Tag

#observability

Tagged “observability”

9 articles
Incident response

Google Cloud publishes a six-phase playbook for its own outages

Google Cloud has documented a five-step Verify to Review loop, plus a Prepare phase, for how customers should respond when its own services fall over. The guidance tells you which of its consoles to check first.

Sep 17, 2026 · Maya Okonkwo
Runners & infrastructure

Kubernetes 1.37 turns native histograms on by default for its own metrics

The Prometheus native-histogram format graduates to beta in Kubernetes 1.37 and ships enabled by default, promising smaller time-series and sharper latency percentiles for the metrics the control plane already emits.

Sep 12, 2026 · Maya Okonkwo
Platform engineering

Apica bolts a natural-language agent and MCP server onto Ascent 3.0

Ascent 3.0 lets operators reshape telemetry pipelines in English via an agent called Venn, and opens the platform's own AI to outside tools through an MCP server. Destructive edits still gate on admin approval.

Sep 9, 2026 · Maya Okonkwo
CI observability

GitHub Actions tracing without editing a single workflow file

A CNCF write-up walks through wiring the OpenTelemetry Collector's alpha githubreceiver to org-level Actions webhooks, turning workflow_run and workflow_job events straight into OTLP spans.

Sep 8, 2026 · Priya Nair
Developer experience

Looker starts running the 'which slice moved' hunt before you open the dashboard

Looker Agentic Workflows, a preview capability available in Looker version 26.08 and later, wires a background agent to a business metric and, when a threshold is crossed, runs a Key Driver Analysis over the underlying model and posts the diagnostic summary to Slack or email with a link back to Conversational Analytics. Setup is by natural-language prompt, gated on the chat_with_agent and create_alerts permissions.

Aug 3, 2026 · Priya Nair
Incident response

Building your own AI SRE moves the toil; it does not remove it

Chronosphere leaders argued in The New Stack that engineering teams should build their own AI SRE to map systems, investigate incidents and support reliable delivery at scale. The operational read is narrower: an in-house AI SRE is a second production system with its own on-call, its own change-management story and its own place in the audit trail.

Aug 2, 2026 · Maya Okonkwo
Platform engineering

GitHub lets enterprises pin Copilot's OpenTelemetry endpoint

A new enterprise-managed setting mandates where the Copilot Chat extension in VS Code and Copilot CLI send OpenTelemetry data. The managed value overrides environment variables and user settings.

Jul 12, 2026 · Maya Okonkwo
Platform engineering

GitHub streams Copilot agent sessions to the SIEM, in preview

GitHub put Copilot agent session streaming into public preview on July 2, exposing prompts, responses and tool calls from cloud agents and IDE clients to enterprise event collectors. For CI-adjacent workflows it turns Copilot from a black box into something an on-call can reconstruct after the fact.

Jul 7, 2026 · Maya Okonkwo
Code quality & testing

Lightrun brings production impact into the PR review, and the merge gate gets more interesting

Lightrun added a capability that assesses a pull request's likely production impact before merge, joining a small category of tools that pull runtime context into code review. A step forward for PR review as a real signal, with some honest caveats about what 'runtime-aware' actually delivers today.

Jul 6, 2026 · Priya Nair