Observability & SRE

Metrics, logs, tracing, error budgets, and reliability engineering culture.

  • 19 Tracked terms
  • Last 30 days Feed window

What this topic collects on

An article joins this feed when it matches these terms. Each one is also a search of its own.

Latest in Observability & SRE


dev.to > calderhayes9638 > game-tenant-dns-cutover-pre-change-ttl-scheduling-and-auditable-restore-gpj

Game Tenant DNS Cutover — Pre-change TTL Scheduling and Auditable Restore

46+ min ago   (158+ words) Short answer: model pre-change TTL lowering, DNS cutover, and TTL restoration as separate idempotent states, each with its own deadline, evidence, and retry policy. That distinction matters because a DNS write succeeding is not the same event as every recursive…...


usatoday.com > press-release > story > 41590 > visibility-code-releases-knowledge-engineering-framework-for-visibility-experts

Visibility Code Releases Knowledge Engineering Framework for Visibility Experts

3+ hour, 7+ min ago   (373+ words) The Visibility Code has released a new public framework for Knowledge Engineering for Answer Engines, an emerging publishing discipline focused on making authoritative knowledge easier for AI-mediated systems to retrieve, resolve, represent accurately, attribute, and maintain. Titled “The Visibility Code:…...


dev.to > causely > how-to-check-an-agents-diagnosis-before-it-touches-production-fp6

How to Check an Agent's Diagnosis Before It Touches Production

2+ hour, 2+ min ago   (100+ words) Originally posted to causely.ai by Ben Yemini TL;DR When teams get ready to add agents to... Tagged with ai, devops, kubernetes, observability....


blogs.oracle.com > autonomous-ai-database > turn-any-sql-query-into-an-oci-monitoring-metric-from-inside-autonomous-ai-database

Turn Any SQL Query Into an OCI Monitoring Metric — From Inside Autonomous AI Database

4+ hour, 41+ min ago   (1158+ words) Information, tips, tricks and sample code for data management in an autonomous, cloud-native, data-driven world No agent. No function. No compute instance. Just a SELECT and a scheduler job. Autonomous AI Database gives you more than forty metrics in OCI…...


dev.to > martin_bammer_6838a4d3b65 > fastlogging-rs-high-performance-logging-for-many-different-programming-languages-772

fastlogging-rs: High-Performance Logging for many different Programming Languages

2+ hour, 36+ min ago   (183+ words) Release, 0.9.0, is available with the following features: Compared to Python’s built-in logging module: When your application logs millions of messages, these speedups can turn minutes into seconds. fastlogging-rs is written in Rust and comes with thin wrappers for your favorite…...


gooddata.ai > blog > ai-observability-tools

AI Observability Tools: Do You Need a Separate One?

8+ hour, 31+ min ago   (17+ words) AI Observability Tools: Do You Need a Separate One, or Is It Already in Your Platform? GoodData...


dev.to > kharesam > the-three-events-that-turn-your-logs-into-an-slo-dashboard-1gie

The Three Events That Turn Your Logs Into an SLO Dashboard

3+ hour, 11+ min ago   (306+ words) That gap between "the infrastructure looks healthy" and "the system is doing its job" is where most observability budget quietly goes to waste. And closing it doesn't take a new tool. It takes logging the right three events per request,…...


dev.to > humzakt > four-stops-instead-of-thirty-rebuilding-the-dashboard-and-retiring-hand-rolled-ui-1a35

Four Stops Instead of Thirty: Rebuilding the Dashboard and Retiring Hand-Rolled UI

4+ hour, 42+ min ago   (1695+ words) An editor's real complaint about the main video-generation service's internal tool wasn't any single bug. It was the shape of using it: thirty confirmation stops on an attended run, a dashboard that buried the one thing you actually owed a…...


dev.to > npayyappilly > from-devops-to-sre-a-practitioners-roadmap-for-enterprise-transformation-252h

From DevOps to SRE: A Practitioner's Roadmap for Enterprise Transformation

6+ hour, 55+ min ago   (859+ words) This post maps the journey from DevOps foundations to SRE practice at the enterprise scale. It is not a theoretical model. It is a phased roadmap with specific exit gates for each phase, derived from the observable signals that distinguish…...


dev.to > tacio_souza > laboratorio-kubernetes-multi-no-no-home-lab-a-base-para-ckne-e-kubestronaut-156j

Laboratório Kubernetes Multi-Nó no Home Lab: A Base para CKNE e Kubestronaut

10+ hour, 8+ min ago   (206+ words) Subi meu laboratório de estudos para a jornada Kubernetes e não poderia estar mais animado com o que vem por aí. Na imagem, você vê a base de tudo: um cluster Kubernetes rodando em 3 VMs (1 Control Plane e 2 Workers) gerenciadas…...