{"slug": "from-telemetry-to-automated-incident-response-my-devops-observability-journey", "title": "From Telemetry to Automated Incident Response: My DevOps & Observability Journey", "summary": "A developer built an observability and automated incident-response pipeline for a FastAPI order-tracking application using OpenTelemetry, Prometheus, Loki, Tempo, Grafana, and Codex CLI. The system routes OTLP telemetry through an OpenTelemetry Collector to separate metrics, logs, and traces stores, fires Grafana alerts only on 5xx responses, and forwards incidents to a dedicated responder service that gathers evidence and hands investigation context to a coding agent. The developer's stated lesson is that observability must connect detection, investigation, remediation, and verification rather than just collect data.", "body_md": "*Description: A hands-on learning journey building observability and automated incident response into a FastAPI order-tracking application using OpenTelemetry, Prometheus, Loki, Tempo, Grafana, and Codex CLI.*\n\nAs part of the **DataTalksClub AI Dev Tools Zoomcamp**, I worked on a small but practical challenge: turning a simple FastAPI application into an observable system with automated incident response.\n\nThe project started as a basic **Order Tracker** application backed by SQLite.\n\nBy the end, it had evolved into a system that could:\n\nThe most valuable lesson for me was that observability is not just about collecting data.\n\nIt is about connecting **detection → investigation → remediation → verification**.\n\nA basic application can work perfectly during development and still fail in production.\n\nFor an order-tracking API, a failure might look simple:\n\n```\nGET /api/orders/express-1002\n```\n\nreturns a server error.\n\nWithout observability, the workflow often becomes:\n\n```\nSomething is broken\n        ↓\nCheck application logs\n        ↓\nSearch for the error\n        ↓\nTry to reproduce it\n        ↓\nFind the problematic code\n        ↓\nApply a fix\n        ↓\nRestart the application\n        ↓\nTest again\n```\n\nThis works, but it becomes increasingly difficult as the system grows.\n\nI wanted to build a workflow where telemetry could help connect these steps.\n\nThe final architecture introduced several components:\n\nThis architecture separates telemetry collection, storage, visualization, alerting, and incident response.\n\nThe first step was adding OpenTelemetry instrumentation to the FastAPI application.\n\nI wanted each order lookup to produce useful telemetry, including:\n\nFor example:\n\n```\nGET /api/orders/standard-1001\n→ 200\n```\n\nand:\n\n```\nGET /api/orders/standard-1002\n→ 404\n```\n\nThe important part was not simply knowing that a request happened.\n\nThe telemetry needed enough context to answer:\n\nWhat endpoint failed, what was the HTTP status, and what happened during the request?\n\nThe next step was connecting the application to an observability stack:\n\nThe OpenTelemetry Collector became the central telemetry pipeline.\n\nConceptually:\n\n```\nApplication\n    │\n    │ OTLP\n    ▼\nOpenTelemetry Collector\n    │\n    ├── Metrics → Prometheus\n    │\n    ├── Logs    → Loki\n    │\n    └── Traces  → Tempo\n```\n\nGrafana then provides a single interface for exploring these signals.\n\nThis was one of the important lessons from the project:\n\nMetrics tell you that something is happening.\n\nLogs help explain what happened.\n\nTraces help show where it happened.\n\nUsing the three together provides much more useful operational context than relying on a single signal.\n\nAfter telemetry was available, I configured Grafana alerting for server-side failures.\n\nThe goal was to detect **5xx responses**.\n\nA 404 response such as:\n\n```\nGET /api/orders/standard-1002\n→ 404\n```\n\nshould not trigger the 5xx incident workflow.\n\nThis distinction was useful because it forced me to think about alert conditions rather than simply alerting on every failed request.\n\nThe alert also needed to provide enough information for the next stage of the workflow.\n\nThe next step was adding a dedicated incident-response service.\n\nThe service exposes:\n\n```\nPOST /alerts\n```\n\nWhen Grafana sends an alert, the responder stores the incident information and gathers useful evidence.\n\nThe idea is to transform:\n\n```\nGrafana Alert\n```\n\ninto:\n\n```\nIncident\n├── endpoint\n├── status\n├── logs\n├── traces\n└── investigation context\n```\n\nThis creates a bridge between observability and automated remediation.\n\nThe most interesting part of the exercise was connecting the incident workflow to a coding agent.\n\nInstead of simply notifying a developer:\n\n```\n🚨 500 error detected\n```\n\nthe system can provide the coding agent with context about the incident.\n\nThe intended workflow becomes:\n\n```\n5xx detected\n     ↓\nGrafana alert\n     ↓\nIncident-response service\n     ↓\nCollect evidence\n     ↓\nCoding agent investigates\n     ↓\nIdentify root cause\n     ↓\nApply fix\n     ↓\nRestart application\n     ↓\nVerify behavior\n```\n\nThis changes the role of AI from simply generating code to participating in an operational workflow.\n\nThe final exercise intentionally exposed a bug in the Express order lookup.\n\nThe problematic implementation calculated the estimated delivery date using:\n\n```\nestimated_at = placed_at.replace(day=placed_at.day + 2)\n```\n\nThis looks reasonable at first glance.\n\nHowever, it assumes that the resulting day exists in the same month.\n\nFor an order created near the end of a month, that assumption can fail.\n\nThe fix was to use date arithmetic instead:\n\n```\nestimated_at = placed_at + timedelta(days=2)\n```\n\nThis allows Python's datetime implementation to correctly handle month boundaries.\n\nThe important part was not only finding the line of code.\n\nThe observability and incident-response workflow provided the path from:\n\n```\nHTTP failure\n    ↓\nTelemetry\n    ↓\nAlert\n    ↓\nIncident evidence\n    ↓\nRoot-cause investigation\n    ↓\nCode change\n    ↓\nVerification\n```\n\nBefore this project, it was easy to think of observability as:\n\n```\n\"Add some logs.\"\n```\n\nThe project changed that perspective.\n\nA useful observability system combines multiple signals and makes them actionable.\n\nAn alert saying:\n\n```\nSomething went wrong\n```\n\nis not particularly useful.\n\nAn actionable incident should provide enough context to start an investigation.\n\n```\nEndpoint\nHTTP status\nLogs\nTrace\nTime window\nDashboard context\n```\n\nAutomatically changing code is not enough.\n\nA remediation workflow should also verify that the application works after the change.\n\nThat makes the workflow closer to:\n\n```\nDetect → Investigate → Fix → Verify\n```\n\nrather than simply:\n\n```\nDetect → Fix\n```\n\nA coding agent becomes much more useful when it receives evidence from the system instead of being asked to investigate blindly.\n\nTelemetry can provide the context needed for the agent to understand what actually happened.\n\nThe project brought several tools together:\n\n| Area | Technology | \n|---|---|\n| API | FastAPI | \n| Runtime | Python | \n| Database | SQLite | \n| Telemetry | OpenTelemetry | \n| Telemetry Pipeline | OpenTelemetry Collector | \n| Metrics | Prometheus | \n| Logs | Loki | \n| Traces | Tempo | \n| Visualization | Grafana | \n| Containers | Docker / Docker Compose | \n| AI-assisted remediation | Codex CLI | \n\nThe completed workflow was validated through the homework tasks.\n\n| Step | Scenario | Result | \n|---|---|---|\n| Q1 | Application health check | `200` | \n| Q2 | Existing order lookup | `200` | \n| Q3 | Missing order lookup | `404` | \n| Q4 | 5xx alert behavior | Normal for 404 | \n| Q5 | Incident-response workflow | Completed | \n| Q6 | Express order incident | Root cause identified and fixed | \n\nThe important outcome was not just that each individual component worked.\n\nThe complete workflow worked as a chain.\n\nThe biggest lesson from this project was that **observability becomes much more valuable when it is connected to action**.\n\nA mature workflow can look like:\n\n```\nApplication\n     ↓\nTelemetry\n     ↓\nObservability\n     ↓\nAlerting\n     ↓\nIncident Response\n     ↓\nAI-assisted Investigation\n     ↓\nRemediation\n     ↓\nVerification\n```\n\nThis project gave me hands-on experience connecting these pieces together rather than learning them as isolated technologies.\n\nFor me, that's the real value of **Learning in Public**: documenting not only what I built, but also the problems I encountered, the root causes I found, and what changed in my understanding along the way.\n\nThis project was completed as part of the **DataTalksClub AI Dev Tools Zoomcamp — Homework 4: DevOps and Observability for AI-Built Apps**.\n\nThanks to the DataTalksClub community for creating practical exercises that connect software development, DevOps, observability, and AI-assisted engineering.\n\nGithub: [order-tracker-v2](https://github.com/ketut-garjita/order-tracker-v2)", "url": "https://wpnews.pro/news/from-telemetry-to-automated-incident-response-my-devops-observability-journey", "canonical_source": "https://dev.to/esadata/from-telemetry-to-automated-incident-response-my-devops-observability-journey-2ob5", "published_at": "2026-10-03 03:49:54+00:00", "updated_at": "2026-10-03 04:08:13.574967+00:00", "lang": "en", "topics": ["ai-agents", "mlops", "developer-tools", "ai-tools"], "entities": ["FastAPI", "OpenTelemetry", "Prometheus", "Loki", "Tempo", "Grafana", "Codex CLI", "DataTalksClub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/from-telemetry-to-automated-incident-response-my-devops-observability-journey", "markdown": "https://wpnews.pro/news/from-telemetry-to-automated-incident-response-my-devops-observability-journey.md", "text": "https://wpnews.pro/news/from-telemetry-to-automated-incident-response-my-devops-observability-journey.txt", "jsonld": "https://wpnews.pro/news/from-telemetry-to-automated-incident-response-my-devops-observability-journey.jsonld"}}