From Telemetry to Automated Incident Response: My DevOps & Observability Journey A developer built an observability and automated incident-response pipeline for a FastAPI order-tracking application using OpenTelemetry, Prometheus, Loki, Tempo, Grafana, and Codex CLI. The system routes OTLP telemetry through an OpenTelemetry Collector to separate metrics, logs, and traces stores, fires Grafana alerts only on 5xx responses, and forwards incidents to a dedicated responder service that gathers evidence and hands investigation context to a coding agent. The developer's stated lesson is that observability must connect detection, investigation, remediation, and verification rather than just collect data. Description: A hands-on learning journey building observability and automated incident response into a FastAPI order-tracking application using OpenTelemetry, Prometheus, Loki, Tempo, Grafana, and Codex CLI. As part of the DataTalksClub AI Dev Tools Zoomcamp , I worked on a small but practical challenge: turning a simple FastAPI application into an observable system with automated incident response. The project started as a basic Order Tracker application backed by SQLite. By the end, it had evolved into a system that could: The most valuable lesson for me was that observability is not just about collecting data. It is about connecting detection → investigation → remediation → verification . A basic application can work perfectly during development and still fail in production. For an order-tracking API, a failure might look simple: GET /api/orders/express-1002 returns a server error. Without observability, the workflow often becomes: Something is broken ↓ Check application logs ↓ Search for the error ↓ Try to reproduce it ↓ Find the problematic code ↓ Apply a fix ↓ Restart the application ↓ Test again This works, but it becomes increasingly difficult as the system grows. I wanted to build a workflow where telemetry could help connect these steps. The final architecture introduced several components: This architecture separates telemetry collection, storage, visualization, alerting, and incident response. The first step was adding OpenTelemetry instrumentation to the FastAPI application. I wanted each order lookup to produce useful telemetry, including: For example: GET /api/orders/standard-1001 → 200 and: GET /api/orders/standard-1002 → 404 The important part was not simply knowing that a request happened. The telemetry needed enough context to answer: What endpoint failed, what was the HTTP status, and what happened during the request? The next step was connecting the application to an observability stack: The OpenTelemetry Collector became the central telemetry pipeline. Conceptually: Application │ │ OTLP ▼ OpenTelemetry Collector │ ├── Metrics → Prometheus │ ├── Logs → Loki │ └── Traces → Tempo Grafana then provides a single interface for exploring these signals. This was one of the important lessons from the project: Metrics tell you that something is happening. Logs help explain what happened. Traces help show where it happened. Using the three together provides much more useful operational context than relying on a single signal. After telemetry was available, I configured Grafana alerting for server-side failures. The goal was to detect 5xx responses . A 404 response such as: GET /api/orders/standard-1002 → 404 should not trigger the 5xx incident workflow. This distinction was useful because it forced me to think about alert conditions rather than simply alerting on every failed request. The alert also needed to provide enough information for the next stage of the workflow. The next step was adding a dedicated incident-response service. The service exposes: POST /alerts When Grafana sends an alert, the responder stores the incident information and gathers useful evidence. The idea is to transform: Grafana Alert into: Incident ├── endpoint ├── status ├── logs ├── traces └── investigation context This creates a bridge between observability and automated remediation. The most interesting part of the exercise was connecting the incident workflow to a coding agent. Instead of simply notifying a developer: 🚨 500 error detected the system can provide the coding agent with context about the incident. The intended workflow becomes: 5xx detected ↓ Grafana alert ↓ Incident-response service ↓ Collect evidence ↓ Coding agent investigates ↓ Identify root cause ↓ Apply fix ↓ Restart application ↓ Verify behavior This changes the role of AI from simply generating code to participating in an operational workflow. The final exercise intentionally exposed a bug in the Express order lookup. The problematic implementation calculated the estimated delivery date using: estimated at = placed at.replace day=placed at.day + 2 This looks reasonable at first glance. However, it assumes that the resulting day exists in the same month. For an order created near the end of a month, that assumption can fail. The fix was to use date arithmetic instead: estimated at = placed at + timedelta days=2 This allows Python's datetime implementation to correctly handle month boundaries. The important part was not only finding the line of code. The observability and incident-response workflow provided the path from: HTTP failure ↓ Telemetry ↓ Alert ↓ Incident evidence ↓ Root-cause investigation ↓ Code change ↓ Verification Before this project, it was easy to think of observability as: "Add some logs." The project changed that perspective. A useful observability system combines multiple signals and makes them actionable. An alert saying: Something went wrong is not particularly useful. An actionable incident should provide enough context to start an investigation. Endpoint HTTP status Logs Trace Time window Dashboard context Automatically changing code is not enough. A remediation workflow should also verify that the application works after the change. That makes the workflow closer to: Detect → Investigate → Fix → Verify rather than simply: Detect → Fix A coding agent becomes much more useful when it receives evidence from the system instead of being asked to investigate blindly. Telemetry can provide the context needed for the agent to understand what actually happened. The project brought several tools together: | Area | Technology | |---|---| | API | FastAPI | | Runtime | Python | | Database | SQLite | | Telemetry | OpenTelemetry | | Telemetry Pipeline | OpenTelemetry Collector | | Metrics | Prometheus | | Logs | Loki | | Traces | Tempo | | Visualization | Grafana | | Containers | Docker / Docker Compose | | AI-assisted remediation | Codex CLI | The completed workflow was validated through the homework tasks. | Step | Scenario | Result | |---|---|---| | Q1 | Application health check | 200 | | Q2 | Existing order lookup | 200 | | Q3 | Missing order lookup | 404 | | Q4 | 5xx alert behavior | Normal for 404 | | Q5 | Incident-response workflow | Completed | | Q6 | Express order incident | Root cause identified and fixed | The important outcome was not just that each individual component worked. The complete workflow worked as a chain. The biggest lesson from this project was that observability becomes much more valuable when it is connected to action . A mature workflow can look like: Application ↓ Telemetry ↓ Observability ↓ Alerting ↓ Incident Response ↓ AI-assisted Investigation ↓ Remediation ↓ Verification This project gave me hands-on experience connecting these pieces together rather than learning them as isolated technologies. For me, that's the real value of Learning in Public : documenting not only what I built, but also the problems I encountered, the root causes I found, and what changed in my understanding along the way. This project was completed as part of the DataTalksClub AI Dev Tools Zoomcamp — Homework 4: DevOps and Observability for AI-Built Apps . Thanks to the DataTalksClub community for creating practical exercises that connect software development, DevOps, observability, and AI-assisted engineering. Github: order-tracker-v2 https://github.com/ketut-garjita/order-tracker-v2