Debugging LLMs Without Guesswork: A Practical Langfuse Tutorial A practical Langfuse tutorial demonstrates how to add LLM observability to a LangChain-based support-ticket routing application, tracing each run through ChatPromptTemplate and ChatOpenAI to expose prompt, model call, response, token usage, latency and metadata instead of only the final answer. The walkthrough builds a small app that routes tickets such as a heating complaint or a double-charge complaint to the correct support team, then connects it to Langfuse to inspect runs, add metadata and tags, and compare multiple traces. The stated goal is understanding how observability and API tracing work so the same approach can be applied to more complex RAG, scientific, bioinformatics or production AI workflows. When I first started building small LLM-powered applications, everything felt quite straightforward. I would write a prompt, send it to a model through an API, get a response back, and move on with a feeling of being in cloud 9. As long as the answer looked reasonable, I did not think too much about what was happening in between. That changed very quickly once my projects became even slightly more complex. I started connecting prompts through LangChain, testing different inputs, adding application logic, and thinking about how these workflows could eventually grow into larger scientific or bioinformatics tools. Then debugging became much less pleasant. Sometimes a response was unexpected. Sometimes one run behaved differently from another. Sometimes I wanted to know whether the issue came from the prompt, the model call, the input itself, or something else in the workflow. And the frustrating part was that I could see the final output, but I could not easily see the complete story behind it. Coming from a bioinformatics background, this felt strangely uncomfortable. In computational biology, we normally want to understand every important stage of an analysis pipeline — the input data, transformations, parameters, intermediate outputs, and final results. We care deeply about reproducibility and traceability. Yet with an LLM application, I was often looking only at the final answer. That was when I started exploring the idea of LLM observability . While learning more about tracing and monitoring AI applications, I came across Langfuse . What immediately caught my attention was the possibility of seeing an LLM request as an actual execution trace rather than just a response printed in the terminal. Instead of: Input → Answer I could begin to see something closer to: Input ↓Prompt ↓Model Call ↓Response ↓Token Usage ↓Latency ↓Metadata That simple change in visibility made the whole idea of debugging LLM applications much more intuitive to me. So I decided to build a very small practice project for myself to understand it properly from the ground up that I would like to share here. In this small project, I will walk through that exact exercise step by step. We will create a small LangChain-based support-ticket application, connect it to Langfuse, trace the model calls, inspect what happened inside each run, add useful metadata and tags, and finally compare multiple traces. The application itself is intentionally simple. The real goal is to understand how observability works , why API tracing matters, and how the same ideas can later be applied to more complex RAG, scientific, bioinformatics, or production AI workflows. Let us start with the problem that made tracing necessary in the first place. Imagine that we have a simple application that receives a customer-support ticket. For example: Ticket 1: My apartment heating is not working. OR Ticket 2: I was charged twice for the same payment. Our LLM application reads the ticket and decides which support team should receive it. The workflow looks simple: User ticket → Prompt → LLM API → Response But suppose something goes wrong. The application sends the heating complaint to the billing department. Now we need to understand: If we only print the final answer in the terminal, answering these questions becomes difficult. This is the fundamental problem that observability tools such as Langfuse help solve. The selected ticket is passed through a small LangChain pipeline. The application then asks the model to determine which team should handle the request. The architecture is intentionally simple as you can see in the app: User Ticket ↓ChatPromptTemplate ↓ChatOpenAI ↓Routing Response LangChain itself is not the main subject of this walkthrough here. Its role here is simply to connect the prompt to the model and provide a convenient callback mechanism that Langfuse can observe. The application uses: python from langchain core.prompts import ChatPromptTemplatefrom langchain openai import ChatOpenAI The model is created using: model = ChatOpenAI model="gpt-4o-mini" A small prompt is created: prompt = ChatPromptTemplate.from messages "system", "You route support tickets to the correct team." , "user", "Ticket: {ticket}" Then the prompt and model are connected: chain = prompt | model Conceptually: Ticket ↓Prompt Template ↓Prepared Model Messages ↓OpenAI Model ↓Response This small workflow is perfect for learning tracing because we can clearly compare the code with what later appears inside Langfuse. python from dotenv import load dotenvfrom langchain core.prompts import ChatPromptTemplatefrom langchain openai import ChatOpenAIfrom openai import APIConnectionError, AuthenticationError, RateLimitErrorload dotenv MODEL = "gpt-4o-mini"EXAMPLES = { "1": { "label": "Heating failure", "ticket": "My heating has stopped working and the apartment is cold.", "ticket type": "maintenance", }, "2": { "label": "Double charge", "ticket": "I was charged twice for this month's rent.", "ticket type": "billing", },}def build chain : """Create the small LangChain prompt to model pipeline.""" prompt = ChatPromptTemplate.from messages "system", "You route support tickets to the correct team. " "Choose one team: maintenance, billing, account-access, or general-support. " "Return exactly two short lines: TEAM: ... and REASON: ...", , "user", "Ticket: {ticket}" , model = ChatOpenAI model=MODEL return prompt | modeldef route ticket ticket, ticket type, build config : """Route one support ticket with optional tracing config.""" chain = build chain config = build config ticket type if config is None: config = {} result = chain.invoke {"ticket": ticket}, config=config, return result.contentdef choose ticket : """Ask the learner which prepared ticket to run.""" print "Choose a support ticket:" print "1. Heating failure" print "2. Double charge" choice = input "Enter 1 or 2: " .strip return choice, EXAMPLES.get choice def run app build config : """Run the walkthrough app with the supplied config builder.""" choice, example = choose ticket if example is None: print "\nINVALID CHOICE" print "--------------" print f"'{choice}' is not one of the available options." print "Choose 1 or 2." return try: response = route ticket example "ticket" , example "ticket type" , build config, except AuthenticationError: print "\nMODEL AUTHENTICATION FAILED" print "---------------------------" print "The model client could not authenticate." return except RateLimitError: print "\nMODEL RATE LIMIT HIT" print "--------------------" print "The model service rejected the request because of rate limits." return except APIConnectionError: print "\nMODEL CONNECTION FAILED" print "-----------------------" print "The app could not connect to the model service." return print "\nMODEL ROUTING RESULT" print "--------------------" print response Go to the Langfuse Cloud interface https://cloud.langfuse.com/ and create an account or sign in. Then: 1. Create a new project. 2. Open the project settings. 3. Generate API keys. 4. Copy the required Langfuse credentials. Give your project a name, mine is llm-bio: Then create your own Langfuse API key: For this walkthrough, we need values corresponding to: LANGFUSE PUBLIC KEY The public key identifies the project . LANGFUSE SECRET KEY The secret key allows our application to send tracing information . LANGFUSE HOST The host tells the SDK where the trace data should be sent . These credentials are Langfuse credentials, not the API key used to call the LLM. Initially I was making the same mistake in my undestanding. Store Credentials Safely. Instead of placing credentials directly inside the Python script, keep them inside a .env file: OPENAI API KEY="sk-xxx"LANGFUSE SECRET KEY="sk-xxx"LANGFUSE PUBLIC KEY="pk-lf-xxx"LANGFUSE BASE URL="https://cloud.langfuse.com" Now we connect our LangChain application to Langfuse. The important import is: python from langfuse.langchain import CallbackHandler The CallbackHandler acts as the tracing hook. When LangChain executes the chain, the callback receives information about the run and sends the relevant tracing data to Langfuse. LangChain Execution ↓CallbackHandler ↓Langfuse ↓Trace Dashboard Next, I did create a small configuration function. python from langfuse.langchain import CallbackHandlerfrom app import run appTRACE NAME = "LangFuse Walkthrough Ticket"def build config ticket type : """Create run config. Add the LangFuse callback here.""" langfuse handler = CallbackHandler return { "callbacks": langfuse handler , "run name": TRACE NAME, "metadata": { "ticket type": ticket type, }, "tags": "langfuse-walkthrough" , }if name == " main ": run app build config This configuration adds four useful pieces of information to every run. Callback "callbacks": langfuse handler This tells LangChain that the Langfuse tracing callback should observe the execution. Run name "run name": TRACE NAME The trace receives a recognizable name. Metadata "metadata": { "ticket type": ticket type} Metadata adds contextual information to the trace. ticket type = maintenance or: ticket type = billing Tags "tags": "langfuse-walkthrough" Tags make groups of related traces easier to find later. When the application invokes the LangChain pipeline, the tracing configuration is passed into the execution. result = chain.invoke {"ticket": ticket}, config=build config ticket type , This is an important moment in the workflow. The callback becomes attached to the chain execution. Now the application is still performing the same task as before, but Langfuse is observing what happens. Ticket ↓chain.invoke ↓Prompt ↓LLM API ↓Response ↓Langfuse Trace The terminal output may look completely normal. The interesting change happens behind the scenes. A trace is being created. Run the application from PyCharm. Choose one of the available tickets. Heating failure The application should return the model’s routing decision. At this point, our application has completed its task. But now we have something additional: a recorded execution trace. Inside the Langfuse dashboard, navigate to: Observability ↓Tracing The latest application run should appear in the trace list. Open the trace. Instead of only seeing the final answer, we can now inspect the full execution story. The trace is effectively the evidence record for this particular request. Read the Trace as a Story When opening the Langfuse dashboard for the first time, there can be a lot of information on the screen. Instead of trying to understand everything immediately, I find it easier to ask a few simple questions. Did my run arrive? Look at the trace list. If a new trace appeared immediately after running the application, the tracing integration is working. What entered the application? Inspect the input. In this example, we should see the selected support ticket. What did the model return? Inspect the output. We should see the routing answer produced by the model. Which model handled the request? Open the model-call section. The trace should indicate the model used for the request. How expensive was the request? Look at the token-usage information. How long did it take? Check the latency. What happened first, second, and third? Inspect the execution timeline. That timeline allows us to reconstruct the application run. The practical signals highlighted in the walkthrough include the model input, model response, model name, token consumption, latency, and trace timeline. One of the most useful exercises is comparing the Langfuse trace with the application code. Our source code follows: ticket ↓prompt template ↓model call ↓model response The Langfuse trace should tell the same story. That is important. Tracing is most useful when we can move smoothly between: Code ↕Runtime behaviour ↕Trace If a problem appears in the trace, we can then return to the corresponding part of the code. This creates a much more systematic debugging workflow. Now the trace becomes genuinely useful. Instead of just looking at it, we can investigate it. Questions I like to ask include: What exactly was sent to the model? What response did the model return? Which model processed the request? How long did the call take? How many tokens were consumed? Did the execution succeed? Does the output match the expected behaviour? Would this trace give me enough information to investigate a user complaint? This last question is particularly important. Imagine a user says: “Your AI routed my support request incorrectly.” Without observability, we may have little evidence beyond the complaint. With Langfuse, we can inspect the exact request that produced the result. The walkthrough describes this nicely: the trace allows the developer to inspect the real run instead of relying only on the final response or trying to reconstruct what happened later. One of my favourite parts of this exercise is metadata. Our trace configuration includes: "metadata": { "ticket type": ticket type} Suppose we run the application many times. Some traces correspond to: maintenance while others correspond to: billing Without metadata, we may have to open every trace individually. With metadata, we can search or filter by operational context. For example: Show maintenance traces. Show billing traces. Show traces from a particular environment. Show traces from a specific test case. This becomes incredibly useful once an application generates hundreds or thousands of traces. Metadata should contain useful operational labels rather than private or sensitive user information. Examples include: ticket type environment test case experiment pipeline version The original exercise specifically uses safe metadata and tags to make traces searchable and easier to organize. Now run the application again. This time choose the second ticket: Double charge The application should produce another response and another Langfuse trace. We now have two executions: Trace 1 → Heating failure Trace 2 → Double charge This is where observability becomes much more interesting. Instead of inspecting one run in isolation, we can compare behaviour across runs. Open the maintenance trace and the billing trace. Now compare them. Ask: Did both calls use the same model? This confirms that both requests followed the expected application path. Ask: Was the correct metadata attached? We should see something like: maintenance for one trace and: billing for the other. Ask: Did latency change? One request may have taken longer. That could become important when investigating performance problems. Ask: Did token usage change? Different prompts or responses may consume different numbers of tokens. At scale, this directly affects application cost. Ask: Was the routing decision correct? Now we are beginning to move from simple tracing toward systematic evaluation. The practice session uses exactly this comparison to demonstrate how metadata, latency, token usage, and outputs can be compared across runs. Our support-ticket example is intentionally simple. But imagine applying the same tracing concepts to a production RAG system. A run might look like: User Question ↓Query Processing ↓Embedding Model ↓Vector Search ↓Retrieved Documents ↓Prompt Construction ↓LLM API ↓Generated Answer Now imagine the answer is wrong. With tracing, we can investigate whether the problem came from: retrieval context quality prompt construction model behaviour token truncation latency API failure Without observability, all of those possibilities are hidden behind one final response. One concept that became clearer to me while doing this exercise is that tracing is not simply a debugging feature. It also contributes to reproducibility . After working through the exercise, this is the mental model I find easiest: Application ↓LLM Run ↓Trace ↓Observe ↓Compare ↓Debug ↓Evaluate ↓Improve The first goal is simply visibility. Once we can see what the system is doing, we can start asking better questions about quality, performance, reliability, and cost. The most valuable idea in this walkthrough is not the few lines of integration code. It is the shift from: “The model gave me this answer.” to: “I can inspect exactly how my application produced this answer.” That distinction becomes increasingly important as AI systems move from notebooks and experiments into applications used by researchers, clinicians, developers, companies, and eventually end users. This walkthrough captures this transition very well: metadata and filtering transform a dashboard from a simple list of model calls into something that can actually be investigated. This is a small example, but the same principle scales naturally toward much larger applications. Langfuse gives developers and engineers something that becomes increasingly valuable as LLM applications grow: visibility . In this small experiment, I only traced two support tickets. But the same concept can be extended to RAG systems, scientific assistants, biomedical knowledge platforms, agentic applications, drug-discovery tools, and production AI services. For me, the key lesson is simple: “Do not treat an LLM call as a black box.” Once we can trace the inputs, prompts, models, outputs, latency, token usage, metadata, and execution sequence, debugging becomes more systematic and application behaviour becomes much easier to understand. And that is the first step toward building AI systems that are not only intelligent, but also observable, reproducible, and maintainable . Debugging LLMs Without Guesswork: A Practical Langfuse Tutorial https://pub.towardsai.net/debugging-llms-without-guesswork-a-practical-langfuse-tutorial-03cb26358f5d was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.