{"slug": "debugging-llms-without-guesswork-a-practical-langfuse-tutorial", "title": "Debugging LLMs Without Guesswork: A Practical Langfuse Tutorial", "summary": "A practical Langfuse tutorial demonstrates how to add LLM observability to a LangChain-based support-ticket routing application, tracing each run through ChatPromptTemplate and ChatOpenAI to expose prompt, model call, response, token usage, latency and metadata instead of only the final answer. The walkthrough builds a small app that routes tickets such as a heating complaint or a double-charge complaint to the correct support team, then connects it to Langfuse to inspect runs, add metadata and tags, and compare multiple traces. The stated goal is understanding how observability and API tracing work so the same approach can be applied to more complex RAG, scientific, bioinformatics or production AI workflows.", "body_md": "When I first started building small LLM-powered applications, everything felt quite straightforward. I would write a prompt, send it to a model through an API, get a response back, and move on with a feeling of being in cloud 9. As long as the answer looked reasonable, I did not think too much about what was happening in between.\n\nThat changed very quickly once my projects became even slightly more complex.\n\nI started connecting prompts through LangChain, testing different inputs, adding application logic, and thinking about how these workflows could eventually grow into larger scientific or bioinformatics tools. Then debugging became much less pleasant. Sometimes a response was unexpected. Sometimes one run behaved differently from another. Sometimes I wanted to know whether the issue came from the prompt, the model call, the input itself, or something else in the workflow.\n\n*And the frustrating part was that I could see the final output, but I could not easily see the complete story behind it.*\n\nComing from a bioinformatics background, this felt strangely uncomfortable. In computational biology, we normally want to understand every important stage of an analysis pipeline — the input data, transformations, parameters, intermediate outputs, and final results. We care deeply about reproducibility and traceability. Yet with an LLM application, I was often looking only at the final answer.\n\nThat was when I started exploring the idea of **LLM observability**. While learning more about tracing and monitoring AI applications, I came across **Langfuse**. What immediately caught my attention was the possibility of seeing an LLM request as an actual execution trace rather than just a response printed in the terminal.\n\nInstead of:\n\n```\nInput → Answer\n```\n\nI could begin to see something closer to:\n\n```\nInput  ↓Prompt  ↓Model Call  ↓Response  ↓Token Usage  ↓Latency  ↓Metadata\n```\n\nThat simple change in visibility made the whole idea of debugging LLM applications much more intuitive to me. So I decided to build a very small practice project for myself to understand it properly from the ground up that I would like to share here.\n\nIn this small project, I will walk through that exact exercise step by step. We will create a small LangChain-based support-ticket application, connect it to Langfuse, trace the model calls, inspect what happened inside each run, add useful metadata and tags, and finally compare multiple traces.\n\nThe application itself is intentionally simple. The real goal is to understand **how observability works**, why API tracing matters, and how the same ideas can later be applied to more complex RAG, scientific, bioinformatics, or production AI workflows.\n\nLet us start with the problem that made tracing necessary in the first place.\n\nImagine that we have a simple application that receives a customer-support ticket.\n\nFor example:\n\n**Ticket 1:** *My apartment heating is not working. (OR*)\n\n**Ticket 2:** *I was charged twice for the same payment.*\n\nOur LLM application reads the ticket and decides which support team should receive it. The workflow looks simple:\n\n**User ticket → Prompt → LLM API → Response**\n\nBut suppose something goes wrong. The application sends the heating complaint to the billing department.\n\nNow we need to understand:\n\nIf we only print the final answer in the terminal, answering these questions becomes difficult. This is the fundamental problem that observability tools such as Langfuse help solve.\n\nThe selected ticket is passed through a small LangChain pipeline. The application then asks the model to determine which team should handle the request.\n\nThe architecture is intentionally simple as you can see in the app:\n\n```\nUser Ticket     ↓ChatPromptTemplate     ↓ChatOpenAI     ↓Routing Response\n```\n\nLangChain itself is not the main subject of this walkthrough here. Its role here is simply to connect the prompt to the model and provide a convenient callback mechanism that Langfuse can observe.\n\nThe application uses:\n\n``` python\nfrom langchain_core.prompts import ChatPromptTemplatefrom langchain_openai import ChatOpenAI\n```\n\nThe model is created using:\n\n```\nmodel = ChatOpenAI(model=\"gpt-4o-mini\")\n```\n\nA small prompt is created:\n\n```\nprompt = ChatPromptTemplate.from_messages([    (\"system\", \"You route support tickets to the correct team.\"),    (\"user\", \"Ticket: {ticket}\")])\n```\n\nThen the prompt and model are connected:\n\n```\nchain = prompt | model\n```\n\nConceptually:\n\n```\nTicket  ↓Prompt Template  ↓Prepared Model Messages  ↓OpenAI Model  ↓Response\n```\n\nThis small workflow is perfect for learning tracing because we can clearly compare the code with what later appears inside Langfuse.\n\n``` python\nfrom dotenv import load_dotenvfrom langchain_core.prompts import ChatPromptTemplatefrom langchain_openai import ChatOpenAIfrom openai import APIConnectionError, AuthenticationError, RateLimitErrorload_dotenv()MODEL = \"gpt-4o-mini\"EXAMPLES = {    \"1\": {        \"label\": \"Heating failure\",        \"ticket\": \"My heating has stopped working and the apartment is cold.\",        \"ticket_type\": \"maintenance\",    },    \"2\": {        \"label\": \"Double charge\",        \"ticket\": \"I was charged twice for this month's rent.\",        \"ticket_type\": \"billing\",    },}def build_chain():    \"\"\"Create the small LangChain prompt to model pipeline.\"\"\"    prompt = ChatPromptTemplate.from_messages([        (            \"system\",            \"You route support tickets to the correct team. \"            \"Choose one team: maintenance, billing, account-access, or general-support. \"            \"Return exactly two short lines: TEAM: ... and REASON: ...\",        ),        (\"user\", \"Ticket: {ticket}\"),    ])    model = ChatOpenAI(model=MODEL)    return prompt | modeldef route_ticket(ticket, ticket_type, build_config):    \"\"\"Route one support ticket with optional tracing config.\"\"\"    chain = build_chain()    config = build_config(ticket_type)    if config is None:        config = {}    result = chain.invoke(        {\"ticket\": ticket},        config=config,    )    return result.contentdef choose_ticket():    \"\"\"Ask the learner which prepared ticket to run.\"\"\"    print(\"Choose a support ticket:\")    print(\"1. Heating failure\")    print(\"2. Double charge\")    choice = input(\"Enter 1 or 2: \").strip()    return choice, EXAMPLES.get(choice)def run_app(build_config):    \"\"\"Run the walkthrough app with the supplied config builder.\"\"\"    choice, example = choose_ticket()    if example is None:        print(\"\\nINVALID CHOICE\")        print(\"--------------\")        print(f\"'{choice}' is not one of the available options.\")        print(\"Choose 1 or 2.\")        return    try:        response = route_ticket(            example[\"ticket\"],            example[\"ticket_type\"],            build_config,        )    except AuthenticationError:        print(\"\\nMODEL AUTHENTICATION FAILED\")        print(\"---------------------------\")        print(\"The model client could not authenticate.\")        return    except RateLimitError:        print(\"\\nMODEL RATE LIMIT HIT\")        print(\"--------------------\")        print(\"The model service rejected the request because of rate limits.\")        return    except APIConnectionError:        print(\"\\nMODEL CONNECTION FAILED\")        print(\"-----------------------\")        print(\"The app could not connect to the model service.\")        return    print(\"\\nMODEL ROUTING RESULT\")    print(\"--------------------\")    print(response)\n```\n\nGo to the [*Langfuse Cloud interface*](https://cloud.langfuse.com/) and create an account or sign in.\n\nThen:\n\n1. Create a new project.\n\n2. Open the project settings.\n\n3. Generate API keys.\n\n4. Copy the required Langfuse credentials.\n\nGive your project a name, mine is llm-bio:\n\nThen create your own Langfuse API key:\n\nFor this walkthrough, we need values corresponding to:\n\nLANGFUSE_PUBLIC_KEY (The public key identifies the project).\n\nLANGFUSE_SECRET_KEY (The secret key allows our application to send tracing information).\n\nLANGFUSE_HOST (The host tells the SDK where the trace data should be sent).\n\nThese credentials are Langfuse credentials, not the API key used to call the LLM. Initially I was making the same mistake in my undestanding.\n\nStore Credentials Safely. Instead of placing credentials directly inside the Python script, keep them inside a .env file:\n\n```\nOPENAI_API_KEY=\"sk-xxx\"LANGFUSE_SECRET_KEY=\"sk-xxx\"LANGFUSE_PUBLIC_KEY=\"pk-lf-xxx\"LANGFUSE_BASE_URL=\"https://cloud.langfuse.com\"\n```\n\nNow we connect our LangChain application to Langfuse. The important import is:\n\n``` python\nfrom langfuse.langchain import CallbackHandler\n```\n\nThe CallbackHandler acts as the tracing hook.\n\nWhen LangChain executes the chain, the callback receives information about the run and sends the relevant tracing data to Langfuse.\n\n```\nLangChain Execution        ↓CallbackHandler        ↓Langfuse        ↓Trace Dashboard\n```\n\nNext, I did create a small configuration function.\n\n``` python\nfrom langfuse.langchain import CallbackHandlerfrom app import run_appTRACE_NAME = \"LangFuse Walkthrough Ticket\"def build_config(ticket_type):    \"\"\"Create run config. Add the LangFuse callback here.\"\"\"    langfuse_handler = CallbackHandler()    return {        \"callbacks\": [langfuse_handler],        \"run_name\": TRACE_NAME,        \"metadata\": {            \"ticket_type\": ticket_type,        },        \"tags\": [\"langfuse-walkthrough\"],    }if __name__ == \"__main__\":    run_app(build_config)\n```\n\nThis configuration adds four useful pieces of information to every run.\n\nCallback\n\n```\n\"callbacks\": [langfuse_handler]\n```\n\nThis tells LangChain that the Langfuse tracing callback should observe the execution.\n\nRun name\n\n```\n\"run_name\": TRACE_NAME\n```\n\nThe trace receives a recognizable name.\n\nMetadata\n\n```\n\"metadata\": {    \"ticket_type\": ticket_type}\n```\n\nMetadata adds contextual information to the trace.\n\n```\nticket_type = maintenance\n```\n\nor:\n\n```\nticket_type = billing\n```\n\nTags\n\n```\n\"tags\": [\"langfuse-walkthrough\"]\n```\n\nTags make groups of related traces easier to find later.\n\nWhen the application invokes the LangChain pipeline, the tracing configuration is passed into the execution.\n\n```\nresult = chain.invoke(    {\"ticket\": ticket},    config=build_config(ticket_type),)\n```\n\nThis is an important moment in the workflow.\n\nThe callback becomes attached to the chain execution.\n\nNow the application is still performing the same task as before, but Langfuse is observing what happens.\n\n```\nTicket   ↓chain.invoke()   ↓Prompt   ↓LLM API   ↓Response   ↓Langfuse Trace\n```\n\nThe terminal output may look completely normal.\n\nThe interesting change happens behind the scenes.\n\nA trace is being created.\n\nRun the application from PyCharm. Choose one of the available tickets.\n\n```\nHeating failure\n```\n\nThe application should return the model’s routing decision. At this point, our application has completed its task. But now we have something additional:\n\n**a recorded execution trace.**\n\nInside the Langfuse dashboard, navigate to:\n\n```\nObservability      ↓Tracing\n```\n\nThe latest application run should appear in the trace list. Open the trace. Instead of only seeing the final answer, we can now inspect the full execution story. The trace is effectively the evidence record for this particular request.\n\n*Read the Trace as a Story*\n\nWhen opening the Langfuse dashboard for the first time, there can be a lot of information on the screen. Instead of trying to understand everything immediately, I find it easier to ask a few simple questions.\n\n*Did my run arrive?*\n\nLook at the trace list. If a new trace appeared immediately after running the application, the tracing integration is working.\n\n*What entered the application?*\n\nInspect the input. In this example, we should see the selected support ticket.\n\n*What did the model return?*\n\nInspect the output. We should see the routing answer produced by the model.\n\n*Which model handled the request?*\n\nOpen the model-call section. The trace should indicate the model used for the request.\n\n*How expensive was the request?*\n\nLook at the token-usage information.\n\n*How long did it take?*\n\nCheck the latency.\n\n*What happened first, second, and third?*\n\nInspect the execution timeline. That timeline allows us to reconstruct the application run.\n\nThe practical signals highlighted in the walkthrough include the model input, model response, model name, token consumption, latency, and trace timeline.\n\nOne of the most useful exercises is comparing the Langfuse trace with the application code. Our source code follows:\n\n```\nticket  ↓prompt template  ↓model call  ↓model response\n```\n\nThe Langfuse trace should tell the same story. That is important.\n\nTracing is most useful when we can move smoothly between:\n\n```\nCode ↕Runtime behaviour ↕Trace\n```\n\nIf a problem appears in the trace, we can then return to the corresponding part of the code. This creates a much more systematic debugging workflow.\n\nNow the trace becomes genuinely useful. Instead of just looking at it, we can investigate it. Questions I like to ask include:\n\n```\nWhat exactly was sent to the model?\nWhat response did the model return?\nWhich model processed the request?\nHow long did the call take?\nHow many tokens were consumed?\nDid the execution succeed?\nDoes the output match the expected behaviour?\nWould this trace give me enough information to investigate a user complaint?\n```\n\nThis last question is particularly important. Imagine a user says:\n\n*“Your AI routed my support request incorrectly.”*\n\nWithout observability, we may have little evidence beyond the complaint.\n\nWith Langfuse, we can inspect the exact request that produced the result.\n\nThe walkthrough describes this nicely: the trace allows the developer to inspect the real run instead of relying only on the final response or trying to reconstruct what happened later.\n\nOne of my favourite parts of this exercise is metadata. Our trace configuration includes:\n\n```\n\"metadata\": {    \"ticket_type\": ticket_type}\n```\n\nSuppose we run the application many times. Some traces correspond to:\n\n```\nmaintenance\n```\n\nwhile others correspond to:\n\n```\nbilling\n```\n\nWithout metadata, we may have to open every trace individually. With metadata, we can search or filter by operational context. For example:\n\n```\nShow maintenance traces.\nShow billing traces.\nShow traces from a particular environment.\nShow traces from a specific test case.\n```\n\nThis becomes incredibly useful once an application generates hundreds or thousands of traces. Metadata should contain useful operational labels rather than private or sensitive user information.\n\nExamples include:\n\n```\nticket_type\nenvironment\ntest_case\nexperiment\npipeline_version\n```\n\nThe original exercise specifically uses safe metadata and tags to make traces searchable and easier to organize.\n\nNow run the application again. This time choose the second ticket:\n\n```\nDouble charge\n```\n\nThe application should produce another response and another Langfuse trace.\n\nWe now have two executions:\n\n```\nTrace 1 → Heating failure\nTrace 2 → Double charge\n```\n\nThis is where observability becomes much more interesting. Instead of inspecting one run in isolation, we can compare behaviour across runs.\n\nOpen the maintenance trace and the billing trace. Now compare them.\n\nAsk: *Did both calls use the same model?*\n\nThis confirms that both requests followed the expected application path.\n\nAsk: *Was the correct metadata attached?*\n\nWe should see something like:\n\n```\nmaintenance\n```\n\nfor one trace and:\n\n```\nbilling\n```\n\nfor the other.\n\nAsk: *Did latency change?*\n\nOne request may have taken longer. That could become important when investigating performance problems.\n\nAsk: *Did token usage change?*\n\nDifferent prompts or responses may consume different numbers of tokens. At scale, this directly affects application cost.\n\nAsk: *Was the routing decision correct?*\n\nNow we are beginning to move from simple tracing toward systematic evaluation.\n\nThe practice session uses exactly this comparison to demonstrate how metadata, latency, token usage, and outputs can be compared across runs.\n\nOur support-ticket example is intentionally simple. But imagine applying the same tracing concepts to a production RAG system. A run might look like:\n\n```\nUser Question      ↓Query Processing      ↓Embedding Model      ↓Vector Search      ↓Retrieved Documents      ↓Prompt Construction      ↓LLM API      ↓Generated Answer\n```\n\nNow imagine the answer is wrong. With tracing, we can investigate whether the problem came from:\n\n```\nretrieval\ncontext quality\nprompt construction\nmodel behaviour\ntoken truncation\nlatency\nAPI failure\n```\n\nWithout observability, all of those possibilities are hidden behind one final response.\n\nOne concept that became clearer to me while doing this exercise is that tracing is not simply a debugging feature.\n\nIt also contributes to **reproducibility**.\n\nAfter working through the exercise, this is the mental model I find easiest:\n\n```\nApplication    ↓LLM Run    ↓Trace    ↓Observe    ↓Compare    ↓Debug    ↓Evaluate    ↓Improve\n```\n\nThe first goal is simply visibility. Once we can see what the system is doing, we can start asking better questions about quality, performance, reliability, and cost.\n\nThe most valuable idea in this walkthrough is not the few lines of integration code.\n\nIt is the shift from: *“The model gave me this answer.”*\n\nto:\n\n*“I can inspect exactly how my application produced this answer.”*\n\nThat distinction becomes increasingly important as AI systems move from notebooks and experiments into applications used by researchers, clinicians, developers, companies, and eventually end users. This walkthrough captures this transition very well: metadata and filtering transform a dashboard from a simple list of model calls into something that can actually be investigated.\n\nThis is a small example, but the same principle scales naturally toward much larger applications.\n\nLangfuse gives developers and engineers something that becomes increasingly valuable as LLM applications grow: **visibility**. In this small experiment, I only traced two support tickets. But the same concept can be extended to RAG systems, scientific assistants, biomedical knowledge platforms, agentic applications, drug-discovery tools, and production AI services.\n\nFor me, the key lesson is simple:\n\n*“Do not treat an LLM call as a black box.”*\n\nOnce we can trace the inputs, prompts, models, outputs, latency, token usage, metadata, and execution sequence, debugging becomes more systematic and application behaviour becomes much easier to understand. And that is the first step toward building AI systems that are not only intelligent, but also **observable, reproducible, and maintainable**.\n\n[Debugging LLMs Without Guesswork: A Practical Langfuse Tutorial](https://pub.towardsai.net/debugging-llms-without-guesswork-a-practical-langfuse-tutorial-03cb26358f5d) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/debugging-llms-without-guesswork-a-practical-langfuse-tutorial", "canonical_source": "https://pub.towardsai.net/debugging-llms-without-guesswork-a-practical-langfuse-tutorial-03cb26358f5d?source=rss----98111c9905da---4", "published_at": "2026-10-03 12:39:35+00:00", "updated_at": "2026-10-03 13:07:27.535071+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "developer-tools", "mlops", "ai-agents"], "entities": ["Langfuse", "LangChain", "ChatPromptTemplate", "ChatOpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/debugging-llms-without-guesswork-a-practical-langfuse-tutorial", "markdown": "https://wpnews.pro/news/debugging-llms-without-guesswork-a-practical-langfuse-tutorial.md", "text": "https://wpnews.pro/news/debugging-llms-without-guesswork-a-practical-langfuse-tutorial.txt", "jsonld": "https://wpnews.pro/news/debugging-llms-without-guesswork-a-practical-langfuse-tutorial.jsonld"}}