{"slug": "building-reliable-data-analytics-agents-lessons-from-the-kdd-cup", "title": "Building Reliable Data Analytics Agents: Lessons from the KDD Cup", "summary": "NVIDIA's KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around making the agent harness smaller, clearer, and easier to verify, the team reported. The competition required agents to answer natural-language questions across databases, CSV and JSON files, prose documents, PDFs, and briefing videos using a small, fixed LLM (Qwen3.5-35B-A3B), which made the harness the main optimization surface. KGMON's two guiding principles were constraining the action space — unifying CSV and JSON into one SQLite-backed SQL interface with schema() and sql(query) functions — and making every attempt inspectable through execution traces and trajectory inspection.", "body_md": "The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a simple idea of making an agent’s harness smaller, clearer, and easier to verify.\n\nThe competition asked [agents](https://www.nvidia.com/en-us/glossary/ai-agents/) to answer natural-language questions over heterogeneous data sources, including databases, CSV and JSON files, prose documents, PDFs, and briefing videos. Every task required more than [retrieval](https://www.nvidia.com/en-us/glossary/retrieval-augmented-generation/), with the agent inspecting available data, choosing the right tools, reasoning across sources, producing a final answer file, and handling traps that appear in analytical workflows.\n\nKDD also required teams to work with a [small, fixed LLM](https://huggingface.co/Qwen/Qwen3.5-35B-A3B) to power the agent, making the harness the main optimization surface. The team aimed to make the task easier for the available model through preprocessing, constrained tools, persistent state, and evaluation. These techniques are especially useful when building reliable systems around smaller [open models](https://www.nvidia.com/en-us/glossary/open-models/).\n\nThis post shares the playbook behind their result. It isn’t a general recipe for every data science agent, but the broader lesson applies to many agent systems. Reliability often comes less from making the model more open-ended and more from building the right harness around it.\n\n## Two principles behind the system\n\nTwo principles shaped the system.\n\nFirst, constrain the action space. Too many ways to inspect data, call tools, write files, or recover from mistakes can lead to failures. KGMON unified data access, exposed a small tool set, and required final answers to follow a single output path.\n\nSecond, make every attempt inspectable. A run can fail because of a bad tool call, an incorrect join, a missed document rule, or an answer-format issue. Execution traces, repeated attempts, and trajectory inspection helped the team identify failures and refine the harness. Figure 1 shows the overall workflow.\n\nThe following techniques offer practical guidance for building reliable data analytics agents.\n\n## 1. Convert structured sources into one query surface\n\nEach KDD task could combine SQL databases, CSV and JSON files, documents, and videos. KGMON converted CSV and JSON files into tables in the existing SQLite database, giving the agent one SQL interface for all structured data.\n\nKGMON’s custom, persistent Python environment exposed two built-in functions for inspecting and querying that unified data surface:\n\n```\nschema()\nsql(query)\n```\n\nThese custom functions gave the agent a single way to discover and query structured data. They were built into KGMON’s environment rather than provided by Python or SQLite.\n\nThe single SQL interface reduced routing failures and wasted turns, leaving the fixed model more room to reason over the data.\n\n**Apply this approach:** Normalize access to structured data before the agent starts. A temporary SQLite database, a virtualized query layer, or a governed warehouse interface can provide a stable, narrow, and well-documented query surface.\n\n## 2. Seed the agent with schema context up front\n\nKGMON added a schema-scouting step before the main [reasoning](https://www.nvidia.com/en-us/glossary/ai-reasoning/) loop. The system inspected tables, columns, possible join keys, duplicate names, look-alike fields, units, null patterns, and row-grain issues.\n\nThe agent received this schema context at the start of each task, saving an early discovery turn.\n\nSchema scouting reduced errors caused by wrong columns, missed joins, or confusion about what each answer row should represent. Figure 2 shows the unified SQL interface and the schema context provided to the agent.\n\n**Apply this approach:** Add a read-only preflight step that briefs the agent on available tables, likely joins, key columns, units, suspicious fields, and known ambiguities. This can be more useful than another retrieval tool.\n\n## 3. Build a small, opinionated tool harness\n\nKGMON limited the agent to helpers for schema inspection, SQL queries, document lookup, prose extraction, and answer writing.\n\nThe environment exposed the following functions:\n\n```\nschema()       \t# inspect tables and columns\nsql(query)     \t# query context.db\nwrite_answer(df)   # write the final answer atomically\nprose_helper() \t# answer from prose or extract prose into SQL\n```\n\nMiddleware repaired malformed tool calls so one bad call did not end an attempt.\n\nA stateful Python environment retained variables between tool calls, letting the agent reuse intermediate results. Predefined functions such as `schema()`, `sql(query)`, and `write_answer(df)` reduced boilerplate, syntax errors, and file-handling mistakes. This saved turns and kept the small LLM focused on analysis. Figure 3 shows the environment’s functions, persistent state, and error handling.\n\nShort, valid attempts left more turns for additional runs, evaluation, and ensembling—combining results from multiple attempts.\n\n**Apply this approach:** Design tools around workflow requirements and remove duplicate paths. A single approved way to run SQL or write an answer reduces opportunities to corrupt state or produce invalid outputs.\n\n## 4. Treat prose as a first-class input, but keep it separate from structured data\n\nAnalytics workflows often include PDFs, Markdown files, documents, policy text, instructions, and reports. In some KDD tasks, these sources contained thresholds, rules, definitions, and table-like records needed to answer questions.\n\nLarge documents can consume an LLM’s context window. KGMON blocked direct reads of entire files through Python’s `open()` function or `.read()` method, providing tools that limited previews and searches by character count or regular expression (regex) matches.\n\nAfter finding a relevant section, the agent could invoke `prose_helper`, a custom tool that passed document chunks to a separate LLM call with temperature set to 0 and reasoning disabled. It returned answers or extracted tables, keeping raw document contents out of the main agent’s context.\n\n```\nFigure 4 shows the document-inspection workflow. KGMON used prose_helper in two modes:\nmode=\"answer\"  # extract a rule, threshold, or short answer\nmode=\"table\"   # extract repeating records into a SQL table\n```\n\nExtracting rules from prose let the agent apply them in SQL analysis while keeping its working context focused.\n\nImportant caveat: Table extraction suited competition tasks that embedded structured information in documents. Production systems may need only targeted prose lookup; table extraction can remain optional.\n\n**Apply this approach:** Provide a document-inspection tool that answers targeted questions and cites sources without loading entire documents into the main agent’s context. Add table extraction when documents contain repeated records that need joins or filters.\n\n## 5. Preprocess video when it is part of the task\n\nSome KDD tasks included briefing videos. To avoid the compute cost of processing video inside the agent loop, KGMON extracted keyframes, transcribed audio, aligned transcript segments with frames, and supplied the resulting evidence to the agent.\n\nEach task included at most one video, often containing slide-based constraints or distractor values. Transcript-aligned keyframes connected spoken context to the correct visual evidence. Figure 5 shows the preprocessing steps.\n\nImportant caveat: Preprocessing suited the competition setup. Applications with many videos may benefit from an on-demand helper, similar to `prose_helper`.\n\n**Apply this approach:** Preprocess a limited set of videos before the agent loop. For larger collections, provide a retrieval or inspection tool that queries video evidence on demand.\n\n## 6. Log traces so another agent or a human can inspect failures\n\nEach attempt logged prompts, tool calls, SQL queries, intermediate results, errors, repairs, document lookups, and final answers.\n\nA specialized inspector agent could review failed trajectories, categorize mistakes, and surface recurring failures to help the team prioritize improvements.\n\nTraces showed whether a wrong answer came from schema confusion, an incorrect join, missed prose evidence, output formatting, or a brittle prompt rule.\n\n**Apply this approach:** Build trace inspection into the development workflow. A subagent or evaluation script can categorize failures from recent runs and recommend harness changes. Figure 6 shows how a trace reveals the first wrong decision.\n\n## 7. Evaluate multiple attempts, but watch the cost\n\nKGMON improved coverage and reliability through repeated attempts and answer selection. It grouped attempts by answer values rather than column names and could give contested tasks additional runs.\n\nMultiple attempts helped distinguish stable answers from one-off mistakes under the leaderboard’s exact value-level scoring.\n\nImportant caveat: Repeated attempts increase token usage, latency, and compute cost. In production, reserve ensembling for tasks whose value, uncertainty, or risk justifies the cost.\n\n**Apply this approach:** Start with single-run evaluation and trace inspection. Use confidence, disagreement, or validation failures to decide when additional attempts justify their cost.\n\n## 8. Use improvement loops carefully\n\nIn autonomous improvement loops, an agent uses evaluation feedback to refine prompts, tools, postprocessing, and evaluation logic. These changes can improve performance but also overfit to the benchmark.\n\nKGMON saw several risks:\n\n- Hardcoding training examples into prompts\n- Accumulating contradictory instructions\n- Adding brittle postprocessing rules\n- Improving one benchmark split while hurting generalization\n\nThe team needed rapid improvements without building a harness that memorized the benchmark.\n\n**Apply this approach:** Require held-out tasks, prompt audits, trace review, and human approval before promoting changes. Build a review gate into the improvement loop.\n\n## 9. Keep humans in the right part of the loop\n\nThe team combined human guidance with experimentation at agent scale.\n\nPeople defined the task requirements, audited traces, guided early agent behavior, rejected brittle changes, and selected improvements to retain in the harness.\n\nTargeted human interventions could improve future runs without requiring oversight of every tool call.\n\n**Apply this approach:** Review task design, evaluation criteria, failure analysis, and proposed reusable skills. Let agents execute and explore, while people decide which improvements to retain. Figure 7 shows where human review fits into the loop.\n\n## Top lessons for building a data science agent\n\nThe exact KDD solution was shaped by the benchmark: a fixed model, no internet access, heterogeneous task bundles, value-level scoring, and a long-run budget. Not every design choice should be copied directly into production.\n\nYou can apply these practices to other agent systems:\n\n- Normalize data access before the agent starts.\n- Give the agent a small, reliable tool surface.\n- Make every attempt traceable.\n- Evaluate both answers and trajectories.\n- Use repeated attempts selectively.\n- Preserve reusable knowledge.\n- Promote improvements only after validation.\n\nThese practices help an agent use data reliably.\n\n## Getting started\n\nStart with one repeatable analytics workflow and build the smallest harness that can complete it. Follow these steps:\n\n- Define the question and required answer format.\n- Give the agent a narrow set of tools for inspecting and querying data.\n- Record every tool call and intermediate result.\n- Create a small evaluation set that includes both successful cases and likely failure modes.\n- Review the agent’s trajectories, then refine its tools, prompts, and validation checks.\n\nOnce the foundation is reliable, add document inspection, shared knowledge, reviewed improvements, and selective evaluation of multiple attempts.\n\nFor additional architecture examples and implementation ideas, explore the [KDD Cup Data Agents presentation archive](https://dataagent.top/presentations), which includes the recorded presentation session and slides from eight featured teams. These materials are intended as design references rather than a step-by-step reimplementation of KGMON’s solution.\n\nAgent reliability depends on the model and its harness. Together, they must support analysis that is inspectable, repeatable, and useful.", "url": "https://wpnews.pro/news/building-reliable-data-analytics-agents-lessons-from-the-kdd-cup", "canonical_source": "https://developer.nvidia.com/blog/building-reliable-data-analytics-agents-lessons-from-the-kdd-cup/", "published_at": "2026-10-08 18:30:02+00:00", "updated_at": "2026-10-08 18:50:25.884679+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "large-language-models", "ai-research", "developer-tools"], "entities": ["NVIDIA", "KGMON", "KDD Cup 2026", "Qwen3.5-35B-A3B", "SQLite"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-reliable-data-analytics-agents-lessons-from-the-kdd-cup", "markdown": "https://wpnews.pro/news/building-reliable-data-analytics-agents-lessons-from-the-kdd-cup.md", "text": "https://wpnews.pro/news/building-reliable-data-analytics-agents-lessons-from-the-kdd-cup.txt", "jsonld": "https://wpnews.pro/news/building-reliable-data-analytics-agents-lessons-from-the-kdd-cup.jsonld"}}