Building Reliable Data Analytics Agents: Lessons from the KDD Cup NVIDIA's KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around making the agent harness smaller, clearer, and easier to verify, the team reported. The competition required agents to answer natural-language questions across databases, CSV and JSON files, prose documents, PDFs, and briefing videos using a small, fixed LLM (Qwen3.5-35B-A3B), which made the harness the main optimization surface. KGMON's two guiding principles were constraining the action space — unifying CSV and JSON into one SQLite-backed SQL interface with schema() and sql(query) functions — and making every attempt inspectable through execution traces and trajectory inspection. The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a simple idea of making an agent’s harness smaller, clearer, and easier to verify. The competition asked agents https://www.nvidia.com/en-us/glossary/ai-agents/ to answer natural-language questions over heterogeneous data sources, including databases, CSV and JSON files, prose documents, PDFs, and briefing videos. Every task required more than retrieval https://www.nvidia.com/en-us/glossary/retrieval-augmented-generation/ , with the agent inspecting available data, choosing the right tools, reasoning across sources, producing a final answer file, and handling traps that appear in analytical workflows. KDD also required teams to work with a small, fixed LLM https://huggingface.co/Qwen/Qwen3.5-35B-A3B to power the agent, making the harness the main optimization surface. The team aimed to make the task easier for the available model through preprocessing, constrained tools, persistent state, and evaluation. These techniques are especially useful when building reliable systems around smaller open models https://www.nvidia.com/en-us/glossary/open-models/ . This post shares the playbook behind their result. It isn’t a general recipe for every data science agent, but the broader lesson applies to many agent systems. Reliability often comes less from making the model more open-ended and more from building the right harness around it. Two principles behind the system Two principles shaped the system. First, constrain the action space. Too many ways to inspect data, call tools, write files, or recover from mistakes can lead to failures. KGMON unified data access, exposed a small tool set, and required final answers to follow a single output path. Second, make every attempt inspectable. A run can fail because of a bad tool call, an incorrect join, a missed document rule, or an answer-format issue. Execution traces, repeated attempts, and trajectory inspection helped the team identify failures and refine the harness. Figure 1 shows the overall workflow. The following techniques offer practical guidance for building reliable data analytics agents. 1. Convert structured sources into one query surface Each KDD task could combine SQL databases, CSV and JSON files, documents, and videos. KGMON converted CSV and JSON files into tables in the existing SQLite database, giving the agent one SQL interface for all structured data. KGMON’s custom, persistent Python environment exposed two built-in functions for inspecting and querying that unified data surface: schema sql query These custom functions gave the agent a single way to discover and query structured data. They were built into KGMON’s environment rather than provided by Python or SQLite. The single SQL interface reduced routing failures and wasted turns, leaving the fixed model more room to reason over the data. Apply this approach: Normalize access to structured data before the agent starts. A temporary SQLite database, a virtualized query layer, or a governed warehouse interface can provide a stable, narrow, and well-documented query surface. 2. Seed the agent with schema context up front KGMON added a schema-scouting step before the main reasoning https://www.nvidia.com/en-us/glossary/ai-reasoning/ loop. The system inspected tables, columns, possible join keys, duplicate names, look-alike fields, units, null patterns, and row-grain issues. The agent received this schema context at the start of each task, saving an early discovery turn. Schema scouting reduced errors caused by wrong columns, missed joins, or confusion about what each answer row should represent. Figure 2 shows the unified SQL interface and the schema context provided to the agent. Apply this approach: Add a read-only preflight step that briefs the agent on available tables, likely joins, key columns, units, suspicious fields, and known ambiguities. This can be more useful than another retrieval tool. 3. Build a small, opinionated tool harness KGMON limited the agent to helpers for schema inspection, SQL queries, document lookup, prose extraction, and answer writing. The environment exposed the following functions: schema inspect tables and columns sql query query context.db write answer df write the final answer atomically prose helper answer from prose or extract prose into SQL Middleware repaired malformed tool calls so one bad call did not end an attempt. A stateful Python environment retained variables between tool calls, letting the agent reuse intermediate results. Predefined functions such as schema , sql query , and write answer df reduced boilerplate, syntax errors, and file-handling mistakes. This saved turns and kept the small LLM focused on analysis. Figure 3 shows the environment’s functions, persistent state, and error handling. Short, valid attempts left more turns for additional runs, evaluation, and ensembling—combining results from multiple attempts. Apply this approach: Design tools around workflow requirements and remove duplicate paths. A single approved way to run SQL or write an answer reduces opportunities to corrupt state or produce invalid outputs. 4. Treat prose as a first-class input, but keep it separate from structured data Analytics workflows often include PDFs, Markdown files, documents, policy text, instructions, and reports. In some KDD tasks, these sources contained thresholds, rules, definitions, and table-like records needed to answer questions. Large documents can consume an LLM’s context window. KGMON blocked direct reads of entire files through Python’s open function or .read method, providing tools that limited previews and searches by character count or regular expression regex matches. After finding a relevant section, the agent could invoke prose helper , a custom tool that passed document chunks to a separate LLM call with temperature set to 0 and reasoning disabled. It returned answers or extracted tables, keeping raw document contents out of the main agent’s context. Figure 4 shows the document-inspection workflow. KGMON used prose helper in two modes: mode="answer" extract a rule, threshold, or short answer mode="table" extract repeating records into a SQL table Extracting rules from prose let the agent apply them in SQL analysis while keeping its working context focused. Important caveat: Table extraction suited competition tasks that embedded structured information in documents. Production systems may need only targeted prose lookup; table extraction can remain optional. Apply this approach: Provide a document-inspection tool that answers targeted questions and cites sources without loading entire documents into the main agent’s context. Add table extraction when documents contain repeated records that need joins or filters. 5. Preprocess video when it is part of the task Some KDD tasks included briefing videos. To avoid the compute cost of processing video inside the agent loop, KGMON extracted keyframes, transcribed audio, aligned transcript segments with frames, and supplied the resulting evidence to the agent. Each task included at most one video, often containing slide-based constraints or distractor values. Transcript-aligned keyframes connected spoken context to the correct visual evidence. Figure 5 shows the preprocessing steps. Important caveat: Preprocessing suited the competition setup. Applications with many videos may benefit from an on-demand helper, similar to prose helper . Apply this approach: Preprocess a limited set of videos before the agent loop. For larger collections, provide a retrieval or inspection tool that queries video evidence on demand. 6. Log traces so another agent or a human can inspect failures Each attempt logged prompts, tool calls, SQL queries, intermediate results, errors, repairs, document lookups, and final answers. A specialized inspector agent could review failed trajectories, categorize mistakes, and surface recurring failures to help the team prioritize improvements. Traces showed whether a wrong answer came from schema confusion, an incorrect join, missed prose evidence, output formatting, or a brittle prompt rule. Apply this approach: Build trace inspection into the development workflow. A subagent or evaluation script can categorize failures from recent runs and recommend harness changes. Figure 6 shows how a trace reveals the first wrong decision. 7. Evaluate multiple attempts, but watch the cost KGMON improved coverage and reliability through repeated attempts and answer selection. It grouped attempts by answer values rather than column names and could give contested tasks additional runs. Multiple attempts helped distinguish stable answers from one-off mistakes under the leaderboard’s exact value-level scoring. Important caveat: Repeated attempts increase token usage, latency, and compute cost. In production, reserve ensembling for tasks whose value, uncertainty, or risk justifies the cost. Apply this approach: Start with single-run evaluation and trace inspection. Use confidence, disagreement, or validation failures to decide when additional attempts justify their cost. 8. Use improvement loops carefully In autonomous improvement loops, an agent uses evaluation feedback to refine prompts, tools, postprocessing, and evaluation logic. These changes can improve performance but also overfit to the benchmark. KGMON saw several risks: - Hardcoding training examples into prompts - Accumulating contradictory instructions - Adding brittle postprocessing rules - Improving one benchmark split while hurting generalization The team needed rapid improvements without building a harness that memorized the benchmark. Apply this approach: Require held-out tasks, prompt audits, trace review, and human approval before promoting changes. Build a review gate into the improvement loop. 9. Keep humans in the right part of the loop The team combined human guidance with experimentation at agent scale. People defined the task requirements, audited traces, guided early agent behavior, rejected brittle changes, and selected improvements to retain in the harness. Targeted human interventions could improve future runs without requiring oversight of every tool call. Apply this approach: Review task design, evaluation criteria, failure analysis, and proposed reusable skills. Let agents execute and explore, while people decide which improvements to retain. Figure 7 shows where human review fits into the loop. Top lessons for building a data science agent The exact KDD solution was shaped by the benchmark: a fixed model, no internet access, heterogeneous task bundles, value-level scoring, and a long-run budget. Not every design choice should be copied directly into production. You can apply these practices to other agent systems: - Normalize data access before the agent starts. - Give the agent a small, reliable tool surface. - Make every attempt traceable. - Evaluate both answers and trajectories. - Use repeated attempts selectively. - Preserve reusable knowledge. - Promote improvements only after validation. These practices help an agent use data reliably. Getting started Start with one repeatable analytics workflow and build the smallest harness that can complete it. Follow these steps: - Define the question and required answer format. - Give the agent a narrow set of tools for inspecting and querying data. - Record every tool call and intermediate result. - Create a small evaluation set that includes both successful cases and likely failure modes. - Review the agent’s trajectories, then refine its tools, prompts, and validation checks. Once the foundation is reliable, add document inspection, shared knowledge, reviewed improvements, and selective evaluation of multiple attempts. For additional architecture examples and implementation ideas, explore the KDD Cup Data Agents presentation archive https://dataagent.top/presentations , which includes the recorded presentation session and slides from eight featured teams. These materials are intended as design references rather than a step-by-step reimplementation of KGMON’s solution. Agent reliability depends on the model and its harness. Together, they must support analysis that is inspectable, repeatable, and useful.