{"slug": "the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm", "title": "The AMD Mini-Challenge 3: citation-by-necessity RAG on ROCm", "summary": "A developer built a retrieval-augmented generation container for Match 3 of the AMD AI League that scored 10/10 (200/200) on the public sample questions with exact citation sets, running on an AMD Instinct MI300X under ROCm. The container uses a persistent Unix-socket server to load models once, a separate worker process to isolate parser failures, and deterministic post-generation citation checks, with models loading in 7.7 seconds and peak VRAM around 20 GB.", "body_md": "*Notes from building a retrieval-augmented generation container for the AMD AI League (Match 3), on an AMD Instinct MI300X.*\n\nMatch 3 of the AMD AI League sounds simple: build a container that answers questions about a folder of documents. The catch is in the scoring. Each answer only counts if **the value is right and the list of cited files is exactly right**. No partial credit, no \"close enough\". Cite one file too many and a correct answer scores zero.\n\nMy container scored **10/10 (200/200)** on the public sample questions, with exact citation sets, in 1–2 seconds per question. Here is how it works, and the traps that nearly cost me points.\n\nThe grader copies a folder of mixed documents into the container and calls one script in two ways:\n\n```\npython3 /app/app.py --index /app/corpus\npython3 /app/app.py --corpus /app/corpus --query-id query_01 --query \"What is the maximum junction temperature?\"\n```\n\nEach query must write `/app/output/query_01_output.json`:\n\n```\n{\"answer\": \"94\", \"citations\": [\"specs/tq40_datasheet_r2.pdf\"], \"confidence\": 0.9}\n```\n\nThe corpus is deliberately messy:\n\nLimits: 10 minutes for startup (including indexing), 30 seconds per question, 1–48 GiB of VRAM, a 60 GiB image, and **no network** during grading.\n\nThe grader starts a **new process for every question**. If you load an 8-billion-parameter model inside `app.py`, you load it ten times and blow the 30-second budget on every question.\n\nSo the container runs two pieces:\n\n`server.py`:` CMD`. It loads the models once, keeps the index in memory, and listens on a Unix socket.`app.py`:` citations` field scores zero for that question.\nTwo models, both baked into the image:\n\nOn the MI300X: models load in 7.7 s, the sample corpus indexes in 3.6 s, and peak VRAM is about 20 GB.\n\nThe challenge spells it out: a corpus walk that crashes on the first unreadable file indexes nothing after it, so *which* files you lose depends on alphabetical order. That is how a solution passes locally and fails on the graded set.\n\nI went one step further and parse every file in a **separate plain-Python worker process**. A file that raises, hangs or crashes the parser costs only that file:\n\n`PermissionError` is caught, skipped and logged.`_WITHDRAWN`) or by phrases like \"superseded by\", and kept out of the model's context. A detail that matters: the Spreadsheets and CSV rows are indexed one row at a time, with the column headers attached (`Part Number: ORR-FAN-2214-B | Description: Fan assembly, field-replaceable | ...`). That way a single row is a self-contained, retrievable fact.\n\nSome questions need two files. For example: *\"The production log shows a thermal throttle incident. Which firmware release fixed the underlying defect?\"* The log gives an identifier, and the bug database maps that identifier to a fix version. Neither file alone answers it.\n\nRetrieval fuses identifier-aware BM25 (so `ORR-1847`, `E7731` and `THERM_ALERT#` survive tokenisation) with dense embeddings. Then it takes one extra step: **rare identifiers** found in the best chunks, which the question didn't mention, pull in the chunks that define them. That puts both links of the chain in front of the model.\n\nThe rule from the challenge brief is: *cite a file only if removing it would make your answer impossible.*\n\nAsking the model nicely is not enough, so the citations go through deterministic checks after the model answers:\n\nTo compare values the way the grader does, I first stripped separators and searched for the value as a substring. Then `4.3.2` became `432`, which \"matched\" inside a log line where `latency_ms=243` was followed by a date starting `2026`. The log looked like a source of the answer, and my citations broke.\n\nThe fix was a token-boundary regex. Separators are still optional (`Q3 FY27` matches `Q3FY27`, and `v4.3.2` matches `4.3.2`), but the value must stand on its own. Small detail, whole question's worth of points.\n\n**1. pip can silently replace ROCm torch with a CUDA build.** Installing `transformers` can pull a CUDA torch over the base image's ROCm build. The error then shows up somewhere unrelated. My Dockerfile pins every torch package to the version already in the base image, installs against those constraints, and fails the build if `torch.__version__` no longer contains `rocm`.\n\n**2. Test the unreadable file the way the grader does.** As root, `chmod 000` doesn't stop you reading a file. The grader drops the `DAC_OVERRIDE` capability, so test with:\n\n```\ndocker run --network none --cap-drop DAC_OVERRIDE --device=/dev/kfd --device=/dev/dri ...\n```\n\n**A bonus warning about the official self-check.** It starts the container without GPU devices. My server couldn't load the model, the client waited out its timeouts, wrote valid empty answers, and every check said **PASS**. The self-check checks the *shape* of the output, not the answers. Always score the sample questions yourself on a real GPU run.\n\n| Check | Result | \n|---|---|\n| Sample questions | 10/10, exact citation sets (200/200) | \n| Per question | 1.0–1.9 s (limit 30 s) | \n| Startup: model load + index | about 12 s (limit 10 min) | \n| Peak VRAM | 20 GB (limit 48 GiB) | \n| Image size | 45.8 GiB uncompressed (limit 60 GiB) | \n\nThe graded corpus is larger and harder than the sample, so these numbers are a starting point, not a final score.\n\nThe code is open source: **[https://github.com/DataGuy-Eterniti/knight-eterniti](https://github.com/DataGuy-Eterniti/knight-eterniti)** (`mc3-rag/`).\n\nThanks to **@lablab.ai** and **[@amd](https://dev.to/amd)** for the AMD AI League and the MI300X access through AMD Developer Cloud. On to Match 4.\n\n*#AMD #ROCm #lablab #RAG #MachineLearning #AMDAILeague*", "url": "https://wpnews.pro/news/the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm", "canonical_source": "https://dev.to/eterniti/the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm-4ljm", "published_at": "2026-10-08 19:32:33+00:00", "updated_at": "2026-10-08 19:49:11.902948+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-infrastructure", "mlops", "ai-chips"], "entities": ["AMD", "AMD AI League", "AMD Instinct MI300X", "ROCm"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm", "markdown": "https://wpnews.pro/news/the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm.md", "text": "https://wpnews.pro/news/the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm.txt", "jsonld": "https://wpnews.pro/news/the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm.jsonld"}}