# The AMD Mini-Challenge 3: citation-by-necessity RAG on ROCm

> Source: <https://dev.to/eterniti/the-amd-mini-challenge-3-citation-by-necessity-rag-on-rocm-4ljm>
> Published: 2026-10-08 19:32:33+00:00

*Notes from building a retrieval-augmented generation container for the AMD AI League (Match 3), on an AMD Instinct MI300X.*

Match 3 of the AMD AI League sounds simple: build a container that answers questions about a folder of documents. The catch is in the scoring. Each answer only counts if **the value is right and the list of cited files is exactly right**. No partial credit, no "close enough". Cite one file too many and a correct answer scores zero.

My container scored **10/10 (200/200)** on the public sample questions, with exact citation sets, in 1–2 seconds per question. Here is how it works, and the traps that nearly cost me points.

The grader copies a folder of mixed documents into the container and calls one script in two ways:

```
python3 /app/app.py --index /app/corpus
python3 /app/app.py --corpus /app/corpus --query-id query_01 --query "What is the maximum junction temperature?"
```

Each query must write `/app/output/query_01_output.json`:

```
{"answer": "94", "citations": ["specs/tq40_datasheet_r2.pdf"], "confidence": 0.9}
```

The corpus is deliberately messy:

Limits: 10 minutes for startup (including indexing), 30 seconds per question, 1–48 GiB of VRAM, a 60 GiB image, and **no network** during grading.

The grader starts a **new process for every question**. If you load an 8-billion-parameter model inside `app.py`, you load it ten times and blow the 30-second budget on every question.

So the container runs two pieces:

`server.py`:` CMD`. It loads the models once, keeps the index in memory, and listens on a Unix socket.`app.py`:` citations` field scores zero for that question.
Two models, both baked into the image:

On the MI300X: models load in 7.7 s, the sample corpus indexes in 3.6 s, and peak VRAM is about 20 GB.

The challenge spells it out: a corpus walk that crashes on the first unreadable file indexes nothing after it, so *which* files you lose depends on alphabetical order. That is how a solution passes locally and fails on the graded set.

I went one step further and parse every file in a **separate plain-Python worker process**. A file that raises, hangs or crashes the parser costs only that file:

`PermissionError` is caught, skipped and logged.`_WITHDRAWN`) or by phrases like "superseded by", and kept out of the model's context. A detail that matters: the Spreadsheets and CSV rows are indexed one row at a time, with the column headers attached (`Part Number: ORR-FAN-2214-B | Description: Fan assembly, field-replaceable | ...`). That way a single row is a self-contained, retrievable fact.

Some questions need two files. For example: *"The production log shows a thermal throttle incident. Which firmware release fixed the underlying defect?"* The log gives an identifier, and the bug database maps that identifier to a fix version. Neither file alone answers it.

Retrieval fuses identifier-aware BM25 (so `ORR-1847`, `E7731` and `THERM_ALERT#` survive tokenisation) with dense embeddings. Then it takes one extra step: **rare identifiers** found in the best chunks, which the question didn't mention, pull in the chunks that define them. That puts both links of the chain in front of the model.

The rule from the challenge brief is: *cite a file only if removing it would make your answer impossible.*

Asking the model nicely is not enough, so the citations go through deterministic checks after the model answers:

To compare values the way the grader does, I first stripped separators and searched for the value as a substring. Then `4.3.2` became `432`, which "matched" inside a log line where `latency_ms=243` was followed by a date starting `2026`. The log looked like a source of the answer, and my citations broke.

The fix was a token-boundary regex. Separators are still optional (`Q3 FY27` matches `Q3FY27`, and `v4.3.2` matches `4.3.2`), but the value must stand on its own. Small detail, whole question's worth of points.

**1. pip can silently replace ROCm torch with a CUDA build.** Installing `transformers` can pull a CUDA torch over the base image's ROCm build. The error then shows up somewhere unrelated. My Dockerfile pins every torch package to the version already in the base image, installs against those constraints, and fails the build if `torch.__version__` no longer contains `rocm`.

**2. Test the unreadable file the way the grader does.** As root, `chmod 000` doesn't stop you reading a file. The grader drops the `DAC_OVERRIDE` capability, so test with:

```
docker run --network none --cap-drop DAC_OVERRIDE --device=/dev/kfd --device=/dev/dri ...
```

**A bonus warning about the official self-check.** It starts the container without GPU devices. My server couldn't load the model, the client waited out its timeouts, wrote valid empty answers, and every check said **PASS**. The self-check checks the *shape* of the output, not the answers. Always score the sample questions yourself on a real GPU run.

| Check | Result | 
|---|---|
| Sample questions | 10/10, exact citation sets (200/200) | 
| Per question | 1.0–1.9 s (limit 30 s) | 
| Startup: model load + index | about 12 s (limit 10 min) | 
| Peak VRAM | 20 GB (limit 48 GiB) | 
| Image size | 45.8 GiB uncompressed (limit 60 GiB) | 

The graded corpus is larger and harder than the sample, so these numbers are a starting point, not a final score.

The code is open source: **[https://github.com/DataGuy-Eterniti/knight-eterniti](https://github.com/DataGuy-Eterniti/knight-eterniti)** (`mc3-rag/`).

Thanks to **@lablab.ai** and **[@amd](https://dev.to/amd)** for the AMD AI League and the MI300X access through AMD Developer Cloud. On to Match 4.

*#AMD #ROCm #lablab #RAG #MachineLearning #AMDAILeague*
