# We catalogued 55+ AI-agent failures in low-level code — and shipped 124 verified skills to fix them

> Source: <https://dev.to/trothbyte/we-catalogued-55-ai-agent-failures-in-low-level-code-and-shipped-124-verified-skills-to-fix-them-369l>
> Published: 2026-08-16 22:26:06+00:00

✨

Disclosure:this article was drafted with AI assistance. Every technical claim in it is source-traced in the linked repository (`registry/claims.yaml`

, 177 primary sources). The failure classes below come from real, documented incidents — not vibes.

AI coding agents are excellent at boilerplate and unreliable at low-level code — and the failures are not random. They cluster into **predictable classes** with a recognizable signature: the code *looks* correct, compiles, and still does the wrong thing.

We spent three research passes collecting **55+ documented failures** (source-traced, not anecdotes) and turned them into **124 verified engineering skills**. Here's what we learned.

The fix isn't "be more careful". It's **mechanical gates**: assemble → disassemble → compare bytes; measure real parallelism; check the API actually exists; make your harness *able to fail*.

Agents invent instructions that do not exist. One generated CDC COMPASS pseudo-ops (`JOB`

, `SST`

, `OCT`

) for a program that "looked like assembly". Another produced ** movqad** — no such instruction.

The dangerous failures don't error — they silently corrupt:

```
; What the agent wrote            ; What actually assembled
imul eax, eax, 38                  ; 69 c0 00 00 00 00  (the 38 is DROPPED)
mov (%rax), %eax                   ; 8b 00              (fine — but one byte vs. eax?)
```

`imul eax, eax, 38`

assembles to `69 c0 00 00 00 00`

— the immediate is silently discarded by the parser. This is exactly the bug class behind **BBoeOS PR#584**. Add AT&T/Intel operand inversion, missing size hints (`inc [counter]`

vs `inc qword [counter]`

), and "AX is 8-bit" claims, and you have a reliable failure generator.

**The gate:**

```
gcc -c sample.s && objdump -d sample.o       # AT&T
gcc -c -masm=intel sample.s && objdump -d -M intel sample.o   # Intel
```

If the disassembly doesn't match what you wrote — same mnemonic, same operands, same size — you didn't write that instruction. Calibration: LLM disassembly gets exact matches right ~14% of the time; decompiler "fixes" are correct ~37%.

Models produce `ConcurrentHashMap`

/atomics that look thread-safe but execute everything on one thread. The **CONCUR benchmark** (arXiv:2603.03683) catches deadlocks and races that linear benchmarks cannot — because the code isn't actually concurrent.

Our gate is mechanical, not syntactic:

| Program | Live threads | Wall-clock | Verdict |
|---|---|---|---|
| "thread-safe" demo | 1 |
1.206 s | fake parallelism |
real `pthread` split |
4 |
0.304 s | real parallelism |

Count live threads and measure wall-clock scaling. Thread-safe *syntax* is not parallelism.

**RustEvo²** (arXiv:2503.16922) is the clearest dataset: models nail *stabilized* APIs at **65.8%**, but *behavioral* changes (same signature, different semantics) at only **38%**. Performance collapses from 56.1% to 32.5% for APIs added after the training cutoff. RAG helps (+13.5%) — but you can't RAG what doesn't exist yet.

And the supply chain angle is worse: agents hallucinate crates that **do not exist but resemble real ones** — a typosquatting risk (`serde-json`

vs `serde_json`

). Studies report **5.2% (commercial) to 21.7% (open-source)** package hallucination rates.

```
cargo info serde-json    # exit 101 — does not exist
cargo info serde_json    # the real crate
```

**The gate:** verify existence (`cargo search`

/crates.io API), pin the toolchain, and treat behavioral changes as the most dangerous class. In crypto specifically, only **23.3%** of generated Rust compiles and **57%** of that is vulnerable — with nonce reuse the leading cause (arXiv:2604.27001).

The worst class of all: a "passing" test that doesn't test the target.

`allclose`

oracle `mmap`

without ever calling `munmap`

— an agent was the

The Iron Law:a harness that cannot fail is not evidence. The ablation test is:break the target — does your test catch it?If it still passes, your test is decoration.

`gcc -O2`

, 500k × 256-byte compares — early-exit `memcmp`

0.054 s vs constant-time ~0 s.Each skill is a compact `SKILL.md`

that answers five questions before a line is written:

| Question | Why it matters |
|---|---|
When to use / when not to |
the agent loads the right tool, not everything |
What the agent often gets wrong |
the named failure classes above |
How to reason correctly |
the positive process, not just "don't" |
What to verify / how |
executable gates, not vibes |
Where the knowledge comes from |
every claim → primary source |

**Verified, not asserted:** 65 of 124 skills were validated by actually running examples on real toolchains (GCC 16.1, rustc 1.97, GDB, objdump, CMake/Ninja). The remaining 59 are honestly marked `researched`

with the exact command that would verify them.

```
git clone https://github.com/TrothByte/low-level-skills-trothbyte
python tools/validate.py     # 124 skills + registry + 177 sources, all gate in seconds
```

Also installable via `npx skills add TrothByte/low-level-skills-trothbyte`

or as a Claude Code plugin marketplace.

The agents aren't broken — **our expectations are**. "It compiles" was never the bar for low-level code. The bar is: *assemble → disassemble → compare bytes; measure real parallelism; check the API exists; make the test able to fail.* The 55+ catalogued failures become 124 skills that encode exactly those gates.

If you write, review, or debug C, C++, Rust, assembly, kernels, or firmware with an AI — you'll recognize these failures. The library is free and MIT-licensed: **github.com/TrothByte/low-level-skills-trothbyte**.

*Surveys with full source traces are in the repo's research/ folder. Found a failure we haven't catalogued? Open an issue — new skills must be source-traced and differentiated from the existing 124.*
