We catalogued 55+ AI-agent failures in low-level code — and shipped 124 verified skills to fix them A developer catalogued over 55 documented AI-agent failures in low-level code, revealing predictable failure classes where code looks correct, compiles, and still does the wrong thing. The project turned these into 124 verified engineering skills, emphasizing mechanical gates such as assemble-disassemble-compare, measuring real parallelism, and verifying API existence. The findings highlight that LLM disassembly matches exact instructions only ~14% of the time, and package hallucination rates range from 5.2% to 21.7%. ✨ Disclosure:this article was drafted with AI assistance. Every technical claim in it is source-traced in the linked repository registry/claims.yaml , 177 primary sources . The failure classes below come from real, documented incidents — not vibes. AI coding agents are excellent at boilerplate and unreliable at low-level code — and the failures are not random. They cluster into predictable classes with a recognizable signature: the code looks correct, compiles, and still does the wrong thing. We spent three research passes collecting 55+ documented failures source-traced, not anecdotes and turned them into 124 verified engineering skills . Here's what we learned. The fix isn't "be more careful". It's mechanical gates : assemble → disassemble → compare bytes; measure real parallelism; check the API actually exists; make your harness able to fail . Agents invent instructions that do not exist. One generated CDC COMPASS pseudo-ops JOB , SST , OCT for a program that "looked like assembly". Another produced movqad — no such instruction. The dangerous failures don't error — they silently corrupt: ; What the agent wrote ; What actually assembled imul eax, eax, 38 ; 69 c0 00 00 00 00 the 38 is DROPPED mov %rax , %eax ; 8b 00 fine — but one byte vs. eax? imul eax, eax, 38 assembles to 69 c0 00 00 00 00 — the immediate is silently discarded by the parser. This is exactly the bug class behind BBoeOS PR 584 . Add AT&T/Intel operand inversion, missing size hints inc counter vs inc qword counter , and "AX is 8-bit" claims, and you have a reliable failure generator. The gate: gcc -c sample.s && objdump -d sample.o AT&T gcc -c -masm=intel sample.s && objdump -d -M intel sample.o Intel If the disassembly doesn't match what you wrote — same mnemonic, same operands, same size — you didn't write that instruction. Calibration: LLM disassembly gets exact matches right ~14% of the time; decompiler "fixes" are correct ~37%. Models produce ConcurrentHashMap /atomics that look thread-safe but execute everything on one thread. The CONCUR benchmark arXiv:2603.03683 catches deadlocks and races that linear benchmarks cannot — because the code isn't actually concurrent. Our gate is mechanical, not syntactic: | Program | Live threads | Wall-clock | Verdict | |---|---|---|---| | "thread-safe" demo | 1 | 1.206 s | fake parallelism | real pthread split | 4 | 0.304 s | real parallelism | Count live threads and measure wall-clock scaling. Thread-safe syntax is not parallelism. RustEvo² arXiv:2503.16922 is the clearest dataset: models nail stabilized APIs at 65.8% , but behavioral changes same signature, different semantics at only 38% . Performance collapses from 56.1% to 32.5% for APIs added after the training cutoff. RAG helps +13.5% — but you can't RAG what doesn't exist yet. And the supply chain angle is worse: agents hallucinate crates that do not exist but resemble real ones — a typosquatting risk serde-json vs serde json . Studies report 5.2% commercial to 21.7% open-source package hallucination rates. cargo info serde-json exit 101 — does not exist cargo info serde json the real crate The gate: verify existence cargo search /crates.io API , pin the toolchain, and treat behavioral changes as the most dangerous class. In crypto specifically, only 23.3% of generated Rust compiles and 57% of that is vulnerable — with nonce reuse the leading cause arXiv:2604.27001 . The worst class of all: a "passing" test that doesn't test the target. allclose oracle mmap without ever calling munmap — an agent was the The Iron Law:a harness that cannot fail is not evidence. The ablation test is:break the target — does your test catch it?If it still passes, your test is decoration. gcc -O2 , 500k × 256-byte compares — early-exit memcmp 0.054 s vs constant-time ~0 s.Each skill is a compact SKILL.md that answers five questions before a line is written: | Question | Why it matters | |---|---| When to use / when not to | the agent loads the right tool, not everything | What the agent often gets wrong | the named failure classes above | How to reason correctly | the positive process, not just "don't" | What to verify / how | executable gates, not vibes | Where the knowledge comes from | every claim → primary source | Verified, not asserted: 65 of 124 skills were validated by actually running examples on real toolchains GCC 16.1, rustc 1.97, GDB, objdump, CMake/Ninja . The remaining 59 are honestly marked researched with the exact command that would verify them. git clone https://github.com/TrothByte/low-level-skills-trothbyte python tools/validate.py 124 skills + registry + 177 sources, all gate in seconds Also installable via npx skills add TrothByte/low-level-skills-trothbyte or as a Claude Code plugin marketplace. The agents aren't broken — our expectations are . "It compiles" was never the bar for low-level code. The bar is: assemble → disassemble → compare bytes; measure real parallelism; check the API exists; make the test able to fail. The 55+ catalogued failures become 124 skills that encode exactly those gates. If you write, review, or debug C, C++, Rust, assembly, kernels, or firmware with an AI — you'll recognize these failures. The library is free and MIT-licensed: github.com/TrothByte/low-level-skills-trothbyte . Surveys with full source traces are in the repo's research/ folder. Found a failure we haven't catalogued? Open an issue — new skills must be source-traced and differentiated from the existing 124.