cd /news/artificial-intelligence/we-catalogued-55-ai-agent-failures-i… · home topics artificial-intelligence article
[ARTICLE · art-99135] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

We catalogued 55+ AI-agent failures in low-level code — and shipped 124 verified skills to fix them

A developer catalogued over 55 documented AI-agent failures in low-level code, revealing predictable failure classes where code looks correct, compiles, and still does the wrong thing. The project turned these into 124 verified engineering skills, emphasizing mechanical gates such as assemble-disassemble-compare, measuring real parallelism, and verifying API existence. The findings highlight that LLM disassembly matches exact instructions only ~14% of the time, and package hallucination rates range from 5.2% to 21.7%.

read4 min views1 publishedAug 16, 2026

Disclosure:this article was drafted with AI assistance. Every technical claim in it is source-traced in the linked repository (registry/claims.yaml

, 177 primary sources). The failure classes below come from real, documented incidents — not vibes.

AI coding agents are excellent at boilerplate and unreliable at low-level code — and the failures are not random. They cluster into predictable classes with a recognizable signature: the code looks correct, compiles, and still does the wrong thing.

We spent three research passes collecting 55+ documented failures (source-traced, not anecdotes) and turned them into 124 verified engineering skills. Here's what we learned.

The fix isn't "be more careful". It's mechanical gates: assemble → disassemble → compare bytes; measure real parallelism; check the API actually exists; make your harness able to fail.

Agents invent instructions that do not exist. One generated CDC COMPASS pseudo-ops (JOB

, SST

, OCT

) for a program that "looked like assembly". Another produced ** movqad** — no such instruction.

The dangerous failures don't error — they silently corrupt:

; What the agent wrote            ; What actually assembled
imul eax, eax, 38                  ; 69 c0 00 00 00 00  (the 38 is DROPPED)
mov (%rax), %eax                   ; 8b 00              (fine — but one byte vs. eax?)

imul eax, eax, 38

assembles to 69 c0 00 00 00 00

— the immediate is silently discarded by the parser. This is exactly the bug class behind BBoeOS PR#584. Add AT&T/Intel operand inversion, missing size hints (inc [counter]

vs inc qword [counter]

), and "AX is 8-bit" claims, and you have a reliable failure generator.

The gate:

gcc -c sample.s && objdump -d sample.o       # AT&T
gcc -c -masm=intel sample.s && objdump -d -M intel sample.o   # Intel

If the disassembly doesn't match what you wrote — same mnemonic, same operands, same size — you didn't write that instruction. Calibration: LLM disassembly gets exact matches right ~14% of the time; decompiler "fixes" are correct ~37%.

Models produce ConcurrentHashMap

/atomics that look thread-safe but execute everything on one thread. The CONCUR benchmark (arXiv:2603.03683) catches deadlocks and races that linear benchmarks cannot — because the code isn't actually concurrent.

Our gate is mechanical, not syntactic:

Program Live threads Wall-clock Verdict
"thread-safe" demo 1
1.206 s fake parallelism
real pthread split
4
0.304 s real parallelism

Count live threads and measure wall-clock scaling. Thread-safe syntax is not parallelism.

RustEvo² (arXiv:2503.16922) is the clearest dataset: models nail stabilized APIs at 65.8%, but behavioral changes (same signature, different semantics) at only 38%. Performance collapses from 56.1% to 32.5% for APIs added after the training cutoff. RAG helps (+13.5%) — but you can't RAG what doesn't exist yet.

And the supply chain angle is worse: agents hallucinate crates that do not exist but resemble real ones — a typosquatting risk (serde-json

vs serde_json

). Studies report 5.2% (commercial) to 21.7% (open-source) package hallucination rates.

cargo info serde-json    # exit 101 — does not exist
cargo info serde_json    # the real crate

The gate: verify existence (cargo search

/crates.io API), pin the toolchain, and treat behavioral changes as the most dangerous class. In crypto specifically, only 23.3% of generated Rust compiles and 57% of that is vulnerable — with nonce reuse the leading cause (arXiv:2604.27001).

The worst class of all: a "passing" test that doesn't test the target.

allclose

oracle mmap

without ever calling munmap

— an agent was the

The Iron Law:a harness that cannot fail is not evidence. The ablation test is:break the target — does your test catch it?If it still passes, your test is decoration.

gcc -O2

, 500k × 256-byte compares — early-exit memcmp

0.054 s vs constant-time ~0 s.Each skill is a compact SKILL.md

that answers five questions before a line is written:

Question Why it matters
When to use / when not to
the agent loads the right tool, not everything
What the agent often gets wrong
the named failure classes above
How to reason correctly
the positive process, not just "don't"
What to verify / how
executable gates, not vibes
Where the knowledge comes from
every claim → primary source

Verified, not asserted: 65 of 124 skills were validated by actually running examples on real toolchains (GCC 16.1, rustc 1.97, GDB, objdump, CMake/Ninja). The remaining 59 are honestly marked researched

with the exact command that would verify them.

git clone https://github.com/TrothByte/low-level-skills-trothbyte
python tools/validate.py     # 124 skills + registry + 177 sources, all gate in seconds

Also installable via npx skills add TrothByte/low-level-skills-trothbyte

or as a Claude Code plugin marketplace.

The agents aren't broken — our expectations are. "It compiles" was never the bar for low-level code. The bar is: assemble → disassemble → compare bytes; measure real parallelism; check the API exists; make the test able to fail. The 55+ catalogued failures become 124 skills that encode exactly those gates.

If you write, review, or debug C, C++, Rust, assembly, kernels, or firmware with an AI — you'll recognize these failures. The library is free and MIT-licensed: github.com/TrothByte/low-level-skills-trothbyte.

Surveys with full source traces are in the repo's research/ folder. Found a failure we haven't catalogued? Open an issue — new skills must be source-traced and differentiated from the existing 124.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @bboeos 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/we-catalogued-55-ai-…] indexed:0 read:4min 2026-08-16 ·