Letting an LLM write executables byte by byte (hello world to playable Tetris) An LLM generated working x86-64 ELF executables byte by byte with no compiler, assembler or linker, producing a 165-byte hello world and a 439-byte HTTP server, according to an experiment run with claude-sonnet-5-5. The raw-bytes approach was unreliable: across five runs per case, the hello world passed 4/5 and the HTTP server 2/5, versus 5/5 for both at the C and assembly levels. The same loop also produced a playable Tetris in assembly (X11, no libc, about 8 KB), verified by comparing screenshots against a reference at four moments. No source language. No compiler. No assembler. No linker. Give an LLM a natural-language spec and it returns the bytes of the executable : ELF header, program header and machine code, by hand. The only thing between the model and the binary is xxd -r . A 439-byte HTTP server, generated from a one-paragraph spec. That's the real output of the generated binary. I've been running this experiment every time a new model comes out for a couple of years, and this is the first time the results are good enough to share. You could say it isn't a compiler in the classical sense, but a model acting as the entire compiler analysis, code generation, assembling and linking in a single step, inside a loop that corrects it. The loop is a bash script of about 100 lines. It hands the spec to the model claude -p , turns whatever comes back into a binary, and tests it for real: it runs it, sends it a curl if it's a server, or compares a screenshot against a reference if it's a game. If it fails, the error goes back to the model exactly as it came out exit code, output, what readelf sees in the file , and it starts over until the test passes or the attempts run out. Nobody is in the middle. php flowchart TD S Natural-language spec -- M LLM M -- |ELF bytes| X xxd -r X -- T{Real test} T -- pass -- OK Binary T -- fail -- E Real error