10M Batch LLM Inference at $0 Cloud Cost: O(1) Memory Clamped Architecture A developer documented running a 10,000,000-record batch LLM inference workload entirely on a local HP Z4 G4 workstation at zero cloud cost, using jemalloc with background thread decay to prevent heap fragmentation, Polars streaming sinks to keep memory independent of row count, and an aiosqlite WAL pipeline with BEGIN IMMEDIATE checkpoints every 50,000 items for crash consistency. The writeup argues that high cloud API costs and OOM failures in large-scale pipelines are architectural defects rather than hardware limits, and that air-gapped local execution avoids HIPAA and GDPR exposure from sending unmasked enterprise data to external APIs. High cloud API costs and Out-Of-Memory OOM failures in large-scale data pipelines are architectural defects, not hardware constraints. This report documents the performance of a 10,000,000-record batch LLM inference workload executed locally on an HP Z4 G4 workstation Intel Xeon W-2155, 128GB ECC DDR4 RAM, NVIDIA RTX 5060 Ti 16GB VRAM, NVMe PCIe Gen4 SSD . PRAGMA integrity check: ok glibc malloc exhibits heap fragmentation under sustained allocation and deallocation cycles. Injecting libjemalloc2 via LD PRELOAD with active background thread decay MALLOC CONF="background thread:true,dirty decay ms:2000,muzzy decay ms:2000" prevents process RSS growth and maintains a predictable host RAM footprint. In-memory array construction introduces O N spatial memory complexity. Streaming chunked outputs directly to disk using Polars collect engine="streaming" and sink parquet ensures memory usage remains independent of row count. State recovery and record tracking are managed via an aiosqlite pipeline using Write-Ahead Logging PRAGMA journal mode=WAL; PRAGMA synchronous=NORMAL; . Explicit BEGIN IMMEDIATE transaction blocks are executed every 50,000 items, with periodic WAL truncation to guarantee system crash consistency and database integrity. Transmitting unmasked enterprise datasets to external APIs introduces regulatory compliance risks under HIPAA and GDPR. Executing high-throughput LLM workloads on local, air-gapped infrastructure --network none prevents telemetry transmission and external data exposure. Reproducible Code & Repository: https://github.com/Matsubara-CEO https://github.com/Matsubara-CEO