{"slug": "10m-batch-llm-inference-at-0-cloud-cost-o-1-memory-clamped-architecture", "title": "10M Batch LLM Inference at $0 Cloud Cost: O(1) Memory Clamped Architecture", "summary": "A developer documented running a 10,000,000-record batch LLM inference workload entirely on a local HP Z4 G4 workstation at zero cloud cost, using jemalloc with background thread decay to prevent heap fragmentation, Polars streaming sinks to keep memory independent of row count, and an aiosqlite WAL pipeline with BEGIN IMMEDIATE checkpoints every 50,000 items for crash consistency. The writeup argues that high cloud API costs and OOM failures in large-scale pipelines are architectural defects rather than hardware limits, and that air-gapped local execution avoids HIPAA and GDPR exposure from sending unmasked enterprise data to external APIs.", "body_md": "High cloud API costs and Out-Of-Memory (OOM) failures in large-scale data pipelines are architectural defects, not hardware constraints.\n\nThis report documents the performance of a 10,000,000-record batch LLM inference workload executed locally on an HP Z4 G4 workstation (Intel Xeon W-2155, 128GB ECC DDR4 RAM, NVIDIA RTX 5060 Ti 16GB VRAM, NVMe PCIe Gen4 SSD).\n\n`PRAGMA integrity_check: ok`)\nglibc malloc exhibits heap fragmentation under sustained allocation and deallocation cycles. Injecting `libjemalloc2` via `LD_PRELOAD` with active background thread decay (`MALLOC_CONF=\"background_thread:true,dirty_decay_ms:2000,muzzy_decay_ms:2000\"`) prevents process RSS growth and maintains a predictable host RAM footprint.\n\nIn-memory array construction introduces O(N) spatial memory complexity. Streaming chunked outputs directly to disk using Polars (`collect(engine=\"streaming\")` and `sink_parquet`) ensures memory usage remains independent of row count.\n\nState recovery and record tracking are managed via an `aiosqlite` pipeline using Write-Ahead Logging (`PRAGMA journal_mode=WAL; PRAGMA synchronous=NORMAL;`). Explicit `BEGIN IMMEDIATE` transaction blocks are executed every 50,000 items, with periodic WAL truncation to guarantee system crash consistency and database integrity.\n\nTransmitting unmasked enterprise datasets to external APIs introduces regulatory compliance risks under HIPAA and GDPR. Executing high-throughput LLM workloads on local, air-gapped infrastructure (`--network none`) prevents telemetry transmission and external data exposure.\n\nReproducible Code & Repository:\n\n[https://github.com/Matsubara-CEO](https://github.com/Matsubara-CEO)", "url": "https://wpnews.pro/news/10m-batch-llm-inference-at-0-cloud-cost-o-1-memory-clamped-architecture", "canonical_source": "https://dev.to/shotamatsubara/10m-batch-llm-inference-at-0-cloud-cost-o1-memory-clamped-architecture-1g62", "published_at": "2026-10-08 14:13:31+00:00", "updated_at": "2026-10-08 14:20:37.062789+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-tools"], "entities": ["HP Z4 G4", "Intel Xeon W-2155", "NVIDIA RTX 5060 Ti", "Polars", "aiosqlite", "jemalloc", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/10m-batch-llm-inference-at-0-cloud-cost-o-1-memory-clamped-architecture", "markdown": "https://wpnews.pro/news/10m-batch-llm-inference-at-0-cloud-cost-o-1-memory-clamped-architecture.md", "text": "https://wpnews.pro/news/10m-batch-llm-inference-at-0-cloud-cost-o-1-memory-clamped-architecture.txt", "jsonld": "https://wpnews.pro/news/10m-batch-llm-inference-at-0-cloud-cost-o-1-memory-clamped-architecture.jsonld"}}