cd /news/large-language-models/10m-batch-llm-inference-at-0-cloud-c… · home › topics › large-language-models › article
[ARTICLE · art-147619] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

10M Batch LLM Inference at $0 Cloud Cost: O(1) Memory Clamped Architecture

A developer documented running a 10,000,000-record batch LLM inference workload entirely on a local HP Z4 G4 workstation at zero cloud cost, using jemalloc with background thread decay to prevent heap fragmentation, Polars streaming sinks to keep memory independent of row count, and an aiosqlite WAL pipeline with BEGIN IMMEDIATE checkpoints every 50,000 items for crash consistency. The writeup argues that high cloud API costs and OOM failures in large-scale pipelines are architectural defects rather than hardware limits, and that air-gapped local execution avoids HIPAA and GDPR exposure from sending unmasked enterprise data to external APIs.

by read1 min views1 publishedOct 8, 2026

High cloud API costs and Out-Of-Memory (OOM) failures in large-scale data pipelines are architectural defects, not hardware constraints.

This report documents the performance of a 10,000,000-record batch LLM inference workload executed locally on an HP Z4 G4 workstation (Intel Xeon W-2155, 128GB ECC DDR4 RAM, NVIDIA RTX 5060 Ti 16GB VRAM, NVMe PCIe Gen4 SSD).

PRAGMA integrity_check: ok) glibc malloc exhibits heap fragmentation under sustained allocation and deallocation cycles. Injecting libjemalloc2 via LD_PRELOAD with active background thread decay (MALLOC_CONF="background_thread:true,dirty_decay_ms:2000,muzzy_decay_ms:2000") prevents process RSS growth and maintains a predictable host RAM footprint.

In-memory array construction introduces O(N) spatial memory complexity. Streaming chunked outputs directly to disk using Polars (collect(engine="streaming") and sink_parquet) ensures memory usage remains independent of row count.

State recovery and record tracking are managed via an aiosqlite pipeline using Write-Ahead Logging (PRAGMA journal_mode=WAL; PRAGMA synchronous=NORMAL;). Explicit BEGIN IMMEDIATE transaction blocks are executed every 50,000 items, with periodic WAL truncation to guarantee system crash consistency and database integrity.

Transmitting unmasked enterprise datasets to external APIs introduces regulatory compliance risks under HIPAA and GDPR. Executing high-throughput LLM workloads on local, air-gapped infrastructure (--network none) prevents telemetry transmission and external data exposure.

Reproducible Code & Repository:

https://github.com/Matsubara-CEO

── more in #large-language-models 4 stories · sorted by recency
── more on @hp z4 g4 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/10m-batch-llm-infere…] indexed:0 read:1min 2026-10-08 · —