High cloud API costs and Out-Of-Memory (OOM) failures in large-scale data pipelines are architectural defects, not hardware constraints.
This report documents the performance of a 10,000,000-record batch LLM inference workload executed locally on an HP Z4 G4 workstation (Intel Xeon W-2155, 128GB ECC DDR4 RAM, NVIDIA RTX 5060 Ti 16GB VRAM, NVMe PCIe Gen4 SSD).
PRAGMA integrity_check: ok)
glibc malloc exhibits heap fragmentation under sustained allocation and deallocation cycles. Injecting libjemalloc2 via LD_PRELOAD with active background thread decay (MALLOC_CONF="background_thread:true,dirty_decay_ms:2000,muzzy_decay_ms:2000") prevents process RSS growth and maintains a predictable host RAM footprint.
In-memory array construction introduces O(N) spatial memory complexity. Streaming chunked outputs directly to disk using Polars (collect(engine="streaming") and sink_parquet) ensures memory usage remains independent of row count.
State recovery and record tracking are managed via an aiosqlite pipeline using Write-Ahead Logging (PRAGMA journal_mode=WAL; PRAGMA synchronous=NORMAL;). Explicit BEGIN IMMEDIATE transaction blocks are executed every 50,000 items, with periodic WAL truncation to guarantee system crash consistency and database integrity.
Transmitting unmasked enterprise datasets to external APIs introduces regulatory compliance risks under HIPAA and GDPR. Executing high-throughput LLM workloads on local, air-gapped infrastructure (--network none) prevents telemetry transmission and external data exposure.
Reproducible Code & Repository:
https://github.com/Matsubara-CEO