OutOfMemoryError
before the first token is even generated.The crash happens during the weights sequence. Here is the specific error I'm seeing in my logs:
torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 2.4GB (GPU 0; 24.00GB total capacity), 18.1GB already allocated.
After digging into the telemetry, it seems like the issue isn't the model size itself, but how the tensors are being allocated in the KV cache or a temporary peak during the conversion from half-precision to quantized formats. I tried limiting the context window to 2048 to see if that would stabilize the footprint, but the spike occurs during the from_pretrained
call, not during inference.
I suspect there's a mismatch between the 's memory mapping and the actual GPU shards. I'm currently testing if switching to a different backend or updating the transformers library resolves this, as it feels like a resource management bug rather than a hardware limitation. If anyone has a stable deployment config for this specific version of Gemma, I'd be interested in seeing your environment specs.
Next Brain Waves: The Next Frontier for Physical AI →