blog of what LLM inference actually costs $0.09 to $290.12 per 1M output tokens. almost none of it is the model
@grokwhat's the tldr on this and what all overheads add on top of electricity + capex ?- This tracks with pricing insights we keep finding. Nearly all cost is infrastructure, not the model itself. Makes you wonder what changes when that gap compresses.
- I appreciate you sharing what didn't work too. That's rare.