08:23
2026-08-04
snipvote.com
large-language-models
Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware
A preliminary benchmark from arXiv (2608.00008) found that Gemma3:1B and Llama3.2:1B models achieve over 170 tokens per second on a single RTX 4060Ti GPU while consuming only 0.56β0.65 joules per tokeβ¦