{"slug": "triple-bottom-line-sustainability-of-language-models-for-edge-ai-a-comparison", "title": "Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs", "summary": "A new arXiv study introduces a Holistic Sustainability Score (HSS) to compare small language models (SLMs) and quantized large language models (LLMs) for edge AI, finding that Qwen3-30B-A3B/GGUF Q4 ranks first with a score of 93.38, followed by Mistral-Small-24B/GGUF Q4 at 92.40, while the highest-ranked SLM, Phi-4-mini/BF16, scores 89.49. The study concludes that native SLMs are not universally the most sustainable edge choice, as optimized quantized LLMs can win overall, though SLMs remain competitive with lower resource demand.", "body_md": "arXiv:2609.00665v1 Announce Type: new\nAbstract: Edge-AI model selection is commonly driven by one isolated metric - accuracy, latency, memory, energy, or safety, even though a deployable language model must balance all five. Our work focuses on answering the question whether na- tively trained small language models (SLMs) or large language models (LLMs) compressed through post-training quantization offer the more sustainable edge- deployment trade-off. We introduce a reproducible Holistic Sustainability Score (HSS) organized around the triple bottom line: an economic pillar for capability and systems efficiency, an environmental pillar for operational GPU energy and a social pillar for harmful-prompt robustness. Five BF16 SLMs and five LLMs under different quantization approaches - BF16, INT8, NF4 4-bit, GPTQ 4-bit, and GGUF Q4 produce 30 measured configurations. Capability is assessed on five zero-shot benchmarks; efficiency uses latency, throughput, peak VRAM and energy; and safety is approximated by attack success rate on five harmful prompts. Qwen3-30B-A3B/GGUF Q4 ranks first in the combined pool (93.38), followed by Mistral-Small-24B/GGUF Q4 (92.40), while Phi-4-mini/BF16 is the highest- ranked SLM in that pool (89.49). Thus, the hypothesis that native SLMs must be the most sustainable edge choice is not supported universally; optimized quantized LLMs can win overall, while SLMs remain competitive through lower resource demand. Quantization is a systems-level choice rather than a monotonic precision- efficiency trade-off and HSS remains relative to its comparison pool and proxy definitions.", "url": "https://wpnews.pro/news/triple-bottom-line-sustainability-of-language-models-for-edge-ai-a-comparison", "canonical_source": "https://www.machinebrief.com/news/triple-bottom-line-sustainability-of-language-models-for-edg-8gkt", "published_at": "2026-09-02 04:00:00+00:00", "updated_at": "2026-09-02 04:25:15.233276+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["arXiv", "Qwen3-30B-A3B", "Mistral-Small-24B", "Phi-4-mini"], "alternates": {"html": "https://wpnews.pro/news/triple-bottom-line-sustainability-of-language-models-for-edge-ai-a-comparison", "markdown": "https://wpnews.pro/news/triple-bottom-line-sustainability-of-language-models-for-edge-ai-a-comparison.md", "text": "https://wpnews.pro/news/triple-bottom-line-sustainability-of-language-models-for-edge-ai-a-comparison.txt", "jsonld": "https://wpnews.pro/news/triple-bottom-line-sustainability-of-language-models-for-edge-ai-a-comparison.jsonld"}}