cd/entity/Qwen3-32B· home entities Qwen3-32B
grep -l @qwen3-32b /news/*.json | wc -l → 54

Qwen3-32B

mentions 54 type Organization page 2/3 feed RSS

// recent coverage 54 mentions

10:43
2026-06-26
dev.to
large-language-models

I Wish I Knew About This OpenAI Swap Sooner — Full Breakdown

An engineer at a company using OpenAI's GPT-4o for LLM inference discovered they were overpaying by up to 40x compared to alternatives like DeepSeek V4 Flash served through Global API. After benchmark…

03:38
2026-06-24
dev.to
large-language-models

Line AI Chatbot In Production: A CTO's Honest Breakdown

A CTO cut inference costs by 40-65% by replacing GPT-4o with a mix of DeepSeek, Qwen, and GLM models via the Line AI Chatbot framework, which uses a model-agnostic API to avoid vendor lock-in. The sys…

11:20
2026-06-21
dev.to
artificial-intelligence

The CTO Playbook for AI Agent Data Analysis on a Budget

A startup CTO cut AI agent data analysis costs by 40-65% by replacing GPT-4o with cheaper models like GLM-4 Plus for 85% of traffic, using a routing layer that classifies queries and dispatches to app…

08:06
2026-06-21
dev.to
artificial-intelligence

I Built an AI Tutor in 48 Hours and Heres What Blew My Mind

A developer built an AI tutoring app in 48 hours using the Global API, which provides access to 184 models. By benchmarking models, they found that GLM-4 Plus at $0.80 per million output tokens and De…

15:11
2026-06-19
dev.to
large-language-models

How I Slashed AI API Costs 60% as a Cloud Architect

A cloud architect rebuilt their inference layer to slash AI API costs by 60% while maintaining sub-2-second p99 latency. By implementing a tiered model routing system that directs simple queries to ch…

11:59
2026-06-19
dev.to
large-language-models

How I Compared Context Windows Across 184 LLM Models in 2026

A developer compared context windows across 184 LLM models in 2026, finding that matching window size to workload can reduce costs by 40-65%. Switching from a 128K model to a smarter routing strategy …

09:56
2026-06-19
dev.to
artificial-intelligence

What I Learned Running Airtable AI Across Three Regions at p99

An engineer at Airtable shared lessons from deploying Airtable AI across three regions with p99 latency under 1.8 seconds and 99.94% uptime. By routing queries to different models based on complexity,…

15:05
2026-06-17
dev.to
large-language-models

DeepSeek vs Gemini 2.0 Pro: Which AI API Actually Wins in 2026?

A developer evaluated DeepSeek V4 Flash and Gemini 2.0 Pro for a high-volume ranking system processing 12 million inference calls daily, finding DeepSeek V4 Flash achieved a p99 latency of 1.18 second…

06:33
2026-06-17
arxiv.org
machine-learning

Fearless Concurrency on the GPU

Researchers introduced cuTile Rust, a tile-based system for safe, idiomatic GPU kernel authoring in Rust that extends Rust's ownership discipline to GPU kernels. On the NVIDIA B200 GPU, cuTile Rust ac…

00:53
2026-06-17
dev.to
large-language-models

How I Cut Costs 65% Migrating LangChain to DeepSeek

A developer cut inference costs by 65% by migrating a LangChain pipeline from GPT-4o to DeepSeek models via Global API. The switch, which took about ten minutes, reduced monthly costs from $1,000 to $…

20:21
2026-06-16
dev.to
artificial-intelligence

Notion AI's Pricing Trap: Why I Went Open Source Instead

A developer abandoned Notion AI after its pricing ballooned, opting for open-source alternatives. Benchmarking showed Notion AI's optimized 2026 stack offered 40-65% cost reduction but relied on commu…

16:03
2026-06-16
dev.to
artificial-intelligence

From Walled Garden to Open Road: A DeepSeek API Nestjs Story

A developer built a NestJS-based inference layer using DeepSeek's open API after receiving a $4,200 invoice from a proprietary AI vendor. The setup provides access to 184 models at prices ranging from…

← prev page 2 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics