{"slug": "i-tested-cloudflare-workers-ai-free-tier-and-found-24-powerful-llms-you-can-use", "title": "I Tested Cloudflare Workers AI Free Tier and Found 24 Powerful LLMs You Can Use Today", "summary": "A developer tested Cloudflare Workers AI's free tier and found 24 open-source LLMs that accept inference requests through a single API endpoint, including OpenAI's gpt-oss-120b, Qwen2.5-Coder-32B, and DeepSeek-R1-Distill-Qwen-32B. The developer reported that the catalog lists over 300 models but several newer ones, such as GLM 5.3 and Kimi K2.6, return a 5035 error stating they are unavailable on the Workers Free plan. GPT-OSS-120B was ranked the best overall coding model, with Qwen2.5-Coder-32B and DeepSeek-R1-Distill-Qwen-32B placed second and third.", "body_md": "Most developers know about OpenAI, Anthropic, Groq, and Cerebras.\n\nWhat many developers don't realize is that Cloudflare Workers AI provides access to a large collection of open-source models through a single API endpoint.\n\nThe interesting part?\n\nMany of these models are usable on Cloudflare's free tier.\n\nI recently decided to test Cloudflare Workers AI to answer three questions:\n\nWhich models are actually available on the free plan?\n\nWhich models are best for coding?\n\nHow do Cloudflare's limits compare to Groq, Cerebras, and other providers?\n\nThe results were honestly surprising.\n\nThe Experiment\n\nCloudflare exposes a model catalog API:\n\nInvoke-RestMethod \n\n  -Uri \"https://api.cloudflare.com/client/v4/accounts/$AccountId/ai/models/search\"\n\n  -Headers $Headers\n\nMy account returned more than 300 models.\n\nInitially, I assumed that every model listed would be available.\n\nI was wrong.\n\nSome models appeared in the catalog but were blocked when inference requests were executed.\n\nFor example:\n\nzai-org/glm-5.3\n\nzai-org/glm-5.3-flash\n\nzai-org/glm-5.2\n\nmoonshotai/kimi-k2.6\n\nmoonshotai/kimi-k2.7-code\n\nreturned:\n\n{\n\n  \"code\": 5035,\n\n  \"message\": \"Model is not available on the Workers Free plan\"\n\n}\n\nThis led me to create a discovery script that:\n\nEnumerated all available models\n\nExecuted a test inference\n\nRecorded successful responses\n\nRecorded paid-plan restrictions\n\nModels That Worked on the Free Tier\n\nAfter testing, these models successfully accepted inference requests.\n\nOpenAI\n\nopenai/gpt-oss-20b\n\nopenai/gpt-oss-120b\n\nQwen\n\nqwen/qwen2.5-coder-32b-instruct\n\nqwen/qwen3-30b-a3b-fp8\n\nqwen/qwen3.8-27b\n\nqwen/qwq-32b\n\nDeepSeek\n\ndeepseek-ai/deepseek-r1-distill-qwen-32b\n\nMeta Llama\n\nmeta/llama-3.1-8b-instruct-fp8\n\nmeta/llama-3.2-1b-instruct\n\nmeta/llama-3.2-3b-instruct\n\nmeta/llama-3.3-70b-instruct-fp8-fast\n\nmeta/llama-4-scout-17b-16e-instruct\n\nGoogle Gemma\n\ngoogle/gemma-2b-it-lora\n\ngoogle/gemma-7b-it-lora\n\ngoogle/gemma-4-26b-a4b-it\n\naisingapore/gemma-sea-lion-v4-27b-it\n\nMistral\n\nmistral/mistral-7b-instruct-v0.2-lora\n\nmistralai/mistral-small-3.1-24b-instruct\n\nZAI\n\n[@cf](https://dev.to/cf)/zai-org/glm-4.7-flash\n\nIBM\n\nibm-granite/granite-4.0-h-micro\n\nNVIDIA\n\nnvidia/nemotron-3-120b-a12b\n\nBiggest Surprise\n\nI initially started testing because I wanted access to GLM 5.3 Flash.\n\nThe result?\n\nGLM 4.7 Flash ✅\n\nGLM 5.2 ❌\n\nGLM 5.3 ❌\n\nGLM 5.3 Flash ❌\n\nGLM 4.7 Flash was available on the free tier while all newer GLM models required a paid plan.\n\nBest Models For Coding\n\nAfter testing many providers over the last year, this is how I would rank the Cloudflare free models for software development.\n\n🥇 GPT-OSS-120B\n\n[@cf](https://dev.to/cf)/openai/gpt-oss-120b\n\nBest overall coding model.\n\nExcellent at:\n\nTerraform\n\nKubernetes\n\nDevOps\n\nAWS Architecture\n\nRefactoring\n\nMulti-file repositories\n\nIf I could pick only one model, it would be GPT-OSS-120B.\n\n🥈 Qwen2.5-Coder-32B\n\n[@cf](https://dev.to/cf)/qwen/qwen2.5-coder-32b-instruct\n\nPurpose-built coding model.\n\nExcellent at:\n\nGenerating code\n\nReviewing pull requests\n\nFixing bugs\n\nUnderstanding repositories\n\nContinue.dev\n\nCline\n\nThis model consistently performs above its size class.\n\n🥉 DeepSeek-R1-Distill-Qwen-32B\n\n[@cf](https://dev.to/cf)/deepseek-ai/deepseek-r1-distill-qwen-32b\n\nBest reasoning model.\n\nExcellent at:\n\nRoot cause analysis\n\nComplex debugging\n\nArchitecture reviews\n\nAgent workflows\n\nThis is the model I would use when a pipeline breaks at 2 AM and nobody knows why.\n\nHonorable Mentions\n\nQWQ-32B\n\n[@cf](https://dev.to/cf)/qwen/qwq-32b\n\nVery strong reasoning.\n\nLlama 4 Scout\n\n[@cf](https://dev.to/cf)/meta/llama-4-scout-17b-16e-instruct\n\nExcellent balance of speed and intelligence.\n\nLlama 3.3 70B\n\n[@cf](https://dev.to/cf)/meta/llama-3.3-70b-instruct-fp8-fast\n\nGreat for reviews, documentation, and system design.\n\nUnderstanding Cloudflare Limits\n\nOne thing I wanted to understand was whether Cloudflare behaves like Groq.\n\nThe answer is no.\n\nGroq commonly applies limits per model.\n\nFor example:\n\nModel A → X RPM\n\nModel B → Y RPM\n\nCloudflare works differently.\n\nCloudflare's free tier provides:\n\n10,000 Neurons per day\n\nThis is effectively a daily AI budget shared across your account.\n\nCloudflare also documents:\n\n300 requests per minute\n\nfor text generation workloads.\n\nSo the practical model looks like:\n\nAccount Level Daily Budget\n\n+\n\nText Generation RPM\n\n+\n\nModel-Specific Overrides\n\nThis is much closer to an account-level quota system than Groq's model-centric approach.\n\nWhy This Matters\n\nMany developers assume they need:\n\nOpenAI Subscription\n\nAnthropic Subscription\n\nGPU Server\n\nRunPod\n\nAWS Inference Endpoint\n\nbefore they can build an AI product.\n\nFor many projects, that's no longer true.\n\nA single free Cloudflare account can already provide access to:\n\nGPT-OSS-120B\n\nQwen Coder 32B\n\nDeepSeek R1\n\nQWQ 32B\n\nLlama 4 Scout\n\nGemma 4\n\nthrough a single API.\n\nThat's enough to power:\n\nCoding assistants\n\nAI agents\n\nInternal copilots\n\nRAG applications\n\nStartup MVPs\n\nDeveloper tools\n\nwithout touching a GPU.\n\nFinal Thoughts\n\nI started this experiment trying to use GLM 5.3 Flash on the Cloudflare free plan.\n\nInstead, I discovered something much more valuable.\n\nCloudflare Free currently provides access to some genuinely powerful models, including GPT-OSS-120B, Qwen Coder 32B, DeepSeek R1 Distill, QWQ 32B, and Llama 4 Scout.\n\nFor indie hackers, startup founders, and developers building agents, Cloudflare Workers AI might be one of the most underrated free inference platforms available today.\n\nIf you're experimenting with AI tooling, it is absolutely worth testing before paying for another inference provider.\n\nWhat I Plan To Test Next\n\nHow many real coding requests fit within 10,000 free neurons?\n\nWhich free model provides the best cost-to-quality ratio?\n\nCloudflare vs Groq vs Cerebras benchmarks\n\nUsing Cloudflare Workers AI with Continue.dev\n\nBuilding AI agents entirely on free infrastructure\n\nStay tuned.", "url": "https://wpnews.pro/news/i-tested-cloudflare-workers-ai-free-tier-and-found-24-powerful-llms-you-can-use", "canonical_source": "https://dev.to/interro-ai/i-tested-cloudflare-workers-ai-free-tier-and-found-24-powerful-llms-you-can-use-today-511h", "published_at": "2026-09-28 05:16:12+00:00", "updated_at": "2026-09-28 05:48:04.102276+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-tools", "developer-tools"], "entities": ["Cloudflare", "Cloudflare Workers AI", "OpenAI", "Qwen", "DeepSeek", "Meta", "Groq", "Cerebras"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-tested-cloudflare-workers-ai-free-tier-and-found-24-powerful-llms-you-can-use", "markdown": "https://wpnews.pro/news/i-tested-cloudflare-workers-ai-free-tier-and-found-24-powerful-llms-you-can-use.md", "text": "https://wpnews.pro/news/i-tested-cloudflare-workers-ai-free-tier-and-found-24-powerful-llms-you-can-use.txt", "jsonld": "https://wpnews.pro/news/i-tested-cloudflare-workers-ai-free-tier-and-found-24-powerful-llms-you-can-use.jsonld"}}