# I Tested Cloudflare Workers AI Free Tier and Found 24 Powerful LLMs You Can Use Today

> Source: <https://dev.to/interro-ai/i-tested-cloudflare-workers-ai-free-tier-and-found-24-powerful-llms-you-can-use-today-511h>
> Published: 2026-09-28 05:16:12+00:00

Most developers know about OpenAI, Anthropic, Groq, and Cerebras.

What many developers don't realize is that Cloudflare Workers AI provides access to a large collection of open-source models through a single API endpoint.

The interesting part?

Many of these models are usable on Cloudflare's free tier.

I recently decided to test Cloudflare Workers AI to answer three questions:

Which models are actually available on the free plan?

Which models are best for coding?

How do Cloudflare's limits compare to Groq, Cerebras, and other providers?

The results were honestly surprising.

The Experiment

Cloudflare exposes a model catalog API:

Invoke-RestMethod 

  -Uri "https://api.cloudflare.com/client/v4/accounts/$AccountId/ai/models/search"

  -Headers $Headers

My account returned more than 300 models.

Initially, I assumed that every model listed would be available.

I was wrong.

Some models appeared in the catalog but were blocked when inference requests were executed.

For example:

zai-org/glm-5.3

zai-org/glm-5.3-flash

zai-org/glm-5.2

moonshotai/kimi-k2.6

moonshotai/kimi-k2.7-code

returned:

{

  "code": 5035,

  "message": "Model is not available on the Workers Free plan"

}

This led me to create a discovery script that:

Enumerated all available models

Executed a test inference

Recorded successful responses

Recorded paid-plan restrictions

Models That Worked on the Free Tier

After testing, these models successfully accepted inference requests.

OpenAI

openai/gpt-oss-20b

openai/gpt-oss-120b

Qwen

qwen/qwen2.5-coder-32b-instruct

qwen/qwen3-30b-a3b-fp8

qwen/qwen3.8-27b

qwen/qwq-32b

DeepSeek

deepseek-ai/deepseek-r1-distill-qwen-32b

Meta Llama

meta/llama-3.1-8b-instruct-fp8

meta/llama-3.2-1b-instruct

meta/llama-3.2-3b-instruct

meta/llama-3.3-70b-instruct-fp8-fast

meta/llama-4-scout-17b-16e-instruct

Google Gemma

google/gemma-2b-it-lora

google/gemma-7b-it-lora

google/gemma-4-26b-a4b-it

aisingapore/gemma-sea-lion-v4-27b-it

Mistral

mistral/mistral-7b-instruct-v0.2-lora

mistralai/mistral-small-3.1-24b-instruct

ZAI

[@cf](https://dev.to/cf)/zai-org/glm-4.7-flash

IBM

ibm-granite/granite-4.0-h-micro

NVIDIA

nvidia/nemotron-3-120b-a12b

Biggest Surprise

I initially started testing because I wanted access to GLM 5.3 Flash.

The result?

GLM 4.7 Flash ✅

GLM 5.2 ❌

GLM 5.3 ❌

GLM 5.3 Flash ❌

GLM 4.7 Flash was available on the free tier while all newer GLM models required a paid plan.

Best Models For Coding

After testing many providers over the last year, this is how I would rank the Cloudflare free models for software development.

🥇 GPT-OSS-120B

[@cf](https://dev.to/cf)/openai/gpt-oss-120b

Best overall coding model.

Excellent at:

Terraform

Kubernetes

DevOps

AWS Architecture

Refactoring

Multi-file repositories

If I could pick only one model, it would be GPT-OSS-120B.

🥈 Qwen2.5-Coder-32B

[@cf](https://dev.to/cf)/qwen/qwen2.5-coder-32b-instruct

Purpose-built coding model.

Excellent at:

Generating code

Reviewing pull requests

Fixing bugs

Understanding repositories

Continue.dev

Cline

This model consistently performs above its size class.

🥉 DeepSeek-R1-Distill-Qwen-32B

[@cf](https://dev.to/cf)/deepseek-ai/deepseek-r1-distill-qwen-32b

Best reasoning model.

Excellent at:

Root cause analysis

Complex debugging

Architecture reviews

Agent workflows

This is the model I would use when a pipeline breaks at 2 AM and nobody knows why.

Honorable Mentions

QWQ-32B

[@cf](https://dev.to/cf)/qwen/qwq-32b

Very strong reasoning.

Llama 4 Scout

[@cf](https://dev.to/cf)/meta/llama-4-scout-17b-16e-instruct

Excellent balance of speed and intelligence.

Llama 3.3 70B

[@cf](https://dev.to/cf)/meta/llama-3.3-70b-instruct-fp8-fast

Great for reviews, documentation, and system design.

Understanding Cloudflare Limits

One thing I wanted to understand was whether Cloudflare behaves like Groq.

The answer is no.

Groq commonly applies limits per model.

For example:

Model A → X RPM

Model B → Y RPM

Cloudflare works differently.

Cloudflare's free tier provides:

10,000 Neurons per day

This is effectively a daily AI budget shared across your account.

Cloudflare also documents:

300 requests per minute

for text generation workloads.

So the practical model looks like:

Account Level Daily Budget

+

Text Generation RPM

+

Model-Specific Overrides

This is much closer to an account-level quota system than Groq's model-centric approach.

Why This Matters

Many developers assume they need:

OpenAI Subscription

Anthropic Subscription

GPU Server

RunPod

AWS Inference Endpoint

before they can build an AI product.

For many projects, that's no longer true.

A single free Cloudflare account can already provide access to:

GPT-OSS-120B

Qwen Coder 32B

DeepSeek R1

QWQ 32B

Llama 4 Scout

Gemma 4

through a single API.

That's enough to power:

Coding assistants

AI agents

Internal copilots

RAG applications

Startup MVPs

Developer tools

without touching a GPU.

Final Thoughts

I started this experiment trying to use GLM 5.3 Flash on the Cloudflare free plan.

Instead, I discovered something much more valuable.

Cloudflare Free currently provides access to some genuinely powerful models, including GPT-OSS-120B, Qwen Coder 32B, DeepSeek R1 Distill, QWQ 32B, and Llama 4 Scout.

For indie hackers, startup founders, and developers building agents, Cloudflare Workers AI might be one of the most underrated free inference platforms available today.

If you're experimenting with AI tooling, it is absolutely worth testing before paying for another inference provider.

What I Plan To Test Next

How many real coding requests fit within 10,000 free neurons?

Which free model provides the best cost-to-quality ratio?

Cloudflare vs Groq vs Cerebras benchmarks

Using Cloudflare Workers AI with Continue.dev

Building AI agents entirely on free infrastructure

Stay tuned.
