# Free GitHub Agent Frameworks I Ship With

> Source: <https://dev.to/lamingsrb/free-github-agent-frameworks-i-ship-with-343>
> Published: 2026-09-28 06:34:34+00:00

Last month I rebuilt a piece of my content pipeline that had been running on a hand-rolled agent loop for about a year. The rewrite took a weekend because I finally leaned on open-source frameworks instead of maintaining my own scaffolding. This post is the honest tour: which free GitHub repos I actually run in production, where each one earns its keep, and where I've been burned.

Everything here is Apache-2.0 or MIT. No paid tier required to ship. The only money you spend is on model tokens and infrastructure.

If you want the TL;DR before the details: **LangGraph** for anything with branching, retries, or human-in-the-loop; **CrewAI** for role-based content and research swarms; **AutoGen** for conversational multi-agent reasoning and code generation; **Pydantic AI** or **llama-index** agents when you want the smallest surface area possible; and **smolagents** from Hugging Face when you need code-writing agents that stay under 1,000 lines of dependencies.

Here is how I actually decide, on a real project:

| Framework | GitHub | Best for | State handling | My verdict | 
|---|---|---|---|---|
| LangGraph | langchain-ai/langgraph | Deterministic workflows with LLM steps | Explicit graph + checkpointer | My default for production | 
| CrewAI | crewAIInc/crewAI | Role-based content/research crews | Task passing | Great DX, watch the abstractions | 
| AutoGen | microsoft/autogen | Chat-driven multi-agent, code exec | Conversation history | Strong for R&D, heavier for prod | 
| Pydantic AI | pydantic/pydantic-ai | Typed tool-calling agents | Minimal, you own it | Underrated, boring in a good way | 
| smolagents | huggingface/smolagents | Code-writing agents, tiny footprint | In-memory | Perfect for narrow tools | 

The rest of the post is what I wish someone had told me before I picked one.

LangGraph (github.com/langchain-ai/langgraph) is what I use for the orchestration layer in my BizFlowAI ContentStudio pipeline. It is a graph runtime: nodes are functions (usually LLM calls or tools), edges are transitions, and state is a typed dict that flows through. That model matches how production agent work actually behaves. You are not chatting with a magic entity, you are moving a piece of state through a series of decisions and side effects.

What makes it stick for real systems:

`SqliteSaver` and `PostgresSaver` let you resume a run after a crash, replay from any node, or hand control to a human and come back later. In my content pipeline, if the "publish" node fails because a WordPress endpoint is down, the graph resumes exactly there on the next scheduled run. No re-running the $0.40 of research.`add_conditional_edges` and the failure modes are visible.
The gotcha I hit: **do not put your entire application state in one giant TypedDict.** Split state per subgraph. When I had 22 fields flowing through 14 nodes, prompt debugging became painful because I could not tell which node mutated which field. Now I use small subgraphs with their own state, composed into a parent graph.

A minimum viable node looks like this:

``` python
from langgraph.graph import StateGraph, END
from typing import TypedDict

class State(TypedDict):
    topic: str
    draft: str
    approved: bool

def research(state: State) -> State:
    # call your LLM, return partial state update
    return {"draft": llm_draft(state["topic"])}

def review(state: State) -> State:
    return {"approved": llm_review(state["draft"])}

g = StateGraph(State)
g.add_node("research", research)
g.add_node("review", review)
g.add_edge("research", "review")
g.add_conditional_edges("review", lambda s: END if s["approved"] else "research")
g.set_entry_point("research")
app = g.compile()
```

That is 20 lines and you already have retry logic, resumability (once you add a checkpointer), and observability via LangSmith or your own logger.

CrewAI (github.com/crewAIInc/crewAI) leans into the "give each agent a role, a goal, and a backstory" metaphor. I was skeptical at first because that sounded like anthropomorphized fluff. It turned out to be a useful abstraction for content workflows specifically, because SEO content really is a small team: researcher, outliner, writer, editor, SEO reviewer.

I use CrewAI for one specific sub-pipeline: **long-form article generation with three specialized roles**. It ships tomorrow, not next month, because the framework does the boring parts (task chaining, output parsing, tool binding) with about 40 lines of YAML or Python.

Real numbers from my setup:

`response_format`
Where CrewAI hurts: **state passing between tasks is loose.** If task B needs a specific field from task A, you often end up parsing the previous task's freeform output. For anything with real branching, I graduate to LangGraph. CrewAI is where I start, not where I end.

Also: pin your version. The API surface has moved several times. I keep `crewai==0.` pinned exactly and read the changelog before bumping.

Microsoft's AutoGen (github.com/microsoft/autogen) treats multi-agent work as a conversation. Agents talk, a group chat manager decides who speaks next, and you can drop a code-executor agent in the middle. The v0.4 rewrite made it more production-friendly with an async event-driven core, but it is still the heaviest of the three big frameworks.

I use AutoGen for one thing in production: a **research and synthesis loop** where a critic agent challenges a writer agent until the writer produces something with sourced claims. The back-and-forth genuinely improves output for research-heavy pieces. For pure content generation it is overkill.

Trade-off I've measured on the same input topic:

| Framework | Tokens used | Wall clock | Output quality (my rubric) | 
|---|---|---|---|
| LangGraph, single pass | ~4k | 12s | 7/10 | 
| CrewAI, 4 roles | ~9k | 90s | 8/10 | 
| AutoGen, critic loop | ~18k | 140s | 8.5/10 | 

That 0.5 quality bump costs 2x the tokens of CrewAI and 4x LangGraph. For a landing page hero, worth it. For a programmatic SEO page, absolutely not. Match the framework to the unit economics.

The frameworks above are opinionated. Sometimes you want almost nothing between you and the model.

**Pydantic AI** (github.com/pydantic/pydantic-ai) is my pick when I need a typed tool-calling agent inside an existing FastAPI service. It is written by the Pydantic team, so the validation story is airtight. Tool definitions are Python functions with type hints. Output types are Pydantic models. That is the whole framework. If your agent is really just "LLM plus a few tools plus structured output", this saves you from importing 400 MB of dependencies.

**smolagents** (github.com/huggingface/smolagents) from Hugging Face is worth studying even if you do not adopt it. Its core idea is that agents should write code, not JSON, to call tools. In practice this means fewer schema errors and more expressive multi-step reasoning in a single generation. I use it for a narrow internal tool that scrapes and normalizes data. Under 1,000 lines of framework code. You can read the entire source in an afternoon.

These are the ones that cost me real time. In no particular order.

`ChatAnthropic` and `ChatOpenAI` behave differently around streaming, tool calls, and system prompts. If you swap providers, test each tool call path. I once shipped a bug where a tool worked fine on Claude but silently returned a stringified null on OpenAI because the function-calling shape differed.`asyncio.Semaphore(3)` for Anthropic and `Semaphore(8)` for OpenAI, tuned to my tier.
If you asked me to greenfield a multi-agent system this week, here is what I would use:

That combination has kept my own content pipeline running 24/7 with almost no intervention. When something breaks, the checkpointer plus Langfuse traces tell me exactly where within a couple of minutes.

Pick LangGraph. Not because it is the most exciting, but because it forces you to think in states and transitions, which is how you have to reason about production agent systems anyway. Add CrewAI only when you have a genuine "team of roles" problem. Reach for AutoGen when a critic loop measurably improves output on your specific task. Keep Pydantic AI in your back pocket for the small stuff.

Do not adopt a framework because a tutorial made it look pretty. Adopt it because you can name the specific failure mode you are trying to prevent. In my experience the failures that matter are: **losing state on retry, silent tool-call errors, unbounded token spend, and lack of observability.** Every framework I named above solves at least three of those. Some solve all four, if you configure them right.

Free open-source frameworks are where I do 90% of my agent work. The paid platforms make sense at a specific scale and for specific compliance stories, but you can ship real revenue-generating systems with the repos above and a Postgres database.

If you are building something along these lines and want a second pair of eyes from someone who runs this stack in production, I take a small number of engagements each quarter. You can reach me at [lazar-milicevic.com/#contact](https://lazar-milicevic.com/#contact), or read more posts on the blog if you want to see how the pieces fit together on real projects.
