cd /news/ai-agents/the-assistants-api-is-gone-here-s-ho… · home › topics › ai-agents › article
[ARTICLE · art-143946] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

The Assistants API is Gone. Here's How to Build Agents Now.

OpenAI shut down its Assistants API on August 26, 2026, ending the managed persistent-thread model for building stateful conversational agents. Developers must now migrate to the stateless Responses and Conversations APIs, where they manage conversation history themselves, and use the new file_search tool with vector stores for retrieval-augmented generation instead of the old Retrieval tool. The shift trades convenience for more control, better performance and more predictable costs.

by read4 min views4 publishedOct 2, 2026

The OpenAI Assistants API was officially shut down on August 26, 2026. This deprecation forces a fundamental shift in how we build stateful agents, moving from a managed, persistent-thread model to a more direct, stateless approach with the new Responses and Conversations APIs. This change simplifies some parts of the stack but places the responsibility for state management squarely back on you, the developer.

The original Assistants API was an attempt to abstract away the complexity of building conversational AI. It provided a stateful environment with persistent "threads," allowing developers to build agents that could maintain context over long interactions without managing the conversation history themselves. It also bundled powerful tools like Code Interpreter and a retrieval system.

However, this abstraction came with trade-offs. The API could feel clunky, and its stateful nature led to unpredictable costs, as the entire conversation thread might be re-processed on every turn. Performance was also a concern, as developers had to poll for updates rather than using a real-time stream.

The new model, centered on the Responses API, is a return to a more primitive, stateless paradigm. The core idea is a direct request/response flow. Persistent threads are replaced by Conversation objects, which you must create and manage. The responsibility for maintaining context between turns now falls to your application code. This is a significant architectural change, but one that offers more control, better performance, and more predictable costs.

For developers building Retrieval-Augmented Generation (RAG) systems, the most important change is the introduction of the file_search tool within the Responses API. This is the successor to the old Retrieval tool and is now the standard way to have models access knowledge from your private documents.

The workflow is straightforward: you create a vector_store, upload your files to it, and OpenAI's backend handles the entire pipeline of chunking, embedding, and indexing. You no longer have to build and manage your own embedding and retrieval logic.

When you make a call to the Responses API, you can make the file_search tool available. The model then intelligently decides when to use it based on the user's query. It performs a semantic search against your vector store, retrieves the most relevant passages, and incorporates them into its response, complete with citations.

Here is a conceptual look at what an API call might look like in Python, enabling the tool for a specific vector store:

from openai import OpenAI
client = OpenAI()

vector_store_id = "vs_123abc"

conversation = client.conversations.create()
client.conversations.messages.create(
    conversation_id=conversation.id,
    role="user",
    content="What were the key findings in the Q3 financial report?"
)

response = client.responses.create(
    model="gpt-6.1-sol",
    conversation_id=conversation.id,
    tools=[{
        "type": "file_search",
        "file_search": {
            "vector_store_ids": [vector_store_id]
        }
    }]
)

print(response.output.text.value)

The primary gain from this migration is control. The move to a stateless API gives you direct authority over conversation history and state management. This means more predictable performance and costs, as you are no longer subject to the black box of the old persistent-thread system. The Responses API also consolidates complex workflows into a single API call, simplifying the overall interaction model.

The most significant loss is convenience. The hand-holding of the managed, infinite-context thread is gone. If your application was architected around the assumption that OpenAI would manage the state, you now have a non-trivial migration project ahead of you. OpenAI does not provide an automatic tool for migrating old Threads to new Conversations.

There is also a key technical trade-off with the managed file_search tool. While it removes the burden of building a RAG pipeline, it also takes away your control over the chunking strategy. For highly structured or complex documents, the automated chunking might not be optimal, which can impact the quality of retrieval. This is a critical limitation to be aware of when deciding whether to use the built-in tool or roll your own retrieval system.

This shift marks a maturation of the AI developer stack. The initial, heavily-abstracted Assistants API has been replaced by more powerful and primitive building blocks. It’s a move away from providing a magic box and toward giving engineers direct access to the core components of RAG and agentic systems. For builders, this means more responsibility, but also more power. The era of the fully managed agent is over; the era of building your own is here.

── more in #ai-agents 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-assistants-api-i…] indexed:0 read:4min 2026-10-02 · —