{"slug": "deleting-rag-from-our-web-agent-made-it-2-3x-faster", "title": "Deleting RAG from our web agent made it 2.3x faster", "summary": "Skyvern-AI rewrote its open source browser automation agent Skyvern without retrieval augmented generation (RAG), producing a version it says is 2.3x faster and scores 90.5% on the Odysseys benchmark. The original RAG-based Skyvern parsed and condensed each page's interactable elements before sending them to an LLM, but Skyvern-AI said DOM parsing of Shadow DOMs, iframes, select2 dropdowns, Kendo UI widgets and Canvas elements became an endless battle and bottlenecked agent performance regardless of model quality. The new agent starts with minimal context and uses tools to scan HTML with model-generated JavaScript, take rustwright-supported actions such as Click, Type, Scroll and Switch tabs, take screenshots, and mark runs final.", "body_md": "# Deleting RAG from our web agent made it 2.3x faster\n\n**TL;DR - We rewrote Skyvern from scratch using** **Pi** **as inspiration, making it 2.3x faster and score 90.5% on** **Odysseys benchmark****. Try it out the via** **Open Source** **or** **Cloud****.**\n\n💡 **Recap: What is Skyvern?** It’s [an open source](https://github.com/Skyvern-AI/Skyvern?ref=skyvern.com) framework that helps non-technical teams prompt to build browser automations. Skyvern is model agnostic, and can be run with closed models like Claude, or open ones like Deepseek. We have helped thousands of companies automate things like interacting with healthcare portals, fetching invoices or utility bills, filling out government forms, applying to jobs, and more.\n\n## RAG? What are you talking about?\n\nFor those unfamiliar with agents, RAG stands for Retrieval Augmented Generation. It’s a way to pass information to an LLM so they make smarter decisions. For example, before asking questions to an LLM about a document, you could parse that document and feed it into an LLM. You can read more about it [here](https://pub.towardsai.net/retrieval-augmented-generation-aka-rag-how-does-it-work-1ef69e3654dd?ref=skyvern.com).\n\nThe first version of Skyvern was built using RAG. We would load up a website, screenshot and annotate it, condense it into a smaller list of elements, and send it to an LLM to take actions on.\n\nThis worked great if you could identify every interactable element on a website with 100% accuracy. We thought we could eventually get there. We were wrong.\n\nExcept.. the web is messy. Accurately identifying every element on a website proved to be a nightmare. What started off as a [nice little utility to identify interactable elements](https://github.com/Skyvern-AI/skyvern/blob/539dc5d14e86fb760f7ab63e9667a01d1dd1a406/skyvern/webeye/scraper/domUtils.js?ref=skyvern.com#L1113) essentially turned into building a new DOM interpreter.\n\nParsing things like Shadow DOMs, iframes, select2 dropdowns, Kendo UI widgets, Canvas elements, …. this became an endless battle.\n\nThe worst part was that agent performance did not improve with better LLMs. Instead, it was bottlenecked by this DOM parsing logic. If it failed to correctly identify an element, the entire agent run failed.\n\n**This is a sign that we need to give the LLM more freedom.**\n\n## Give the LLM freedom to do what it wants\n\nRAG systems, if you aren’t careful, can assume that the pre-processing layer always happens ahead of any LLM calls, and don’t let the LLM decide when it is necessary. This was a good idea when LLMs weren’t good at following instructions, which meant that you couldn’t trust them to reliably pull the information when needed, but instruction following hasn’t been a big problem since the models released after Opus 4.5.\n\nWe noticed 2 moves in the market that highlighted this change in thinking:\n\n1. Coding agents like Claude code were moving away from embedding style code search to grep style code search because the models were increasingly getting better at using primitive tools to fetch the context they needed\n2. Minimalist coding agents like Pi were becoming increasingly popular, with extensions of Pi like [Prime-intellect](https://www.primeintellect.ai/?ref=skyvern.com) hitting SOTA on the[ARC-AGI-3 benchmark](https://x.com/PrimeIntellect/status/2085087000764568010?ref=skyvern.com) by tuning the harness and letting the agent generate and execute python code as needed\n\nSo we asked ourselves: **would the same approach work for browser agents?** What if the agent started with the minimum context required, and gave it the ability to pull additional context as needed?\n\nIn practice, it means giving the agent the following tools:\n\n1. Scan the HTML with model-generated javascript\n2. Take any [rustwright](https://github.com/Skyvern-AI/rustwright?ref=skyvern.com) -supported action on the website, for example Click, Type, Scroll, Switch tabs, etc\n3. Take a screenshot\n4. Mark the run as final (success, or failed with reason)\n\n## The results were better than we expected\n\nOur hypothesis was that we would see an increase in performance, in exchange for added cost and a slower processing speed. The new architecture should result more calls to the LLM to get the same task done.\n\nTo test this hypothesis, we split the test up into two branches:\n\n1. The leading public browser agent benchmark: Odysseys to test for accuracy\n2. A live A/B test across ~100,000 customer runs to test for speed / cost\n\n### Odysseys benchmark results\n\nSkyvern 3.0 topped the [Odysseys benchmark](https://odysseysbench.com/leaderboard?ref=skyvern.com). It reported a **90.5% perfect rubric success rate,** pretty much exactly as we expected. What surprised us though was that it was highly efficient in its runs, at only 65.4 average steps.\n\n### Production A/B Test results\n\nWe ran the production A/B test on approx 100,000 runs and measured the impact\n\n|  | Skyvern 2.0 | Skyvern 3.0 | Improvement | \n|---|---|---|---|\n| Mean run time | 593 s | 262s | 2.3x | \n| Total cost per run | $0.039 | $0.0302 | -22.5% | \n\nThe production A/B test results were surprising. The new architecture was.. faster and cheaper? We expected at least one of these dimensions to be a net negative\n\nWhy was that the case? Digging into it, we found a few interesting things:\n\n1. **Screenshot utilization dropped by 90% (27% —> 2.7% of all LLM calls)** . LLM calls with images go through a vision encoder to help the LLM “see” the image, which[tend to be 1.7x slower than text-only encoders](https://arxiv.org/html/2607.23373v1?ref=skyvern.com)\n2. **Token consumption increased by 52% (188K tokens / run —> 286K tokens per run)** Average number of turns (ie LLM Calls) increased by 3.59x (6.4 —> 23), but tokens per turn decreased by 59% (29K tokens —> 12K tokens)\n3. **Prompt Cache hit rates increased by 195% (27% —> 79.7%)** because we stopped loading the screenshot or HTML on every turn, which meant a larger % of the prompt was cacheable\n\nDespite the increased token usage, cached tokens carry a 70% discount, which explains the net reduction in cost.\n\n## So… are browser agents solved?\n\nThe answer is still **no**. They can still be improved across all 3 dimensions (cost, speed, accuracy).\n\nOur team is now exploring the role of [”system one models” like Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev?ref=skyvern.com) for browser agents. The current generation is still highly limited: 64K context window, no image processing, which causes unacceptable losses in accuracy, but we expect this to change over the next 6-12 months 🙂\n\n## Give Skyvern a try!\n\nRun it via our [Skyvern Open Source](https://github.com/Skyvern-AI/Skyvern?ref=skyvern.com) or [Skyvern Cloud](https://app.skyvern.com/?ref=skyvern.com) versions and let us know what you think!", "url": "https://wpnews.pro/news/deleting-rag-from-our-web-agent-made-it-2-3x-faster", "canonical_source": "https://www.skyvern.com/blog/deleting-rag-from-our-web-agent-made-it-2-3x-faster/", "published_at": "2026-10-07 13:26:56+00:00", "updated_at": "2026-10-07 13:50:19.889222+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "ai-tools", "developer-tools"], "entities": ["Skyvern", "Skyvern-AI", "Pi", "Claude", "Deepseek", "Prime-intellect", "rustwright", "Odysseys benchmark"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/deleting-rag-from-our-web-agent-made-it-2-3x-faster", "markdown": "https://wpnews.pro/news/deleting-rag-from-our-web-agent-made-it-2-3x-faster.md", "text": "https://wpnews.pro/news/deleting-rag-from-our-web-agent-made-it-2-3x-faster.txt", "jsonld": "https://wpnews.pro/news/deleting-rag-from-our-web-agent-made-it-2-3x-faster.jsonld"}}