cd /news/ai-agents/browser-ai-agents-dont-need-screensh… · home › topics › ai-agents › article
[ARTICLE · art-146008] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Browser AI Agents Don’t Need Screenshots for Everything

A developer built Jev Ultrafast Agent, a Chrome extension that lets browser AI agents act on an indexed, text-based representation of a page's interactive elements instead of processing screenshots at every step. The architecture feeds the model a compact action space (CLICK, TYPE_TEXT, SELECT, SCROLL, WAIT, DONE, BLOCKED) and executes one typed decision per observe-decide-execute loop, which the developer says cuts token usage and removes coordinate ambiguity from visually similar elements. The extension supports Claude and OpenAI-compatible APIs and requires no Python environment or local automation server.

by read4 min views4 publishedOct 6, 2026

Browser agents have become surprisingly capable.

Give an AI agent a task like “find this product, fill in the form and submit it”, and it can often navigate a website almost like a human.

But there is a problem with how many browser agents understand what is happening on the screen.

They look at it.

Literally.

A common browser-agent loop looks something like this:

Screenshot → model → visual reasoning → action → new screenshot → repeat

This approach is powerful because it works almost anywhere. If a human can see an interface, a multimodal model can potentially understand it too.

But does an AI agent really need to see a button to know that it can click it?

That question led us to experiment with a different architecture.

A web page isn't just pixels.

The browser already has access to structured information about the elements on the page: buttons, inputs, links, selects and other interactive components.

Instead of converting the interface into an image and asking the model to visually interpret it, we can give the model a simplified representation of the available actions.

For example, instead of a screenshot of a login page, the model could receive something conceptually similar to:

[1] Email — input
[2] Password — input
[3] Remember me — checkbox
[4] Sign in — button
[5] Forgot password? — link

Now the model doesn't need to determine coordinates or visually locate the button.

It can simply decide:

TYPE_TEXT 1 "user@example.com"

Then:

TYPE_TEXT 2 "password"

And finally:

CLICK 4

The browser executes the action, generates an updated representation of the page and sends it back to the model.

The loop becomes:

Page → indexed action space → model → typed action → updated page

No screenshot is required for many interactions.

The first obvious reason is token usage.

Images can introduce considerably more information into an agent loop than a compact representation of the interactive state of a page.

And browser agents don't make one decision.

A relatively simple task can require dozens of steps.

If every step requires processing another screenshot, the cost accumulates.

A compact action representation gives the model only the information it needs to make the next decision.

There is another benefit: less ambiguity.

Imagine that a page contains three visually similar buttons.

A vision-based agent needs to determine which button it wants and where that button is located.

With an indexed representation, the decision can simply become:

CLICK 17

The execution layer already knows what element 17 refers to.

We took this idea further while building Jev Ultrafast Agent.

The agent receives the indexed state of the page and returns one typed decision per step.

The action space includes operations such as:

CLICK
TYPE_TEXT
SELECT
SCROLL
WAIT
DONE
BLOCKED

The extension executes the decision and observes the page again.

So instead of asking the model to generate an elaborate plan and then attempting to execute that plan against a changing website, the system operates in a short feedback loop:

observe
↓
decide
↓
execute
↓
observe again

This matters because websites are dynamic.

Clicking something can open a modal, trigger validation, load new content or completely change the DOM.

A long pre-generated sequence of actions can therefore become outdated after the first interaction.

A one-action-per-step architecture lets the model reconsider the environment after every meaningful change.

Another decision we made was to move the implementation directly into the browser.

Jev is distributed as a Chrome extension.

That means the basic architecture doesn't require users to set up a Python environment or run a separate local automation server.

The extension handles the interaction layer while the user chooses the model provider.

The current implementation supports Claude and OpenAI-compatible APIs, among other options.

The idea is essentially:

your browser + your model + your API key.

For experimentation with browser agents, this makes the barrier to entry considerably lower.

No.

And this is probably the most interesting part of the experiment.

There are interfaces where the DOM or accessibility information isn't enough.

Canvas-based applications are an obvious example. So are certain visual editors, maps, unusual custom controls and tasks where understanding the actual appearance of something is necessary.

If the task is:

Click the Submit button.

Structured information is probably enough.

But if the task is:

Choose the product photo where the jacket is dark green.

Now vision actually matters.

This suggests a potentially more efficient architecture:

structured representation by default, vision when necessary.

Instead of making vision the primary way an agent understands every webpage, it becomes another tool the agent can invoke when the structured representation doesn't contain enough information.

We packaged this approach into a Chrome extension called Jev Ultrafast Agent.

It is still an early project, and we're particularly interested in finding cases where the indexed approach performs poorly compared with vision-first agents.

👉 Try Jev Ultrafast Agent on the Chrome Web Store:

https://chromewebstore.google.com/detail/jev-ultrafast-agent/ibimdlcpofepgagkmapaodcbdmnccogo

The larger question is more interesting than the extension itself:

Should browser agents really use vision as their default interface to the web?

Our current hypothesis is that they shouldn't.

For a large percentage of ordinary browser interactions, the browser already contains a cleaner representation of what an agent needs to know.

Vision can handle the rest.

── more in #ai-agents 4 stories · sorted by recency
── more on @jev ultrafast agent 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/browser-ai-agents-do…] indexed:0 read:4min 2026-10-06 · —