Browser agents have become surprisingly capable.
Give an AI agent a task like “find this product, fill in the form and submit it”, and it can often navigate a website almost like a human.
But there is a problem with how many browser agents understand what is happening on the screen.
They look at it.
Literally.
A common browser-agent loop looks something like this:
Screenshot → model → visual reasoning → action → new screenshot → repeat
This approach is powerful because it works almost anywhere. If a human can see an interface, a multimodal model can potentially understand it too.
But does an AI agent really need to see a button to know that it can click it?
That question led us to experiment with a different architecture.
A web page isn't just pixels.
The browser already has access to structured information about the elements on the page: buttons, inputs, links, selects and other interactive components.
Instead of converting the interface into an image and asking the model to visually interpret it, we can give the model a simplified representation of the available actions.
For example, instead of a screenshot of a login page, the model could receive something conceptually similar to:
[1] Email — input
[2] Password — input
[3] Remember me — checkbox
[4] Sign in — button
[5] Forgot password? — link
Now the model doesn't need to determine coordinates or visually locate the button.
It can simply decide:
TYPE_TEXT 1 "user@example.com"
Then:
TYPE_TEXT 2 "password"
And finally:
CLICK 4
The browser executes the action, generates an updated representation of the page and sends it back to the model.
The loop becomes:
Page → indexed action space → model → typed action → updated page
No screenshot is required for many interactions.
The first obvious reason is token usage.
Images can introduce considerably more information into an agent loop than a compact representation of the interactive state of a page.
And browser agents don't make one decision.
A relatively simple task can require dozens of steps.
If every step requires processing another screenshot, the cost accumulates.
A compact action representation gives the model only the information it needs to make the next decision.
There is another benefit: less ambiguity.
Imagine that a page contains three visually similar buttons.
A vision-based agent needs to determine which button it wants and where that button is located.
With an indexed representation, the decision can simply become:
CLICK 17
The execution layer already knows what element 17 refers to.
We took this idea further while building Jev Ultrafast Agent.
The agent receives the indexed state of the page and returns one typed decision per step.
The action space includes operations such as:
CLICK
TYPE_TEXT
SELECT
SCROLL
WAIT
DONE
BLOCKED
The extension executes the decision and observes the page again.
So instead of asking the model to generate an elaborate plan and then attempting to execute that plan against a changing website, the system operates in a short feedback loop:
observe
↓
decide
↓
execute
↓
observe again
This matters because websites are dynamic.
Clicking something can open a modal, trigger validation, load new content or completely change the DOM.
A long pre-generated sequence of actions can therefore become outdated after the first interaction.
A one-action-per-step architecture lets the model reconsider the environment after every meaningful change.
Another decision we made was to move the implementation directly into the browser.
Jev is distributed as a Chrome extension.
That means the basic architecture doesn't require users to set up a Python environment or run a separate local automation server.
The extension handles the interaction layer while the user chooses the model provider.
The current implementation supports Claude and OpenAI-compatible APIs, among other options.
The idea is essentially:
your browser + your model + your API key.
For experimentation with browser agents, this makes the barrier to entry considerably lower.
No.
And this is probably the most interesting part of the experiment.
There are interfaces where the DOM or accessibility information isn't enough.
Canvas-based applications are an obvious example. So are certain visual editors, maps, unusual custom controls and tasks where understanding the actual appearance of something is necessary.
If the task is:
Click the Submit button.
Structured information is probably enough.
But if the task is:
Choose the product photo where the jacket is dark green.
Now vision actually matters.
This suggests a potentially more efficient architecture:
structured representation by default, vision when necessary.
Instead of making vision the primary way an agent understands every webpage, it becomes another tool the agent can invoke when the structured representation doesn't contain enough information.
We packaged this approach into a Chrome extension called Jev Ultrafast Agent.
It is still an early project, and we're particularly interested in finding cases where the indexed approach performs poorly compared with vision-first agents.
👉 Try Jev Ultrafast Agent on the Chrome Web Store:
https://chromewebstore.google.com/detail/jev-ultrafast-agent/ibimdlcpofepgagkmapaodcbdmnccogo
The larger question is more interesting than the extension itself:
Should browser agents really use vision as their default interface to the web?
Our current hypothesis is that they shouldn't.
For a large percentage of ordinary browser interactions, the browser already contains a cleaner representation of what an agent needs to know.
Vision can handle the rest.