$ cat posts/building-brw.md

Building a browser interface for agents

What a failed click taught us about browser automation, and how brw reduces repeated work while checking that actions actually succeed.

19 min read#brw#browser-automation#agents#mcp

I recently watched an agent try to add someone to a chat space, find the Add button and click it, only to get a response saying that the page had changed while the dialogue remained open. After watching it try again, I eventually clicked the button myself to finish the job, because the automation was reporting actions without establishing whether the person had actually been added.

In our 1 October logs, we found a messaging workflow that had taken eight minutes and 37 seconds and made 78 browser calls, yet their combined browser handling time was only about 22 seconds. That doesn’t account for where every second went, but it does suggest that we need to look beyond the individual browser commands to understand why getting a fairly ordinary job done takes so long.

That’s the problem we’ve been working on in brw, the open-source browser tooling we build at Don Works. It’s a Go service that controls Chromium directly through the Chrome DevTools Protocol, or through an extension in a browser you’re already signed into. Agents reach it through MCP, the protocol many agent hosts use to expose tools, or through its HTTP API and command-line interface.

The work in brw 0.20.0 follows that task from giving the model a useful view of the page, through carrying out steps that are already decided, to checking what the application actually did. Along the way we’ve found improvements in fairly ordinary data structures, as well as problems that are harder to resolve, such as a browser losing its connection just after submitting something important.

Giving the model the right controls

To fill in a form, a model needs to know which fields belong to the task, what state they’re in and how to refer to them on the next call, without having to work through a description of every other element on the page.

brw gives controls short references such as e17, alongside their role, name and state. The agent can ask to fill that ref or click it, rather than inventing a CSS selector or finding the same button in another screenshot. Screenshots are still useful when meaning lives in the pixels, but a labelled textbox already tells us quite a lot about how to use it.

Those references need to remain useful when a framework replaces a button, changes its text or inserts another element ahead of it. We stamp a ref on the element and keep a recovery identity built from attributes such as its ID, name, link destination and accessibility attributes, which lets a replacement recover the same ref where those attributes still identify it. Anonymous controls need more positional information, while navigation to a new document still requires a fresh observation.

Having identified the controls, we also need to select the ones the agent is likely to need, so our default snapshot returns a ranked view of up to 40 controls, which we call the frontier. Previously, keyboard focus remaining on the page behind an open dialogue could favour background controls enough to bury the dialogue’s own buttons. We now prioritise the active dialogue, including useful controls below its scroll fold, so the agent can discover them before scrolling to them.

Some sites make a checkbox’s native input transparent and draw the visible box in its label, even though the input still holds the checked state and receives the interaction. Filtering that input out because of its opacity left the agent looking at the decoration without the state it needed, so we now retain it when its visible label makes it actionable.

We found similar opportunities to reduce irrelevant information outside the page, with five tab listings accounting for roughly 22% of the returned data in one earlier live sample. A compact listing with search, ownership filters and a result limit saves the agent from repeatedly working through every open tab. It also reports how many tabs matched and whether the result was truncated, so the agent can tell when it has only part of the answer.

Sending changes without forgetting the page

Once the agent has seen a form, it can ask for changes since a snapshot version rather than receiving the whole form again after every action. The response contains added or changed controls, along with refs that have disappeared, but making that comparison work means keeping track of which view the agent is asking about.

An agent might inspect one field, click a button, then ask what changed from the broader snapshot it started with, in which case keeping only the last snapshot would replace its intended baseline with a different view. We retain up to eight baselines per document, bounded by 2 MiB of serialised state, and record the query, role, limit and other options that produced each one. If the requested baseline is missing or the options don’t match, brw returns a full snapshot and explains why.

When these change responses are chained, the saved baseline still needs to describe the complete selected view. If only the Plan field changes from free to pro, the reply can contain that one field, but we must also retain Email and Save in the baseline so the next comparison doesn’t report those unchanged controls as new.

Send the change, keep the full view

  1. 1 · First observationThe agent sees the form
    EmailPlan = freeSave
  2. 2 · Next observationOnly Plan has changed
    EmailPlan = proSave
  3. 3 · Reply to the agentSend just the change
    Plan = pro
    Email and Save are unchanged, so leave them out of this reply.

Saved in the browser: both full viewsEach saved view, called a baseline, includes Email, Plan and Save. The next comparison can use the complete new view, even though this reply contains only Plan.

The browser retains the complete selected view for the next comparison, even when the reply contains only one changed field. If that saved view is missing or the selection options differ, brw sends a full snapshot instead.

On one recorded form, a compact initial snapshot followed by four compact change responses totalled 1,884 characters, compared with 13,865 for five full JSON observations. That 86% reduction combines compact formatting with sending changes, so it measures the size of the results rather than model tokens or inference time. The benchmark notes retain the fixture and exact comparison.

An empty change response only tells us that nothing changed within the selected view, which is why the result includes candidate counts and truncation information. The agent may still need to widen its query to find out what happened elsewhere on the page.

Counting the same siblings once

Looking at the cost of producing that smaller view led to my favourite optimisation in this work, which has very little to do with AI. We found an unnecessarily expensive loop in the code that walks the browser’s document tree, or DOM, repeatedly counting elements it had already visited.

To recover a control, that code sometimes needs its position among siblings of the same tag, such as button:nth-of-type(27). The old implementation found the position by walking backwards through the preceding siblings for each element, so on a wide row of buttons the first had almost nothing to count while the last scanned almost the whole row. Across all the buttons, those repeated walks made the work grow quadratically.

The replacement builds an index for each parent as it visits the children, recording each node’s position and keeping a separate count per tag. Later lookups reuse the recorded position, with the scan moving forwards only when necessary. Because the index lives for one snapshot, we don’t have to keep a persistent DOM cache correct between calls.

Find each element’s position without counting again

Before · Count backwards each timeRevisit earlier siblings

Each dashed square is an earlier sibling checked again. The numbered square is the element being located.

Find 1
Find 2
Find 3
Find 4
Find 5
Find 6
For these six siblings: 0 + 1 + 2 + 3 + 4 + 5 = 15 repeated checks of earlier elements.
After · Scan forwards onceRemember each position
123456
Saved positionsFirst element → 1
Second element → 2
… and so on through 6
Later lookups reuse these positions instead of walking backwards again.
These six elements share a parent and a tag, so each position can be recorded as the scan advances through them. The index keeps those positions for the rest of this snapshot, advancing further only when needed.

The DOM can still change during a synchronous walk, because stamping a ref writes an attribute and a custom element can respond immediately through attributeChangedCallback. If that callback inserts a sibling before we finish observing the page, the positions we’ve already recorded can become stale within the same call.

We therefore invalidate the index after stamping a custom element, with a regression test that inserts a sibling from the callback to check this behaviour. Other tests cover mixed tags, shadow roots, frames and reordered elements, while the performance regression counts sibling accesses so it can catch repeated work without depending on how busy the test machine is.

Paired measurements on my M4 Max with Chrome 154 showed the effect:

Sibling controls Previous walker Indexed walker
100 1.0 ms 0.8 ms
1,000 17.5 ms 6.1 ms
5,000 301.8 ms 32.4 ms

These are median times for the complete in-page walker, from nine alternating pairs after two warmup pairs, excluding transport, page loading and the model. Although the largest case improved by about 9.3 times, our small 35-command browser fixture showed no demonstrated overall latency improvement, so the benefit we’ve measured is specific to the CPU work on dense pages.

We can avoid more of that work by rejecting irrelevant elements earlier: if the caller only wants textboxes, checking the role first saves us calculating names, recovery paths and geometry for elements we’ll discard. Both changes and their raw measurements are in the benchmark notes.

Letting the browser finish known steps

Once the model knows which fields to fill, a batch can carry out those actions with checks between them and return one final observation, saving another model turn until there’s something new to decide. If a step fails, the batch stops there and returns the information the model needs to work out how to proceed.

In a twenty-control form fixture, running five fills and five clicks as a batch returned 743 bytes of result JSON, compared with 4,951 bytes for the same ten steps made separately. The checks still ran in the browser, but nine intermediate observations no longer had to cross the boundary.

Keeping the checks close to the action also lets us wait for a particular result, such as a field reaching the expected value, instead of guessing when the page has settled. JavaScript can change an input’s value without producing the DOM mutation that a generic waiting heuristic expects. The direct-browser batch path can use the batch’s next check as its wait condition, provided that check was false before the action and we’re still in the same document. Those conditions keep it from accepting an old result or applying the check to a replacement page.

On a controlled field-update fixture, this brought the median from 122.2 ms to 27.4 ms. It relies on having a suitable check for that flow; the measurement notes cover delayed updates, overlays and replacement documents.

Batches were also paying the human-pacing delay for passive waits, assertions and observations, even though that delay is meant to space out inputs. By keeping a clock for the last actual input, we can run those passive steps immediately while the next fill or click still respects the configured spacing.

With human pacing enabled, a local batch containing two fills, two waits and two checks changed as follows:

Browser connection Previous pacing Shared action clock
Direct browser connection 4,141 ms 1,110 ms
Extension 4,380 ms 726 ms

Those medians come from eight runs per case, including the final observation but excluding setup, navigation and the model, with the expected values and input events still matching. The improvement in these small samples comes from removing delays on passive work while retaining the configured spacing between real inputs. The timing notes include the fixtures and exclusions.

For any of these waits to be useful, its deadline has to survive a stalled page. A timer started after awaiting a JavaScript condition can’t help if that condition never resolves, while a timer inside the page may itself stall when the renderer stops responding. We moved enforcement to the host where needed, reject late results, and check that a failed wait prevents the next batch action from running.

Using an operation the page already provides

Where a site provides an operation through WebMCP, an agent can sometimes avoid several interactions with the page altogether. A booking interface exposing find_available_slots, for example, saves the agent from reconstructing availability by clicking through a calendar.

On the Revitt booking page, we compared that interface with the ordinary page controls across three runs per path, stopping before submitting a booking in every case. Page tools took three calls and returned about 7 KB of text, while the DOM path took eleven calls and returned about 55 KB. The tool also returned several days of availability at once, whereas the form showed one day at a time.

Measured tool time was similar, at about a second on either path, and there was no model in that timing loop, so we can’t treat this as a measured agent speedup. It does give us an example of how a site’s own operation can reduce the interaction needed to answer a question, while the browser interface remains necessary for sites that don’t provide one.

Checking that the action worked

Returning to the Add-button failure, reducing the time and data involved would have been little comfort while I was still waiting for the person to be added. We needed to establish both whether the input had reached the control and whether the application had produced the intended result.

In our local reproductions, script-generated events weren’t trusted input, and Chrome could acknowledge a command aimed at an inactive extension tab without delivering the expected interaction. brw now checks that the target is where it is painted and is not disabled, inert or covered, then dispatches input through the browser. It checks again after moving the pointer, because that movement can bring up an overlay, while an inactive extension tab requires an explicit focus action so we don’t quietly move the user’s working tab.

In touch emulation, a mouse command could stall until the 30-second deadline in our checkbox fixture, whereas native touch start and end events completed the flow. We haven’t instrumented the underlying cause inside Chromium, but across direct and extension connections, headed and headless browsers, and desktop and touch modes, all 24 seeded local runs produced the same semantic observation and expected final state. That gives us evidence of consistency on this fixture, without establishing whether other sites behave the same way or how the connections compare for speed.

A delivered click is only part of the job

  1. 1 · brw checks the controlCan Add be clicked?Check that the target is visible and can receive input.
  2. 2 · brw delivers inputClick AddSend mouse or touch input through the browser.
  3. 3 · Agent or recipe checks the taskDid the member appear?Read the application’s result, not just the click acknowledgement.

What did the task check establish?

Yes · Member confirmedThis step is completeContinue from the observed result.
Not confirmed · Outcome unclearStop and inspectFind out what happened before deciding whether another click is safe.
The agent or recipe supplies a check for the expected result, since brw’s click response alone cannot establish whether a member was added. This shows how that check fits into the workflow, although the original Add operation has yet to be retested live.

After delivering the input, we still need a check tied to the task, such as the expected member appearing or the dialogue closing with a confirmation. A generic report that something changed on the page leaves that question unanswered, and if input delivery fails ambiguously, brw stops rather than automatically replaying it.

We also checked the installed release in Google Chat, where the default snapshot exposed the space-setup dialogue and its styled external-members checkbox. Clicking changed its checked state and Cancel closed the dialogue, after which we closed the test tab without creating a space or sending invitations. Because we haven’t repeated the original membership-changing Add operation live, our evidence is still limited to the local reproductions and this narrower check.

Recovering after an interrupted write

If the browser submits a form and the daemon dies before seeing the confirmation, we have the additional problem of working out whether the site accepted the request or ever received it. Repeating the click before resolving that uncertainty risks doing the operation twice, which is why brw recipes can record an attempt before carrying it out.

Recipes are versioned browser workflows with declared targets and checks for success, pinned by ID, version and content digest so execution uses the version that was selected. With a receipt-capable HTTP provider, a recipe can also record an in-flight receipt before dispatching an action that changes remote state.

We derive the receipt key from the recipe digest, step ID, exact origin and declared inputs, so the same attempted operation produces the same key after a restart. Input names are sorted, line endings are normalised, missing and empty values remain distinct, and each field includes its length so ambiguous boundaries can’t make different inputs share an encoding. That gives us a reproducible key without relying on the old clock, host or session.

If two runners try to start the same operation, a lookup followed by a separate write would leave both able to see no receipt and proceed. The provider’s begin operation must therefore report whether this caller created the receipt, so only the caller that did so dispatches the browser action.

We record completion only after the recipe’s success check passes, so another run finding an in-flight receipt must first read the remote state using the recipe’s verification step. If that confirms the result, it can finish the receipt without another click; otherwise the workflow stops for a human to reconcile the outcome.

After a lost connection, check before submitting again

Normal run · With a provider that stores durable receipts
  1. 1 · Save the attemptRecord “pending”The provider saves an attempt record, called a receipt, before the browser submits.
  2. 2 · Act on the websiteSubmit onceOnly the run that created this record may submit the action.
  3. 3 · Confirm successCheck, then mark completedUpdate the saved record only after the website’s result passes the recipe’s success check.

If the connection is lost before completion is recorded

On restart, read the pending record and check the websiteThe saved record tells us an attempt started, but leaves us to check whether the website accepted it.

Did the website already complete this action?

Yes · Result verifiedMark the record completedDo not submit the action again.
Not confirmed · Still uncertainStop for human reviewKeep the record pending while a human checks the outcome, without submitting again.
Both recovery branches avoid another submission while preserving what we know about the attempt across a restart. The receipt helps us recover from an interruption, but cannot guarantee that an arbitrary website executes an action exactly once.

Because the receipt is on our side of the network, it can’t guarantee exactly-once execution on an arbitrary website. It gives us a record of uncertainty that survives a restart and prevents an unresolved attempt from turning into an automatic retry, provided the recipe uses a provider that supports durable receipts. Local directory recipes don’t have that support; the recipe documentation describes the provider contract.

When several agents share a browser, we also need to keep ownership of a tab while an agent reads a result, thinks and decides what to do next. Locking only during the click would leave the tab available to another agent between those steps, so brw leases it to the session across the whole sequence, including reads and verification. Another session trying to use it gets a contention error, and a change to the browser’s active tab doesn’t silently change the agent’s working tab.

Measuring the whole journey

Across all this work, the eight-minute workflow is a useful reminder that faster snapshots and smaller responses only help if they contribute to getting a correct result sooner. Our local usage reporting records input and output sizes and timings at separate CLI, HTTP and MCP boundaries, with character-based token estimates labelled as such because they don’t represent the host model’s complete context or the provider’s billing count.

Those boundaries need to stay separate because HTTP and MCP records can describe the same operation, and several calls can overlap. In the 1 October workflow, the browser’s roughly 22 seconds of handling sat inside a larger execution trace containing 43 groups of tool execution, with a median gap of 12 seconds between their start times. Subtracting one total from another wouldn’t tell us how much time the model spent on inference, because host orchestration, waiting and other work need their own trace. The live investigation keeps those measurements distinct.

Comparing the measurements also led us to leave out an optimisation that reused more action scripts: although it sent 2.46% fewer bytes on our 35-command fixture, it increased browser protocol calls from 237 to 242 without demonstrating lower latency. The smaller payload hadn’t translated into a benefit for the task, which made the extra calls difficult to justify.

The next experiment is to join these browser measurements to the host model’s timing and token usage, then follow the complete sequence from an observation through a decision to a verified result. We haven’t yet measured that loop against a human, and these component benchmarks don’t establish that an agent can navigate faster than one.

For the chat-space task, I’d want that comparison to measure whether the person was added and how much attention I had to give it along the way. Helping the model find the right control and avoiding needless back-and-forth are useful insofar as they let it finish the job and verify the outcome without me having to watch every step. brw and the measurement fixtures are on GitHub, and the next comparison will need to account for that whole process.