Building an autonomous browser agent that can navigate web applications—booking flights, filling forms, and managing dashboards—seems straightforward until you inspect modern web pages. A modern web application is a sprawling jungle of 20,000 DOM nodes, nested <div> wrappers, tracking scripts, inline CSS styling, and SVG graphics spanning 2 to 5 megabytes of raw HTML.
The Context Waste of Raw HTML
If you dump raw DOM HTML into a language model's context window, 95% of the tokens consist of layout wrappers, CSS classes, and hidden JavaScript tags. The model is forced to search for interactive buttons amidst a mountain of visual noise, frequently hallucinating non-existent selectors.
[Raw DOM: 50,000 Tokens of Nested Divs and Script Noise][Accessibility Tree (AXTree): 800 Tokens of Clean Semantic Primitives] [id=14] Button "Submit" (clickable) [id=15] Combobox "Departure City" value="San Francisco" (expanded) [id=16] Textbox "Departure Date" (required)The Semantic Purity of the Accessibility Tree
Web browsers already maintain a specialized, highly curated data structure specifically designed for screen readers: the Accessibility Tree (AXTree).
The AXTree strips away all visual layout styling, CSS classes, and tracking scripts, exposing only the pure semantic interactive elements: role (button, textbox, combobox), accessible name, current state (focused, checked, disabled), and unique element IDs.
The Systems Result
By feeding the AXTree rather than the raw DOM, browser agents consume 90% fewer tokens, execute actions with near 100% selector reliability, and navigate complex single-page web applications with human-like precision.
Reference Paper / Context: Mind2Web: Towards a Generalist Agent for the Web (Deng et al., Ohio State) — Read source ↗