← Back to all stories

The Frustration of Web Scraping: Why Accessibility Trees Beat Raw HTML for Browser Agents

Building an autonomous browser agent that can navigate web applications—booking flights, filling forms, and managing dashboards—seems straightforward until you inspect modern web pages. A modern web application is a sprawling jungle of 20,000 DOM nodes, nested <div> wrappers, tracking scripts, inline CSS styling, and SVG graphics spanning 2 to 5 megabytes of raw HTML.

The Context Waste of Raw HTML

If you dump raw DOM HTML into a language model's context window, 95% of the tokens consist of layout wrappers, CSS classes, and hidden JavaScript tags. The model is forced to search for interactive buttons amidst a mountain of visual noise, frequently hallucinating non-existent selectors.

[Raw DOM: 50,000 Tokens of Nested Divs and Script Noise]
[Accessibility Tree (AXTree): 800 Tokens of Clean Semantic Primitives] [id=14] Button "Submit" (clickable) [id=15] Combobox "Departure City" value="San Francisco" (expanded) [id=16] Textbox "Departure Date" (required)

The Semantic Purity of the Accessibility Tree

Web browsers already maintain a specialized, highly curated data structure specifically designed for screen readers: the Accessibility Tree (AXTree).

The AXTree strips away all visual layout styling, CSS classes, and tracking scripts, exposing only the pure semantic interactive elements: role (button, textbox, combobox), accessible name, current state (focused, checked, disabled), and unique element IDs.

The Systems Result

By feeding the AXTree rather than the raw DOM, browser agents consume 90% fewer tokens, execute actions with near 100% selector reliability, and navigate complex single-page web applications with human-like precision.

Reference Paper / Context: Mind2Web: Towards a Generalist Agent for the Web (Deng et al., Ohio State) — Read source ↗
About the Author

Vikram Samal is an AI systems architect focusing on test-time reasoning, high-throughput inference runtimes, and distributed agent infrastructure. Writing weekly architectural stories on Sundays.

Previous
← ColBERT and the Magic of Late Interaction: The Sweet Spot of Information Retrieval
Next
Scaling Beyond RAM: The Architecture of DiskANN and SSD-Resident Vector Search →