Browser control at DOM speed, built for Codex.
Codex Web DOM Bridge is a lightweight browser-control layer that makes Codex dramatically faster and more reliable when working on the web.
Instead of forcing Codex to inspect screenshots, guess coordinates, move the mouse, click pixels, and wait through slow visual feedback loops, this bridge lets Codex interact directly with the page's live DOM. It exposes a compact set of browser tools for observing page structure, clicking buttons, typing into fields, waiting for page changes, extracting content, searching pages, and running multi-step workflows in a single browser hop.
The result is a faster, cleaner, and more accurate way for Codex to control websites.
Watch/download the short product promo: codex-dom-bridge-product-promo.mp4.
Codex can use the bridge through MCP tools such as web_observe, web_act, web_wait, web_extract, web_search, and web_run. The bridge maps live page elements into stable temporary handles, then resolves later actions against the actual DOM node. That means Codex can work more like a browser-native automation agent instead of a remote human trying to click around visually.
This is especially useful for:
- Agentic browser workflows
- Web research automation
- Form filling
- Search and extraction
- Browser-based testing
- Codex-controlled web apps
- Faster local AI tool execution
- Reducing failures from bad clicks, missed elements, and visual ambiguity
At its core, this project keeps Codex out of the pixel business whenever the web page already provides better structure. The fast path is run: compose the obvious DOM steps and let the browser content script perform them locally in one efficient workflow.
Codex CLI/tool calls
-> local bridge server
-> browser content script
-> live DOM
The public contract is intentionally tiny:
observereturns a compact, accessibility-flavored map of links, buttons, inputs, headings, and forms.actclicks, types, checks, selects, focuses, or presses against a DOM handle.waitwaits for text, selectors, readiness, or quiet DOM.extractpulls text, HTML, links, fields, or repeated records out of the page.runexecutes a tiny multi-step DOM workflow in one browser hop.searchfinds the page's search input, submits a query, waits for quiet DOM, and extracts results.
Start the bridge:
npm startOpen the demo page:
open http://127.0.0.1:8797/demo/The demo injects the bridge runtime automatically. For arbitrary sites, load the unpacked extension in extension/.
List connected pages:
node bin/codex-web.mjs clientsObserve the active page:
node bin/codex-web.mjs observeType into an element and click a button:
node bin/codex-web.mjs act active e1 type "lightning fast web control"
node bin/codex-web.mjs act active e2 clickExtract visible page text:
node bin/codex-web.mjs extract active --textRun a full DOM workflow in one command:
node bin/codex-web.mjs run '[
{"act":{"selector":"#query","action":"type","value":"speed"}},
{"act":{"selector":"#submit-search","action":"click"}},
{"wait":{"for":"quiet","quietMs":250}},
{"extract":{"selector":"#results","text":true}}
]'Search the current page using its own search box:
node bin/codex-web.mjs search "structured" --results '#results'- Start the server with
npm start. - Open Chrome or a Chromium browser.
- Go to
chrome://extensions. - Enable Developer Mode.
- Choose "Load unpacked" and select the
extension/directory.
The extension content script connects to http://127.0.0.1:8797 by default.
Codex should call this through MCP, not by prompt convention alone. The MCP wrapper exposes the bridge as tools named web_status, web_observe, web_act, web_wait, web_extract, web_search, and web_run.
First install dependencies:
npm installThen add the MCP server to ~/.codex/config.toml:
[mcp_servers.codex-web]
command = "node"
args = ["/Volumes/EXT/Applications/Webtool/bin/codex-web-mcp.mjs"]
[mcp_servers.codex-web.env]
CODEX_DOM_BRIDGE_URL = "http://127.0.0.1:8797"For the current demo server on port 8798, use:
[mcp_servers.codex-web]
command = "node"
args = ["/Volumes/EXT/Applications/Webtool/bin/codex-web-mcp.mjs"]
[mcp_servers.codex-web.env]
CODEX_DOM_BRIDGE_URL = "http://127.0.0.1:8798"Restart Codex after editing the config. In a new Codex turn, ask it to use web_status; it should report the bridge URL and connected browser clients. From there:
Use web_search with query "structured" and resultsSelector "#results".
The browser page still needs the content script: use the demo page, or load the unpacked extension in extension/ for arbitrary sites.
With the demo open, run:
CODEX_DOM_BRIDGE_URL=http://127.0.0.1:8798 npm run benchmark -- --url http://127.0.0.1:8798 --iterations 9The benchmark reports:
bridge.search: one tool call for search, wait, and extraction.bridge.run: one tool call for a scripted DOM workflow.bridge.split: the older multi-call pattern.
Current demo baseline numbers live in benchmarks/demo-baseline.md.
All command endpoints accept active in place of a client id.
GET /api/clients
POST /api/:clientId/observe
POST /api/:clientId/act
POST /api/:clientId/wait
POST /api/:clientId/extract
POST /api/:clientId/run
POST /api/:clientId/searchExample:
curl -s http://127.0.0.1:8797/api/active/observe \
-H 'content-type: application/json' \
-d '{"include":["inputs","buttons"]}'One-hop workflow:
curl -s http://127.0.0.1:8797/api/active/run \
-H 'content-type: application/json' \
-d '{"steps":[{"search":{"query":"codex dom","resultsSelector":"#results"}}]}'The bridge keeps Codex out of the pixel business whenever the page gives us better structure. A browser client assigns stable temporary handles like e12, returns concise element descriptors, and resolves later actions against the real DOM node.
Fallback order:
- Element id from
observe. - CSS selector.
- Accessibility role and name.
- Browser surface clicking, only when the DOM path is blocked.
The fast path is run: compose the obvious DOM steps and let the content script do them locally. That removes the slow screenshot/coordinate loop and avoids paying a server round trip for every keystroke, click, wait, and extraction.