← 开源
platonai

Browser4

Browser4 — an AI-native browser engine for autonomous agents, intelligent extraction, and large-scale web automation.

InfrastructureMake world agent-readyKotlin
在 GitHub 打开
增长势头
+124 小时新增 Star+0.1%
1.15k
Star
152
Fork
—
本周
7
贡献者
创建于 2018-03-12 · 更新于 2026-10-05 · 今日第 4499 名
主要开发者
README

🤖 Browser4

CI build Release Stars Maven Central npm Website Top language License


English | 简体中文 | 中国镜像

Table of Contents

🌟 Introduction

💖 Browser4 — an AI-native browser engine for autonomous agents, intelligent extraction, and large-scale web automation. 💖

✨ Key Capabilities

  • 🤖 Agent Browser — AI agents and humans drive real browsers via a Rust CLI, MCP, and an agentic backend: navigate, click, fill, snapshot, batch, and loop.
  • 🧬 Zero-Token Extraction — X-SQL + CSS selectors for deterministic extraction from live pages or stored HTML snapshots; WebMiner ML clustering turns HTML corpora into spreadsheet and report views with no LLM tokens.
  • 🧠 Hybrid Intelligence — Combine LLM extraction, ML clustering, X-SQL, and a progressive experience store that reuses learned selectors and blockers.
  • ⚡ High-Performance Runtime — Coroutine-safe, CDP-native engine designed for 100k–200k complex page visits per machine per day via swarm/crawl scale-out.
  • 📦 Enterprise-Scale Automation — Swarm crawling, batch/loop jobs, stateful sessions, plugins, runtime skills, browser extension, and MCP-over-HTTP.
  • 🛠️ Programming-Agent Kernel — 50+ coding.* tools (sandboxed shell/fs, scaffolding, validation, self-development) for agents building Browser4 artifacts — or Browser4 itself.

Quick Start

Paste the following instruction to your favorite AI agent like dsh, claude, codex, workbuddy or openclaw and run it:

Read https://browser4.io/SKILL.md, install or upgrade browser4-cli for browser automation, perform the following task:

1. Open the browser in headed mode (`open --headed`) so the window is visible — this is a human-facing demo
2. go to amazon.com
3. search for pens to draw on whiteboards
4. compare the first 4 ones
5. write the result to a markdown file

DeepSeek Harness integration

https://github.com/platonai/dsh-browser4

dsh plugin --profile web add dsh-browser4                  # npm registry
dsh plugin --profile web add github:platonai/dsh-browser4  # GitHub

🧭 Tool Selection Guide

Choosing the right tool for your task:

How to Interact with a Page

Use snapshot -i --boxes to see clickable/typeable elements with refs like e15, then click , fill "", type/press, select, hover/drag/scroll, and wait to drive the page. All interaction commands accept CSS selectors too. Chain multiple steps efficiently with batch.

Content embedded in ``s (payment forms, editors, widgets) is reached with the built-in frame switching: frames lists the frame tree, frame "" scopes subsequent element commands into that frame (same-origin iframes fully supported), and frame main returns to the main document — no manual contentDocument eval needed.

Typical interactive flow:

# Humans usually want to see the browser — open it headed
browser4-cli open --headed https://example.com/login
browser4-cli snapshot -i --boxes
browser4-cli fill e3 "[email protected]"
browser4-cli fill e4 "secret" --submit
browser4-cli wait --load networkidle
browser4-cli snapshot -i
# iframe-heavy page:
browser4-cli frame "#pay-frame"
browser4-cli fill "#card-number" "4111 1111 1111 1111"
browser4-cli frame main

How to Extract Data

Need to extract data from a page?
├─ Interactive page (click, fill, scroll first)? → snapshot + refs, then extract
├─ Static page, one field? → htmlsnapshot get text ""
├─ Static page, all matches of one field? → htmlsnapshot get all text ""
├─ Static page, multiple correlated fields (title+price+url per item)?
│  → htmlsnapshot query --sql @query.sql
├─ Live JS / complex DOM logic? → eval --json
├─ Natural language ("find the product price")? → extract (needs LLM key)
└─ High volume, many pages? → crawl or swarm with --sql

How to Process at Scale

Need to process multiple pages?
├─ Single list page (search results)? → htmlsnapshot query with DOM_LOAD_AND_SELECT
├─ List of known URLs (in a file)? → crawl --seed-file urls.txt --depth 0 --sql @query.sql
├─ Crawl from a start URL (follow links)? → crawl  --out-link-selector "..." --depth N
├─ Need parallel execution (high throughput)? → swarm create → swarm query --seed-file ...
├─ Repeated monitoring (check every hour)? → loop -i 3600 -- eval "..."
└─ Just a few URLs in a shell script?
   → browser4-cli open --headed "https://first-url"   # humans: open once, visibly
   → for url in ...; do browser4-cli goto "$url"; ... done

How to Turn HTML into Spreadsheets — Zero Tokens

WebMiner runs ML clustering on downloaded HTML files to produce structured spreadsheets and interactive reports — no LLM tokens, everything runs locally. webminer is a first-class Browser4 CLI citizen: browser4-cli webminer install + browser4-cli webminer all runs the whole pipeline without PowerShell.

Have HTML files and want structured data — without tokens?
├─ < 20 pages? → browser4-cli crawl --seed-file urls.txt --depth 0 --sql @query.sql
├─ < 1,000 pages (small to medium)? → WebMiner Free (SMILE ML engine)
│  browser4-cli webminer install
│  browser4-cli webminer all ./pages/
│  → Interactive HTML report + Excel spreadsheets — local, zero cost
├─ > 1,000 pages (production scale)? → WebMiner Commercial (Apache Spark ML)
│  Same encode → cluster → views pipeline, distributed across machines
└─ Need to acquire pages first?
   ├─ Single pages: browser4-cli htmlsnapshot export
   ├─ Bulk download: browser4-cli crawl --seed-file urls.txt --depth 0
   └─ High throughput: browser4-cli swarm create → swarm query --seed-file ...
       Then feed the HTML directory to WebMiner

Pipeline: encode (HTML → feature vectors → CSV) → cluster (KMeans, auto-detected K) → views (HTML report + Excel). Free tier uses the SMILE ML library for single-machine clustering (< 1,000 pages). Requires JDK 17+ (auto-detected). See web-miner and browser4-cli help webminer for usage.


📦 Installation

Manually installation is optional since your AI agent is smart enough to install it after reading the SKILL.

Install browser4-cli globally using npm (requires Node.js):

npm install -g browser4-cli
browser4-cli install

Or bootstrap the native binary directly with a single command. The scripts also install the Browser4 backend (runtime bundle) afterwards — running browser4-cli install on a fresh machine, or browser4-cli upgrade when a backend is already present (pass --skip-backend to skip this step):

Windows (PowerShell):

irm https://browser4.oss-cn-beijing.aliyuncs.com/scripts/install-browser4-cli.ps1 | iex

Linux / macOS (bash):

curl -fsSL https://browser4.oss-cn-beijing.aliyuncs.com/scripts/install-browser4-cli.sh | bash

💡 CLI Guide for Humans

browser4-cli is a human-usable browser automation shell, not just an agent backend. You can drive a real browser, inspect state, extract structured data, run X-SQL, orchestrate crawl/swarm jobs, manage server plugins and skills, and hand long-running work to built-in AI features.

If you want the embedded agent-facing instructions, see skills/browser4-cli/SKILL.md. This section is the human reference.

Quick start

# Open a visible browser session (humans usually want to see the window;
# agents should prefer the default headless mode — see SKILL.md)
browser4-cli open --headed https://browser4.io

# Inspect the page and get element refs
browser4-cli snapshot --boxes

# Interact using a ref from the snapshot
browser4-cli click e15
browser4-cli fill e16 "Browser4" --submit

# Extract data from the live page
browser4-cli get text "h1"

# Capture the page into the store, then extract from that copy
browser4-cli htmlsnapshot
browser4-cli htmlsnapshot get text "#main-content"
browser4-cli htmlsnapshot query --sql @query.sql

# Save output
browser4-cli screenshot --full-page --filename page.jpg
browser4-cli pdf --filename page.pdf

Mental model

  1. Session-oriented: commands work against the current browser session; use -s for isolated named sessions.
  2. Two page views: snapshot is for interactive work with element refs like e15; htmlsnapshot is for DOM/X-SQL extraction with CSS selectors.
  3. Interactive vs static extraction: use click, fill, type, press, wait when the page must be manipulated first; use htmlsnapshot query when you need structured extraction from the DOM.
  4. Synchronous vs async jobs: agent, swarm, crawl, and async chat-style commands return task IDs you poll later.

Global options

These flags can appear before any command.

Flag Meaning
-h, --help [command|category] Show top-level help, category help, or detailed command help
--help-json Emit the machine-readable command reference
-v, --version Print the CLI version
-s, --session Use a named session instead of the default session. A global flag: place it before the command (browser4-cli -s job-42 snapshot) — after the command it is rejected as a positional argument
--server Override the Browser4 server URL
--timeout Override the HTTP timeout for the current command
--proxy Proxy used for runtime downloads/install operations
--json Emit machine-readable JSON only
--pretty Pretty-print JSON output
-q, --quiet Suppress normal human-readable output
-tip, --show-tip Show a relevant tip on stderr after commands

Key concepts before the command list

Element refs vs CSS selectors

  • snapshot returns accessibility-tree refs such as e5, e12, e42
  • most interaction commands accept either a snapshot ref or a CSS selector
  • htmlsnapshot commands use CSS selectors, not accessibility refs

snapshot vs htmlsnapshot

Tool Best for Input model Output model
snapshot clicking, typing, finding interactive elements live accessibility tree refs like e15
htmlsnapshot DOM inspection, CSS extraction, X-SQL a fresh snapshot of the active page (every command captures the tab first, then reads; --expires reads the stored snapshot instead) CSS selectors and query results

LLM configuration

AI-powered commands such as extract, summarize, chat, agent run, and X-SQL llm_* functions require an LLM provider key.

Provider Environment variables
DeepSeek DEEPSEEK_API_KEY
OpenRouter OPENROUTER_API_KEY, OPENROUTER_MODEL_NAME, OPENROUTER_BASE_URL
Volcengine VOLCENGINE_API_KEY, VOLCENGINE_MODEL_NAME, VOLCENGINE_BASE_URL
OpenAI-compatible OPENAI_API_KEY, OPENAI_MODEL_NAME, OPENAI_BASE_URL
Aliyun Qwen OPENAI_API_KEY, OPENAI_MODEL_NAME, OPENAI_BASE_URL
export DEEPSEEK_API_KEY=sk-xxxxxxxxxxxxxxxxxxxxxxxx

Configure one provider. When several provider keys are present, the first one in the built-in detection order wins — openai comes after the dedicated providers such as deepseek, so a leftover DEEPSEEK_API_KEY silently wins over OPENAI_*. browser4-cli doctor prints the key that won (Selected key) and the model the backend will request (Active model). See LLM configuration.

Complete command reference

Session lifecycle and server administration

Command Description
open [url] Open a browser session or reconnect to an existing one. Headless by default. Supports --headed (visible window), --headless, --profile , --profile-mode , --interact-level . Note: SYSTEM_DEFAULT is deprecated and unsupported on Chrome ≥ 143 — use attach + state-save/state-load to reuse system browser state (see browser-state-import.md).
attach Attach to an existing browser via CDP or the Browser4 extension. Supports --cdp and remote endpoint options. After a successful attach the CLI prints the browser that actually connected (Connected browser: … / Attached to … at …) and shows a ⚠ warning when it conflicts with the requested channel (e.g. requested msedge but Chrome connected) — verify it before driving the session.
close Close the active browser session.
list List browser sessions with their status and next-open behavior. The Connection column prefers the backend-reported actual browser and annotates channel conflicts (e.g. requested msedge · actual Google Chrome). Supports --all.
session-default Make a named session become the default unnamed session.
close-all Close all sessions without stopping the backend.
kill-all Force-stop the backend and Browser4-managed browser processes.
stop Gracefully stop the Browser4 server.
status Show server version, port, health, and the web status panel URL (http://:18182/status). When a session is active it also prints a current-session block: Name / Session ID / Status / Connection / Next open.
doctor Run diagnostics: build info, LLM status (including the configuration file it reads, the winning provider key and the active model), stale daemon cleanup, optional repair. Supports --verbose and --fix; --fix also writes a commented LLM config template.
doctor log [name] List, view, tail, or grep backend log files. Supports --tail, grep-style flags, and doctor log grep .
doctor metrics [filter] List, filter, or grep backend metrics. Supports doctor metrics grep .
doctor status [--section ] [--verbose] Print the aggregated status panel report in the terminal: summary layer by default, full detail with --verbose, one report with --section (health, build, runtime, llm, sessions, pulsar-sessions, swarm, url-pool, browsers, drivers, privacy, plugins, skills, metrics, logs), machine-readable JSON with --json.
delete-data Delete session data.
install Install the Browser4 runtime bundle. Supports --tag and --force.
upgrade Upgrade the CLI/runtime bundle. Supports --tag and --force.
uninstall Remove global installs and runtime data. Supports -y, --yes, and --dry-run.
browser4-cli open --headed https://example.com
browser4-cli attach --cdp chrome
browser4-cli doctor --verbose
browser4-cli doctor log server.log --tail
browser4-cli doctor metrics grep request
browser4-cli doctor status --section skills --verbose

Web status panel: open http://127.0.0.1:18182/status in a browser for a live dashboard (health, version, JVM/runtime, LLM config, sessions, Pulsar sessions — SDK identity, context and main-loop state, swarm — swarm session plus task summary, URL pool — queued/real-time/delay counts per priority cache, browsers & open tabs — per-session browser/driver binding and tab counts, with on-demand live tab details via GET /api/system/tabs — driver pools, plugins: load/enable state and SDK compatibility, skills: registered skills with origin (classpath/filesystem/programmatic), metrics, log files; auto-refreshes, set ?refresh= to change the interval). The panel is backed by the aggregated GET /api/system/status endpoint; the individual endpoints (/api/system/health, /api/system/build, /api/doctor/llm-status, /api/doctor/metrics, /api/doctor/log-files, /api/plugins, /api/skills) remain available. browser4-cli plugin-list also reports load/enable state and SDK version for every installed plugin; the same reports can be read from the terminal with browser4-cli doctor status.

Page screenshots: open http://127.0.0.1:18182/pages.html for a grid of every open page across sessions. The active tab of each session is captured automatically (click a screenshot to re-capture it); inactive tabs show a placeholder that captures on click. Swarm sessions only show placeholders. Screenshots load asynchronously — the backend captures in the background (202 Accepted with Retry-After while capturing, cached image/png when ready), so the panel never blocks on a capture. Backed by GET /api/pages and GET /api/pages/{sessionId}/{guid}/screenshot.png (?refresh=1 forces a new capture).

Navigation

Command Description
goto Navigate to a URL; auto-opens/reconnects a session if needed.
go-back Go back in browser history.
go-forward Go forward in browser history.
reload Reload the current page.

Core interaction

All interaction commands accept a snapshot ref such as e15 or a CSS selector unless noted otherwise. Most of them also support --no-snapshot to skip the automatic post-action accessibility snapshot.

Command Description
click [button] Click an element. Supports --modifiers, --follow, --auto-dismiss-dialogs.
dblclick [button] Double-click an element. Supports --modifiers, --follow, --auto-dismiss-dialogs.
hover Hover over an element.
fill Clear and fill text into an editable field. Supports --submit, --verify.
type [ref] Type text into the focused element or a target element. Supports --submit, --verify, --focus, --interactable-timeout, and --method auto|chars|exec (requires a target ref; auto types short text per-character and bulk-inserts long/multi-line text on textarea/contenteditable via a single execCommand('insertText'), chars forces per-character typing, exec forces the bulk insert).
press [ref] Press a key on the focused element or a target element. Supports --verify, --follow.
select Select a dropdown value. Supports --verify.
check Check a checkbox or radio button.
uncheck Uncheck a checkbox or radio button.
drag Drag and drop from one element to another.
wait [target] Wait for a selector/ref, duration, text, URL pattern, page-load state, or JavaScript expression. Supports --timeout, --text, --url, --load, --fn.
upload [file...] Upload one or more local files to a page file input. The target must be an `` (anything else errors); the paths must be readable by the browser process — local mode: the same machine, and empty/missing files are rejected with a hint; remote backend: paths resolve on the backend host. Supports --no-snapshot.

wait --load accepts domcontentloaded, load, and networkidle. networkidle only proves the network went quiet — it does not mean late-rendered results exist; for result pages poll the element instead: wait "".

browser4-cli click e8 --follow
browser4-cli fill e4 "[email protected]" --submit
browser4-cli type "Browser4" e7 --verify
browser4-cli wait --text "Success"
browser4-cli wait --load networkidle

Keyboard and mouse

Command Description
keydown Press and hold a key.
keyup Release a held key.
mousemove Move the mouse to screen/page coordinates.
mousedown [button] Press a mouse button.
mouseup [button] Release a mouse button.
mousewheel Scroll using a wheel delta.
scroll Scroll the page up, down, left, or right.

Page inspection and live extraction

Command Description
snapshot Capture an accessibility-tree snapshot. Supports --boxes/--no-boxes, -i/--interactive, -u/--urls, -c/--compact, --no-compact, -d/--depth, -l/--limit, -s/--selector, --raw, --stdout, -vp/--viewport, --filename. --stdout/--raw output is paginated at 2000 lines/page by default — when truncated, stdout (if piped) gets a # … output truncated: showing N of M lines … hint and the full footer goes to stderr; use --all or --page-size 0 for the complete tree, and bound very large pages with -v N/--depth/--selector/--no-boxes.
snapshot grep Search saved/current snapshot YAML with grep-style flags such as -i, -v, -c, -l, -F, -w, -A, -B, -C, --selector, --page, --page-size, --all.
snapshot list List saved snapshot files with timestamps and sizes.
snapshot clean Remove old snapshot files. Supports --dry-run.
get [name] Extract text, html, box, styles, property, or attr from a live page element.
eval [expression] [ref] Evaluate JavaScript on the page or an element. Supports --file, --stdin, --base64, --await, --wait-selector, --json.
console [min-level] List browser console messages. Supports --clear.
cdp Send an arbitrary Chrome DevTools Protocol command. Supports --json .
generate-locator Generate the best CSS selector for a snapshot ref or existing selector.
resize Resize the browser window.
dialog-accept [prompt] Accept an alert/confirm/prompt dialog, optionally filling the prompt.
dialog-dismiss Dismiss an alert/confirm/prompt dialog.

get modes:

Mode Meaning Example
text visible inner text browser4-cli get text ".price"
html inner HTML browser4-cli get html "#main"
box bounding box browser4-cli get box "#hero"
styles computed styles browser4-cli get styles e9
property DOM property value browser4-cli get property "input" value
attr HTML attribute value browser4-cli get attr "a" href
browser4-cli snapshot -i --boxes
browser4-cli snapshot grep -C 2 "button"
browser4-cli eval "document.title"
browser4-cli eval --file script.js --await
browser4-cli console warn
browser4-cli cdp Runtime.evaluate --json '{"expression":"document.title"}'

HTML snapshot and X-SQL extraction

htmlsnapshot captures a raw DOM snapshot of the page the active tab is showing and is the center of Browser4's structured extraction workflow. Every htmlsnapshot command works on a fresh snapshot of the active page: it captures the tab first (serializing the document the tab already shows, without navigating) and then operates on that snapshot — capture returns the metadata, while the reads consume it, so a read already sees the live document (form results, SPA updates, eval mutations). A read can also be told to serve the stored snapshot instead: --expires (default 0s, the live page) serves the stored copy of the active page while it is younger than the window — htmlsnapshot get text ".price" --expires 1d reads the previous snapshot version and never touches the tab. A command aimed at another URL (readability , query --url ) reads that URL's stored copy, or an independent read-only load when the store has nothing — a URL the tab does not show is never captured, so the tab's document is never filed under it.

Command Description
htmlsnapshot Short form of htmlsnapshot capture.
htmlsnapshot capture Capture and store a static HTML snapshot with metadata about the page and interactive elements.
htmlsnapshot get [selector] [name] Extract the first matching text, textcontent, html, or attr from a fresh snapshot of the active page. Supports --expires .
htmlsnapshot get all [selector] [name] Extract all matching values from a fresh snapshot of the active page. Supports --offset, --limit, and --expires .
htmlsnapshot query [url] Run X-SQL against a fresh snapshot of the active page, or against an explicit URL's stored page. Supports --sql , --sql-stdin, --sql-base64, --expires , result pagination, and extraction-focused output flags.
htmlsnapshot export Export a fresh snapshot's HTML to a file. Supports positional file path or --file plus --clean and --expires .
htmlsnapshot summary [--algorithm ] Generate a compressed Web Page Summary Index (WPSI) from a fresh snapshot of the active page; plugins may contribute additional algorithm ids. Supports --expires .
htmlsnapshot algorithms List installed summary algorithms (built-in wpsi plus plugin-contributed ids). No page required.
htmlsnapshot grep Search a fresh snapshot's HTML with grep-style flags. Supports --expires .
htmlsnapshot inspect [selector] Discover recurring DOM patterns and selector candidates in a fresh snapshot of the active page. Supports --max, --depth, --stdin, --selector-base64, --expires .
htmlsnapshot readability [url] Extract the main article content with a Readability-style heuristic — no LLM, no tokens. Without a URL: the active page; with one: that URL's own stored copy. Supports --text-only, --expires , and pagination.

Important rules:

  • use snapshot when you need refs and interaction
  • use htmlsnapshot when you need repeated DOM extraction — every command captures the tab first, so simply rerunning it sees the page as it is now
  • use --expires on a read when you want the snapshot already in the store (--expires 1d) instead of the live page (--expires 0s, the default) — the stored copy is left untouched and the tab is not serialized
  • htmlsnapshot query --sql @query.sql is the recommended way to avoid shell quoting issues
  • for correlated list extraction, prefer htmlsnapshot query over repeated get all
  • for one-step article extraction (no selectors needed), use htmlsnapshot readability
browser4-cli htmlsnapshot
browser4-cli htmlsnapshot get text "#productTitle"
browser4-cli htmlsnapshot get all text ".result-title" --offset 10 --limit 5
browser4-cli htmlsnapshot inspect ".s-result-item" --depth 6 --max 20
browser4-cli htmlsnapshot export --file page.html --clean
browser4-cli htmlsnapshot query --sql @query.sql
browser4-cli htmlsnapshot readability --text-only --all

For deep X-SQL usage, see skills/browser4-cli/references/htmlsnapshot.md and skills/browser4-cli/references/x-sql-dom-load-select.md.

Screenshots and PDF

Command Description
screenshot [ref] Take a page or element screenshot. Supports --filename, --full-page, --viewport.
pdf Save the current page as PDF. Supports --filename.

Tabs

Command Description
tab-list List open tabs with indexes, titles, and URLs; use --json for full GUIDs.
tab-new [url] Open a new tab, optionally navigating to a URL.
tab-close [index] Close a tab by index; supports --guid .
tab-select Switch to a tab by index; supports --guid .

Browser storage and local page data

Command Description
state-save [filename] Save cookies and localStorage to a JSON file.
state-load Restore cookies and localStorage from a JSON file.
cookie-list List cookies. Supports --domain, --path.
cookie-get Get a cookie by name.
cookie-set Set a cookie. Supports --domain, --path, --expires, --httpOnly, --secure, --sameSite.
cookie-delete Delete a cookie by name. Supports --domain, --path.
cookie-clear Clear all cookies.
localstorage-list List localStorage entries.
localstorage-get Read a localStorage key.
localstorage-set Set a localStorage key.
localstorage-delete Delete a localStorage key.
localstorage-clear Clear localStorage.
sessionstorage-list List sessionStorage entries.
sessionstorage-get Read a sessionStorage key.
sessionstorage-set Set a sessionStorage key.
sessionstorage-delete Delete a sessionStorage key.
sessionstorage-clear Clear sessionStorage.
webdb export Export pages from the Browser4 web database to a local directory.
webdb normalize Normalize a URL into the web database key format.

AI extraction, chat, and autonomous agent tasks

These commands require an LLM key.

Command Description
extract Extract structured data from the current page. Supports --schema , --filename, --raw, --stdout.
summarize [instruction] Summarize the current page. Supports --selector, --filename, --raw, --stdout.
chat Send a plain AI chat request without auto-appended browser context.
chat-result Retrieve the result of an async chat task.
agent run Submit an autonomous browser task and immediately receive a task ID. Supports --wait (block for the result) and --wait-timeout (default 600).
agent status Check a running task.
agent result Fetch a completed result.
agent list List tracked agent tasks and their status.
browser4-cli extract "product name, price, rating"
browser4-cli extract "contacts" --schema @schema.json
browser4-cli summarize --selector "#reviews"
browser4-cli agent run "Go to amazon.com, compare the first 3 keyboards, write a summary"
browser4-cli agent status agent-task-1

Batch and loop automation

Command Description
batch [command...] Execute multiple commands in one invocation. Supports --bail and --json for stdin-driven command arrays.
loop [task] Run a task repeatedly. Supports --name, -i/--interval, -n/--count, -t/--timeout, --shell, --list, --pause, --resume, --pause-all, --resume-all, --stop, --stop-all, --status, --history, --keep-state.

Batch-compatible commands:

goto  go-back  go-forward  reload  press  type  keydown  keyup
click  dblclick  hover  fill  select  check  uncheck  drag
mousemove  mousedown  mouseup  mousewheel  scroll  wait
get  eval  snapshot  screenshot  pdf  dialog-accept  dialog-dismiss
resize  tab-list  tab-new  tab-close  tab-select
browser4-cli batch --bail "goto https://example.com" "snapshot" "screenshot"
browser4-cli loop "load https://example.com and extract the title" -i 300 -n 10
browser4-cli loop --shell "curl -s https://api.example.com/health" -i 60
browser4-cli loop --list

Network inspection, HAR recording & request routing

Inspect what the page actually loaded (XHR/fetch calls, status codes, headers, response bodies), record a HAR 1.2 archive importable by Chrome DevTools, and route (mock/abort) matching requests. See skills/browser4-cli/references/network.md for the full guide.

Command Description
network requests List tracked requests. Supports --filter, --type, --method, --status (200, 2xx, 400-499), --clear.
network request Full detail of one request: headers, timing, and the response body (fetched on demand).
network har start [--content ] Start a HAR recording. Content mode: none, text, or all (binary base64).
network har stop [path] Stop recording and print the HAR JSON, or write it to a .har file when a path is given.
network route --body |--abort Intercept matching requests (mock response or fail them) via CDP Fetch. Supports --content-type, --resource-type.
network unroute [pattern] Remove routes; without a pattern, disable interception entirely.
browser4-cli network requests --filter api --status 2xx
browser4-cli network har start --content text
browser4-cli network har stop ./capture.har
browser4-cli network route "**/api/users" --body '{"users":[]}' --content-type application/json

Swarm and crawl for scale

The co prefix is accepted as an alias for swarm.

Command Description
swarm create Create a parallel scraping session. Supports --profile-mode, --max-open-tabs, --max-browser-contexts, --display-mode.
swarm submit [url] Submit URLs or X-SQL payloads as jobs. Supports --seed-file, --sql, --deadline, --expires, --refresh, --parse.
swarm query Run an X-SQL extraction job against one or more loaded pages. Supports --sql, --seed-file, --deadline, --expires, --refresh.
swarm status Check a swarm task status.
swarm result Fetch a completed swarm result.
swarm list List tracked swarm tasks.
swarm close Close the swarm session and release browser resources.
crawl [url] Crawl from a URL or seed file. Supports --seed-file, --sql, --sql-stdin, --sql-base64, --format, --output, -d/--depth, -ol/--out-link-selector, -olp/--out-link-pattern, -tl/--top-links, -a/--args, --refresh, --parse, --expires, -p/--priority, --page-load-timeout, --ignore-url-query, --no-norm, --readonly, -bg/--background.
crawl status Check crawl task status.
crawl result Fetch crawl results.
crawl cancel Cancel a running crawl (the checkpoint is kept, so it can be resumed).
crawl resume Continue an interrupted crawl from its checkpoint: same task id, no repeat requests for URLs that already succeeded. --retry-failed re-fetches terminal failures; automatic resume at startup is off by default (crawl.autoResume). See Crawl checkpoint & resume.
crawl clear Remove terminal-state crawl tasks; supports force-style cleanup options. Resumable checkpoints are kept until crawl clear --all.
crawl list List tracked crawl tasks (--status interrupted shows the resumable ones).
browser4-cli swarm create --max-open-tabs 12 --display-mode HEADLESS
browser4-cli swarm query --seed-file urls.txt --sql @query.sql --refresh
browser4-cli crawl "https://example.com" --depth 2 --out-link-selector "a[href]"
browser4-cli crawl list
browser4-cli crawl resume     # continue an interrupted crawl

Bundled skill files vs installed runtime skills

Browser4 has two different "skill" surfaces:

  1. skills ... manages bundled, embedded skill documents that ship with the CLI.
  2. skill-* manages installed runtime skills exposed by the backend.
Bundled CLI skills
Command Description
skills List bundled skill names.
skills list Same as skills.
skills get Print a skill's SKILL.md. Supports --full and --all.
skills path [name] Print the bundled skill directory path.
skills unpack [dest] Unpack bundled skill files to a directory.
Installed runtime skills
Command Description
skill-list List installed backend skills.
skill-info Show detailed skill metadata.
skill-install Install a skill from a directory containing SKILL.md. Supports --overwrite.
skill-uninstall Remove a skill by ID.
skill-reload Reload a skill from its source directory.

Progressive experience memory

These commands operate on Browser4's learned experience store.

Command Description
experience save Save a task execution trace. Supports --outcome, --intent, --task-type (canonical types including publish_post), and --facts — merges retrospective knowledge (selectors / interaction_hints / known_blockers / anti_patterns, camelCase or snake_case keys) into the (domain, intent) facts entry; the merge is refused when that entry is VERIFIED (immutable).
experience query Query known selectors, blockers, and hints for a URL/domain. Supports --intent.
experience list List stored experience entries. Supports --filter, --intent-filter, --page, --page-size.
experience deep-learn Run deeper analysis on stored traces. Supports --force.

Plugins

Plugins are server-side JARs that extend Browser4.

Command Description
plugin list List installed plugins.
plugin info Show plugin details.
plugin install Install a plugin from a local JAR file. Supports --replace.
plugin remove Remove a plugin. Supports -y, --yes.

Advanced and currently hidden commands

These commands exist in the CLI but are intentionally kept out of the default public help.

Command Description
act Experimental natural-language action translator that turns plain text into a browser command and runs it.

Timeout environment variables

Variable Default Used for
BROWSER4_CLI_HTTP_TIMEOUT_SECS 30 most commands
BROWSER4_CLI_INPUT_TIMEOUT_SECS 90 type, fill, and other slower input workflows
BROWSER4_CLI_NAVIGATION_TIMEOUT_SECS 120 goto, reload, go-back, go-forward
export BROWSER4_CLI_INPUT_TIMEOUT_SECS=180
export BROWSER4_CLI_NAVIGATION_TIMEOUT_SECS=300

State persistence

CLI state lives under ~/.browser4 unless overridden.

  • default session: ~/.browser4/cli-state.json
  • named sessions: ~/.browser4/sessions/.json
  • loops: ~/.browser4/loops/.json

The runtime bundle is stored separately in a platform-conventional application-data directory, so clearing session state does not force a re-download of Browser4 itself.


🚀 Build from Source

Prerequisites: Git, JDK 25+ (Eclipse Temurin), Chrome/Chromium, and PowerShell 7 (Linux/macOS only). For the full prerequisites table, platform-specific tools, and Chrome auto-detection paths, see Build from Source.

  1. Clone the repository

    git clone https://github.com/platonai/Browser4.git
    cd Browser4
    
  2. Configure your LLM API key

    Edit application.properties and add your API key, or set environment variables. See LLM Configuration for supported providers and variable names.

  3. Build the project

    ./mvnw -DskipTests
    
  4. Build and run the CLI (from source)

    # Build the Rust CLI (requires Rust toolchain)
    cd cli/browser4-cli && cargo build --release
    
    # Or run directly without installing:
    cargo run --manifest-path cli/browser4-cli/Cargo.toml -- --help
    
    # Add --quiet to suppress Cargo build-status output:
    cargo run --quiet --manifest-path cli/browser4-cli/Cargo.toml -- 
    
    # Or install globally:
    cd cli/browser4-cli && cargo install --path .
    

    On Windows, prefix the command with chcp 65001 >nul && for proper UTF-8 output. See Build from Source for full platform-specific instructions.

    Dev-mode wrappers (no install needed): The repo root provides shell wrappers that auto-build from source. Use ./b4w.ps1 (PowerShell), ./b4w.sh (Git Bash / Linux / macOS), or ./b4w.bat (CMD) — all accept the same arguments as the installed browser4-cli binary. Dev mode only ever serves the checked-out code: a local runtime bundle built from a different project version makes the command refuse to start (naming the reason, the bundle's versions and build time, and the rebuild command) rather than silently testing an outdated backend; a checkout whose sources merely look newer than the bundle jars is rebuilt before the server starts — see CLI install & upgrade.


🎬 YouTube: Watch the video

📺 Bilibili: https://www.bilibili.com/video/BV1kM2rYrEFC


Architecture

browser4-cli (Rust) ──MCP over HTTP──▶ browser4-rest (Kotlin/Spring) ──▶ PulsarWebDriver (Kotlin/CDP)
  • CLI (cli/browser4-cli) — native Rust binary, talks to the backend via MCP tool calls
  • Backend (browser4-rest) — Spring Boot server, dispatches MCP tools to browser drivers
  • Browser driver (browser4-core/browser4-browser) — wraps Chrome DevTools Protocol
  • Agent tools (browser4-agentic) — maps MCP tool names to browser automation methods
  • Programming kernel (browser4-coding) — dependency-light agent toolkit (sandboxed shell/filesystem, scaffolding, validation, self-development tools) — see below

📦 Modules Overview

Module Description
cli/browser4-cli Rust CLI — fast, native binary for browser automation
skills/browser4-cli AI agent skill definitions (SKILL.md)
browser4-core Core engine: sessions, scheduling, DOM, browser control
browser4-dependencies BOM and dependency version alignment
browser4-tools Operational tools and launch helpers
browser4-agentic AI agents, MCP integration, skill registration
browser4-coding Programming-agent kernel — sandboxed shell/fs, artifact scaffolding & validation, self-development tools (47 coding.* tools)
browser4-agent-tools High-level agent tools: scraping, crawling, stateful page interaction
browser4-rest Spring Boot REST layer & command endpoints
browser4-apps/browser4-standalone Product packaging — unified launcher (target/Browser4.jar)
examples/browser4-examples Runnable examples and demos
browser4-tests E2E, integration, and scenario test suites
cdp-protocol Chrome DevTools Protocol JSON definitions
coworker/ Builtin AI coworker

🧩 Programming-Agent Kernel (browser4-coding)

browser4-coding is the dependency-light programming kernel that lets an AI agent create Browser4 artifacts and develop Browser4 itself. It is independent of browser4-agentic and pulsar-common (only SLF4J + Jackson + coroutines), so it can be reused by non-agent hosts. Heavy backends (LSP servers, kotlin-compiler-embeddable) are probed at runtime and never downloaded by default.

The coding domain exposes 47 tools in four groups:

Group Count Highlights
Shell & filesystem 28 sandboxed coding.shell (command whitelist), snapshot-based edit primitives with revert, diff (Myers/Patience), repo-governance protection (coding.protect)
Artifact creation & validation 6 scaffold (plugin/skill/js/script), scaffoldFlow (multi-file dev-flow), scaffoldFromExample (anti-staleness live templates, directory mode + stem-derived renames), validate (incl. repo-consistency)
Self-development 7 mvnBuild (structured diagnostics), ktSymbols/ktReferences/ktInheritance (zero-dep Kotlin analysis), impact + moduleGraph (live pom graph), devTask (AGENTS.md flow + execution), trapCheck (CDP pitfalls)
LSP 4 on-demand diagnostics/symbols/references for ts/js/py/rs (degrades gracefully when a server is missing)

Generic vs project-specific: the kernel is layered by mechanism vs data — diff, sandbox, LSP client, Kotlin analysis, Maven passthrough and the pom-graph scanner are generic and portable; the scaffolds, validators, ModuleMap, CdpTrapCheck and the governance defaults encode Browser4 conventions and are the layer to rewrite when reusing the kernel elsewhere.

  • Full tool reference & workflows: skills/browser4-coding/SKILL.md
  • Developing Browser4 itself: skills/browser4-dev/SKILL.md
  • Four-artifact comparison examples (real vs scaffold output): docs-dev/copilot/examples/
  • Evaluation summary (P1–P7): docs-dev/copilot/browser4-programming-support-eval.md

🧪 Test Fixture Server (MockSite)

Browser4 includes a lightweight MockSite server that serves static HTML pages for testing and demos. Start it from the repository root:

Windows (pwsh — PowerShell 7+): ./bin/test.ps1 mock-site -Dmock.site.port=18080 Linux/macOS: ./bin/test.sh mock-site -Dmock.site.port=18080

Key demo pages are served at http://localhost:18080/generated/. For the full page listing, environment variables, Python fallback, and Maven-based launch, see MockSite. For the test taxonomy and tagging system, see Test Taxonomy.


🤝 Support & Community

Join our community for support, feedback, and collaboration!

  • GitHub Discussions: Engage with developers and users.
  • Issue Tracker: Report bugs or request features.
  • Social Media: Follow us for updates and news.

We welcome contributions! See CONTRIBUTING.md for details.

群聊:dsh-Browser4社区交流3群


📜 Documentation

Comprehensive documentation is available in the docs/ directory and on our GitHub Pages site.


🔧 Proxy Configuration - Unblock Website Access

Set the environment variable PROXY_ROTATION_URL to the rotation URL provided by your proxy service provider:

export PROXY_ROTATION_URL=https://your-proxy-provider.com/rotation-endpoint

Each time you access this rotation URL, it should return a response containing one or more fresh proxy IPs. If you need this type of URL, please contact your proxy service provider.


License

Apache 2.0 License. See LICENSE for details.

Links