Unofficial, community-maintained TypeScript port of the Python browser-use library
A TypeScript-first library for building AI-powered web agents that can autonomously browse, interact with, and extract data from the web using LLMs and Playwright.
Unofficial project. This repository is a community-maintained TypeScript port of the Python browser-use library. It is not developed, maintained, or endorsed by the browser-use maintainers or by Browser Use (the company), and its versions are independent of the Python package.
- Official Python library: github.com/browser-use/browser-use. Official docs and Browser Use Cloud: docs.browser-use.com.
- Questions and bugs about this TypeScript port belong in webllm/browser-use issues, not in the official project.
The port follows the Python library's behavior closely, with a native Node.js experience, full type safety, and support for all major LLM providers.
- 🤖 Autonomous Browser Control — AI-driven navigation, clicking, typing, form filling, scrolling, and tab management
- 🧠 15+ LLM Providers & Adapters — OpenAI, Anthropic, Google Gemini, Azure, AWS Bedrock, Groq, Ollama, DeepSeek, OpenRouter, Mistral, Cerebras, Browser Use, LiteLLM, OCI Raw, Vercel, and custom providers
- 👁️ Vision Support — Screenshot-based understanding for visual web interactions
- 🔧 45+ Built-in Actions — Navigation, element interaction, scrolling, forms, tabs, content extraction, file I/O, and more
- 🧩 Custom Actions — Extensible registry with Zod schema validation, domain restrictions, and page filters
- 🔌 MCP Server — Model Context Protocol support for Claude Desktop and MCP-compatible clients
- ⌨️ CLI Tool — Interactive and one-shot modes for quick browser tasks
- 🔒 Security First — Sensitive data masking, domain restrictions, and Chromium sandboxing
- 📊 Observability — Event system, performance tracing, and session recording (GIF)
- 🐳 Docker Ready — Configurable for containerized and CI/CD environments
| This repository | Official browser-use | |
|---|---|---|
| Language | TypeScript / Node.js | Python |
| Maintained by | The community (webllm) | The Browser Use team |
| Package | browser-use on npm |
browser-use on PyPI |
| Issues and help | webllm/browser-use | browser-use/browser-use |
- The port tracks the Python library's behavior and naming (snake_case options, the same action set) and ports upstream changes periodically. Some upstream features are intentionally not ported; see the architecture decisions.
- Browser Use Cloud and
ChatBrowserUseare hosted services operated by Browser Use. This port only calls their public APIs with your own API key. Browser Use's official TypeScript SDK for its cloud API is the separatebrowser-use-sdkpackage. - This port collects no telemetry or usage analytics.
- "Browser Use" and related names belong to their respective owners and are used here only to describe compatibility.
npm install browser-use
# Playwright browsers are installed automatically via postinstallexport OPENAI_API_KEY=sk-your-api-key
# or ANTHROPIC_API_KEY, GOOGLE_API_KEY, etc.import { Agent } from 'browser-use';
import { ChatOpenAI } from 'browser-use/llm/openai';
const agent = new Agent({
task: 'Go to google.com and search for "TypeScript tutorials"',
llm: new ChatOpenAI({
model: 'gpt-4o',
apiKey: process.env.OPENAI_API_KEY,
}),
});
const history = await agent.run();
console.log('Result:', history.final_result());
console.log('Success:', history.is_successful());npx tsx example.ts# Interactive mode
npx browser-use
# One-shot task
npx browser-use "Go to example.com and extract the page title"
# With specific model
npx browser-use --model claude-opus-5 -p "Search for AI news"
# Headless mode
npx browser-use --headless -p "Check the weather"
# MCP server mode
npx browser-use --mcp
# Minimal direct-browser MCP server for coding agents
npx browser-use --cli-mcp
# Install the bundled coding-agent skill (codex, claude, cursor, openclaw, etc.)
npx browser-use skill install --target codexThe bundled browser-use skill teaches coding
agents to drive the persistent browser through the two-tool CLI MCP server or
the browser-use-direct fallback. Run npx browser-use skill install without
a target to install it for every supported coding agent, or use
npx browser-use skill show to inspect it first.
A second skill, browser-use-ts, is a
reference for writing TypeScript code with this package (Agent, browser
configuration, custom actions, providers, the Actor API, and integrations).
Install it with npx browser-use skill install --skill browser-use-ts, and run
npx browser-use skill list to see every bundled skill.
For multi-step direct control, browser-use-direct script <file|-> runs a
JavaScript file against the persistent browser with helpers such as
goto_url, state, click, js, and print. Scripts run with your Node.js
privileges and are not exposed through the MCP server.
┌─────────────────────────────────────────────────────┐
│ Browser-Use │
├─────────────────────────────────────────────────────┤
│ Agent ← MessageManager ← LLM Providers │
│ ↓ │
│ Controller → Action Registry → BrowserSession │
│ ↓ │
│ DomService │
└─────────────────────────────────────────────────────┘
| Component | Description |
|---|---|
| Agent | Central orchestrator — runs the observe → think → act loop |
| Controller | Manages action registration and execution via Registry |
| BrowserSession | Playwright wrapper — browser lifecycle, tab management, screenshots |
| DomService | Extracts interactive elements with indexed mapping for LLM consumption |
| MessageManager | Manages LLM conversation history with token optimization |
| LLM Providers | Unified BaseChatModel interface across 15+ providers and adapters |
- Agent receives a natural language task
- DomService extracts the current page state (interactive elements + optional screenshot)
- LLM analyzes the state and returns actions to take
- Controller validates and executes actions through the Registry
- Results feed back to the LLM for the next step
- Loop continues until
doneaction ormax_steps
| Provider | Import | Vision | Notes |
|---|---|---|---|
| OpenAI | browser-use/llm/openai |
✅ | Default provider, reasoning models (o1/o3/o4) |
| Codex | browser-use/llm/codex |
✅ | Experimental ChatGPT/Codex OAuth provider |
| Anthropic | browser-use/llm/anthropic |
✅ | Prompt caching support |
| Google Gemini | browser-use/llm/google |
✅ | Extended thinking support |
| Azure OpenAI | browser-use/llm/azure |
✅ | Enterprise deployment |
| AWS Bedrock | browser-use/llm/aws |
✅ | Claude via AWS |
| Groq | browser-use/llm/groq |
❌ | Fastest inference |
| Ollama | browser-use/llm/ollama |
❌ | Local/self-hosted models |
| DeepSeek | browser-use/llm/deepseek |
❌ | Cost-effective |
| OpenRouter | browser-use/llm/openrouter |
Varies | Multi-model routing |
| Mistral | browser-use/llm/mistral |
Varies | Mistral models |
| Cerebras | browser-use/llm/cerebras |
❌ | Fast inference |
| Browser Use | browser-use/llm/browser-use |
Varies | Hosted Browser Use LLM |
| LiteLLM | browser-use/llm/litellm |
Varies | OpenAI-compatible LiteLLM gateway |
| OCI Raw | browser-use/llm/oci-raw |
Varies | Oracle Cloud Generative AI |
| Vercel | browser-use/llm/vercel |
Varies | Vercel AI Gateway / routed models |
Provider examples
// OpenAI
import { ChatOpenAI } from 'browser-use/llm/openai';
const llm = new ChatOpenAI({
model: 'gpt-4o',
apiKey: process.env.OPENAI_API_KEY,
});
// Anthropic
import { ChatAnthropic } from 'browser-use/llm/anthropic';
const llm = new ChatAnthropic({
model: 'claude-opus-5',
apiKey: process.env.ANTHROPIC_API_KEY,
});
// Google Gemini
import { ChatGoogle } from 'browser-use/llm/google';
const llm = new ChatGoogle('gemini-2.5-flash');
// Ollama (local)
import { ChatOllama } from 'browser-use/llm/ollama';
const llm = new ChatOllama('llama3', 'http://localhost:11434');
// OpenAI Reasoning Models
const llm = new ChatOpenAI({ model: 'o3-mini', reasoningEffort: 'medium' });
// Codex OAuth provider (experimental)
// First run: npx browser-use auth codex login
import { ChatCodex } from 'browser-use/llm/codex';
const codexLlm = new ChatCodex({ model: 'gpt-5.5' });
// Browser Use gateway: one BROWSER_USE_API_KEY can route provider/model ids
import { ChatBrowserUse } from 'browser-use/llm/browser-use';
const gatewayLlm = new ChatBrowserUse({
model: 'anthropic/claude-sonnet-4-6',
});const agent = new Agent({
task: `Go to amazon.com, search for "wireless keyboard",
extract the name, price, and rating of the first 5 products as JSON`,
llm,
use_vision: true,
});
const history = await agent.run(30);
console.log(history.final_result());const agent = new Agent({
task: 'Login to the dashboard',
llm,
sensitive_data: {
'*.example.com': {
username: process.env.SITE_USERNAME!,
password: process.env.SITE_PASSWORD!,
},
},
browser_session: new BrowserSession({
browser_profile: new BrowserProfile({
allowed_domains: ['*.example.com'],
}),
}),
});import fs from 'node:fs';
import { Controller, ActionResult } from 'browser-use';
import { z } from 'zod';
const controller = new Controller();
controller.registry.action('Save screenshot to file', {
param_model: z.object({
filename: z.string().describe('Output filename'),
}),
})(async function save_screenshot(params, ctx) {
const screenshot = await ctx.page.screenshot();
fs.writeFileSync(`./screenshots/${params.filename}`, screenshot);
return new ActionResult({
extracted_content: `Screenshot saved as ${params.filename}`,
});
});
const agent = new Agent({ task: '...', llm, controller });const agent = new Agent({
task: 'Navigate to hacker news and summarize the top stories',
llm,
use_vision: true,
vision_detail_level: 'high', // 'auto' | 'low' | 'high'
generate_gif: './session.gif',
});const agent = new Agent({
task: `Compare "Sony WH-1000XM5" prices:
1. Open amazon.com and search for the product
2. Open bestbuy.com in a new tab and search
3. Provide a comparison summary`,
llm,
use_vision: true,
});const agent = new Agent({ task: '...', llm });
agent.eventbus.on('CreateAgentStepEvent', (event) => {
console.log('Step completed:', event.step_id);
});
await agent.run();const agent = new Agent({
task: 'Your task',
llm,
use_vision: true, // Enable screenshot analysis
max_actions_per_step: 5, // Actions per LLM call
max_failures: 3, // Max retries on failure
generate_gif: './recording.gif', // Session recording
validate_output: true, // Strict output validation
use_thinking: true, // Extended thinking prompts
llm_timeout: 60, // LLM call timeout (seconds)
step_timeout: 180, // Step timeout (seconds)
extend_system_message: 'Be concise', // Custom prompt additions
});
const history = await agent.run(50); // Max 50 stepsimport { BrowserProfile, BrowserSession } from 'browser-use';
const profile = new BrowserProfile({
headless: true,
viewport: { width: 1920, height: 1080 },
user_data_dir: './my-profile', // Persistent sessions
allowed_domains: ['*.example.com'], // Domain restrictions
highlight_elements: true, // Visual debugging
proxy: { server: 'http://proxy:8080' },
});
const session = new BrowserSession({ browser_profile: profile });
const agent = new Agent({ task: '...', llm, browser_session: session });| Variable | Description |
|---|---|
OPENAI_API_KEY |
OpenAI API key |
ANTHROPIC_API_KEY |
Anthropic API key |
GOOGLE_API_KEY |
Google API key |
BROWSER_USE_HEADLESS |
Run browser headlessly (true/false) |
BROWSER_USE_LOGGING_LEVEL |
Log level: debug, info, warning, error |
BROWSER_USE_ALLOWED_DOMAINS |
Comma-separated domain allowlist |
See Configuration Guide for the full list.
Browser-Use can run as an MCP server, exposing browser automation as tools for Claude Desktop:
npx browser-use --mcpAdd to your Claude Desktop config (~/Library/Application Support/Claude/claude_desktop_config.json):
{
"mcpServers": {
"browser-use": {
"command": "npx",
"args": ["browser-use", "--mcp"],
"env": {
"OPENAI_API_KEY": "your-api-key"
}
}
}
}Core MCP tools include retry_with_browser_use_agent, browser_navigate, browser_click, browser_type, browser_get_state, browser_extract_content, browser_scroll, browser_go_back, browser_list_tabs, browser_switch_tab, browser_close_tab, browser_list_sessions, browser_close_session, and browser_close_all. The server also exposes registered controller actions as additional MCP tools.
See MCP Server Guide for more details.
Run Anthropic's browser toolset (browser_toolset_20260801) on a Browser Use browser — local Chromium, Browser Use Cloud, or an existing session — together with a bounded Bash tool:
import Anthropic from '@anthropic-ai/sdk';
import {
BrowserUseToolset,
createBashTool,
runBrowserToolsetConversation,
} from 'browser-use/integrations/anthropic';
const toolset = new BrowserUseToolset();
try {
const { message } = await runBrowserToolsetConversation({
client: new Anthropic(),
toolset,
tools: [createBashTool({ outputDir: 'outputs' })],
task: 'Read the first three Hacker News posts and save their titles to hn.md.',
});
console.log(message.content);
} finally {
await toolset.close();
}See Claude Browser Toolset for runtimes, members, and safety controls.
- Sensitive Data Masking — Credentials are automatically masked in logs and LLM context
- Domain Restrictions — Lock browser navigation to trusted domains
- Domain-scoped Secrets — Credentials are only injected on matching domains
- Sensitive Data Warning — Browser-Use warns when
sensitive_datais used withoutallowed_domains - Chromium Sandbox — Enabled by default for production security
const agent = new Agent({
task: 'Login and fetch invoices',
llm,
sensitive_data: {
'*.example.com': {
username: process.env.USERNAME!,
password: process.env.PASSWORD!,
},
},
browser_session: new BrowserSession({
browser_profile: new BrowserProfile({
allowed_domains: ['*.example.com'],
}),
}),
});See Security Guide for production deployment best practices.
| Document | Description |
|---|---|
| Quick Start | Get started in 5 minutes |
| Architecture | System design and component overview |
| API Reference | Complete API documentation |
| Configuration | All configuration options |
| LLM Providers | Provider setup and comparison |
| Actions | Built-in and custom actions |
| MCP Server | MCP integration guide |
| Claude Browser Toolset | Anthropic browser toolset driver |
| Security | Security best practices |
| Examples | More code examples |
| Contributing | Contribution guidelines |
# Install dependencies
pnpm install
# Build
pnpm build
# Run tests
pnpm test
# Lint & format
pnpm lint
pnpm prettier
# Type checking
pnpm typecheck
# Run an example
pnpm exec tsx examples/simple-search.ts- Node.js 20.16+ (Node 20) or 22.3+
- LLM API Key — At least one supported provider
- Playwright — Installed automatically as a dependency