browser-use
Python library that makes web browsers accessible to AI agents; built on Playwright and LangChain. Supports multi-tab, vision + accessibility-tree hybrid mode, custom actions, and a self-correcting agent loop.
Search and filter by type across the catalog
130 results in Evals, Identity, Infrastructure · page 1 of 6
Python library that makes web browsers accessible to AI agents; built on Playwright and LangChain. Supports multi-tab, vision + accessibility-tree hybrid mode, custom actions, and a self-correcting agent loop.
Secure and elastic sandboxes for running AI-generated code. The public repository stopped updating in June 2026, when development moved to a private codebase.
Use as the default runtime when an agent must execute untrusted code or commands — Firecracker microVMs with ~150ms cold start give each run an isolated, disposable computer.
Never stop coding. Free MIT AI gateway: one endpoint, 352 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by 550+ contributors
Contributed by Sentinel
Use as the payments rail when an agent should earn or spend money in code — create customers, prices, payment links, and usage-based billing — the infrastructure behind an agent that funds its own compute.
Open-source agent platform that automates browser-based workflows using LLMs and computer vision — identifies interactive elements via screenshots, handles CAPTCHAs, and supports complex multi-step form flows.
Use when an agent must operate the live web — navigate, act, and extract on real pages — via a cloud browser driven by act/extract/observe primitives, with a local-Chromium escape hatch using the same code.
Gradio web UI on top of the browser-use framework — lets users run AI browser agents interactively, configure LLM providers, watch live recordings, and replay task sessions without writing Python.
trycua/cua open-source computer-use agent framework — Apple Silicon-native, runs lightweight macOS/Linux VMs with sub-second cold starts; provides a unified Python interface for screen capture, click, and type actions.
Use to let a hosted agent reach a private-data MCP server behind your firewall — an outbound tunnel plus a proxy, with per-server OAuth — so internal tools are usable without exposing them to the public internet.
Open-source Chrome extension that runs a multi-agent browser automation system locally — Planner, Navigator, and Validator agents collaborate inside the browser with no external API calls for web tasks.
Browserless.io headless browser service — provides a Docker-deployable or cloud-hosted Chrome endpoint with REST and WebSocket APIs for screenshot, PDF, scraping, and Puppeteer/Playwright remote sessions.
Bytebot open-source computer-use agent — Docker-based Ubuntu desktop with AI-controlled mouse and keyboard; exposes an HTTP API for agents to send click, type, screenshot, and macro commands.
Open-source browser API optimised for AI agents — provides session management, stealth settings, proxy rotation, and a REST/WebSocket interface on top of Chromium for cloud-scale agent browser access.