OpenAI Operator
OpenAI's AI agents that can browse the web for you.
🔥 A list of tools, frameworks, and resources for building AI web agents
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
OpenAI's AI agents that can browse the web for you.
SOTA agent and framework that makes the web LLM-friendly.
Framework to automate browser-based workflows.
A research prototype exploring the future of human-agent interaction, starting with your browser.
A tool for building more deterministic and explainable web agents using semantic geometry on web content.
State-of-the-art AI agent that helps automate complex, cumbersome, multi-step tasks without repetitive manual input.
Vision-enabled web agent.
Agent & framework with HTML DOM distillation.
A general AI agent that can execute long running tasks across tools like browsers, terminals, and text editors.
An AI-powered Chrome extension that understands natural language and takes actions in your browser on your behalf.
MultimodalWebSurfer is a multimodal agent that can search the web and visit web pages.
A generalist multi-agent system for solving complex tasks including surfing the web via Autogen's MultimodalWebSurfer.
An AI-powered Chrome extension & browser agent that understands natural language and takes actions on your behalf.
A multi-agent system that executes browser-based tasks in parallel given a natural language prompt.
AI web agent Chrome extension that autonomously does tasks, scrapes to Sheets, and calls APIs with prompts in your own browser.
An open-source & local-first AI web agent Chrome extension with flexible LLM options and multi-agent system.
An open-source & self-hostable browser automation library for AI agents.
WebAgent for information seeking built by Tongyi Lab, Alibaba Group.
An MIT-licensed, open alternative to Anthropic's Cowork built with Opencode and dev-browser. Supports multiple LLM providers for launching computer-use agents to automate browser workflows.
An AI coworking agent in your browser that automates tasks, navigates pages, and works with files and 2000+ apps from a side panel.
Autonomous research agent that traverses the web to build a knowledge graph, then synthesizes an answer through adversarial redrafting.
Computer use agent that can control your browser.
A framework to enable multimodal models to operate a computer.
Desktop activity layer that helps models understand your workflow and complete tasks faster.
An open-source CLI based agent that can write & execute code as well as control your browser.
A GUI agent model designed to interact seamlessly with GUIs using human-like perception, reasoning, and action capabilities.
Hosted browser agents for SMEs to automate complex workflows.
AI-powered browser automation for data extraction.
Experimental project using GPT-4 Vision to browse the web via the Vimium extension.
An AI browser agent that helps companies maintain up-to-date documentation.
An AI coworker embedding into and controlling your browser.
A free, experimental Chrome extension that leverages Claude Computer Use to automate tasks in your browser.
AI-powered scraping automations that evolve with your target sites.
A no-code browser automation platform that turns screen recordings into reusable automated workflows.
A Chrome extension that enables AI-powered browser automations, allowing users to automate tasks and workflows directly within the browser.
Browser assistant for web task automation.
Browser extension for page summaries and Q&A.
Chrome extension webscraping that can leverage AI for structured data extraction.
An integrated platform that combines a browser, file manager, and AI assistant with browser-level context.
An AI-powered browser by Perplexity. Not much more details out yet.
AI-first web browser envisioned by The Browser Company (Arc).
No-code web data extraction solution using agentic AI.
Chrome extension that uses AI to control and read webpages, including auto summaries, web automation, scraping, and MCP support.
AI browser workflow automation platform that turns recorded web tasks into reusable runs with API triggers, schedules, credentials, logs, and human review.
Open-source headless browser API built specifically for AI agents and apps.
Open-source deep research harness for building cited web research agents with ledger-based coverage audits, pluggable search providers, and Steel-backed browser fetches.
Tool for parsing GUIs for vision based agents.
Framework for natural language web automation.
Toolkit integration with AI agents.
A headless browser API for AI workflows.
AI web browsing framework.
Vision utilities library for web interaction agents.
Remote web agents that execute tasks on any website and return structured JSON via a single API call.
Containerized computer use agent framework with a virtual desktop environment.
Vision-first browser agent with self-healing deterministic replay. Screenshot → model → action loop over CDP, multi-provider (Anthropic, Google, OpenAI), action caching for zero-token reruns.
Open-source API for browser agents to automate repetitive workflows. Works with multiple browsers/LLM providers, and minimizes costs with a self-healing action cache.
Local-first trace viewer for debugging Playwright, Browser Use, Stagehand, and other web-agent runs with redacted shareable exports.
Browser infrastructure for AI agents with managed sessions, an agent runtime, and credential vault and persona authentication primitives.
Playwright wrapper for a stealth-patched Firefox 150 build. Drop-in replacement returning native Playwright Browser objects; spoofing happens in C++ source with no JS-level overrides.
Browser agent framework from Microsoft Research where the agent writes and runs Playwright scripts in a terminal workspace; supports OpenAI, Anthropic, and OpenRouter backends.
Browser automation CLI and skills for AI agents to operate real browsers, manage sessions, support human handoff, and capture screenshots and evidence.
Browser extension that sits between an AI agent and the page, stripping prompt injection, masking PII/credentials, and removing dark patterns before content reaches the model.
Open-source SDK for building browser and computer-use RL environments to evaluate and train web agents, with task-based verifiable rewards runnable as evals or RL training across any model.
Configurable web proxy and browser-as-a-service for deploying and operating AI agents in a sandbox layer on top of any third-party website, using client-side extensions and without source-code access.
Self-learning browser infrastructure for AI agents that records how a site is navigated, then compiles that context into deterministic CLI adapters and reusable sitemap memory for later runs.
Unofficial open-source Chrome extension and local companion that let Hermes Agent control only explicitly attached tabs in a user's Chrome session.
APIs for turning websites into LLM-friendly markdown.
Open-source LLM Friendly Web Crawler & Scraper.
Python scraper based on AI.
The web-browsing agent module of the OpenAgents platform (HKU). Enables autonomous navigation of websites via natural language, as part of a larger multi-modal agent framework.
Turns any website into a type-safe API you can rely on.
Uses LLMs for intelligent scraping and content understanding.
Open-source headless browser engine for AI agents. Compiles HTML to Semantic Object Model (SOM) with 17.5x token compression. 13 MCP tools. First browser tool on the MCP Registry. Rust, Apache-2.0.
Create complex Playwright spiders with natural language prompts.
AI web data agent that plans browser scraping tasks, extracts structured web data, and exposes MCP tools for dataset workflows.
Web APIs and MCP tools for search, scraping, crawling, schema-based extraction, document parsing, monitoring, and batch jobs.
A query language and toolkit that makes the web AI-ready.
Search API that provides Google Search results for your agents.
Performant and cost effective search API that provides Google Search results for your agents.
Neural search platform for web data.
Search engine that indexes 1,750+ agent-first tools ranked by agentic readiness. Available as an MCP server with tools for searching, scoring, and monitoring agent infrastructure.
Web search API for AI agents with five tools (search, news, images, scrape, research); agents pay per call in USDC via the x402 protocol, or use a free API key.
Open-source web search and evidence tool for AI agents with query rewriting, source-domain zoom-in, structured sourced outputs, and MCP and LangGraph integrations.
Unified social and commerce data API covering 50+ platforms as one GET and one JSON schema.
Leaderboard compiling AI agent products and their performance on widely used WebVoyager benchmarks.
A collection of challenges designed for testing general-purpose web-browsing AI agents.
An open-source evaluation framework for web-based AI agents.
A large-scale dataset for generalist web agents.
OpenAI's research paper that introduces World of Bits: a platform where agents complete tasks on the internet by performing low-level keyboard and mouse actions.
A classic suite of 104 mini web browser tasks in a synthetic environment. It is an extension of the OpenAI MiniWoB benchmark.
51-URL benchmark comparing HTML vs Markdown vs SOM representations for AI agents. Measures token efficiency, latency, and accuracy across GPT-4o and Claude Sonnet 4.
A realistic, self-hostable web environment for autonomous agents. Includes official leaderboard tracking agent performance.
An online evaluation framework for dynamic web environments. Tests agents on live websites.
OpenAI's browser-assisted question-answering research project.
A simulated e-commerce shopping environment with 1.18M real Amazon products.
Vision-enabled benchmark for real-world website interaction with large multimodal models.
A suite of 33 browser-based tasks for enterprise "knowledge worker" scenarios.
A gym environment for web task automation.
A benchmark on historical versions of web UI.
283 everyday tasks (V1 153 + V2 130) on 163 live production websites across 15 categories. Two-stage scoring (final HTTP-request interception + LLM judge) blocks only the write request so real sites stay clean. Public leaderboard with 5-layer execution traces (recording, action log, request log,…
Tutorial demonstrating how to build a web navigation agent using LangGraph Agents, Vision Models, and Web Voyager.
Step-by-step guide to create an AI that browses the web using Playwright and the Browser-Use library.
Instructions on installing the open-source Browser-Use agent with a local LLM.
Walks through deploying a Browser-Use web UI agent powered by the DeepSeek model on a cloud VM.
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…