Armory
Source

Browse

Search and filter by type across the catalog

264 results in Evals, Identity, Infrastructure, Hooks · page 1 of 11

InfrastructurePreview

browser-use

Python library that makes web browsers accessible to AI agents; built on Playwright and LangChain. Supports multi-tab, vision + accessibility-tree hybrid mode, custom actions, and a self-correcting agent loop.

99.9111,989 stars · 12,312 forks · 11 mentions · passed install test
Harnessclaudecursorcodexopencodegemini
browserbrowser-use
No one-command install · SourceDetails
InfrastructurePreview

daytona

Secure and elastic sandboxes for running AI-generated code. The public repository stopped updating in June 2026, when development moved to a private codebase.

99.971,846 stars · 5,651 forks · 9 mentions · passed install test
Harnessclaudecursorcodexopencodegemini
infrastructuredev-environments
No one-command install · SourceDetails
InfrastructureExperimental

ggml-org-llama-cpp

LLM inference in C/C++

Contributed by Sentinel

99.9126,728 stars · 22,647 forks · 5 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureExperimental

vllm-project-vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

Contributed by Sentinel

99.990,743 stars · 21,577 forks · 15 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureStable

e2b-sandbox

Use as the default runtime when an agent must execute untrusted code or commands. Firecracker microVMs with ~150ms cold start give each run an isolated, disposable computer.

99.813,635 stars · 1,015 forks · 10 mentions · passed install test
Harnessclaudecodex
sandboxruntimefirecrackermicrovm
No one-command install · SourceDetails
InfrastructureExperimental

diegosouzapw-omniroute

Never stop coding. Free MIT AI gateway: one endpoint, 352 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by 550+ contributors

Contributed by Sentinel

99.860,005 stars · 8,342 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureExperimental

sgl-project-sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

Contributed by Sentinel

99.733,202 stars · 8,472 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureStable

stripe-agent-toolkit

Use as the payments rail when an agent should earn or spend money in code (create customers, prices, payment links, and usage-based billing): the infrastructure behind an agent that funds its own compute.

99.71,785 stars · 329 forks · passed install test
Harnessclaudecodex
paymentsstripebillingfinancial-rails
No one-command install · SourceDetails
EvalsPreview

promptfoo

CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.

99.524,737 stars · 2,255 forks
Harnessclaudecursorcodexopencodegemini
evalsred-teamingcicli
No one-command install · SourceDetails
InfrastructurePreview

skyvern

Open-source agent platform that automates browser-based workflows using LLMs and computer vision. It identifies interactive elements via screenshots, handles CAPTCHAs, and supports complex multi-step form flows.

99.522,907 stars · 2,152 forks
Harnessclaudecursorcodexopencodegemini
browserskyvern
No one-command install · SourceDetails
EvalsPreview

openai-evals

OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.

99.519,509 stars · 3,093 forks
Harnessclaudecursorcodexopencodegemini
evalsregistrybenchmark
No one-command install · SourceDetails
EvalsPreview

lm-evaluation-harness

EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.

99.513,860 stars · 3,533 forks
Harnessclaudecursorcodexopencodegemini
evalsacademicbenchmarkharness
No one-command install · SourceDetails
InfrastructureStable

browserbase-bb

Use when an agent must operate the live web (navigate, act, and extract on real pages) via a cloud browser driven by act/extract/observe primitives, with a local-Chromium escape hatch using the same code.

99.424,125 stars · 1,664 forks · 3 mentions
Harnessclaudecodex
browserweb-automationstagehandbrowserbase
No one-command install · SourceDetails
InfrastructurePreview

browser-use-webui

Gradio web UI on top of the browser-use framework. It lets users run AI browser agents interactively, configure LLM providers, watch live recordings, and replay task sessions without writing Python.

99.416,310 stars · 2,721 forks
Harnessclaudecursorcodexopencodegemini
browserbrowser-use
No one-command install · SourceDetails
InfrastructurePreview

cua-computer-use-agent

trycua/cua open-source computer-use agent framework: Apple Silicon-native, runs lightweight macOS/Linux VMs with sub-second cold starts; provides a unified Python interface for screen capture, click, and type actions.

99.422,092 stars · 1,519 forks
Harnessclaudecursorcodexopencodegemini
browsercomputer-use
No one-command install · SourceDetails
EvalsPreview

deepeval

Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.

99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
No one-command install · SourceDetails
EvalsPreview

ragas

Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.

99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
No one-command install · SourceDetails
InfrastructureStable

mcp-tunnels-cloudflared

Use to let a hosted agent reach a private-data MCP server behind your firewall (an outbound tunnel plus a proxy, with per-server OAuth) so internal tools are usable without exposing them to the public internet.

99.315,474 stars · 1,417 forks
Harnessclaude
tunnelmcpprivate-datacloudflared
No one-command install · SourceDetails
InfrastructurePreview

nanobrowser

Open-source Chrome extension that runs a multi-agent browser automation system locally. Planner, Navigator, and Validator agents collaborate inside the browser with no external API calls for web tasks.

99.313,713 stars · 1,451 forks
Harnessclaudecursorcodexopencodegemini
browsermulti-agent
No one-command install · SourceDetails
InfrastructurePreview

browserless

Browserless.io headless browser service. It provides a Docker-deployable or cloud-hosted Chrome endpoint with REST and WebSocket APIs for screenshot, PDF, scraping, and Puppeteer/Playwright remote sessions.

99.313,654 stars · 1,039 forks
Harnessclaudecursorcodexopencodegemini
browserbrowserless
No one-command install · SourceDetails
InfrastructurePreview

bytebot

Bytebot open-source computer-use agent: a Docker-based Ubuntu desktop with AI-controlled mouse and keyboard; exposes an HTTP API for agents to send click, type, screenshot, and macro commands.

99.311,084 stars · 1,506 forks
Harnessclaudecursorcodexopencodegemini
browsercomputer-use
No one-command install · SourceDetails
EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
InfrastructurePreview

steel-browser

Open-source browser API optimised for AI agents. It provides session management, stealth settings, proxy rotation, and a REST/WebSocket interface on top of Chromium for cloud-scale agent browser access.

99.17,696 stars · 982 forks
Harnessclaudecursorcodexopencodegemini
browsersteel
No one-command install · SourceDetails
HooksPreview

plannotator

Interactive plan review UI that intercepts ExitPlanMode via hooks, letting users visually annotate plans with comments, deletions, and replacements before approving or denying with detailed feedback.

99.18,344 stars · 618 forks
Harnessclaude
hook
No one-command install · SourceDetails
Browse · Armory