Armory
Source

Browse

Search and filter by type across the catalog

852 results in Sub-Agents, Memory, Infrastructure, Evals · page 2 of 36

InfrastructureStable

browserbase-bb

Use when an agent must operate the live web (navigate, act, and extract on real pages) via a cloud browser driven by act/extract/observe primitives, with a local-Chromium escape hatch using the same code.

99.424,125 stars · 1,664 forks · 3 mentions
Harnessclaudecodex
browserweb-automationstagehandbrowserbase
No one-command install · SourceDetails
MemoryExperimental

gibsonai-memori

Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems. Built for enterprise, Memori works with the data infrastructure you already run, no rip-and-replace, and deploys across managed cloud, single-tenant cloud, VPC, and on-premises.

Contributed by Sentinel

99.416,313 stars · 3,283 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
InfrastructurePreview

browser-use-webui

Gradio web UI on top of the browser-use framework. It lets users run AI browser agents interactively, configure LLM providers, watch live recordings, and replay task sessions without writing Python.

99.416,310 stars · 2,721 forks
Harnessclaudecursorcodexopencodegemini
browserbrowser-use
No one-command install · SourceDetails
InfrastructurePreview

cua-computer-use-agent

trycua/cua open-source computer-use agent framework: Apple Silicon-native, runs lightweight macOS/Linux VMs with sub-second cold starts; provides a unified Python interface for screen capture, click, and type actions.

99.422,092 stars · 1,519 forks
Harnessclaudecursorcodexopencodegemini
browsercomputer-use
No one-command install · SourceDetails
EvalsPreview

deepeval

Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.

99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
No one-command install · SourceDetails
EvalsPreview

ragas

Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.

99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
No one-command install · SourceDetails
InfrastructureStable

mcp-tunnels-cloudflared

Use to let a hosted agent reach a private-data MCP server behind your firewall (an outbound tunnel plus a proxy, with per-server OAuth) so internal tools are usable without exposing them to the public internet.

99.315,474 stars · 1,417 forks
Harnessclaude
tunnelmcpprivate-datacloudflared
No one-command install · SourceDetails
MemoryExperimental

semantica-agi-semantica

Use when an agent's stored context needs provenance — you must be able to say where a remembered fact came from.

99.313,484 stars · 1,525 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
InfrastructurePreview

nanobrowser

Open-source Chrome extension that runs a multi-agent browser automation system locally. Planner, Navigator, and Validator agents collaborate inside the browser with no external API calls for web tasks.

99.313,713 stars · 1,451 forks
Harnessclaudecursorcodexopencodegemini
browsermulti-agent
No one-command install · SourceDetails
MemoryExperimental

nevamind-ai-memu

Use when one person's memory should follow them across several different agents.

99.314,387 stars · 1,063 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
InfrastructurePreview

browserless

Browserless.io headless browser service. It provides a Docker-deployable or cloud-hosted Chrome endpoint with REST and WebSocket APIs for screenshot, PDF, scraping, and Puppeteer/Playwright remote sessions.

99.313,654 stars · 1,039 forks
Harnessclaudecursorcodexopencodegemini
browserbrowserless
No one-command install · SourceDetails
InfrastructurePreview

bytebot

Bytebot open-source computer-use agent: a Docker-based Ubuntu desktop with AI-controlled mouse and keyboard; exposes an HTTP API for agents to send click, type, screenshot, and macro commands.

99.311,084 stars · 1,506 forks
Harnessclaudecursorcodexopencodegemini
browsercomputer-use
No one-command install · SourceDetails
MemoryExperimental

evermind-ai-everos

Use when you want the agent's memory to be plain Markdown on your own disk rather than a hosted database.

99.212,757 stars · 911 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

memtensor-memos

Use when memory should persist across tasks and be reused, not just retrieved once per conversation.

99.211,599 stars · 1,060 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
InfrastructurePreview

steel-browser

Open-source browser API optimised for AI agents. It provides session management, stealth settings, proxy rotation, and a REST/WebSocket interface on top of Chromium for cloud-scale agent browser access.

99.17,696 stars · 982 forks
Harnessclaudecursorcodexopencodegemini
browsersteel
No one-command install · SourceDetails
MemoryExperimental

plastic-labs-honcho

Memory library for building stateful agents

Contributed by Sentinel

99.16,980 stars · 865 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
InfrastructurePreview

microsandbox

Use as the OSS self-hosted sandbox when you need to run agent code on your own infra: libkrun-based microVM isolation with no per-sandbox vendor cost, the escape hatch from a managed runtime at high volume.

99.18,436 stars · 454 forks
Harnessclaudecodex
sandboxself-hostedlibkrunmicrovm
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
MemoryExperimental

getzep-zep

Examples, framework integrations and tools for Zep Cloud, Zep's hosted agent memory service; the repository says it is not the product itself. The open-source knowledge-graph engine behind Zep is Graphiti (getzep-graphiti).

Contributed by Sentinel

99.04,882 stars · 651 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
MemoryExperimental

caviraoss-openmemory

Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.

Contributed by Sentinel

98.94,478 stars · 504 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
Browse · Armory