Armory
Source

Browse

Search and filter by type across the catalog

854 results in Evals, Infrastructure, Identity, Sub-Agents · page 2 of 36

InfrastructurePreview

steel-browser

Open-source browser API optimised for AI agents. It provides session management, stealth settings, proxy rotation, and a REST/WebSocket interface on top of Chromium for cloud-scale agent browser access.

99.17,696 stars · 982 forks
Harnessclaudecursorcodexopencodegemini
browsersteel
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
InfrastructurePreview

microsandbox

Use as the OSS self-hosted sandbox when you need to run agent code on your own infra: libkrun-based microVM isolation with no per-sandbox vendor cost, the escape hatch from a managed runtime at high volume.

99.18,436 stars · 454 forks
Harnessclaudecodex
sandboxself-hostedlibkrunmicrovm
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
InfrastructurePreview

claude-managed-agents-selfhost

Use for enterprise hosted-control agents (the agent loop runs at the provider while execution happens on your own infrastructure) when you want a managed control plane but data and code must stay on your machines.

98.93,876 stars · 836 forks
Harnessclaude
managed-agentshosted-controlenterpriseself-hosted-execution
No one-command install · SourceDetails
EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
EvalsPreview

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
No one-command install · SourceDetails
InfrastructurePreview

sysbox

Next-generation container runtime enabling Docker-in-Docker and VM-like isolation without privileged containers.

98.63,847 stars · 228 forks
Harnessclaudecursorcodexopencodegemini
infrastructurecontainers
No one-command install · SourceDetails
EvalsPreview

lighteval

Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.

98.62,533 stars · 553 forks
Harnessclaudecursorcodexopencodegemini
evalshuggingfacebenchmarklightweight
No one-command install · SourceDetails
IdentityExperimental

metapriseai-orgkernel

Use when every agent action must be traceable to a specific agent instance with its own key, scoped token and tamper-evident log.

98.52,699 stars · 247 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
InfrastructurePreview

open-operator

Browserbase Open Operator: an open-source Operator-style web agent built on Stagehand; demonstrates full task decomposition, action planning, and evidence collection using the Browserbase cloud.

98.41,953 stars · 325 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserstagehand
No one-command install · SourceDetails
InfrastructurePreview

notte

Notte open-source web agent environment. It converts browser sessions into a Markov Decision Process with structured observation/action spaces, making browsers first-class RL and LLM agent environments.

98.32,003 stars · 181 forks
Harnessclaudecursorcodexopencodegemini
browserrl
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
IdentityExperimental

infisical-agent-vault

A HTTP credential proxy and vault for AI agents like Claude Code, OpenClaw, Hermes, custom agents + harnesses, and more.

Contributed by Sentinel

98.32,171 stars · 141 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
EvalsPreview

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructurePreview

e2b-desktop

E2B Desktop Sandbox: a cloud virtual desktop (Ubuntu + VNC) with Python SDK for screenshot, mouse, keyboard, and process control; designed for AI agents that need a full GUI environment in an isolated VM.

98.11,461 stars · 179 forks
Harnessclaudecursorcodexopencodegemini
browsere2b
No one-command install · SourceDetails
InfrastructurePreview

agent-e

Emergence AI agent-E browser agent: hierarchical LLM-based web automation that uses DOM distillation and action abstraction layers to achieve significantly higher benchmark accuracy than prior browser agents.

98.01,249 stars · 190 forks
Harnessclaudecursorcodexopencodegemini
browseremergence
No one-command install · SourceDetails
Browse · Armory