Armory
Source

Browse

Search and filter by type across the catalog

816 results in Infrastructure, Sub-Agents, Evals · page 3 of 34

InfrastructurePreview

webvoyager

Research browser agent from Zhejiang University and HKU. It uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).

97.81,127 stars · 124 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserresearch
No one-command install · SourceDetails
InfrastructurePreview

surf-computer-use

E2B Surf: a Stagehand-powered computer-use interface layer for E2B Firecracker sandboxes; connects the act/extract/observe primitives directly to microVM display output for lightweight headless computer use.

97.6856 stars · 141 forks
Harnessclaudecursorcodexopencodegemini
browsere2b
No one-command install · SourceDetails
EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars · 107 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureExperimental

modal-labs-modal-client

Modal — serverless GPU/CPU containers for running agents, sandboxes and model inference from Python (this is the client SDK).

Contributed by Sentinel

97.0518 stars · 132 forks · 11 mentions
Harnessclaudecodexcursorgeminiopencode
infrastructure
No one-command install · SourceDetails
EvalsExperimental

stonybrooknlp-appworld

🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars · 78 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars · 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install · SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars · 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars · 62 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install · SourceDetails
EvalsPreview

vellum-evals

Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.

90.582 stars · 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
No one-command install · SourceDetails
EvalsPreview

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

82.327 stars · 12 forks
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingsdk
No one-command install · SourceDetails
EvalsPreview

galileo-evaluate

Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.

80.322 stars · 11 forks
Harnessclaudecursorcodexopencodegemini
evalshallucinationobservabilitysdk
No one-command install · SourceDetails
InfrastructurePreview

e2b

Open-source secure cloud sandboxes (Firecracker microVMs) for running AI-generated code. ~150ms cold start.

80.0passed install test
Harnessclaudecursorcodexopencodegemini
infrastructuresandbox
No one-command install · SourceDetails
InfrastructurePreview

modal

Serverless cloud platform for running Python functions, containers, and AI workloads with zero infra management.

80.0passed install test
Harnessclaudecursorcodexopencodegemini
infrastructureserverless
No one-command install · SourceDetails
InfrastructurePreview

railway

Zero-config cloud platform for deploying agent backends, databases, and services from a Git push.

80.0passed install test
Harnessclaudecursorcodexopencodegemini
infrastructuredeploy
No one-command install · SourceDetails
InfrastructurePreview

vercel

Frontend cloud platform with serverless functions and AI SDK integrations for deploying agent-facing UIs.

80.0passed install test
Harnessclaudecursorcodexopencodegemini
infrastructuredeploy
No one-command install · SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars · 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install · SourceDetails
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install · SourceDetails
Sub-AgentsPreview

3d-artist

3D art and asset creation specialist for game development. Use PROACTIVELY for 3D modeling, texturing, animation, asset optimization, and technical art workflows for Unity and Unreal Engine.

UnrankedNo signals yet
Harnessclaude
game-developmentsubagents
Details
Sub-AgentsPreview

4-1-beast

An agent definition that sets GPT-4.1 up as a coding agent.

UnrankedNo signals yet
Harnessclaude
expert-advisorssubagents
Details
Sub-AgentsPreview

a11y-architect

Accessibility Architect specializing in WCAG 2.2 compliance for Web and Native platforms. Use PROACTIVELY when designing UI components, establishing design systems, or auditing code for inclusive user experiences.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
subagent
Details
Sub-AgentsPreview

ab-test-analysis

Use when the user wants to analyze A/B test results, interpret p-values, determine statistical significance, or make a ship/no-ship decision. Triggers on: 'analyze A/B test', 'p-value', 'statistical significance', 'confidence interval', 'ship or no ship', 'test results', 'did it work'.

UnrankedNo signals yet
Harnessclaude
research-analysissubagent
Details
Sub-AgentsPreview

academic-research-synthesizer

Academic research synthesis specialist. Use PROACTIVELY for comprehensive research on academic topics, literature reviews, technical investigations, and well-cited analysis combining multiple sources.

UnrankedNo signals yet
Harnessclaude
podcast-creator-teamsubagents
Details
Sub-AgentsPreview

academic-researcher

Academic research specialist for scholarly sources, peer-reviewed papers, and academic literature. Use PROACTIVELY for research paper analysis, literature reviews, citation tracking, and academic methodology evaluation.

UnrankedNo signals yet
Harnessclaude
deep-research-teamsubagents
Details
Browse · Armory