Armory
Source

Browse

Search and filter by type across the catalog

1,535 results in Sub-Agents, Evals, Workflows, Infrastructure · page 3 of 64

EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
WorkflowsExperimental

noahshinn-reflexion

[NeurIPS 2023] Reflexion: Language Agents with Verbal Reinforcement Learning

Contributed by Sentinel

98.73,251 stars · 317 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
No one-command install · SourceDetails
InfrastructurePreview

sysbox

Next-generation container runtime enabling Docker-in-Docker and VM-like isolation without privileged containers.

98.63,847 stars · 228 forks
Harnessclaudecursorcodexopencodegemini
infrastructurecontainers
No one-command install · SourceDetails
EvalsPreview

lighteval

Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.

98.62,533 stars · 553 forks
Harnessclaudecursorcodexopencodegemini
evalshuggingfacebenchmarklightweight
No one-command install · SourceDetails
WorkflowsPreview

ralph-orchestrator

Ralph Orchestrator implements the simple but effective "Ralph Wiggum" technique for autonomous task completion, continuously running an AI agent against a prompt file until the task is marked as complete or limits are reached. This implementation provides a robust, well-tested, and feature-complete orchestration system for AI-driven development. Also cited in the Anthropic Ralph plugin documentation.

98.63,120 stars · 292 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
InfrastructurePreview

open-operator

Browserbase Open Operator: an open-source Operator-style web agent built on Stagehand; demonstrates full task decomposition, action planning, and evidence collection using the Browserbase cloud.

98.41,953 stars · 325 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserstagehand
No one-command install · SourceDetails
WorkflowsPreview

claude-codepro

A development environment for Claude Code with a spec-driven workflow, TDD enforcement, cross-session memory, semantic search, quality hooks and modular rules. Large, with wide coverage.

98.32,063 stars · 176 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
InfrastructurePreview

notte

Notte open-source web agent environment. It converts browser sessions into a Markov Decision Process with structured observation/action spaces, making browsers first-class RL and LLM agent environments.

98.32,003 stars · 181 forks
Harnessclaudecursorcodexopencodegemini
browserrl
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
EvalsPreview

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructurePreview

e2b-desktop

E2B Desktop Sandbox: a cloud virtual desktop (Ubuntu + VNC) with Python SDK for screenshot, mouse, keyboard, and process control; designed for AI agents that need a full GUI environment in an isolated VM.

98.11,461 stars · 179 forks
Harnessclaudecursorcodexopencodegemini
browsere2b
No one-command install · SourceDetails
InfrastructurePreview

agent-e

Emergence AI agent-E browser agent: hierarchical LLM-based web automation that uses DOM distillation and action abstraction layers to achieve significantly higher benchmark accuracy than prior browser agents.

98.01,249 stars · 190 forks
Harnessclaudecursorcodexopencodegemini
browseremergence
No one-command install · SourceDetails
WorkflowsPreview

the-ralph-playbook

A detailed guide to the Ralph Wiggum technique for autonomous coding loops, with the reasoning behind it and practical guidelines.

97.91,029 stars · 266 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
EvalsPreview

wandb-weave-evals

Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.

97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
No one-command install · SourceDetails
EvalsPreview

langtrace

Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.

97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
No one-command install · SourceDetails
InfrastructurePreview

webvoyager

Research browser agent from Zhejiang University and HKU. It uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).

97.81,127 stars · 124 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserresearch
No one-command install · SourceDetails
WorkflowsPreview

claude-code-documentation-mirror

A mirror of Anthropic's documentation pages for Claude Code, updated every few hours.

97.7983 stars · 138 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
InfrastructurePreview

surf-computer-use

E2B Surf: a Stagehand-powered computer-use interface layer for E2B Firecracker sandboxes; connects the act/extract/observe primitives directly to microVM display output for lightweight headless computer use.

97.6856 stars · 141 forks
Harnessclaudecursorcodexopencodegemini
browsere2b
No one-command install · SourceDetails
EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
Browse · Armory