Armory
Source

Browse

Search and filter by type across the catalog

248 results in CLIs & Tools, Evals, Infrastructure · page 6 of 11

CLIs & ToolsExperimental

letta-ai-letta-code

Stateful agents that are like people, with memory, identity, and the ability to learn and adapt

Contributed by Sentinel

98.73,185 stars · 385 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
No one-command install · SourceDetails
InfrastructurePreview

sysbox

Next-generation container runtime enabling Docker-in-Docker and VM-like isolation without privileged containers.

98.63,847 stars · 228 forks
Harnessclaudecursorcodexopencodegemini
infrastructurecontainers
No one-command install · SourceDetails
EvalsPreview

lighteval

Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.

98.62,533 stars · 553 forks
Harnessclaudecursorcodexopencodegemini
evalshuggingfacebenchmarklightweight
No one-command install · SourceDetails
CLIs & ToolsPreview

crystal

A full-fledged desktop application for orchestrating, monitoring, and interacting with Claude Code agents.

98.53,114 stars · 197 forks
Harnessclaude
clientcli
No one-command install · SourceDetails
CLIs & ToolsPreview

omnara

A command center for AI agents that syncs Claude Code sessions across terminal, web, and mobile. Allows for remote monitoring, human-in-the-loop interaction, and team collaboration.

98.52,780 stars · 213 forks · 4 mentions
Harnessclaude
claude-codealternative-clients
No one-command install · SourceDetails
CLIs & ToolsExperimental

primeintellect-ai-prime-rl

Agentic RL Training at Scale

Contributed by Sentinel

98.52,007 stars · 421 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructurePreview

open-operator

Browserbase Open Operator: an open-source Operator-style web agent built on Stagehand; demonstrates full task decomposition, action planning, and evidence collection using the Browserbase cloud.

98.41,953 stars · 325 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserstagehand
No one-command install · SourceDetails
CLIs & ToolsPreview

tweakcc

Command-line tool to customize your Claude Code styling.

98.42,476 stars · 199 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
InfrastructurePreview

notte

Notte open-source web agent environment. It converts browser sessions into a Markov Decision Process with structured observation/action spaces, making browsers first-class RL and LLM agent environments.

98.32,003 stars · 181 forks
Harnessclaudecursorcodexopencodegemini
browserrl
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-tools

Well-crafted toolset for session continuity, featuring skills/commands to avoid compaction and recover context across sessions with cross-agent handoff between Claude Code and Codex CLI. Includes a fast Rust/Tantivy-powered full-text session search (TUI for humans, skill/CLI for agents), tmux-cli skill + command for interacting with scripts and CLI agents, and safety hooks to block dangerous commands.

98.21,989 stars · 132 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsPreview

cc-sessions

An opinionated approach to productive development with Claude Code

98.11,551 stars · 191 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsPreview

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructurePreview

e2b-desktop

E2B Desktop Sandbox: a cloud virtual desktop (Ubuntu + VNC) with Python SDK for screenshot, mouse, keyboard, and process control; designed for AI agents that need a full GUI environment in an isolated VM.

98.11,461 stars · 179 forks
Harnessclaudecursorcodexopencodegemini
browsere2b
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-ide-el

claude-code-ide.el integrates Claude Code with Emacs, like Anthropic’s VS Code/IntelliJ extensions. It shows ediff-based code suggestions, pulls LSP/flymake/flycheck diagnostics, and tracks buffer context. It adds an extensible MCP tool support for symbol refs/defs, project metadata, and tree-sitter AST queries.

98.01,658 stars · 112 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
InfrastructurePreview

agent-e

Emergence AI agent-E browser agent: hierarchical LLM-based web automation that uses DOM distillation and action abstraction layers to achieve significantly higher benchmark accuracy than prior browser agents.

98.01,249 stars · 190 forks
Harnessclaudecursorcodexopencodegemini
browseremergence
No one-command install · SourceDetails
CLIs & ToolsPreview

rulesync

A Node.js CLI tool that automatically generates configs (rules, ignore files, MCP servers, commands, and subagents) for various AI coding agents. Rulesync can convert configs between Claude Code and other AI agents in both directions.

98.01,373 stars · 143 forks
Harnessclaude
toolingcli
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-nvim

A seamless integration between Claude Code AI assistant and Neovim.

97.92,096 stars · 70 forks
Harnessclaude
toolingcliide-integrations
No one-command install · SourceDetails
EvalsPreview

wandb-weave-evals

Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.

97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
No one-command install · SourceDetails
EvalsPreview

langtrace

Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.

97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
No one-command install · SourceDetails
InfrastructurePreview

webvoyager

Research browser agent from Zhejiang University and HKU. It uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).

97.81,127 stars · 124 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserresearch
No one-command install · SourceDetails
Browse · Armory