Armory
Source

Browse

Search and filter by type across the catalog

231 results in Evals, Identity, CLIs & Tools · page 5 of 10

EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-devtools

A desktop app that shows your Claude Code sessions by reading their logs: context use per turn across categories, compaction, sub-agent execution trees and custom notification triggers.

98.73,893 stars · 298 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
CLIs & ToolsExperimental

letta-ai-letta-code

Stateful agents that are like people, with memory, identity, and the ability to learn and adapt

Contributed by Sentinel

98.73,185 stars · 385 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
No one-command install · SourceDetails
EvalsPreview

lighteval

Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.

98.62,533 stars · 553 forks
Harnessclaudecursorcodexopencodegemini
evalshuggingfacebenchmarklightweight
No one-command install · SourceDetails
IdentityExperimental

metapriseai-orgkernel

Use when every agent action must be traceable to a specific agent instance with its own key, scoped token and tamper-evident log.

98.52,699 stars · 247 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
CLIs & ToolsPreview

crystal

A full-fledged desktop application for orchestrating, monitoring, and interacting with Claude Code agents.

98.53,114 stars · 197 forks
Harnessclaude
clientcli
No one-command install · SourceDetails
CLIs & ToolsPreview

omnara

A command center for AI agents that syncs Claude Code sessions across terminal, web, and mobile. Allows for remote monitoring, human-in-the-loop interaction, and team collaboration.

98.52,780 stars · 213 forks · 4 mentions
Harnessclaude
claude-codealternative-clients
No one-command install · SourceDetails
CLIs & ToolsExperimental

primeintellect-ai-prime-rl

Agentic RL Training at Scale

Contributed by Sentinel

98.52,007 stars · 421 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsPreview

tweakcc

Command-line tool to customize your Claude Code styling.

98.42,476 stars · 199 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
IdentityExperimental

infisical-agent-vault

A HTTP credential proxy and vault for AI agents like Claude Code, OpenClaw, Hermes, custom agents + harnesses, and more.

Contributed by Sentinel

98.32,171 stars · 141 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-tools

Well-crafted toolset for session continuity, featuring skills/commands to avoid compaction and recover context across sessions with cross-agent handoff between Claude Code and Codex CLI. Includes a fast Rust/Tantivy-powered full-text session search (TUI for humans, skill/CLI for agents), tmux-cli skill + command for interacting with scripts and CLI agents, and safety hooks to block dangerous commands.

98.21,989 stars · 132 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsPreview

cc-sessions

An opinionated approach to productive development with Claude Code

98.11,551 stars · 191 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsPreview

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-ide-el

claude-code-ide.el integrates Claude Code with Emacs, like Anthropic’s VS Code/IntelliJ extensions. It shows ediff-based code suggestions, pulls LSP/flymake/flycheck diagnostics, and tracks buffer context. It adds an extensible MCP tool support for symbol refs/defs, project metadata, and tree-sitter AST queries.

98.01,658 stars · 112 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsPreview

rulesync

A Node.js CLI tool that automatically generates configs (rules, ignore files, MCP servers, commands, and subagents) for various AI coding agents. Rulesync can convert configs between Claude Code and other AI agents in both directions.

98.01,373 stars · 143 forks
Harnessclaude
toolingcli
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-nvim

A seamless integration between Claude Code AI assistant and Neovim.

97.92,096 stars · 70 forks
Harnessclaude
toolingcliide-integrations
No one-command install · SourceDetails
EvalsPreview

wandb-weave-evals

Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.

97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
No one-command install · SourceDetails
Browse · Armory