Armory
Source

Browse

Search and filter by type across the catalog

1,260 results in Sub-Agents, Evals, Identity, CLAUDE.md / Rules · page 1 of 53

CLAUDE.md / RulesStable

karpathy-coding-discipline

Drop into CLAUDE.md/AGENTS.md as the first behavior norm a coding agent ingrains: think before coding, prefer the simplest solution, change only what you own, and execute toward the stated goal.

99.9209,417 stars · 21,315 forks
Harnessclaudecodexcursorgeminiopencode
disciplinecodingconstitutionbehavior-norm
Details
Sub-AgentsExperimental

wshobson-agents

Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity

Contributed by Sentinel

99.739,338 stars · 4,195 forks
Harnessclaudecodexcursorgeminiopencode
subagent
No one-command install · SourceDetails
Sub-AgentsExperimental

voltagent-awesome-claude-code-subagents

A collection of 100+ specialized Claude Code subagents covering a wide range of development use cases

Contributed by Sentinel

99.524,793 stars · 2,870 forks
Harnessclaudecodexcursorgeminiopencode
subagent
No one-command install · SourceDetails
CLAUDE.md / RulesExperimental

humanlayer-12-factor-agents

What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?

Contributed by Sentinel

99.525,642 stars · 1,951 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

promptfoo

CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.

99.524,737 stars · 2,255 forks
Harnessclaudecursorcodexopencodegemini
evalsred-teamingcicli
No one-command install · SourceDetails
EvalsPreview

openai-evals

OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.

99.519,509 stars · 3,093 forks
Harnessclaudecursorcodexopencodegemini
evalsregistrybenchmark
No one-command install · SourceDetails
EvalsPreview

lm-evaluation-harness

EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.

99.513,860 stars · 3,533 forks
Harnessclaudecursorcodexopencodegemini
evalsacademicbenchmarkharness
No one-command install · SourceDetails
EvalsPreview

deepeval

Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.

99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
No one-command install · SourceDetails
EvalsPreview

ragas

Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.

99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
No one-command install · SourceDetails
EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
EvalsPreview

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
No one-command install · SourceDetails
EvalsPreview

lighteval

Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.

98.62,533 stars · 553 forks
Harnessclaudecursorcodexopencodegemini
evalshuggingfacebenchmarklightweight
No one-command install · SourceDetails
IdentityExperimental

metapriseai-orgkernel

Use when every agent action must be traceable to a specific agent instance with its own key, scoped token and tamper-evident log.

98.52,699 stars · 247 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
IdentityExperimental

infisical-agent-vault

A HTTP credential proxy and vault for AI agents like Claude Code, OpenClaw, Hermes, custom agents + harnesses, and more.

Contributed by Sentinel

98.32,171 stars · 141 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
Browse · Armory