1,394 results in Sub-Agents, CLAUDE.md / Rules, Hooks, Evals, Identity · page 1 of 59
Harness Claude Code Cursor Codex Gemini OpenCode karpathy-coding-discipline Drop into CLAUDE.md/AGENTS.md as the first behavior norm a coding agent ingrains: think before coding, prefer the simplest solution, change only what you own, and execute toward the stated goal.
99.9 209,417 stars · 21,315 forks
Harness claude codex cursor gemini opencode
discipline coding constitution behavior-norm
wshobson-agents Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity
Contributed by Sentinel
99.7 39,338 stars · 4,195 forks
Harness claude codex cursor gemini opencode
subagent
voltagent-awesome-claude-code-subagents A collection of 100+ specialized Claude Code subagents covering a wide range of development use cases
Contributed by Sentinel
99.5 24,793 stars · 2,870 forks
Harness claude codex cursor gemini opencode
subagent
CLAUDE.md / Rules Experimental humanlayer-12-factor-agents What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?
Contributed by Sentinel
99.5 25,642 stars · 1,951 forks · 4 mentions
Harness claude codex cursor gemini opencode
promptfoo CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.
99.5 24,737 stars · 2,255 forks
Harness claude cursor codex opencode gemini
evals red-teaming ci cli
openai-evals OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.
99.5 19,509 stars · 3,093 forks
Harness claude cursor codex opencode gemini
evals registry benchmark
lm-evaluation-harness EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.
99.5 13,860 stars · 3,533 forks
Harness claude cursor codex opencode gemini
evals academic benchmark harness
deepeval Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.
99.4 18,041 stars · 1,890 forks
Harness claude cursor codex opencode gemini
evals metrics rag ci
ragas Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.
99.4 15,853 stars · 1,727 forks · 1 mention · failed install test
Harness claude cursor codex opencode gemini
evals rag metrics
phoenix Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.
99.2 11,286 stars · 1,086 forks · 2 mentions
Harness claude cursor codex opencode gemini
evals observability rag agents
plannotator Interactive plan review UI that intercepts ExitPlanMode via hooks, letting users visually annotate plans with comments, deletions, and replacements before approving or denying with detailed feedback.
99.1 8,344 stars · 618 forks
Harness claude
hook
swe-bench SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
99.1 5,762 stars · 957 forks · 9 mentions
Harness claude cursor codex opencode gemini
evals code benchmark agents
giskard Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.0 5,838 stars · 542 forks
Harness claude cursor codex opencode gemini
evals safety vulnerability scan
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
big-bench Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.8 3,247 stars · 618 forks
Harness claude cursor codex opencode gemini
evals benchmark google academic
trulens Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.7 3,530 stars · 335 forks · 1 mention
Harness claude cursor codex opencode gemini
evals rag tracking dashboard
xlang-ai-osworld [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.7 3,117 stars · 530 forks · 7 mentions
Harness claude codex cursor gemini opencode
inspect-ai UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.7 2,683 stars · 685 forks
Harness claude cursor codex opencode gemini
evals safety aisi government
helm Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.7 2,898 stars · 412 forks
Harness claude cursor codex opencode gemini
evals benchmark academic stanford
lighteval Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.
98.6 2,533 stars · 553 forks
Harness claude cursor codex opencode gemini
evals huggingface benchmark lightweight
metapriseai-orgkernel Use when every agent action must be traceable to a specific agent instance with its own key, scoped token and tamper-evident log.
98.5 2,699 stars · 247 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
tdd-guard A hooks-driven system that monitors file operations in real-time and blocks changes that violate TDD principles.
98.4 2,324 stars · 185 forks
Harness claude
hook
evalplus Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.
98.3 1,819 stars · 208 forks
Harness claude cursor codex opencode gemini
evals code-generation humaneval benchmark
Page 1 of 59 Next
Browse · Armory