1,222 results in CLAUDE.md / Rules, Sub-Agents, Evals · page 2 of 51
Harness Claude Code Cursor Codex Gemini OpenCode wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
langtrace Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.8 1,228 stars · 127 forks
Harness claude cursor codex opencode gemini
evals observability opentelemetry tracing
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
braintrust Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
82.3 27 stars · 12 forks
Harness claude cursor codex opencode gemini
evals experiment-tracking sdk
galileo-evaluate Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.
80.3 22 stars · 11 forks
Harness claude cursor codex opencode gemini
evals hallucination observability sdk
humanloop-evals Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.
69.0 12 stars · 3 forks
Harness claude cursor codex opencode gemini
evals human-eval dataset sdk
pre-commit-hooks A repository of pre-commit hooks whose CLAUDE.md and .claude/ documentation is a thorough, compact example of project instructions for Claude Code.
57.3 5 stars · 3 forks
Harness claude
claude-code claude-md-files
angular-coding-style angular rule: apply when working on angular and you need Angular Coding Style.
16.3 1 mention
Harness claude codex cursor gemini opencode
rules angular
agentbench Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.
Unranked No signals yet
Harness claude codex
eval benchmark scoring harness
3d-artist 3D art and asset creation specialist for game development. Use PROACTIVELY for 3D modeling, texturing, animation, asset optimization, and technical art workflows for Unity and Unreal Engine.
Unranked No signals yet
Harness claude
game-development subagents
4-1-beast An agent definition that sets GPT-4.1 up as a coding agent.
Unranked No signals yet
Harness claude
expert-advisors subagents
a11y-architect Accessibility Architect specializing in WCAG 2.2 compliance for Web and Native platforms. Use PROACTIVELY when designing UI components, establishing design systems, or auditing code for inclusive user experiences.
Unranked No signals yet
Harness claude codex cursor gemini opencode
subagent
ab-test-analysis Use when the user wants to analyze A/B test results, interpret p-values, determine statistical significance, or make a ship/no-ship decision. Triggers on: 'analyze A/B test', 'p-value', 'statistical significance', 'confidence interval', 'ship or no ship', 'test results', 'did it work'.
Unranked No signals yet
Harness claude
research-analysis subagent
academic-research-synthesizer Academic research synthesis specialist. Use PROACTIVELY for comprehensive research on academic topics, literature reviews, technical investigations, and well-cited analysis combining multiple sources.
Unranked No signals yet
Harness claude
podcast-creator-team subagents
academic-researcher Academic research specialist for scholarly sources, peer-reviewed papers, and academic literature. Use PROACTIVELY for research paper analysis, literature reviews, citation tracking, and academic methodology evaluation.
Unranked No signals yet
Harness claude
deep-research-team subagents
accessibility Expert assistant for web accessibility (WCAG 2.1/2.2), inclusive UX, and a11y testing
Unranked No signals yet
Harness claude
web-tools subagents
accessibility-expert Expert accessibility specialist ensuring WCAG compliance, inclusive design, and assistive technology compatibility. Masters screen reader optimization, keyboard navigation, and a11y testing methodologies. Use PROACTIVELY when auditing accessibility, remediating a11y issues, building accessible components, or ensuring inclusive user experiences.
Unranked No signals yet
Harness claude
agents subagents
accessibility-tester Use this agent when you need comprehensive accessibility testing, WCAG compliance verification, or assessment of assistive technology support.
Unranked No signals yet
Harness claude
quality-security subagent
Previous Page 2 of 51 Next
Browse · Armory