Armory
Source

Browse

Search and filter by type across the catalog

761 results in Sub-Agents, Evals 路 page 2 of 32

EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars 路 81 forks 路 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 馃洅WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars 路 107 forks 路 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
EvalsExperimental

stonybrooknlp-appworld

馃實 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars 路 78 forks 路 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars 路 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install 路 SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars 路 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install 路 SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars 路 62 forks 路 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install 路 SourceDetails
EvalsPreview

vellum-evals

Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.

90.582 stars 路 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
No one-command install 路 SourceDetails
EvalsPreview

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

82.327 stars 路 12 forks
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingsdk
No one-command install 路 SourceDetails
EvalsPreview

galileo-evaluate

Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.

80.322 stars 路 11 forks
Harnessclaudecursorcodexopencodegemini
evalshallucinationobservabilitysdk
No one-command install 路 SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars 路 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install 路 SourceDetails
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install 路 SourceDetails
Sub-AgentsPreview

3d-artist

3D art and asset creation specialist for game development. Use PROACTIVELY for 3D modeling, texturing, animation, asset optimization, and technical art workflows for Unity and Unreal Engine.

UnrankedNo signals yet
Harnessclaude
game-developmentsubagents
Details
Sub-AgentsPreview

4-1-beast

An agent definition that sets GPT-4.1 up as a coding agent.

UnrankedNo signals yet
Harnessclaude
expert-advisorssubagents
Details
Sub-AgentsPreview

a11y-architect

Accessibility Architect specializing in WCAG 2.2 compliance for Web and Native platforms. Use PROACTIVELY when designing UI components, establishing design systems, or auditing code for inclusive user experiences.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
subagent
Details
Sub-AgentsPreview

ab-test-analysis

Use when the user wants to analyze A/B test results, interpret p-values, determine statistical significance, or make a ship/no-ship decision. Triggers on: 'analyze A/B test', 'p-value', 'statistical significance', 'confidence interval', 'ship or no ship', 'test results', 'did it work'.

UnrankedNo signals yet
Harnessclaude
research-analysissubagent
Details
Sub-AgentsPreview

academic-research-synthesizer

Academic research synthesis specialist. Use PROACTIVELY for comprehensive research on academic topics, literature reviews, technical investigations, and well-cited analysis combining multiple sources.

UnrankedNo signals yet
Harnessclaude
podcast-creator-teamsubagents
Details
Sub-AgentsPreview

academic-researcher

Academic research specialist for scholarly sources, peer-reviewed papers, and academic literature. Use PROACTIVELY for research paper analysis, literature reviews, citation tracking, and academic methodology evaluation.

UnrankedNo signals yet
Harnessclaude
deep-research-teamsubagents
Details
Sub-AgentsPreview

accessibility

Expert assistant for web accessibility (WCAG 2.1/2.2), inclusive UX, and a11y testing

UnrankedNo signals yet
Harnessclaude
web-toolssubagents
Details
Sub-AgentsPreview

accessibility-expert

Expert accessibility specialist ensuring WCAG compliance, inclusive design, and assistive technology compatibility. Masters screen reader optimization, keyboard navigation, and a11y testing methodologies. Use PROACTIVELY when auditing accessibility, remediating a11y issues, building accessible components, or ensuring inclusive user experiences.

UnrankedNo signals yet
Harnessclaude
agentssubagents
Details
Sub-AgentsPreview

accessibility-tester

Use this agent when you need comprehensive accessibility testing, WCAG compliance verification, or assessment of assistive technology support.

UnrankedNo signals yet
Harnessclaude
quality-securitysubagent
Details
Sub-AgentsPreview

accessibility-tester-2

Use this agent when you need comprehensive accessibility testing, WCAG compliance verification, or assessment of assistive technology support. Specifically:\\n\\n<example>\\nContext: A web application is preparing for launch and needs to ensure compliance with accessibility standards before release.\\nuser: \"Can you test our web app for accessibility issues? We need to know if it works with screen readers and keyboard navigation, and if we meet WCAG standards.\"\\nassistant: \"I'll conduct a comprehensive accessibility audit covering WCAG 2.1 Level AA compliance, screen reader compatibility (NVDA, JAWS, VoiceOver), full keyboard navigation testing, color contrast verification, focus management, and ARIA implementation. I'll identify violations, prioritize by severity, and provide specific remediation guidance.\"\\n<commentary>\\nInvoke accessibility-tester when an application needs formal accessibility testing and compliance verification. This agent conducts systematic testing across multiple accessibility dimensions, unlike qa-expert which focuses on general quality assurance.\\n</commentary>\\n</example>\\n\\n<example>\\nContext: Users with disabilities reported difficulty using the mobile app, and the team wants to understand and fix accessibility barriers.\\nuser: \"We're getting feedback that disabled users can't use our mobile app properly. Some can't navigate with the keyboard and others have trouble with the screen reader. What's wrong?\"\\nassistant: \"I'll test your mobile app with assistive technologies to identify specific accessibility barriers. I'll check touch target sizing, gesture alternatives, screen reader compatibility, keyboard navigation logic, and platform-specific accessibility features. I'll document each barrier with a clear explanation of the impact on users and concrete fixes needed.\"\\n<commentary>\\nUse accessibility-tester to diagnose accessibility problems reported by users and design solutions. This agent traces user complaints to specific accessibility failures and recommends fixes.\\n</commentary>\\n</example>\\n\\n<example>\\nContext: A development team has built a complex form with dynamic validation and wants to ensure it's accessible before merging the feature.\\nuser: \"We built a new checkout form with real-time validation. Can you check if it's accessible? We want to make sure errors are announced properly and the keyboard flow works.\"\\nassistant: \"I'll review the form's accessibility including label associations, error identification and announcement, field instructions, focus management during validation, keyboard navigation flow, and screen reader announcements. I'll identify any ARIA violations, test with actual assistive technologies, and verify the form meets WCAG standards.\"\\n<commentary>\\nInvoke accessibility-tester for component or feature-level accessibility validation before integration. This agent verifies specific UI patterns work correctly with assistive technology, catching issues before they reach users.\\n</commentary>\\n</example>

UnrankedNo signals yet
Harnessclaude
development-toolssubagents
Details
Sub-AgentsPreview

ad-security-reviewer

Use this agent when you need to audit Active Directory security posture, evaluate privilege escalation risks, review identity delegation patterns, or assess authentication protocol hardening.

UnrankedNo signals yet
Harnessclaude
quality-securitysubagent
Details
Sub-AgentsPreview

address-comments

Address PR comments

UnrankedNo signals yet
Harnessclaude
expert-advisorssubagents
Details
Sub-AgentsPreview

adr-generator

Expert agent for creating comprehensive Architectural Decision Records (ADRs) with structured formatting optimized for AI consumption and human readability.

UnrankedNo signals yet
Harnessclaude
data-aisubagents
Details
Browse 路 Armory