Armory
Source

Browse

Search and filter by type across the catalog

1,893 results in Evals, Sub-Agents, Skills 路 page 3 of 79

EvalsExperimental

stonybrooknlp-appworld

馃實 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars 路 78 forks 路 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
SkillsPreview

web-assets-generator-skill

Easily generate web assets from Claude Code including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn. Handles image resizing, text-to-image generation, emojis, and provides proper HTML meta tags.

96.4490 stars 路 52 forks
Harnessclaude
claude-codeagent-skills
No one-command install 路 SourceDetails
SkillsExperimental

superdesign-skill

The design skill for Claude Code, Cursor and any coding agent. Stop shipping AI-slop UI: turn it into shippable, tasteful frontend. Install: npx skills add superdesigndev/superdesign-skill. Powered by superdesign.dev

Contributed by Sentinel

96.2520 stars 路 38 forks
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install 路 SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars 路 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install 路 SourceDetails
SkillsPreview

cc-devops-skills

Skills for DevOps work that generate and validate infrastructure-as-code with shell scripts and CLI tools, for most deployment platforms. Also useful as documentation.

95.2303 stars 路 34 forks
Harnessclaude
claude-codeagent-skills
No one-command install 路 SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars 路 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install 路 SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars 路 62 forks 路 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install 路 SourceDetails
SkillsPreview

claude-code-agents

Comprehensive E2E development workflow with helpful Claude Code subagent prompts for solo devs. Run multiple auditors in parallel, automate fix cycles with micro-checkpoint protocols, and do browser-based QA. Includes strict protocols to prevent AI going rogue.

91.9147 stars 路 15 forks
Harnessclaude
skill
No one-command install 路 SourceDetails
SkillsPreview

book-factory

A comprehensive pipeline of Skills that replicates traditional publishing infrastructure for nonfiction book creation using specialized Claude skills.

91.6108 stars 路 20 forks
Harnessclaude
skill
No one-command install 路 SourceDetails
EvalsPreview

vellum-evals

Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.

90.582 stars 路 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
No one-command install 路 SourceDetails
EvalsPreview

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

82.327 stars 路 12 forks
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingsdk
No one-command install 路 SourceDetails
EvalsPreview

galileo-evaluate

Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.

80.322 stars 路 11 forks
Harnessclaudecursorcodexopencodegemini
evalshallucinationobservabilitysdk
No one-command install 路 SourceDetails
SkillsPreview

claude-mountaineering-skills

Claude Code skill that automates mountain route research for North American peaks. Aggregates data from 10+ mountaineering sources like Mountaineers.org, PeakBagger.com and SummitPost.com to generate detailed route beta reports with weather, avalanche conditions, and trip reports.

73.333 stars 路 1 fork
Harnessclaude
skill
No one-command install 路 SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars 路 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install 路 SourceDetails
SkillsPreview

read-only-postgres

Read-only PostgreSQL query skill for Claude Code. Executes SELECT/SHOW/EXPLAIN/WITH queries across configured databases with strict validation, timeouts, and row limits. Supports multiple connections with descriptions for database selection.

65.614 stars 路 1 fork
Harnessclaude
claude-codeagent-skills
No one-command install 路 SourceDetails
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install 路 SourceDetails
SkillsPreview

2d-games

2D game development principles. Sprites, tilemaps, physics, camera.

UnrankedNo signals yet
Harnessclaude
2d-gamesskills
Details
Sub-AgentsPreview

3d-artist

3D art and asset creation specialist for game development. Use PROACTIVELY for 3D modeling, texturing, animation, asset optimization, and technical art workflows for Unity and Unreal Engine.

UnrankedNo signals yet
Harnessclaude
game-developmentsubagents
Details
SkillsPreview

3d-games

3D game development principles. Rendering, shaders, physics, cameras.

UnrankedNo signals yet
Harnessclaude
3d-gamesskills
Details
SkillsPreview

3d-web-experience

Expert in building 3D experiences for the web - Three.js, React Three Fiber, Spline, WebGL, and interactive 3D scenes. Covers product configurators, 3D portfolios, immersive websites, and bringing depth to web experiences. Use when: 3D website, three.js, WebGL, react three fiber, 3D experience.

UnrankedNo signals yet
Harnessclaude
3d-web-experienceskills
Details
Sub-AgentsPreview

4-1-beast

An agent definition that sets GPT-4.1 up as a coding agent.

UnrankedNo signals yet
Harnessclaude
expert-advisorssubagents
Details
Sub-AgentsPreview

a11y-architect

Accessibility Architect specializing in WCAG 2.2 compliance for Web and Native platforms. Use PROACTIVELY when designing UI components, establishing design systems, or auditing code for inclusive user experiences.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
subagent
Details
Sub-AgentsPreview

ab-test-analysis

Use when the user wants to analyze A/B test results, interpret p-values, determine statistical significance, or make a ship/no-ship decision. Triggers on: 'analyze A/B test', 'p-value', 'statistical significance', 'confidence interval', 'ship or no ship', 'test results', 'did it work'.

UnrankedNo signals yet
Harnessclaude
research-analysissubagent
Details
SkillsPreview

ab-test-setup

When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," or "hypothesis." For tracking implementation, see analytics-tracking.

UnrankedNo signals yet
Harnessclaude
ab-test-setupskills
Details
Browse 路 Armory