Armory
Source

Browse

Search and filter by type across the catalog

2,027 results in Sub-Agents, Hooks, Evals, Skills 路 page 3 of 85

EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars 路 81 forks 路 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 馃洅WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars 路 107 forks 路 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
EvalsExperimental

stonybrooknlp-appworld

馃實 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars 路 78 forks 路 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
HooksPreview

claude-code-hook-comms-hcom

A lightweight CLI tool for real-time communication between Claude Code sub-agents through hooks, with @-mention targeting, a live monitoring dashboard and no dependencies. It was described as unstable when it was listed.

96.6470 stars 路 70 forks
Harnessclaude
claude-codehooks
No one-command install 路 SourceDetails
SkillsPreview

web-assets-generator-skill

Easily generate web assets from Claude Code including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn. Handles image resizing, text-to-image generation, emojis, and provides proper HTML meta tags.

96.4490 stars 路 52 forks
Harnessclaude
claude-codeagent-skills
No one-command install 路 SourceDetails
SkillsExperimental

superdesign-skill

The design skill for Claude Code, Cursor and any coding agent. Stop shipping AI-slop UI: turn it into shippable, tasteful frontend. Install: npx skills add superdesigndev/superdesign-skill. Powered by superdesign.dev

Contributed by Sentinel

96.2520 stars 路 38 forks
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install 路 SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars 路 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install 路 SourceDetails
SkillsPreview

cc-devops-skills

Skills for DevOps work that generate and validate infrastructure-as-code with shell scripts and CLI tools, for most deployment platforms. Also useful as documentation.

95.2303 stars 路 34 forks
Harnessclaude
claude-codeagent-skills
No one-command install 路 SourceDetails
HooksPreview

claude-hooks

A TypeScript-based system for configuring and customizing Claude Code hooks with a powerful and flexible interface.

95.2389 stars 路 26 forks
Harnessclaude
claude-codehooks
No one-command install 路 SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars 路 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install 路 SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars 路 62 forks 路 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install 路 SourceDetails
HooksPreview

dippy

Auto-approve safe bash commands using AST-based parsing, while prompting for destructive operations. Solves permission fatigue without disabling safety. Supports Claude Code, Gemini CLI, and Cursor.

94.0243 stars 路 21 forks
Harnessclaude
hook
No one-command install 路 SourceDetails
HooksPreview

cc-notify

CCNotify provides desktop notifications for Claude Code, alerting you to input needs or task completion, with one-click jumps back to VS Code and task duration display.

93.9216 stars 路 23 forks
Harnessclaude
claude-codehooks
No one-command install 路 SourceDetails
HooksPreview

typescript-quality-hooks

A quality-check hook for Node.js TypeScript projects: TypeScript compilation, ESLint auto-fixing and Prettier formatting, with SHA256 config caching that keeps validation under 5 ms during editing.

92.4178 stars 路 14 forks
Harnessclaude
claude-codehooks
No one-command install 路 SourceDetails
SkillsPreview

claude-code-agents

Comprehensive E2E development workflow with helpful Claude Code subagent prompts for solo devs. Run multiple auditors in parallel, automate fix cycles with micro-checkpoint protocols, and do browser-based QA. Includes strict protocols to prevent AI going rogue.

91.9147 stars 路 15 forks
Harnessclaude
skill
No one-command install 路 SourceDetails
SkillsPreview

book-factory

A comprehensive pipeline of Skills that replicates traditional publishing infrastructure for nonfiction book creation using specialized Claude skills.

91.6108 stars 路 20 forks
Harnessclaude
skill
No one-command install 路 SourceDetails
HooksPreview

cchooks

A lightweight Python SDK with a clean API and good documentation; simplifies the process of writing hooks and integrating them into your codebase, providing a nice abstraction over the JSON configuration files.

90.7130 stars 路 11 forks
Harnessclaude
hook
No one-command install 路 SourceDetails
EvalsPreview

vellum-evals

Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.

90.582 stars 路 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
No one-command install 路 SourceDetails
HooksPreview

claudio

A small library that plays OS-native sounds for Claude Code events through hooks.

89.2113 stars 路 8 forks
Harnessclaude
claude-codehooks
No one-command install 路 SourceDetails
HooksPreview

claude-code-hooks-sdk

A Laravel-inspired PHP SDK for building Claude Code hook responses with a clean, fluent API. This SDK makes it easy to create structured JSON responses for Claude Code hooks using an expressive, chainable interface.

87.068 stars 路 8 forks
Harnessclaude
hook
No one-command install 路 SourceDetails
EvalsPreview

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

82.327 stars 路 12 forks
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingsdk
No one-command install 路 SourceDetails
EvalsPreview

galileo-evaluate

Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.

80.322 stars 路 11 forks
Harnessclaudecursorcodexopencodegemini
evalshallucinationobservabilitysdk
No one-command install 路 SourceDetails
HooksPreview

parry

Prompt injection scanner for Claude Code hooks. Scans tool inputs and outputs for injection attacks, secrets, and data exfiltration attempts. In early development when it was listed.

75.445 stars 路 1 fork
Harnessclaude
hook
No one-command install 路 SourceDetails
HooksPreview

britfix

Claude outputs American spellings by default, which can have an impact on: professional credibility, compliance, documentation, and more. Britfix converts to British English, with a Claude Code hook for automatic conversion as files are written. Context-aware: handles code files intelligently by only converting comments and docstrings, never identifiers or string literals.

74.518 stars 路 4 forks
Harnessclaude
claude-codehooks
No one-command install 路 SourceDetails
Browse 路 Armory