Armory
Source

Browse

Search and filter by type across the catalog

2,099 results in Observability, Evals, Hooks, Sub-Agents, Identity, Skills · page 1 of 88

SkillsPreview

superpowers

A bundle of skills for software engineering that covers much of the development life cycle: planning, reviewing, testing and debugging. Many consolidate standard engineering practice.

99.9280,507 stars · 25,127 forks · 1 mention
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
SkillsPreview

everything-claude-code

A large collection of resources for Claude Code across core engineering domains. Most resources stand alone, and the author's own workflow is optional.

99.9245,823 stars · 37,094 forks · 1 mention
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
SkillsExperimental

mattpocock-skills

Skills for Real Engineers. Straight from my .agents directory.

Contributed by Sentinel

99.9244,329 stars · 20,767 forks · 5 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
SkillsExperimental

tt-a1i-archify

Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.

Contributed by Sentinel

99.972,120 stars · 4,865 forks · 1 mention · passed install test
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
SkillsExperimental

anthropics-skills

Public repository for Agent Skills

Contributed by Sentinel

99.9178,534 stars · 21,125 forks
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
SkillsExperimental

garrytan-gstack

Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA

Contributed by Sentinel

99.9131,046 stars · 19,660 forks · 11 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
SkillsExperimental

nextlevelbuilder-ui-ux-pro-max-skill

An AI skill that provides design intelligence for building professional UI/UX across multiple platforms.

Contributed by Sentinel

99.9125,818 stars · 13,457 forks
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
ObservabilityPreview

sentry-llm-monitoring

Sentry's error and performance monitoring extended to LLM applications. It captures exceptions, latency, and AI token usage with OpenTelemetry integration.

99.744,709 stars · 4,832 forks · 15 mentions
Harnessclaudecursorcodexopencodegemini
observabilityerrorsapm
No one-command install · SourceDetails
SkillsPreview

claude-scientific-skills

A set of ready-to-use Agent Skills for research, science, engineering, analysis, finance and writing.

99.741,662 stars · 3,831 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
Sub-AgentsExperimental

wshobson-agents

Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity

Contributed by Sentinel

99.739,338 stars · 4,195 forks
Harnessclaudecodexcursorgeminiopencode
subagent
No one-command install · SourceDetails
ObservabilityPreview

mlflow-tracing

MLflow's LLM tracing module instruments model calls, agent steps, and tool invocations, storing them alongside experiment runs for reproducibility.

99.627,768 stars · 6,246 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingexperiment-tracking
No one-command install · SourceDetails
ObservabilityPreview

langfuse

Open-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications.

99.634,067 stars · 3,678 forks · 10 mentions
Harnessclaudecursorcodexopencodegemini
observabilitytracingevals
No one-command install · SourceDetails
SkillsExperimental

virgiliojr94-book-to-skill

Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.

Contributed by Sentinel

99.627,865 stars · 2,878 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
SkillsExperimental

alirezarezvani-claude-skills

380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts) for Claude Code, Codex, Gemini CLI, Cursor, and 8 more coding agents — engineering, marketing, product, compliance, C-level advisory, research, business operations, commercial & finance, and your daily productivity skills.

Contributed by Sentinel

99.525,396 stars · 3,596 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
Sub-AgentsExperimental

voltagent-awesome-claude-code-subagents

A collection of 100+ specialized Claude Code subagents covering a wide range of development use cases

Contributed by Sentinel

99.524,793 stars · 2,870 forks
Harnessclaudecodexcursorgeminiopencode
subagent
No one-command install · SourceDetails
EvalsPreview

promptfoo

CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.

99.524,737 stars · 2,255 forks
Harnessclaudecursorcodexopencodegemini
evalsred-teamingcicli
No one-command install · SourceDetails
SkillsPreview

compound-engineering-plugin

A very pragmatic set of well-designed agents, skills, and commands, built around a discipline of turning past mistakes and errors into lessons and opportunities for future growth and improvement. Good documentation.

99.524,761 stars · 2,047 forks
Harnessclaude
skill
No one-command install · SourceDetails
ObservabilityPreview

claude-hud

A status line for Claude Code that shows context usage, tools, agents, to-dos and more. Highly configurable, and maintained when it was listed.

99.527,778 stars · 1,286 forks
Harnessclaude
statuslineobservability
No one-command install · SourceDetails
EvalsPreview

openai-evals

OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.

99.519,509 stars · 3,093 forks
Harnessclaudecursorcodexopencodegemini
evalsregistrybenchmark
No one-command install · SourceDetails
EvalsPreview

lm-evaluation-harness

EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.

99.513,860 stars · 3,533 forks
Harnessclaudecursorcodexopencodegemini
evalsacademicbenchmarkharness
No one-command install · SourceDetails
ObservabilityPreview

comet-opik

Opik by Comet is an open-source LLM evaluation and tracing platform: log traces, run automated evals, create datasets, and track prompt improvements over time.

99.422,249 stars · 1,827 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
observabilityevalstracing
No one-command install · SourceDetails
ObservabilityPreview

opentelemetry-genai

OpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group.

99.44,897 stars · 3,851 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
EvalsPreview

deepeval

Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.

99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
No one-command install · SourceDetails
EvalsPreview

ragas

Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.

99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
No one-command install · SourceDetails
Browse · Armory