Armory
Source

Browse

Search and filter by type across the catalog

1,239 results in Observability, Evals, Skills, Memory · page 2 of 52

SkillsPreview

compound-engineering-plugin

A very pragmatic set of well-designed agents, skills, and commands, built around a discipline of turning past mistakes and errors into lessons and opportunities for future growth and improvement. Good documentation.

99.524,761 stars · 2,047 forks
Harnessclaude
skill
No one-command install · SourceDetails
ObservabilityPreview

claude-hud

A status line for Claude Code that shows context usage, tools, agents, to-dos and more. Highly configurable, and maintained when it was listed.

99.527,778 stars · 1,286 forks
Harnessclaude
statuslineobservability
No one-command install · SourceDetails
EvalsPreview

openai-evals

OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.

99.519,509 stars · 3,093 forks
Harnessclaudecursorcodexopencodegemini
evalsregistrybenchmark
No one-command install · SourceDetails
EvalsPreview

lm-evaluation-harness

EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.

99.513,860 stars · 3,533 forks
Harnessclaudecursorcodexopencodegemini
evalsacademicbenchmarkharness
No one-command install · SourceDetails
MemoryExperimental

gibsonai-memori

Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems. Built for enterprise, Memori works with the data infrastructure you already run, no rip-and-replace, and deploys across managed cloud, single-tenant cloud, VPC, and on-premises.

Contributed by Sentinel

99.416,313 stars · 3,283 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
ObservabilityPreview

comet-opik

Opik by Comet is an open-source LLM evaluation and tracing platform: log traces, run automated evals, create datasets, and track prompt improvements over time.

99.422,249 stars · 1,827 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
observabilityevalstracing
No one-command install · SourceDetails
ObservabilityPreview

opentelemetry-genai

OpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group.

99.44,897 stars · 3,851 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
EvalsPreview

deepeval

Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.

99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
No one-command install · SourceDetails
EvalsPreview

ragas

Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.

99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
No one-command install · SourceDetails
SkillsExperimental

hardikpandya-stop-slop

A skill file for removing AI tells from prose

Contributed by Sentinel

99.316,882 stars · 1,223 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
MemoryExperimental

semantica-agi-semantica

Use when an agent's stored context needs provenance — you must be able to say where a remembered fact came from.

99.313,484 stars · 1,525 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

nevamind-ai-memu

Use when one person's memory should follow them across several different agents.

99.314,387 stars · 1,063 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
SkillsExperimental

nidhinjs-prompt-master

A Claude skill that writes the accurate prompts for any AI tool. Zero tokens or credits wasted. Full context and memory retention

Contributed by Sentinel

99.312,473 stars · 1,462 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
ObservabilityPreview

portkey-ai-gateway

Open-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load balancing.

99.312,873 stars · 1,280 forks
Harnessclaudecursorcodexopencodegemini
observabilitygatewayproxy
No one-command install · SourceDetails
MemoryExperimental

evermind-ai-everos

Use when you want the agent's memory to be plain Markdown on your own disk rather than a hosted database.

99.212,757 stars · 911 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

memtensor-memos

Use when memory should persist across tasks and be reused, not just retrieved once per conversation.

99.211,599 stars · 1,060 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
SkillsPreview

fullstack-dev-skills

A Claude Code plugin with 65 skills for full-stack development across many frameworks, 9 workflow commands for Jira and Confluence, and a /common-ground command that lists Claude's assumptions about your project.

99.211,283 stars · 1,080 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
ObservabilityPreview

ccstatusline

A highly customizable status line formatter for Claude Code CLI that displays model info, git branch, token usage, and other metrics in your terminal.

99.213,035 stars · 579 forks
Harnessclaude
statuslineobservability
No one-command install · SourceDetails
SkillsExperimental

agricidaniel-claude-ads

Claude-first paid-media operations skill for Claude Code across 12 ad platforms (Google, Meta, YouTube, LinkedIn, TikTok, Microsoft, Apple, Amazon, Reddit, Pinterest, Snapchat, X): source-grounded audits, deterministic scoring, versioned JSON reports, and capability-gated account changes.

Contributed by Sentinel

99.28,670 stars · 1,293 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
ObservabilityPreview

openllmetry

OpenTelemetry-based observability for LLM applications. It auto-instruments OpenAI, Anthropic, LangChain, and 20+ providers with zero code changes.

99.17,452 stars · 1,099 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
MemoryExperimental

plastic-labs-honcho

Memory library for building stateful agents

Contributed by Sentinel

99.16,980 stars · 865 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
SkillsPreview

trail-of-bits-security-skills

A very professional collection of over a dozen security-focused skills for code auditing and vulnerability detection. Includes skills for static analysis with CodeQL and Semgrep, variant analysis across codebases, fix verification, and differential code review.

99.06,939 stars · 597 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
Browse · Armory