Armory
Source

Browse

Search and filter by type across the catalog

1,339 results in Skills, Evals, Memory, Hooks · page 2 of 56

MemoryExperimental

gibsonai-memori

Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems. Built for enterprise, Memori works with the data infrastructure you already run, no rip-and-replace, and deploys across managed cloud, single-tenant cloud, VPC, and on-premises.

Contributed by Sentinel

99.416,313 stars · 3,283 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

deepeval

Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.

99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
No one-command install · SourceDetails
EvalsPreview

ragas

Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.

99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
No one-command install · SourceDetails
SkillsExperimental

hardikpandya-stop-slop

A skill file for removing AI tells from prose

Contributed by Sentinel

99.316,882 stars · 1,223 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
MemoryExperimental

semantica-agi-semantica

Use when an agent's stored context needs provenance — you must be able to say where a remembered fact came from.

99.313,484 stars · 1,525 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

nevamind-ai-memu

Use when one person's memory should follow them across several different agents.

99.314,387 stars · 1,063 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
SkillsExperimental

nidhinjs-prompt-master

A Claude skill that writes the accurate prompts for any AI tool. Zero tokens or credits wasted. Full context and memory retention

Contributed by Sentinel

99.312,473 stars · 1,462 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
MemoryExperimental

evermind-ai-everos

Use when you want the agent's memory to be plain Markdown on your own disk rather than a hosted database.

99.212,757 stars · 911 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

memtensor-memos

Use when memory should persist across tasks and be reused, not just retrieved once per conversation.

99.211,599 stars · 1,060 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
SkillsPreview

fullstack-dev-skills

A Claude Code plugin with 65 skills for full-stack development across many frameworks, 9 workflow commands for Jira and Confluence, and a /common-ground command that lists Claude's assumptions about your project.

99.211,283 stars · 1,080 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
SkillsExperimental

agricidaniel-claude-ads

Claude-first paid-media operations skill for Claude Code across 12 ad platforms (Google, Meta, YouTube, LinkedIn, TikTok, Microsoft, Apple, Amazon, Reddit, Pinterest, Snapchat, X): source-grounded audits, deterministic scoring, versioned JSON reports, and capability-gated account changes.

Contributed by Sentinel

99.28,670 stars · 1,293 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

plastic-labs-honcho

Memory library for building stateful agents

Contributed by Sentinel

99.16,980 stars · 865 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
HooksPreview

plannotator

Interactive plan review UI that intercepts ExitPlanMode via hooks, letting users visually annotate plans with comments, deletions, and replacements before approving or denying with detailed feedback.

99.18,344 stars · 618 forks
Harnessclaude
hook
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
SkillsPreview

trail-of-bits-security-skills

A very professional collection of over a dozen security-focused skills for code auditing and vulnerability detection. Includes skills for static analysis with CodeQL and Semgrep, variant analysis across codebases, fix verification, and differential code review.

99.06,939 stars · 597 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
SkillsPreview

codebase-to-course

A Claude Code skill that turns any codebase into an interactive single-page HTML course for non-technical readers.

99.05,505 stars · 551 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
MemoryExperimental

getzep-zep

Examples, framework integrations and tools for Zep Cloud, Zep's hosted agent memory service; the repository says it is not the product itself. The open-source knowledge-graph engine behind Zep is Graphiti (getzep-graphiti).

Contributed by Sentinel

99.04,882 stars · 651 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
MemoryExperimental

caviraoss-openmemory

Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.

Contributed by Sentinel

98.94,478 stars · 504 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
MemoryExperimental

flowelement-m-flow

Use when recall should follow associations between memories rather than nearest-neighbour similarity alone.

98.84,497 stars · 255 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
Browse · Armory