Armory
Source

Browse

Search and filter by type across the catalog

702 results in Memory, Evals, CLAUDE.md / Rules, Observability, Hooks · page 1 of 30

CLAUDE.md / RulesStable

karpathy-coding-discipline

Drop into CLAUDE.md/AGENTS.md as the first behavior norm a coding agent ingrains: think before coding, prefer the simplest solution, change only what you own, and execute toward the stated goal.

99.9209,417 stars · 21,315 forks
Harnessclaudecodexcursorgeminiopencode
disciplinecodingconstitutionbehavior-norm
Details
MemoryExperimental

thedotmack-claude-mem

Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

Contributed by Sentinel

99.894,727 stars · 8,379 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
cli
No one-command install · SourceDetails
MemoryExperimental

mempalace

Use when you want an agent memory system whose recall quality has actually been benchmarked rather than asserted.

99.858,897 stars · 7,548 forks · 2 mentions
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
ObservabilityPreview

sentry-llm-monitoring

Sentry's error and performance monitoring extended to LLM applications. It captures exceptions, latency, and AI token usage with OpenTelemetry integration.

99.744,709 stars · 4,832 forks · 15 mentions
Harnessclaudecursorcodexopencodegemini
observabilityerrorsapm
No one-command install · SourceDetails
MemoryExperimental

microsoft-graphrag

A modular graph-based Retrieval-Augmented Generation (RAG) system

Contributed by Sentinel

99.635,783 stars · 3,754 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
ObservabilityPreview

mlflow-tracing

MLflow's LLM tracing module instruments model calls, agent steps, and tool invocations, storing them alongside experiment runs for reproducibility.

99.627,768 stars · 6,246 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingexperiment-tracking
No one-command install · SourceDetails
ObservabilityPreview

langfuse

Open-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications.

99.634,067 stars · 3,678 forks · 10 mentions
Harnessclaudecursorcodexopencodegemini
observabilitytracingevals
No one-command install · SourceDetails
MemoryExperimental

volcengine-openviking

Use when an agent's memory, retrieved knowledge and learned skills should live in one store that reorganises itself.

99.635,893 stars · 2,739 forks · 2 mentions
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

garrytan-gbrain

Garry's Opinionated OpenClaw/Hermes Agent Brain

Contributed by Sentinel

99.629,460 stars · 4,394 forks · 16 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

getzep-graphiti

Build Real-Time Knowledge Graphs for AI Agents

Contributed by Sentinel

99.630,509 stars · 3,098 forks · 1 mention
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
MemoryExperimental

topoteretes-cognee

Use when an agent needs long-term memory backed by a knowledge graph you can host yourself.

99.630,556 stars · 3,010 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

supermemoryai-supermemory

Use when memory has to be fast, run locally, and be reachable from an app as well as an agent.

99.629,256 stars · 2,557 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

tencentcloud-tencentdb-agent-memory

TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and code into four reusable memory assets (Chat Memory, Skill, LLM-Wiki, Code-Graph) that are governed, shared, and equipped across agents and frameworks.

Contributed by Sentinel

99.525,635 stars · 2,398 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

letta-ai-letta

Platform for stateful agents: AI with advanced memory that can learn and self-improve over time.

Contributed by Sentinel

99.524,552 stars · 2,609 forks · 19 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
CLAUDE.md / RulesExperimental

humanlayer-12-factor-agents

What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?

Contributed by Sentinel

99.525,642 stars · 1,951 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

promptfoo

CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.

99.524,737 stars · 2,255 forks
Harnessclaudecursorcodexopencodegemini
evalsred-teamingcicli
No one-command install · SourceDetails
ObservabilityPreview

claude-hud

A status line for Claude Code that shows context usage, tools, agents, to-dos and more. Highly configurable, and maintained when it was listed.

99.527,778 stars · 1,286 forks
Harnessclaude
statuslineobservability
No one-command install · SourceDetails
EvalsPreview

openai-evals

OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.

99.519,509 stars · 3,093 forks
Harnessclaudecursorcodexopencodegemini
evalsregistrybenchmark
No one-command install · SourceDetails
EvalsPreview

lm-evaluation-harness

EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.

99.513,860 stars · 3,533 forks
Harnessclaudecursorcodexopencodegemini
evalsacademicbenchmarkharness
No one-command install · SourceDetails
MemoryExperimental

gibsonai-memori

Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems. Built for enterprise, Memori works with the data infrastructure you already run, no rip-and-replace, and deploys across managed cloud, single-tenant cloud, VPC, and on-premises.

Contributed by Sentinel

99.416,313 stars · 3,283 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
ObservabilityPreview

comet-opik

Opik by Comet is an open-source LLM evaluation and tracing platform: log traces, run automated evals, create datasets, and track prompt improvements over time.

99.422,249 stars · 1,827 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
observabilityevalstracing
No one-command install · SourceDetails
ObservabilityPreview

opentelemetry-genai

OpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group.

99.44,897 stars · 3,851 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
EvalsPreview

deepeval

Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.

99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
No one-command install · SourceDetails
EvalsPreview

ragas

Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.

99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
No one-command install · SourceDetails
Browse · Armory