Armory
Source

Browse

Search and filter by type across the catalog

1,239 results in Observability, Evals, Memory, Skills · page 4 of 52

MemoryExperimental

zilliztech-memsearch

Use when several coding agents should share one memory store instead of each keeping its own notes.

98.52,571 stars · 238 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

kingjulio8238-memary

Use when you want a well-known reference implementation of agent memory over a knowledge graph to read or fork.

98.52,644 stars · 205 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

redplanethq-core

Use when one memory graph should serve Claude Code, Codex and your other assistants at once.

98.31,963 stars · 189 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
MemoryExperimental

langchain-ai-langmem

Long-term memory for agents: tools that extract what matters from conversations, refine prompts from feedback and keep memory across sessions, with LangGraph's store built in.

Contributed by Sentinel

98.21,684 stars · 192 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

bai-lab-memoryos

Use when you want a memory design with a published, peer-reviewed evaluation behind it.

98.11,570 stars · 161 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

cortexkit-magic-context

Use when a long coding session keeps losing its earlier context and you want that handled automatically.

98.12,044 stars · 106 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
ObservabilityPreview

openinference

OpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI.

98.11,192 stars · 302 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
SkillsPreview

context-engineering-kit

Hand-crafted collection of advanced context engineering techniques and patterns with minimal token footprint focused on improving agent result quality.

98.01,508 stars · 154 forks
Harnessclaude
skill
No one-command install · SourceDetails
MemoryExperimental

dataojitori-nocturne-memory

Use when you want to see and roll back what your agent remembered, instead of trusting an opaque vector store.

98.01,373 stars · 170 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
ObservabilityPreview

langsmith

LangChain's platform for tracing, evaluating, and monitoring LLM applications: deep integration with LangChain/LangGraph plus a REST API for any stack.

98.01,043 stars · 288 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingevals
No one-command install · SourceDetails
MemoryExperimental

claudiodrews-memory-os

Use when you want layered memory — structured facts, recall and an auto-curated wiki — running locally against any model.

97.91,355 stars · 128 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

wandb-weave-evals

Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.

97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
No one-command install · SourceDetails
SkillsPreview

codex-skill

Enables users to prompt codex from claude code. Unlike the raw codex mcp server, this skill infers parameters such as model, reasoning effort, sandboxing from your prompt or asks you to specify them. It also simplifies continuing prior codex sessions so that codex can continue with the prior context.

97.91,424 stars · 109 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
EvalsPreview

langtrace

Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.

97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
No one-command install · SourceDetails
MemoryExperimental

agiresearch-a-mem

A-MEM: Agentic Memory for LLM Agents

Contributed by Sentinel

97.81,164 stars · 121 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
SkillsPreview

claude-codex-settings

A set of plugins for core developer tasks, covering GitHub, Azure, MongoDB, Tavily, Playwright and more. Also works with a few other providers.

97.71,117 stars · 107 forks
Harnessclaude
skill
No one-command install · SourceDetails
SkillsPreview

agentsys

Workflow automation system for Claude with a group of useful plugins, agents, and skills. Automates task-to-production workflows, PR management, code cleanup, performance investigation, drift detection, and multi-agent code review. Includes agnix for linting agent configurations. Built on thousands of lines of code with thousands of tests. Uses deterministic detection (regex, AST) with LLM judgment for efficiency. Used on many production systems.

97.6981 stars · 113 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
ObservabilityPreview

claude-powerline

A vim-style powerline statusline for Claude Code with real-time usage tracking, git integration, custom themes, and more

97.61,163 stars · 82 forks
Harnessclaude
claude-codestatus-lines
No one-command install · SourceDetails
EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

codeabra-iai-personal-memory-engine

Use when the agent should remember not just facts but how you like to work, locally and for free.

97.5862 stars · 105 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
Browse · Armory