Armory
Source

Browse

Search and filter by type across the catalog

826 results in Evals, Workflows, Observability, Memory · page 4 of 35

MemoryExperimental

zilliztech-memsearch

Use when several coding agents should share one memory store instead of each keeping its own notes.

98.52,571 stars · 238 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

kingjulio8238-memary

Use when you want a well-known reference implementation of agent memory over a knowledge graph to read or fork.

98.52,644 stars · 205 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
WorkflowsPreview

claude-codepro

A development environment for Claude Code with a spec-driven workflow, TDD enforcement, cross-session memory, semantic search, quality hooks and modular rules. Large, with wide coverage.

98.32,063 stars · 176 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
MemoryExperimental

redplanethq-core

Use when one memory graph should serve Claude Code, Codex and your other assistants at once.

98.31,963 stars · 189 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
MemoryExperimental

langchain-ai-langmem

Long-term memory for agents: tools that extract what matters from conversations, refine prompts from feedback and keep memory across sessions, with LangGraph's store built in.

Contributed by Sentinel

98.21,684 stars · 192 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

bai-lab-memoryos

Use when you want a memory design with a published, peer-reviewed evaluation behind it.

98.11,570 stars · 161 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

cortexkit-magic-context

Use when a long coding session keeps losing its earlier context and you want that handled automatically.

98.12,044 stars · 106 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
ObservabilityPreview

openinference

OpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI.

98.11,192 stars · 302 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
MemoryExperimental

dataojitori-nocturne-memory

Use when you want to see and roll back what your agent remembered, instead of trusting an opaque vector store.

98.01,373 stars · 170 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
ObservabilityPreview

langsmith

LangChain's platform for tracing, evaluating, and monitoring LLM applications: deep integration with LangChain/LangGraph plus a REST API for any stack.

98.01,043 stars · 288 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingevals
No one-command install · SourceDetails
WorkflowsPreview

the-ralph-playbook

A detailed guide to the Ralph Wiggum technique for autonomous coding loops, with the reasoning behind it and practical guidelines.

97.91,029 stars · 266 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
MemoryExperimental

claudiodrews-memory-os

Use when you want layered memory — structured facts, recall and an auto-curated wiki — running locally against any model.

97.91,355 stars · 128 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

wandb-weave-evals

Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.

97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
No one-command install · SourceDetails
EvalsPreview

langtrace

Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.

97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
No one-command install · SourceDetails
MemoryExperimental

agiresearch-a-mem

A-MEM: Agentic Memory for LLM Agents

Contributed by Sentinel

97.81,164 stars · 121 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
WorkflowsPreview

claude-code-documentation-mirror

A mirror of Anthropic's documentation pages for Claude Code, updated every few hours.

97.7983 stars · 138 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
ObservabilityPreview

claude-powerline

A vim-style powerline statusline for Claude Code with real-time usage tracking, git integration, custom themes, and more

97.61,163 stars · 82 forks
Harnessclaude
claude-codestatus-lines
No one-command install · SourceDetails
EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

codeabra-iai-personal-memory-engine

Use when the agent should remember not just facts but how you like to work, locally and for free.

97.5862 stars · 105 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
WorkflowsPreview

awesome-ralph

A curated list of resources about Ralph, the AI coding technique that runs AI coding agents in automated loops until specifications are fulfilled.

97.3918 stars · 74 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
Browse · Armory