Armory
Source

Browse

Search and filter by type across the catalog

835 results in Sub-Agents, Identity, Evals, Memory · page 3 of 35

EvalsPreview

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

bai-lab-memoryos

Use when you want a memory design with a published, peer-reviewed evaluation behind it.

98.11,570 stars · 161 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

cortexkit-magic-context

Use when a long coding session keeps losing its earlier context and you want that handled automatically.

98.12,044 stars · 106 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

dataojitori-nocturne-memory

Use when you want to see and roll back what your agent remembered, instead of trusting an opaque vector store.

98.01,373 stars · 170 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

claudiodrews-memory-os

Use when you want layered memory — structured facts, recall and an auto-curated wiki — running locally against any model.

97.91,355 stars · 128 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

wandb-weave-evals

Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.

97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
No one-command install · SourceDetails
EvalsPreview

langtrace

Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.

97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
No one-command install · SourceDetails
MemoryExperimental

agiresearch-a-mem

A-MEM: Agentic Memory for LLM Agents

Contributed by Sentinel

97.81,164 stars · 121 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
IdentityExperimental

letta-ai-agent-file

Use when you need to save, share, version or move a whole agent — its persona, memory and behaviour — as one portable file.

97.81,197 stars · 114 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

codeabra-iai-personal-memory-engine

Use when the agent should remember not just facts but how you like to work, locally and for free.

97.5862 stars · 105 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars · 107 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
IdentityExperimental

unicity-aos-capsule-identity

Use when an agent's identity must be persisted as state and assembled into its system prompt at boot, rather than pasted into a prompt by hand.

96.88,503 stars · 19 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsExperimental

stonybrooknlp-appworld

🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars · 78 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars · 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install · SourceDetails
IdentityExperimental

agntcy-oasf

Use when agents must describe themselves to other systems in a common schema so they can be catalogued, discovered and verified.

95.7332 stars · 47 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

telagod-code-abyss

Use when you want a coding agent to have a consistent, composable personality and voice across Claude Code, Codex, Gemini CLI and OpenClaw.

94.6239 stars · 32 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

agntcy-dir

Use when agents and multi-agent systems need to announce themselves and be found across organisations rather than hardcoded.

94.6184 stars · 55 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars · 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars · 62 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install · SourceDetails
IdentityExperimental

character-card-spec-v2

Use when reading or writing the character-card files that the roleplay-agent ecosystem actually ships, including the PNG-embedded variant.

93.9188 stars · 29 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

cloudflare-web-bot-auth

Use when a website has to be able to tell that a request really came from your agent and not from someone impersonating it.

93.8157 stars · 40 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
MemoryExperimental

nemori-ai-nemori

Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.

93.5207 stars · 20 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
Browse · Armory