Armory
Source

Browse

Search and filter by type across the catalog

931 results in Evals, Hooks, Memory, Sub-Agents · page 2 of 39

HooksPreview

plannotator

Interactive plan review UI that intercepts ExitPlanMode via hooks, letting users visually annotate plans with comments, deletions, and replacements before approving or denying with detailed feedback.

99.18,344 stars · 618 forks
Harnessclaude
hook
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
MemoryExperimental

getzep-zep

Examples, framework integrations and tools for Zep Cloud, Zep's hosted agent memory service; the repository says it is not the product itself. The open-source knowledge-graph engine behind Zep is Graphiti (getzep-graphiti).

Contributed by Sentinel

99.04,882 stars · 651 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
MemoryExperimental

caviraoss-openmemory

Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.

Contributed by Sentinel

98.94,478 stars · 504 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
MemoryExperimental

flowelement-m-flow

Use when recall should follow associations between memories rather than nearest-neighbour similarity alone.

98.84,497 stars · 255 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
MemoryExperimental

mirix-ai-mirix

Use when what the agent should remember is what actually happened on screen, consolidated into structured memories.

98.73,440 stars · 270 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
No one-command install · SourceDetails
EvalsPreview

lighteval

Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.

98.62,533 stars · 553 forks
Harnessclaudecursorcodexopencodegemini
evalshuggingfacebenchmarklightweight
No one-command install · SourceDetails
MemoryExperimental

kayba-ai-agentic-context-engine

Use when an agent should carry forward what it learned from its own successes and failures into later runs.

98.52,564 stars · 307 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

memodb-io-memobase

User Profile-Based Long-Term Memory for AI Chatbot Applications.

Contributed by Sentinel

98.52,875 stars · 232 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
MemoryExperimental

zilliztech-memsearch

Use when several coding agents should share one memory store instead of each keeping its own notes.

98.52,571 stars · 238 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

kingjulio8238-memary

Use when you want a well-known reference implementation of agent memory over a knowledge graph to read or fork.

98.52,644 stars · 205 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
HooksPreview

tdd-guard

A hooks-driven system that monitors file operations in real-time and blocks changes that violate TDD principles.

98.42,324 stars · 185 forks
Harnessclaude
hook
No one-command install · SourceDetails
MemoryExperimental

redplanethq-core

Use when one memory graph should serve Claude Code, Codex and your other assistants at once.

98.31,963 stars · 189 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
MemoryExperimental

langchain-ai-langmem

Long-term memory for agents: tools that extract what matters from conversations, refine prompts from feedback and keep memory across sessions, with LangGraph's store built in.

Contributed by Sentinel

98.21,684 stars · 192 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
Browse · Armory