245 results in Identity, Evals, Memory, Hooks · page 1 of 11
Harness Claude Code Cursor Codex Gemini OpenCode thedotmack-claude-mem Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Contributed by Sentinel
99.8 94,727 stars · 8,379 forks · 6 mentions
Harness claude codex cursor gemini opencode
cli
mempalace Use when you want an agent memory system whose recall quality has actually been benchmarked rather than asserted.
99.8 58,897 stars · 7,548 forks · 2 mentions
Harness claude codex cursor gemini opencode
cp138-seed memory
microsoft-graphrag A modular graph-based Retrieval-Augmented Generation (RAG) system
Contributed by Sentinel
99.6 35,783 stars · 3,754 forks · 3 mentions
Harness claude codex cursor gemini opencode
volcengine-openviking Use when an agent's memory, retrieved knowledge and learned skills should live in one store that reorganises itself.
99.6 35,893 stars · 2,739 forks · 2 mentions
Harness claude codex cursor gemini opencode
cp138-seed memory
garrytan-gbrain Garry's Opinionated OpenClaw/Hermes Agent Brain
Contributed by Sentinel
99.6 29,460 stars · 4,394 forks · 16 mentions
Harness claude codex cursor gemini opencode
getzep-graphiti Build Real-Time Knowledge Graphs for AI Agents
Contributed by Sentinel
99.6 30,509 stars · 3,098 forks · 1 mention
Harness claude codex cursor gemini opencode
memory
topoteretes-cognee Use when an agent needs long-term memory backed by a knowledge graph you can host yourself.
99.6 30,556 stars · 3,010 forks · 4 mentions
Harness claude codex cursor gemini opencode
cp138-seed memory
supermemoryai-supermemory Use when memory has to be fast, run locally, and be reachable from an app as well as an agent.
99.6 29,256 stars · 2,557 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
tencentcloud-tencentdb-agent-memory TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and code into four reusable memory assets (Chat Memory, Skill, LLM-Wiki, Code-Graph) that are governed, shared, and equipped across agents and frameworks.
Contributed by Sentinel
99.5 25,635 stars · 2,398 forks · 3 mentions
Harness claude codex cursor gemini opencode
letta-ai-letta Platform for stateful agents: AI with advanced memory that can learn and self-improve over time.
Contributed by Sentinel
99.5 24,552 stars · 2,609 forks · 19 mentions
Harness claude codex cursor gemini opencode
memory
promptfoo CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.
99.5 24,737 stars · 2,255 forks
Harness claude cursor codex opencode gemini
evals red-teaming ci cli
openai-evals OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.
99.5 19,509 stars · 3,093 forks
Harness claude cursor codex opencode gemini
evals registry benchmark
lm-evaluation-harness EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.
99.5 13,860 stars · 3,533 forks
Harness claude cursor codex opencode gemini
evals academic benchmark harness
gibsonai-memori Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems. Built for enterprise, Memori works with the data infrastructure you already run, no rip-and-replace, and deploys across managed cloud, single-tenant cloud, VPC, and on-premises.
Contributed by Sentinel
99.4 16,313 stars · 3,283 forks
Harness claude codex cursor gemini opencode
memory
deepeval Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.
99.4 18,041 stars · 1,890 forks
Harness claude cursor codex opencode gemini
evals metrics rag ci
ragas Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.
99.4 15,853 stars · 1,727 forks · 1 mention · failed install test
Harness claude cursor codex opencode gemini
evals rag metrics
semantica-agi-semantica Use when an agent's stored context needs provenance — you must be able to say where a remembered fact came from.
99.3 13,484 stars · 1,525 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
nevamind-ai-memu Use when one person's memory should follow them across several different agents.
99.3 14,387 stars · 1,063 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
evermind-ai-everos Use when you want the agent's memory to be plain Markdown on your own disk rather than a hosted database.
99.2 12,757 stars · 911 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
memtensor-memos Use when memory should persist across tasks and be reused, not just retrieved once per conversation.
99.2 11,599 stars · 1,060 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
phoenix Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.
99.2 11,286 stars · 1,086 forks · 2 mentions
Harness claude cursor codex opencode gemini
evals observability rag agents
plastic-labs-honcho Memory library for building stateful agents
Contributed by Sentinel
99.1 6,980 stars · 865 forks · 6 mentions
Harness claude codex cursor gemini opencode
memory
plannotator Interactive plan review UI that intercepts ExitPlanMode via hooks, letting users visually annotate plans with comments, deletions, and replacements before approving or denying with detailed feedback.
99.1 8,344 stars · 618 forks
Harness claude
hook
swe-bench SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
99.1 5,762 stars · 957 forks · 9 mentions
Harness claude cursor codex opencode gemini
evals code benchmark agents
Page 1 of 11 Next
Browse · Armory