145 results in Memory, Observability, Identity, Evals · page 4 of 7
Harness Claude Code Cursor Codex Gemini OpenCode agiresearch-a-mem A-MEM: Agentic Memory for LLM Agents
Contributed by Sentinel
97.8 1,164 stars · 121 forks
Harness claude codex cursor gemini opencode
memory
letta-ai-agent-file Use when you need to save, share, version or move a whole agent — its persona, memory and behaviour — as one portable file.
97.8 1,197 stars · 114 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
claude-powerline A vim-style powerline statusline for Claude Code with real-time usage tracking, git integration, custom themes, and more
97.6 1,163 stars · 82 forks
Harness claude
claude-code status-lines
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
codeabra-iai-personal-memory-engine Use when the agent should remember not just facts but how you like to work, locally and for free.
97.5 862 stars · 105 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
unicity-aos-capsule-identity Use when an agent's identity must be persisted as state and assembled into its system prompt at boot, rather than pasted into a prompt by hand.
96.8 8,503 stars · 19 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
claude-code-statusline Enhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring
95.9 476 stars · 35 forks
Harness claude
statusline observability
phospho Phospho is a text analytics and evaluation platform for LLM apps. It logs sessions, runs clustering, detects failures, and surfaces actionable insights.
95.8 439 stars · 35 forks
Harness claude cursor codex opencode gemini
observability analytics evals
agntcy-oasf Use when agents must describe themselves to other systems in a common schema so they can be catalogued, discovered and verified.
95.7 332 stars · 47 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
telagod-code-abyss Use when you want a coding agent to have a consistent, composable personality and voice across Claude Code, Codex, Gemini CLI and OpenClaw.
94.6 239 stars · 32 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
athina-ai Athina AI provides developer-focused LLM monitoring and eval framework: real-time inference logging, automated evals, and regression detection in CI.
94.6 301 stars · 23 forks
Harness claude cursor codex opencode gemini
observability evals logging
agntcy-dir Use when agents and multi-agent systems need to announce themselves and be found across organisations rather than hardcoded.
94.6 184 stars · 55 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
character-card-spec-v2 Use when reading or writing the character-card files that the roleplay-agent ecosystem actually ships, including the PNG-embedded variant.
93.9 188 stars · 29 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
cloudflare-web-bot-auth Use when a website has to be able to tell that a request really came from your agent and not from someone impersonating it.
93.8 157 stars · 40 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
claude-pace A lightweight Bash + jq statusline for Claude Code that displays rate limit pace delta (burn rate vs. time remaining), 5h/7d usage percentage, context window usage, git branch and diff stats. Compares current consumption rate against time remaining in each rate limit window to indicate whether quota is being used faster or slower than the window allows. Single file with no external dependencies beyond jq.
93.6 229 stars · 19 forks
Harness claude
claude-code status-lines
nemori-ai-nemori Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.
93.5 207 stars · 20 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
highflame-ai-zeroid Use when a fleet of autonomous agents needs issued identities with a lifecycle — created, rotated and revoked.
92.7 163 stars · 18 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
jean-technologies-jean-memory Use when you want mem0-style and graph-style memory combined behind one interface.
92.1 171 stars · 13 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
agentmail-to-agentmail-toolkit Use when an agent needs its own email address so people and systems can reach it, and it can act on what arrives.
91.6 98 stars · 26 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
Previous Page 4 of 7 Next
Browse · Armory