831 results in Observability, Sub-Agents, Memory, Evals · page 4 of 35
Harness Claude Code Cursor Codex Gemini OpenCode agiresearch-a-mem A-MEM: Agentic Memory for LLM Agents
Contributed by Sentinel
97.8 1,164 stars · 121 forks
Harness claude codex cursor gemini opencode
memory
claude-powerline A vim-style powerline statusline for Claude Code with real-time usage tracking, git integration, custom themes, and more
97.6 1,163 stars · 82 forks
Harness claude
claude-code status-lines
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
codeabra-iai-personal-memory-engine Use when the agent should remember not just facts but how you like to work, locally and for free.
97.5 862 stars · 105 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
claude-code-statusline Enhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring
95.9 476 stars · 35 forks
Harness claude
statusline observability
phospho Phospho is a text analytics and evaluation platform for LLM apps. It logs sessions, runs clustering, detects failures, and surfaces actionable insights.
95.8 439 stars · 35 forks
Harness claude cursor codex opencode gemini
observability analytics evals
athina-ai Athina AI provides developer-focused LLM monitoring and eval framework: real-time inference logging, automated evals, and regression detection in CI.
94.6 301 stars · 23 forks
Harness claude cursor codex opencode gemini
observability evals logging
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
claude-pace A lightweight Bash + jq statusline for Claude Code that displays rate limit pace delta (burn rate vs. time remaining), 5h/7d usage percentage, context window usage, git branch and diff stats. Compares current consumption rate against time remaining in each rate limit window to indicate whether quota is being used faster or slower than the window allows. Single file with no external dependencies beyond jq.
93.6 229 stars · 19 forks
Harness claude
claude-code status-lines
nemori-ai-nemori Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.
93.5 207 stars · 20 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
jean-technologies-jean-memory Use when you want mem0-style and graph-style memory combined behind one interface.
92.1 171 stars · 13 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
kyros-ai Use when the agent's memory needs to resolve its own contradictions and forget on a schedule.
82.5 96 stars · 2 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
braintrust Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
82.3 27 stars · 12 forks
Harness claude cursor codex opencode gemini
evals experiment-tracking sdk
claudia-statusline High-performance Rust-based statusline for Claude Code with persistent stats tracking, progress bars, and optional cloud sync. Features SQLite-first persistence, git integration, context progress bars, burn rate calculation, XDG-compliant with theme support (dark/light, NO_COLOR).
81.2 36 stars · 5 forks
Harness claude
claude-code status-lines
galileo-evaluate Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.
80.3 22 stars · 11 forks
Harness claude cursor codex opencode gemini
evals hallucination observability sdk
honeycomb Honeycomb's OpenTelemetry-native observability platform: a high-cardinality event store ideal for tracing LLM pipelines and debugging slow agent traces.
75.1 16 stars · 6 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
baserun Baserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and measuring quality.
74.4 16 stars · 5 forks
Harness claude cursor codex opencode gemini
observability tracing testing
humanloop-evals Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.
69.0 12 stars · 3 forks
Harness claude cursor codex opencode gemini
evals human-eval dataset sdk
wikimem Use to give an agent a queryable wiki knowledge base when memory should be a navigable knowledge graph, not just a flat log: ingest files, folders, and URLs into a linked vault, then search or ask it in natural language.
63.4 7 stars · 4 forks
Harness claude
memory knowledge-base wiki ingest
Previous Page 4 of 35 Next
Browse · Armory