229 results in CLIs & Tools, Memory, Evals · page 6 of 10
Harness Claude Code Cursor Codex Gemini OpenCode letta-ai-letta-code Stateful agents that are like people, with memory, identity, and the ability to learn and adapt
Contributed by Sentinel
98.7 3,185 stars · 385 forks · 6 mentions
Harness claude codex cursor gemini opencode
mirix-ai-mirix Use when what the agent should remember is what actually happened on screen, consolidated into structured memories.
98.7 3,440 stars · 270 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
helm Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.7 2,898 stars · 412 forks
Harness claude cursor codex opencode gemini
evals benchmark academic stanford
lighteval Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.
98.6 2,533 stars · 553 forks
Harness claude cursor codex opencode gemini
evals huggingface benchmark lightweight
kayba-ai-agentic-context-engine Use when an agent should carry forward what it learned from its own successes and failures into later runs.
98.5 2,564 stars · 307 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
memodb-io-memobase User Profile-Based Long-Term Memory for AI Chatbot Applications.
Contributed by Sentinel
98.5 2,875 stars · 232 forks
Harness claude codex cursor gemini opencode
memory
crystal A full-fledged desktop application for orchestrating, monitoring, and interacting with Claude Code agents.
98.5 3,114 stars · 197 forks
Harness claude
client cli
omnara A command center for AI agents that syncs Claude Code sessions across terminal, web, and mobile. Allows for remote monitoring, human-in-the-loop interaction, and team collaboration.
98.5 2,780 stars · 213 forks · 4 mentions
Harness claude
claude-code alternative-clients
primeintellect-ai-prime-rl Agentic RL Training at Scale
Contributed by Sentinel
98.5 2,007 stars · 421 forks · 3 mentions
Harness claude codex cursor gemini opencode
zilliztech-memsearch Use when several coding agents should share one memory store instead of each keeping its own notes.
98.5 2,571 stars · 238 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
kingjulio8238-memary Use when you want a well-known reference implementation of agent memory over a knowledge graph to read or fork.
98.5 2,644 stars · 205 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
tweakcc Command-line tool to customize your Claude Code styling.
98.4 2,476 stars · 199 forks
Harness claude
claude-code tooling
redplanethq-core Use when one memory graph should serve Claude Code, Codex and your other assistants at once.
98.3 1,963 stars · 189 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
evalplus Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.
98.3 1,819 stars · 208 forks
Harness claude cursor codex opencode gemini
evals code-generation humaneval benchmark
webarena WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.
98.2 1,592 stars · 249 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents browser benchmark
langchain-ai-langmem Long-term memory for agents: tools that extract what matters from conversations, refine prompts from feedback and keep memory across sessions, with LangGraph's store built in.
Contributed by Sentinel
98.2 1,684 stars · 192 forks
Harness claude codex cursor gemini opencode
memory
claude-code-tools Well-crafted toolset for session continuity, featuring skills/commands to avoid compaction and recover context across sessions with cross-agent handoff between Claude Code and Codex CLI. Includes a fast Rust/Tantivy-powered full-text session search (TUI for humans, skill/CLI for agents), tmux-cli skill + command for interacting with scripts and CLI agents, and safety hooks to block dangerous commands.
98.2 1,989 stars · 132 forks
Harness claude
claude-code tooling
cc-sessions An opinionated approach to productive development with Claude Code
98.1 1,551 stars · 191 forks
Harness claude
claude-code tooling
tau-bench Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
98.1 1,416 stars · 215 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents tool-use benchmark
harbor-framework-terminal-bench Measuring and evolving with the frontier of agent work
Contributed by Sentinel
98.1 588 stars · 434 forks · 13 mentions
Harness claude codex cursor gemini opencode
bai-lab-memoryos Use when you want a memory design with a published, peer-reviewed evaluation behind it.
98.1 1,570 stars · 161 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
cortexkit-magic-context Use when a long coding session keeps losing its earlier context and you want that handled automatically.
98.1 2,044 stars · 106 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
dataojitori-nocturne-memory Use when you want to see and roll back what your agent remembered, instead of trusting an opaque vector store.
98.0 1,373 stars · 170 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
claude-code-ide-el claude-code-ide.el integrates Claude Code with Emacs, like Anthropic’s VS Code/IntelliJ extensions. It shows ediff-based code suggestions, pulls LSP/flymake/flycheck diagnostics, and tracks buffer context. It adds an extensible MCP tool support for symbol refs/defs, project metadata, and tree-sitter AST queries.
98.0 1,658 stars · 112 forks
Harness claude
claude-code tooling
Previous Page 6 of 10 Next
Browse · Armory