797 results in Sub-Agents, Memory, Evals · page 3 of 34
Harness Claude Code Cursor Codex Gemini OpenCode bai-lab-memoryos Use when you want a memory design with a published, peer-reviewed evaluation behind it.
98.1 1,570 stars · 161 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
cortexkit-magic-context Use when a long coding session keeps losing its earlier context and you want that handled automatically.
98.1 2,044 stars · 106 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
dataojitori-nocturne-memory Use when you want to see and roll back what your agent remembered, instead of trusting an opaque vector store.
98.0 1,373 stars · 170 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
claudiodrews-memory-os Use when you want layered memory — structured facts, recall and an auto-curated wiki — running locally against any model.
97.9 1,355 stars · 128 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
langtrace Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.8 1,228 stars · 127 forks
Harness claude cursor codex opencode gemini
evals observability opentelemetry tracing
agiresearch-a-mem A-MEM: Agentic Memory for LLM Agents
Contributed by Sentinel
97.8 1,164 stars · 121 forks
Harness claude codex cursor gemini opencode
memory
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
codeabra-iai-personal-memory-engine Use when the agent should remember not just facts but how you like to work, locally and for free.
97.5 862 stars · 105 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
nemori-ai-nemori Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.
93.5 207 stars · 20 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
jean-technologies-jean-memory Use when you want mem0-style and graph-style memory combined behind one interface.
92.1 171 stars · 13 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
kyros-ai Use when the agent's memory needs to resolve its own contradictions and forget on a schedule.
82.5 96 stars · 2 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
braintrust Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
82.3 27 stars · 12 forks
Harness claude cursor codex opencode gemini
evals experiment-tracking sdk
galileo-evaluate Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.
80.3 22 stars · 11 forks
Harness claude cursor codex opencode gemini
evals hallucination observability sdk
humanloop-evals Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.
69.0 12 stars · 3 forks
Harness claude cursor codex opencode gemini
evals human-eval dataset sdk
wikimem Use to give an agent a queryable wiki knowledge base when memory should be a navigable knowledge graph, not just a flat log: ingest files, folders, and URLs into a linked vault, then search or ask it in natural language.
63.4 7 stars · 4 forks
Harness claude
memory knowledge-base wiki ingest
agentbench Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.
Unranked No signals yet
Harness claude codex
eval benchmark scoring harness
3d-artist 3D art and asset creation specialist for game development. Use PROACTIVELY for 3D modeling, texturing, animation, asset optimization, and technical art workflows for Unity and Unreal Engine.
Unranked No signals yet
Harness claude
game-development subagents
Previous Page 3 of 34 Next
Browse · Armory