{"note":"The single source for the 11 canonical harness components and their picks (CP138, set 2026-09-07). /c, /c/[component], /stack and /api/stack all read this file, so the shelves and the recommended stack can never disagree. Deliberately holds NO score: a pick's Score and Signals are resolved live from lib/rank.mjs by armoryName, so a re-rank can never leave a stale number here. A pick whose armoryName is null is not a catalog row and renders as Not Indexed, never faked into the table. Our own repositories are picks only when they win on score. `plane` is the capability plane (CP138 axis 2): accounts an agent is provisioned onto, not open-source code, so never scored. A pick's `reason` says in one line why it is listed though rows above it score higher; the pages show it only while the pick is not its shelf's top row.","as_of":"2026-09-07","regenerate":"node scripts/stack-evidence.mjs","stack":{"identity":"karpathy-coding-discipline","memory":"letta-ai-letta","skills":"superpowers","tools":"playwright-cli","hooks":"plannotator","subagents":"wshobson-agents","mcps":"github-mcp","dispatch":"pocketflow","evals":"promptfoo","observability":"langfuse","sandbox":"e2b-sandbox"},"total":11,"components":[{"slug":"identity","label":"Identity","one_line":"Who the agent is, and how the outside world reaches it","aggregates":["identity","rules"],"indexed":499,"ranked":42,"ranked_pct":8.4,"top_score":99.9,"fit":null,"leaderboard":"/leaderboard?component=rules","page":"/c/identity","picks":[{"name":"karpathy-coding-discipline","why":"Behavior norms loaded from CLAUDE.md: think first, simplify, keep the diff small","the_pick":true,"shelf_rank":1,"reason":null,"indexed":true,"armory_name":"karpathy-coding-discipline","url":"https://github.com/multica-ai/andrej-karpathy-skills","detail":"/e/claudemd-rules/karpathy-coding-discipline","type":"claudemd-rules","universal":99.9,"evidence":2,"ours":false,"signals":{"stars":209417,"usage":null,"tested":null,"mentions":null,"forks":21315}}]},{"slug":"memory","label":"Memory","one_line":"What the agent keeps between runs","aggregates":["memory"],"indexed":36,"ranked":36,"ranked_pct":100,"top_score":99.8,"fit":null,"leaderboard":"/leaderboard?component=memory","page":"/c/memory","picks":[{"name":"letta-ai-letta","why":"Stateful agents that edit their own long-term memory","the_pick":true,"shelf_rank":10,"reason":"Most rows above it are memory stores or graph builders outside the agent; in Letta the agent itself decides what to keep.","indexed":true,"armory_name":"letta-ai-letta","url":"https://github.com/letta-ai/letta","detail":"/e/memory/letta-ai-letta","type":"memory","universal":99.5,"evidence":3,"ours":false,"signals":{"stars":24552,"usage":null,"tested":null,"mentions":19,"forks":2609}},{"name":"getzep-graphiti","why":"Knowledge graph memory in which each fact records when it became true and when it was replaced","the_pick":false,"shelf_rank":6,"reason":"Listed as the open-source engine behind Zep Cloud, so the same memory can run on your own machines.","indexed":true,"armory_name":"getzep-graphiti","url":"https://github.com/getzep/graphiti","detail":"/e/memory/getzep-graphiti","type":"memory","universal":99.6,"evidence":3,"ours":false,"signals":{"stars":30509,"usage":null,"tested":null,"mentions":1,"forks":3098}},{"name":"plastic-labs-honcho","why":"Memory library for stateful agents, kept per user and per session","the_pick":false,"shelf_rank":16,"reason":"Listed because it reasons in the background to keep a picture of each person an agent deals with, not only a record of what was said.","indexed":true,"armory_name":"plastic-labs-honcho","url":"https://github.com/plastic-labs/honcho","detail":"/e/memory/plastic-labs-honcho","type":"memory","universal":99.1,"evidence":3,"ours":false,"signals":{"stars":6980,"usage":null,"tested":null,"mentions":6,"forks":865}}]},{"slug":"skills","label":"Skills","one_line":"Procedures loaded on demand, one file each","aggregates":["skill"],"indexed":1132,"ranked":29,"ranked_pct":2.6,"top_score":99.9,"fit":null,"leaderboard":"/leaderboard?component=skill","page":"/c/skills","picks":[{"name":"superpowers","why":"One bundle covering planning, review, testing and the rest of the lifecycle","the_pick":true,"shelf_rank":1,"reason":null,"indexed":true,"armory_name":"superpowers","url":"https://github.com/obra/superpowers","detail":"/e/skills/superpowers","type":"skills","universal":99.9,"evidence":3,"ours":false,"signals":{"stars":280507,"usage":null,"tested":null,"mentions":1,"forks":25127}},{"name":"everything-claude-code","why":"Skills across core engineering domains, one file per competency","the_pick":false,"shelf_rank":2,"reason":"Second by score; the only row above it is the pick.","indexed":true,"armory_name":"everything-claude-code","url":"https://github.com/affaan-m/everything-claude-code","detail":"/e/skills/everything-claude-code","type":"skills","universal":99.9,"evidence":3,"ours":false,"signals":{"stars":245823,"usage":null,"tested":null,"mentions":1,"forks":37094}}]},{"slug":"tools","label":"Tools","one_line":"Terminal, browser and computer commands an agent runs","aggregates":["cli","tool"],"indexed":57,"ranked":54,"ranked_pct":94.7,"top_score":99.9,"fit":{"purpose":"run as commands from a terminal","filed":147,"left_out":90},"leaderboard":"/leaderboard?component=cli","page":"/c/tools","picks":[{"name":"playwright-cli","why":"Codegen, screenshot, PDF and trace commands for headless browser scripting","the_pick":true,"shelf_rank":9,"reason":"Picked to drive a browser from the terminal; most rows above it are terminal apps and other command-line tools, and Puppeteer, the closest match, has no WebKit.","indexed":true,"armory_name":"playwright-cli","url":"https://github.com/microsoft/playwright","detail":"/e/clis-tools/playwright-cli","type":"clis-tools","universal":99.8,"evidence":4,"ours":false,"signals":{"stars":13015,"usage":null,"tested":1,"mentions":1,"forks":704}},{"name":"gh","why":"GitHub from the terminal: issues, pull requests, releases, repository search","the_pick":false,"shelf_rank":1,"reason":null,"indexed":true,"armory_name":"gh","url":"https://github.com/cli/cli","detail":"/e/clis-tools/gh","type":"clis-tools","universal":99.9,"evidence":3,"ours":false,"signals":{"stars":46103,"usage":null,"tested":1,"mentions":null,"forks":8949}},{"name":"crawl4ai","why":"Async web crawler that turns pages into clean Markdown for an agent to read","the_pick":false,"shelf_rank":10,"reason":"Runs locally as a Python library with no account; Firecrawl, higher on this shelf, is mostly used as a hosted service.","indexed":true,"armory_name":"crawl4ai","url":"https://github.com/unclecode/crawl4ai","detail":"/e/clis-tools/crawl4ai","type":"clis-tools","universal":99.8,"evidence":2,"ours":false,"signals":{"stars":84297,"usage":null,"tested":null,"mentions":null,"forks":8717}}]},{"slug":"hooks","label":"Hooks","one_line":"Deterministic code fired on agent lifecycle events","aggregates":["hook"],"indexed":134,"ranked":12,"ranked_pct":9,"top_score":99.1,"fit":null,"leaderboard":"/leaderboard?component=hook","page":"/c/hooks","picks":[{"name":"plannotator","why":"Intercepts ExitPlanMode so a plan is annotated and approved before work starts","the_pick":true,"shelf_rank":1,"reason":null,"indexed":true,"armory_name":"plannotator","url":"https://github.com/backnotprop/plannotator","detail":"/e/hooks/plannotator","type":"hooks","universal":99.1,"evidence":2,"ours":false,"signals":{"stars":8344,"usage":null,"tested":null,"mentions":null,"forks":618}},{"name":"tdd-guard","why":"Blocks file writes that violate test-first order, as they happen","the_pick":false,"shelf_rank":2,"reason":"Second by score; the only row above it is the pick.","indexed":true,"armory_name":"tdd-guard","url":"https://github.com/nizos/tdd-guard","detail":"/e/hooks/tdd-guard","type":"hooks","universal":98.4,"evidence":2,"ours":false,"signals":{"stars":2324,"usage":null,"tested":null,"mentions":null,"forks":185}}]},{"slug":"subagents","label":"Sub-Agents","one_line":"Delegated workers, each with its own context window","aggregates":["subagent"],"indexed":724,"ranked":2,"ranked_pct":0.3,"top_score":99.7,"fit":null,"leaderboard":"/leaderboard?component=subagent","page":"/c/subagents","picks":[{"name":"wshobson-agents","why":"Plugin marketplace of sub-agents for Claude Code, Codex, Cursor and OpenCode","the_pick":true,"shelf_rank":1,"reason":null,"indexed":true,"armory_name":"wshobson-agents","url":"https://github.com/wshobson/agents","detail":"/e/subagents/wshobson-agents","type":"subagents","universal":99.7,"evidence":2,"ours":false,"signals":{"stars":39338,"usage":null,"tested":null,"mentions":null,"forks":4195}},{"name":"voltagent-awesome-claude-code-subagents","why":"100+ specialized Claude Code sub-agents, one file each","the_pick":false,"shelf_rank":2,"reason":"Second by score; the only row above it is the pick.","indexed":true,"armory_name":"voltagent-awesome-claude-code-subagents","url":"https://github.com/VoltAgent/awesome-claude-code-subagents","detail":"/e/subagents/voltagent-awesome-claude-code-subagents","type":"subagents","universal":99.5,"evidence":2,"ours":false,"signals":{"stars":24793,"usage":null,"tested":null,"mentions":null,"forks":2870}}]},{"slug":"mcps","label":"MCPs","one_line":"Servers exposing tools and data over one protocol","aggregates":["mcp"],"indexed":61707,"ranked":28693,"ranked_pct":46.5,"top_score":99.9,"fit":null,"leaderboard":"/leaderboard?component=mcp","page":"/c/mcps","picks":[{"name":"github-mcp","why":"Issues, pull requests, repository search and commits without shelling out to git","the_pick":true,"shelf_rank":3,"reason":"Picked as the server most coding agents need first, published by GitHub itself; the rows above it drive a browser and convert documents.","indexed":true,"armory_name":"github-mcp","url":"https://github.com/github/github-mcp-server","detail":"/e/mcps/github-mcp","type":"mcps","universal":99.9,"evidence":3,"ours":false,"signals":{"stars":32663,"usage":null,"tested":1,"mentions":null,"forks":4883}},{"name":"genai-toolbox","why":"Pre-defined queries against Postgres, MySQL, Spanner and other databases, set up in YAML","the_pick":false,"shelf_rank":5,"reason":"A runner-up for database access; the rows above it drive a browser, convert documents, work with GitHub and follow the news.","indexed":true,"armory_name":"genai-toolbox","url":"https://github.com/googleapis/mcp-toolbox","detail":"/e/mcps/genai-toolbox","type":"mcps","universal":99.8,"evidence":3,"ours":false,"signals":{"stars":16293,"usage":null,"tested":1,"mentions":null,"forks":1706}},{"name":"chrome-devtools","why":"Chrome control through the DevTools protocol: automation, debugging, performance traces","the_pick":false,"shelf_rank":1,"reason":null,"indexed":true,"armory_name":"chrome-devtools","url":"https://github.com/chromedevtools/chrome-devtools-mcp","detail":"/e/mcps/chrome-devtools","type":"mcps","universal":99.9,"evidence":4,"ours":false,"signals":{"stars":52627,"usage":null,"tested":1,"mentions":4,"forks":4864}}]},{"slug":"dispatch","label":"Dispatch","one_line":"Routing between agents, plus scheduled and looping runs","aggregates":["workflow"],"indexed":62,"ranked":21,"ranked_pct":33.9,"top_score":99.9,"fit":{"purpose":"route, schedule or loop agent work","filed":731,"left_out":669},"leaderboard":"/leaderboard?component=workflow","page":"/c/dispatch","picks":[{"name":"pocketflow","why":"100-line graph framework; nodes and edges route work between agents","the_pick":true,"shelf_rank":12,"reason":"Picked as a library you drop into your own code: 100 lines with no dependencies. Above it, n8n and Paperclip are platforms you run, LangGraph, AutoGen, CrewAI, the OpenAI Agents SDK and Mastra are much larger frameworks, and ruflo, Orca, Symphony and Auto-Claude run coding agents for you.","indexed":true,"armory_name":"pocketflow","url":"https://github.com/The-Pocket/PocketFlow","detail":"/e/workflows/pocketflow","type":"workflows","universal":99.2,"evidence":2,"ours":false,"signals":{"stars":11139,"usage":null,"tested":null,"mentions":null,"forks":1217}},{"name":"n8n-io-n8n","why":"Workflow automation with 400+ integrations, built visually or in code; heavier to run","the_pick":false,"shelf_rank":1,"reason":null,"indexed":true,"armory_name":"n8n-io-n8n","url":"https://github.com/n8n-io/n8n","detail":"/e/workflows/n8n-io-n8n","type":"workflows","universal":99.9,"evidence":3,"ours":false,"signals":{"stars":206033,"usage":null,"tested":null,"mentions":26,"forks":60896}}]},{"slug":"evals","label":"Evals","one_line":"Graded tasks that prove a change helped","aggregates":["eval"],"indexed":37,"ranked":32,"ranked_pct":86.5,"top_score":99.5,"fit":null,"leaderboard":"/leaderboard?component=eval","page":"/c/evals","picks":[{"name":"promptfoo","why":"Assertions and red-teaming for prompts and agents, runs in CI","the_pick":true,"shelf_rank":1,"reason":null,"indexed":true,"armory_name":"promptfoo","url":"https://github.com/promptfoo/promptfoo","detail":"/e/evals/promptfoo","type":"evals","universal":99.5,"evidence":2,"ours":false,"signals":{"stars":24737,"usage":null,"tested":null,"mentions":null,"forks":2255}},{"name":"swe-bench","why":"Real GitHub issues drawn from 12 Python repositories, scored on the resulting patch","the_pick":false,"shelf_rank":7,"reason":"Listed as the benchmark coding agents are compared on; the rows above it test prompts, apps and models instead.","indexed":true,"armory_name":"swe-bench","url":"https://github.com/SWE-bench/SWE-bench","detail":"/e/evals/swe-bench","type":"evals","universal":99.1,"evidence":3,"ours":false,"signals":{"stars":5762,"usage":null,"tested":null,"mentions":9,"forks":957}},{"name":"deepeval","why":"Evaluation framework with 14+ metrics such as hallucination and faithfulness, runs in CI","the_pick":false,"shelf_rank":4,"reason":"OpenAI Evals and the LM Evaluation Harness, above it, mainly grade models; this grades your own app.","indexed":true,"armory_name":"deepeval","url":"https://github.com/confident-ai/deepeval","detail":"/e/evals/deepeval","type":"evals","universal":99.4,"evidence":2,"ours":false,"signals":{"stars":18041,"usage":null,"tested":null,"mentions":null,"forks":1890}}]},{"slug":"observability","label":"Observability","one_line":"Traces, cost and errors from each run","aggregates":["observability"],"indexed":34,"ranked":26,"ranked_pct":76.5,"top_score":99.7,"fit":null,"leaderboard":"/leaderboard?component=observability","page":"/c/observability","picks":[{"name":"langfuse","why":"Traces, prompt management, datasets and evals in one self-hostable platform","the_pick":true,"shelf_rank":3,"reason":"The two rows above it are tracing add-ons to Sentry and MLflow; this is a whole platform on its own.","indexed":true,"armory_name":"langfuse","url":"https://github.com/langfuse/langfuse","detail":"/e/observability/langfuse","type":"observability","universal":99.6,"evidence":3,"ours":false,"signals":{"stars":34067,"usage":null,"tested":null,"mentions":10,"forks":3678}},{"name":"comet-opik","why":"Open-source tracing and evaluation: log traces, run evals, track prompt changes","the_pick":false,"shelf_rank":5,"reason":"Built like the pick; above it are two tracing add-ons, the pick and a Claude Code status line.","indexed":true,"armory_name":"comet-opik","url":"https://github.com/comet-ml/opik","detail":"/e/observability/comet-opik","type":"observability","universal":99.4,"evidence":3,"ours":false,"signals":{"stars":22249,"usage":null,"tested":null,"mentions":1,"forks":1827}}]},{"slug":"sandbox","label":"Sandbox","one_line":"Isolated compute an agent executes and deploys in","aggregates":["infra"],"indexed":26,"ranked":14,"ranked_pct":53.8,"top_score":99.9,"fit":{"purpose":"run agent code in isolation","filed":58,"left_out":32},"leaderboard":"/leaderboard?component=infra","page":"/c/sandbox","picks":[{"name":"e2b-sandbox","why":"Firecracker microVM per run, ~150ms cold start, for untrusted code","the_pick":true,"shelf_rank":2,"reason":"The one row above it is Daytona, whose public repository stopped updating in June 2026.","indexed":true,"armory_name":"e2b-sandbox","url":"https://github.com/e2b-dev/E2B","detail":"/e/infrastructure/e2b-sandbox","type":"infrastructure","universal":99.8,"evidence":4,"ours":false,"signals":{"stars":13635,"usage":null,"tested":1,"mentions":10,"forks":1015}},{"name":"daytona","why":"Sandboxes for AI-generated code that start in under 90 ms; its public repository stopped updating in June 2026","the_pick":false,"shelf_rank":1,"reason":null,"indexed":true,"armory_name":"daytona","url":"https://github.com/daytonaio/daytona","detail":"/e/infrastructure/daytona","type":"infrastructure","universal":99.9,"evidence":4,"ours":false,"signals":{"stars":71846,"usage":null,"tested":1,"mentions":9,"forks":5651}},{"name":"microsandbox","why":"Self-hosted microVM sandbox on your own machines, no per-sandbox vendor cost","the_pick":false,"shelf_rank":5,"reason":"Above it, Daytona and E2B are mostly used as hosted services, container-use gives each coding agent a container on its own git branch, and Cua is built for computer-use agents; this one runs untrusted code in tiny virtual machines you host yourself.","indexed":true,"armory_name":"microsandbox","url":"https://github.com/microsandbox/microsandbox","detail":"/e/infrastructure/microsandbox","type":"infrastructure","universal":99.1,"evidence":2,"ours":false,"signals":{"stars":8436,"usage":null,"tested":null,"mentions":null,"forks":454}}]}],"plane":[{"slot":"Computer","pick":"Orgo","access":"ORGO_API_KEY, provisioned through bezalel.sh so the agent never sees the key"},{"slot":"Sandbox","pick":"E2B","access":"E2B_API_KEY, one microVM per run"},{"slot":"Database","pick":"Supabase","access":"URL and anon key to the agent; the service-role key is injected at the tool boundary, never into a prompt or a log"},{"slot":"Email","pick":"AgentMail","access":"API key and a per-inbox webhook secret (starts whsec_); verify the signature, or the route is an open relay"},{"slot":"Phone / SMS / WhatsApp","pick":"Twilio","access":"account SID, auth token and a purchased number; inbound requests are signed with X-Twilio-Signature"},{"slot":"iMessage","pick":"Photon Spectrum","access":"project id and secret, and the owner pairs their own handle with a one-time code (OTP)"},{"slot":"Card / spend","pick":"Agent Card HQ","access":"an identity check (KYC), then a funded balance; the slowest slot, so start it first"},{"slot":"Payments / earn","pick":"Stripe agent-toolkit","access":"a restricted key scoped to the agent, never the account key"},{"slot":"Memory store","pick":"Supermemory","access":"SUPERMEMORY_API_KEY, or self-host"},{"slot":"Connectors","pick":"Composio","access":"a key, plus the owner's OAuth consent per app; the agent never holds the third-party token"},{"slot":"Deploy / hosting","pick":"Vercel","access":"token and a linked project; after a CLI deploy, vercel alias set is a separate step, or nothing ships"},{"slot":"Web / browser","pick":"Browserbase + Stagehand","access":"key and project id, or local Chromium on the owner's profile, read-only"}],"provisioning":"Start the slow slots first: the card needs an identity check (KYC) and iMessage needs the owner's one-time code, and both take days; then buy the phone number, then set up the email inbox and its webhook secret; everything that only needs a key takes a minute."}