833 results in Observability, Evals, Sub-Agents, Identity · page 3 of 35
Harness Claude Code Cursor Codex Gemini OpenCode unicity-aos-capsule-identity Use when an agent's identity must be persisted as state and assembled into its system prompt at boot, rather than pasted into a prompt by hand.
96.8 8,503 stars · 19 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
claude-code-statusline Enhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring
95.9 476 stars · 35 forks
Harness claude
statusline observability
phospho Phospho is a text analytics and evaluation platform for LLM apps. It logs sessions, runs clustering, detects failures, and surfaces actionable insights.
95.8 439 stars · 35 forks
Harness claude cursor codex opencode gemini
observability analytics evals
agntcy-oasf Use when agents must describe themselves to other systems in a common schema so they can be catalogued, discovered and verified.
95.7 332 stars · 47 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
telagod-code-abyss Use when you want a coding agent to have a consistent, composable personality and voice across Claude Code, Codex, Gemini CLI and OpenClaw.
94.6 239 stars · 32 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
athina-ai Athina AI provides developer-focused LLM monitoring and eval framework: real-time inference logging, automated evals, and regression detection in CI.
94.6 301 stars · 23 forks
Harness claude cursor codex opencode gemini
observability evals logging
agntcy-dir Use when agents and multi-agent systems need to announce themselves and be found across organisations rather than hardcoded.
94.6 184 stars · 55 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
character-card-spec-v2 Use when reading or writing the character-card files that the roleplay-agent ecosystem actually ships, including the PNG-embedded variant.
93.9 188 stars · 29 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
cloudflare-web-bot-auth Use when a website has to be able to tell that a request really came from your agent and not from someone impersonating it.
93.8 157 stars · 40 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
claude-pace A lightweight Bash + jq statusline for Claude Code that displays rate limit pace delta (burn rate vs. time remaining), 5h/7d usage percentage, context window usage, git branch and diff stats. Compares current consumption rate against time remaining in each rate limit window to indicate whether quota is being used faster or slower than the window allows. Single file with no external dependencies beyond jq.
93.6 229 stars · 19 forks
Harness claude
claude-code status-lines
highflame-ai-zeroid Use when a fleet of autonomous agents needs issued identities with a lifecycle — created, rotated and revoked.
92.7 163 stars · 18 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
agentmail-to-agentmail-toolkit Use when an agent needs its own email address so people and systems can reach it, and it can act on what arrives.
91.6 98 stars · 26 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
agntcy-identity Use when agents, MCP servers and multi-agent systems all need issued identities that another party can verify.
91.4 101 stars · 20 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
character-card-spec-v3 Use when authoring an agent persona against the current character-card standard, with lorebooks, assets and decorators.
90.2 109 stars · 11 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
agentrhq-authsome Use when agents must stay logged in to third-party services without ever seeing your credentials.
88.6 87 stars · 9 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
gebruder-wirken Use when autonomous agents need one gateway that holds their credentials, isolates them per channel, and logs every session.
88.6 171 stars · 5 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
opena2a-agent-identity-management Use when non-human identities need the same lifecycle a workforce IAM gives people — issue, authorize, audit, revoke.
88.5 59 stars · 18 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
vestauth Use when an agent needs to authenticate as itself to services, without you hand-rolling keys and rotation.
87.5 166 stars · 4 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
vorim-ai-labs-vorim-mcp-server Use when you want agent identity, scoped permissions and an audit trail exposed to the agent as tools it can call.
87.4 73 stars · 8 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
Previous Page 3 of 35 Next
Browse · Armory