799 results in Evals, Identity, Sub-Agents · page 2 of 34
wandb-weave-evals
Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
langtrace
Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
letta-ai-agent-file
Use when you need to save, share, version or move a whole agent — its persona, memory and behaviour — as one portable file.
97.81,197 stars · 114 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
xiaowu0162-longmemeval
Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
princeton-nlp-webshop
[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1589 stars · 107 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
unicity-aos-capsule-identity
Use when an agent's identity must be persisted as state and assembled into its system prompt at boot, rather than pasted into a prompt by hand.
96.88,503 stars · 19 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
stonybrooknlp-appworld
🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7500 stars · 78 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
continuous-eval
Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2517 stars · 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
agntcy-oasf
Use when agents must describe themselves to other systems in a common schema so they can be catalogued, discovered and verified.
95.7332 stars · 47 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
telagod-code-abyss
Use when you want a coding agent to have a consistent, composable personality and voice across Claude Code, Codex, Gemini CLI and OpenClaw.
94.6239 stars · 32 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
agntcy-dir
Use when agents and multi-agent systems need to announce themselves and be found across organisations rather than hardcoded.
94.6184 stars · 55 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
metr-task-standard
METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3192 stars · 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
harbor-framework-terminal-bench-2-1
Terminal-Bench 2.1
Contributed by Sentinel
94.3119 stars · 62 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
character-card-spec-v2
Use when reading or writing the character-card files that the roleplay-agent ecosystem actually ships, including the PNG-embedded variant.
93.9188 stars · 29 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
cloudflare-web-bot-auth
Use when a website has to be able to tell that a request really came from your agent and not from someone impersonating it.
93.8157 stars · 40 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
highflame-ai-zeroid
Use when a fleet of autonomous agents needs issued identities with a lifecycle — created, rotated and revoked.
92.7163 stars · 18 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
agentmail-to-agentmail-toolkit
Use when an agent needs its own email address so people and systems can reach it, and it can act on what arrives.
91.698 stars · 26 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
agntcy-identity
Use when agents, MCP servers and multi-agent systems all need issued identities that another party can verify.
91.4101 stars · 20 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
vellum-evals
Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.582 stars · 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
character-card-spec-v3
Use when authoring an agent persona against the current character-card standard, with lorebooks, assets and decorators.
90.2109 stars · 11 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
agentrhq-authsome
Use when agents must stay logged in to third-party services without ever seeing your credentials.
88.687 stars · 9 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
gebruder-wirken
Use when autonomous agents need one gateway that holds their credentials, isolates them per channel, and logs every session.
88.6171 stars · 5 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
opena2a-agent-identity-management
Use when non-human identities need the same lifecycle a workforce IAM gives people — issue, authorize, audit, revoke.
88.559 stars · 18 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
vestauth
Use when an agent needs to authenticate as itself to services, without you hand-rolling keys and rotation.
87.5166 stars · 4 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
Browse · Armory