111 results in Evals, Memory, Identity · page 5 of 5
Harness Claude Code Cursor Codex Gemini OpenCode clawsouls-soulspec Use when you want one file to define an agent's persistent identity in a way any compatible runtime can load.
69.7 21 stars · 1 fork
Harness claude codex cursor gemini opencode
cp138-seed identity
humanloop-evals Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.
69.0 12 stars · 3 forks
Harness claude cursor codex opencode gemini
evals human-eval dataset sdk
intelliger-ai-oati Use when an agent's authority, and every action it took under that authority, must be verifiable after the fact.
64.8 13 stars · 1 fork
Harness claude codex cursor gemini opencode
cp138-seed identity
wikimem Use to give an agent a queryable wiki knowledge base when memory should be a navigable knowledge graph, not just a flat log: ingest files, folders, and URLs into a linked vault, then search or ask it in natural language.
63.4 7 stars · 4 forks
Harness claude
memory knowledge-base wiki ingest
chrisdbaldwin-masques Use when an agent needs to put on a temporary role — a bundle of intent, context and lens — for one task and take it off afterwards.
59.8 14 stars
Harness claude codex cursor gemini opencode
cp138-seed identity
imphillip-soultavern Use when you have character cards from the roleplay ecosystem and want them as SOUL.md personas an agent runtime can load.
59.0 13 stars
Harness claude codex cursor gemini opencode
cp138-seed identity
sunilp-aip Use when an agent's identity must carry across both MCP and agent-to-agent calls under one delegable scheme.
58.1 6 stars · 2 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
amirf194-wingfoot Use when your agent keeps getting blocked by sites and you need it to present a verifiable bot identity and be told why it failed.
57.2 7 stars · 1 fork
Harness claude codex cursor gemini opencode
cp138-seed identity
paymanai-sigilum Use when an agent's identity has to leave an auditable trail rather than just gate a request.
54.8 9 stars
Harness claude codex cursor gemini opencode
cp138-seed identity
omnidotdev-persona-json Use when a non-human actor needs a portable, machine-readable identity document that travels with it between systems.
38.4 3 stars
Harness claude codex cursor gemini opencode
cp138-seed identity
agentbench Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.
Unranked No signals yet
Harness claude codex
eval benchmark scoring harness
evals-cookbooks OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals cookbook templates openai
gaia-benchmark GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals benchmark agents tool-use
honeyhive LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals experiment-tracking tracing dataset
patronus-ai Automated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals hallucination judge sdk
Previous Page 5 of 5
Browse · Armory