Armory
Source

Browse

Search and filter by type across the catalog

111 results in Memory, Identity, Evals · page 5 of 5

IdentityExperimental

clawsouls-soulspec

Use when you want one file to define an agent's persistent identity in a way any compatible runtime can load.

69.721 stars · 1 fork
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars · 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install · SourceDetails
IdentityExperimental

intelliger-ai-oati

Use when an agent's authority, and every action it took under that authority, must be verifiable after the fact.

64.813 stars · 1 fork
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
MemoryPreview

wikimem

Use to give an agent a queryable wiki knowledge base when memory should be a navigable knowledge graph, not just a flat log: ingest files, folders, and URLs into a linked vault, then search or ask it in natural language.

63.47 stars · 4 forks
Harnessclaude
memoryknowledge-basewikiingest
No one-command install · SourceDetails
IdentityExperimental

chrisdbaldwin-masques

Use when an agent needs to put on a temporary role — a bundle of intent, context and lens — for one task and take it off afterwards.

59.814 stars
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

imphillip-soultavern

Use when you have character cards from the roleplay ecosystem and want them as SOUL.md personas an agent runtime can load.

59.013 stars
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

sunilp-aip

Use when an agent's identity must carry across both MCP and agent-to-agent calls under one delegable scheme.

58.16 stars · 2 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

amirf194-wingfoot

Use when your agent keeps getting blocked by sites and you need it to present a verifiable bot identity and be told why it failed.

57.27 stars · 1 fork
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

paymanai-sigilum

Use when an agent's identity has to leave an auditable trail rather than just gate a request.

54.89 stars
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

omnidotdev-persona-json

Use when a non-human actor needs a portable, machine-readable identity document that travels with it between systems.

38.43 stars
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install · SourceDetails
EvalsPreview

evals-cookbooks

OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalscookbooktemplatesopenai
No one-command install · SourceDetails
EvalsPreview

gaia-benchmark

GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkagentstool-use
No one-command install · SourceDetails
EvalsPreview

honeyhive

LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingtracingdataset
No one-command install · SourceDetails
EvalsPreview

patronus-ai

Automated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalshallucinationjudgesdk
No one-command install · SourceDetails
Browse · Armory