Armory
Source

Browse

Search and filter by type across the catalog

145 results in Evals, Observability, Memory, Identity · page 6 of 7

IdentityExperimental

clawsouls

Use when you would rather install a ready-made agent persona than write one, and want it to work across OpenClaw, Claude Code and Cursor.

72.821 stars · 2 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

kanoniv-agent-auth

Use when authority must be handed down a chain of agents and each hop stays provable.

72.715 stars · 4 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

clawsouls-soulspec

Use when you want one file to define an agent's persistent identity in a way any compatible runtime can load.

69.721 stars · 1 fork
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars · 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install · SourceDetails
IdentityExperimental

intelliger-ai-oati

Use when an agent's authority, and every action it took under that authority, must be verifiable after the fact.

64.813 stars · 1 fork
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
MemoryPreview

wikimem

Use to give an agent a queryable wiki knowledge base when memory should be a navigable knowledge graph, not just a flat log: ingest files, folders, and URLs into a linked vault, then search or ask it in natural language.

63.47 stars · 4 forks
Harnessclaude
memoryknowledge-basewikiingest
No one-command install · SourceDetails
IdentityExperimental

chrisdbaldwin-masques

Use when an agent needs to put on a temporary role — a bundle of intent, context and lens — for one task and take it off afterwards.

59.814 stars
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

imphillip-soultavern

Use when you have character cards from the roleplay ecosystem and want them as SOUL.md personas an agent runtime can load.

59.013 stars
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

sunilp-aip

Use when an agent's identity must carry across both MCP and agent-to-agent calls under one delegable scheme.

58.16 stars · 2 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

amirf194-wingfoot

Use when your agent keeps getting blocked by sites and you need it to present a verifiable bot identity and be told why it failed.

57.27 stars · 1 fork
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

paymanai-sigilum

Use when an agent's identity has to leave an auditable trail rather than just gate a request.

54.89 stars
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

omnidotdev-persona-json

Use when a non-human actor needs a portable, machine-readable identity document that travels with it between systems.

38.43 stars
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install · SourceDetails
ObservabilityPreview

datadog-llm-observability

Datadog's managed LLM Observability product. It traces LLM calls, monitors prompt/completion quality, detects anomalies, and integrates with existing APM.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitymanagedapm
No one-command install · SourceDetails
EvalsPreview

evals-cookbooks

OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalscookbooktemplatesopenai
No one-command install · SourceDetails
ObservabilityPreview

fiddler-ai

Fiddler AI Observability platform monitors LLM applications for hallucinations, toxicity, bias, and drift, with explainability and alerting for production AI.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitymanagedsafety
No one-command install · SourceDetails
EvalsPreview

gaia-benchmark

GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkagentstool-use
No one-command install · SourceDetails
EvalsPreview

honeyhive

LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingtracingdataset
No one-command install · SourceDetails
ObservabilityPreview

honeyhive

HoneyHive is an AI evaluation and observability platform for tracing agent pipelines, running evaluations, and debugging regressions in production.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilityevalstracing
No one-command install · SourceDetails
ObservabilityPreview

langtrace

Open-source, OpenTelemetry-compliant LLM observability tool by Scale3Labs. It traces calls to all major LLM providers and frameworks with a self-hostable UI.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
ObservabilityPreview

literal-ai

Literal AI is an observability and evaluation platform for conversational AI. It captures multi-step threads, scores responses, and integrates with Chainlit.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitytracingevals
No one-command install · SourceDetails
ObservabilityPreview

lunary

Open-source LLM observability and prompt management platform. It tracks conversations, errors, costs, and user feedback for production AI applications.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilityloggingevals
No one-command install · SourceDetails
ObservabilityPreview

maxim-ai

Maxim AI is an evaluation and observability platform for AI agents. It supports multi-step trace analysis, prompt testing, and production quality monitoring.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilityevalsagents
No one-command install · SourceDetails
ObservabilityPreview

new-relic-ai-monitoring

New Relic AI Monitoring instruments LLM calls end-to-end. It traces model invocations, measures token costs, and surfaces anomalies via the New Relic platform.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitymanagedapm
No one-command install · SourceDetails
Browse · Armory