Armory
Source

Browse

Search and filter by type across the catalog

107 results in Evals, Memory, Observability · page 5 of 5

EvalsPreview

evals-cookbooks

OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalscookbooktemplatesopenai
No one-command install · SourceDetails
ObservabilityPreview

fiddler-ai

Fiddler AI Observability platform monitors LLM applications for hallucinations, toxicity, bias, and drift, with explainability and alerting for production AI.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitymanagedsafety
No one-command install · SourceDetails
EvalsPreview

gaia-benchmark

GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkagentstool-use
No one-command install · SourceDetails
EvalsPreview

honeyhive

LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingtracingdataset
No one-command install · SourceDetails
ObservabilityPreview

honeyhive

HoneyHive is an AI evaluation and observability platform for tracing agent pipelines, running evaluations, and debugging regressions in production.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilityevalstracing
No one-command install · SourceDetails
ObservabilityPreview

langtrace

Open-source, OpenTelemetry-compliant LLM observability tool by Scale3Labs. It traces calls to all major LLM providers and frameworks with a self-hostable UI.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
ObservabilityPreview

literal-ai

Literal AI is an observability and evaluation platform for conversational AI. It captures multi-step threads, scores responses, and integrates with Chainlit.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitytracingevals
No one-command install · SourceDetails
ObservabilityPreview

lunary

Open-source LLM observability and prompt management platform. It tracks conversations, errors, costs, and user feedback for production AI applications.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilityloggingevals
No one-command install · SourceDetails
ObservabilityPreview

maxim-ai

Maxim AI is an evaluation and observability platform for AI agents. It supports multi-step trace analysis, prompt testing, and production quality monitoring.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilityevalsagents
No one-command install · SourceDetails
ObservabilityPreview

new-relic-ai-monitoring

New Relic AI Monitoring instruments LLM calls end-to-end. It traces model invocations, measures token costs, and surfaces anomalies via the New Relic platform.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitymanagedapm
No one-command install · SourceDetails
EvalsPreview

patronus-ai

Automated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalshallucinationjudgesdk
No one-command install · SourceDetails
Browse · Armory