Armory
Source

Browse

Search and filter by type across the catalog

831 results in Evals, Memory, Observability, Sub-Agents · page 2 of 35

MemoryExperimental

semantica-agi-semantica

Use when an agent's stored context needs provenance — you must be able to say where a remembered fact came from.

99.313,484 stars · 1,525 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

nevamind-ai-memu

Use when one person's memory should follow them across several different agents.

99.314,387 stars · 1,063 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
ObservabilityPreview

portkey-ai-gateway

Open-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load balancing.

99.312,873 stars · 1,280 forks
Harnessclaudecursorcodexopencodegemini
observabilitygatewayproxy
No one-command install · SourceDetails
MemoryExperimental

evermind-ai-everos

Use when you want the agent's memory to be plain Markdown on your own disk rather than a hosted database.

99.212,757 stars · 911 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

memtensor-memos

Use when memory should persist across tasks and be reused, not just retrieved once per conversation.

99.211,599 stars · 1,060 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
ObservabilityPreview

ccstatusline

A highly customizable status line formatter for Claude Code CLI that displays model info, git branch, token usage, and other metrics in your terminal.

99.213,035 stars · 579 forks
Harnessclaude
statuslineobservability
No one-command install · SourceDetails
ObservabilityPreview

openllmetry

OpenTelemetry-based observability for LLM applications. It auto-instruments OpenAI, Anthropic, LangChain, and 20+ providers with zero code changes.

99.17,452 stars · 1,099 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
MemoryExperimental

plastic-labs-honcho

Memory library for building stateful agents

Contributed by Sentinel

99.16,980 stars · 865 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
ObservabilityPreview

helicone

Open-source LLM observability platform: proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.

99.06,123 stars · 662 forks
Harnessclaudecursorcodexopencodegemini
observabilityproxylogging
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
ObservabilityPreview

grafana-tempo

Grafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM stacks.

99.05,461 stars · 746 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingdistributed
No one-command install · SourceDetails
MemoryExperimental

getzep-zep

Examples, framework integrations and tools for Zep Cloud, Zep's hosted agent memory service; the repository says it is not the product itself. The open-source knowledge-graph engine behind Zep is Graphiti (getzep-graphiti).

Contributed by Sentinel

99.04,882 stars · 651 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
MemoryExperimental

caviraoss-openmemory

Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.

Contributed by Sentinel

98.94,478 stars · 504 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
ObservabilityPreview

pydantic-logfire

Logfire by Pydantic: OpenTelemetry-based structured logging and tracing for Python applications with built-in support for FastAPI, SQLAlchemy, and Anthropic.

98.84,450 stars · 283 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrylogging
No one-command install · SourceDetails
MemoryExperimental

flowelement-m-flow

Use when recall should follow associations between memories rather than nearest-neighbour similarity alone.

98.84,497 stars · 255 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
ObservabilityPreview

langwatch

LangWatch provides real-time LLM analytics, guardrails, and evaluation pipelines with a visual studio for monitoring multi-step agent conversations.

98.73,522 stars · 362 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingguardrails
No one-command install · SourceDetails
EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
Browse · Armory