145 results in Identity, Memory, Observability, Evals · page 2 of 7
Harness Claude Code Cursor Codex Gemini OpenCode portkey-ai-gateway Open-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load balancing.
99.3 12,873 stars · 1,280 forks
Harness claude cursor codex opencode gemini
observability gateway proxy
evermind-ai-everos Use when you want the agent's memory to be plain Markdown on your own disk rather than a hosted database.
99.2 12,757 stars · 911 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
memtensor-memos Use when memory should persist across tasks and be reused, not just retrieved once per conversation.
99.2 11,599 stars · 1,060 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
phoenix Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.
99.2 11,286 stars · 1,086 forks · 2 mentions
Harness claude cursor codex opencode gemini
evals observability rag agents
ccstatusline A highly customizable status line formatter for Claude Code CLI that displays model info, git branch, token usage, and other metrics in your terminal.
99.2 13,035 stars · 579 forks
Harness claude
statusline observability
openllmetry OpenTelemetry-based observability for LLM applications. It auto-instruments OpenAI, Anthropic, LangChain, and 20+ providers with zero code changes.
99.1 7,452 stars · 1,099 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
plastic-labs-honcho Memory library for building stateful agents
Contributed by Sentinel
99.1 6,980 stars · 865 forks · 6 mentions
Harness claude codex cursor gemini opencode
memory
swe-bench SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
99.1 5,762 stars · 957 forks · 9 mentions
Harness claude cursor codex opencode gemini
evals code benchmark agents
helicone Open-source LLM observability platform: proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.
99.0 6,123 stars · 662 forks
Harness claude cursor codex opencode gemini
observability proxy logging
giskard Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.0 5,838 stars · 542 forks
Harness claude cursor codex opencode gemini
evals safety vulnerability scan
grafana-tempo Grafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM stacks.
99.0 5,461 stars · 746 forks
Harness claude cursor codex opencode gemini
observability tracing distributed
getzep-zep Examples, framework integrations and tools for Zep Cloud, Zep's hosted agent memory service; the repository says it is not the product itself. The open-source knowledge-graph engine behind Zep is Graphiti (getzep-graphiti).
Contributed by Sentinel
99.0 4,882 stars · 651 forks · 7 mentions
Harness claude codex cursor gemini opencode
memory
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
caviraoss-openmemory Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
Contributed by Sentinel
98.9 4,478 stars · 504 forks
Harness claude codex cursor gemini opencode
memory
big-bench Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.8 3,247 stars · 618 forks
Harness claude cursor codex opencode gemini
evals benchmark google academic
pydantic-logfire Logfire by Pydantic: OpenTelemetry-based structured logging and tracing for Python applications with built-in support for FastAPI, SQLAlchemy, and Anthropic.
98.8 4,450 stars · 283 forks
Harness claude cursor codex opencode gemini
observability opentelemetry logging
flowelement-m-flow Use when recall should follow associations between memories rather than nearest-neighbour similarity alone.
98.8 4,497 stars · 255 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
langwatch LangWatch provides real-time LLM analytics, guardrails, and evaluation pipelines with a visual studio for monitoring multi-step agent conversations.
98.7 3,522 stars · 362 forks
Harness claude cursor codex opencode gemini
observability tracing guardrails
trulens Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.7 3,530 stars · 335 forks · 1 mention
Harness claude cursor codex opencode gemini
evals rag tracking dashboard
xlang-ai-osworld [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.7 3,117 stars · 530 forks · 7 mentions
Harness claude codex cursor gemini opencode
inspect-ai UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.7 2,683 stars · 685 forks
Harness claude cursor codex opencode gemini
evals safety aisi government
mirix-ai-mirix Use when what the agent should remember is what actually happened on screen, consolidated into structured memories.
98.7 3,440 stars · 270 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
helm Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.7 2,898 stars · 412 forks
Harness claude cursor codex opencode gemini
evals benchmark academic stanford
Previous Page 2 of 7 Next
Browse · Armory