967 results in Hooks, Identity, Evals, Observability, Sub-Agents · page 1 of 41
Harness Claude Code Cursor Codex Gemini OpenCode sentry-llm-monitoring Sentry's error and performance monitoring extended to LLM applications. It captures exceptions, latency, and AI token usage with OpenTelemetry integration.
99.7 44,709 stars · 4,832 forks · 15 mentions
Harness claude cursor codex opencode gemini
observability errors apm
wshobson-agents Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity
Contributed by Sentinel
99.7 39,338 stars · 4,195 forks
Harness claude codex cursor gemini opencode
subagent
mlflow-tracing MLflow's LLM tracing module instruments model calls, agent steps, and tool invocations, storing them alongside experiment runs for reproducibility.
99.6 27,768 stars · 6,246 forks
Harness claude cursor codex opencode gemini
observability tracing experiment-tracking
langfuse Open-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications.
99.6 34,067 stars · 3,678 forks · 10 mentions
Harness claude cursor codex opencode gemini
observability tracing evals
voltagent-awesome-claude-code-subagents A collection of 100+ specialized Claude Code subagents covering a wide range of development use cases
Contributed by Sentinel
99.5 24,793 stars · 2,870 forks
Harness claude codex cursor gemini opencode
subagent
promptfoo CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.
99.5 24,737 stars · 2,255 forks
Harness claude cursor codex opencode gemini
evals red-teaming ci cli
claude-hud A status line for Claude Code that shows context usage, tools, agents, to-dos and more. Highly configurable, and maintained when it was listed.
99.5 27,778 stars · 1,286 forks
Harness claude
statusline observability
openai-evals OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.
99.5 19,509 stars · 3,093 forks
Harness claude cursor codex opencode gemini
evals registry benchmark
lm-evaluation-harness EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.
99.5 13,860 stars · 3,533 forks
Harness claude cursor codex opencode gemini
evals academic benchmark harness
comet-opik Opik by Comet is an open-source LLM evaluation and tracing platform: log traces, run automated evals, create datasets, and track prompt improvements over time.
99.4 22,249 stars · 1,827 forks · 1 mention
Harness claude cursor codex opencode gemini
observability evals tracing
opentelemetry-genai OpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group.
99.4 4,897 stars · 3,851 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
deepeval Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.
99.4 18,041 stars · 1,890 forks
Harness claude cursor codex opencode gemini
evals metrics rag ci
ragas Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.
99.4 15,853 stars · 1,727 forks · 1 mention · failed install test
Harness claude cursor codex opencode gemini
evals rag metrics
portkey-ai-gateway Open-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load balancing.
99.3 12,873 stars · 1,280 forks
Harness claude cursor codex opencode gemini
observability gateway proxy
phoenix Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.
99.2 11,286 stars · 1,086 forks · 2 mentions
Harness claude cursor codex opencode gemini
evals observability rag agents
ccstatusline A highly customizable status line formatter for Claude Code CLI that displays model info, git branch, token usage, and other metrics in your terminal.
99.2 13,035 stars · 579 forks
Harness claude
statusline observability
openllmetry OpenTelemetry-based observability for LLM applications. It auto-instruments OpenAI, Anthropic, LangChain, and 20+ providers with zero code changes.
99.1 7,452 stars · 1,099 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
plannotator Interactive plan review UI that intercepts ExitPlanMode via hooks, letting users visually annotate plans with comments, deletions, and replacements before approving or denying with detailed feedback.
99.1 8,344 stars · 618 forks
Harness claude
hook
swe-bench SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
99.1 5,762 stars · 957 forks · 9 mentions
Harness claude cursor codex opencode gemini
evals code benchmark agents
helicone Open-source LLM observability platform: proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.
99.0 6,123 stars · 662 forks
Harness claude cursor codex opencode gemini
observability proxy logging
giskard Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.0 5,838 stars · 542 forks
Harness claude cursor codex opencode gemini
evals safety vulnerability scan
grafana-tempo Grafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM stacks.
99.0 5,461 stars · 746 forks
Harness claude cursor codex opencode gemini
observability tracing distributed
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
Page 1 of 41 Next
Browse · Armory