570 results in CLAUDE.md / Rules, Observability, Evals, Identity · page 1 of 24
Harness Claude Code Cursor Codex Gemini OpenCode karpathy-coding-discipline Drop into CLAUDE.md/AGENTS.md as the first behavior norm a coding agent ingrains: think before coding, prefer the simplest solution, change only what you own, and execute toward the stated goal.
99.9 209,417 stars · 21,315 forks
Harness claude codex cursor gemini opencode
discipline coding constitution behavior-norm
sentry-llm-monitoring Sentry's error and performance monitoring extended to LLM applications. It captures exceptions, latency, and AI token usage with OpenTelemetry integration.
99.7 44,709 stars · 4,832 forks · 15 mentions
Harness claude cursor codex opencode gemini
observability errors apm
mlflow-tracing MLflow's LLM tracing module instruments model calls, agent steps, and tool invocations, storing them alongside experiment runs for reproducibility.
99.6 27,768 stars · 6,246 forks
Harness claude cursor codex opencode gemini
observability tracing experiment-tracking
langfuse Open-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications.
99.6 34,067 stars · 3,678 forks · 10 mentions
Harness claude cursor codex opencode gemini
observability tracing evals
CLAUDE.md / Rules Experimental humanlayer-12-factor-agents What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?
Contributed by Sentinel
99.5 25,642 stars · 1,951 forks · 4 mentions
Harness claude codex cursor gemini opencode
promptfoo CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.
99.5 24,737 stars · 2,255 forks
Harness claude cursor codex opencode gemini
evals red-teaming ci cli
claude-hud A status line for Claude Code that shows context usage, tools, agents, to-dos and more. Highly configurable, and maintained when it was listed.
99.5 27,778 stars · 1,286 forks
Harness claude
statusline observability
openai-evals OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.
99.5 19,509 stars · 3,093 forks
Harness claude cursor codex opencode gemini
evals registry benchmark
lm-evaluation-harness EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.
99.5 13,860 stars · 3,533 forks
Harness claude cursor codex opencode gemini
evals academic benchmark harness
comet-opik Opik by Comet is an open-source LLM evaluation and tracing platform: log traces, run automated evals, create datasets, and track prompt improvements over time.
99.4 22,249 stars · 1,827 forks · 1 mention
Harness claude cursor codex opencode gemini
observability evals tracing
opentelemetry-genai OpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group.
99.4 4,897 stars · 3,851 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
deepeval Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.
99.4 18,041 stars · 1,890 forks
Harness claude cursor codex opencode gemini
evals metrics rag ci
ragas Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.
99.4 15,853 stars · 1,727 forks · 1 mention · failed install test
Harness claude cursor codex opencode gemini
evals rag metrics
portkey-ai-gateway Open-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load balancing.
99.3 12,873 stars · 1,280 forks
Harness claude cursor codex opencode gemini
observability gateway proxy
phoenix Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.
99.2 11,286 stars · 1,086 forks · 2 mentions
Harness claude cursor codex opencode gemini
evals observability rag agents
ccstatusline A highly customizable status line formatter for Claude Code CLI that displays model info, git branch, token usage, and other metrics in your terminal.
99.2 13,035 stars · 579 forks
Harness claude
statusline observability
openllmetry OpenTelemetry-based observability for LLM applications. It auto-instruments OpenAI, Anthropic, LangChain, and 20+ providers with zero code changes.
99.1 7,452 stars · 1,099 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
swe-bench SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
99.1 5,762 stars · 957 forks · 9 mentions
Harness claude cursor codex opencode gemini
evals code benchmark agents
helicone Open-source LLM observability platform: proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.
99.0 6,123 stars · 662 forks
Harness claude cursor codex opencode gemini
observability proxy logging
giskard Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.0 5,838 stars · 542 forks
Harness claude cursor codex opencode gemini
evals safety vulnerability scan
grafana-tempo Grafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM stacks.
99.0 5,461 stars · 746 forks
Harness claude cursor codex opencode gemini
observability tracing distributed
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
big-bench Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.8 3,247 stars · 618 forks
Harness claude cursor codex opencode gemini
evals benchmark google academic
Page 1 of 24 Next
Browse · Armory