126 results in Evals, Observability, Infrastructure · page 4 of 6
Harness Claude Code Cursor Codex Gemini OpenCode claude-code-statusline Enhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring
95.9 476 stars · 35 forks
Harness claude
statusline observability
phospho Phospho is a text analytics and evaluation platform for LLM apps. It logs sessions, runs clustering, detects failures, and surfaces actionable insights.
95.8 439 stars · 35 forks
Harness claude cursor codex opencode gemini
observability analytics evals
athina-ai Athina AI provides developer-focused LLM monitoring and eval framework: real-time inference logging, automated evals, and regression detection in CI.
94.6 301 stars · 23 forks
Harness claude cursor codex opencode gemini
observability evals logging
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
claude-pace A lightweight Bash + jq statusline for Claude Code that displays rate limit pace delta (burn rate vs. time remaining), 5h/7d usage percentage, context window usage, git branch and diff stats. Compares current consumption rate against time remaining in each rate limit window to indicate whether quota is being used faster or slower than the window allows. Single file with no external dependencies beyond jq.
93.6 229 stars · 19 forks
Harness claude
claude-code status-lines
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
braintrust Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
82.3 27 stars · 12 forks
Harness claude cursor codex opencode gemini
evals experiment-tracking sdk
claudia-statusline High-performance Rust-based statusline for Claude Code with persistent stats tracking, progress bars, and optional cloud sync. Features SQLite-first persistence, git integration, context progress bars, burn rate calculation, XDG-compliant with theme support (dark/light, NO_COLOR).
81.2 36 stars · 5 forks
Harness claude
claude-code status-lines
galileo-evaluate Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.
80.3 22 stars · 11 forks
Harness claude cursor codex opencode gemini
evals hallucination observability sdk
e2b Open-source secure cloud sandboxes (Firecracker microVMs) for running AI-generated code. ~150ms cold start.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure sandbox
modal Serverless cloud platform for running Python functions, containers, and AI workloads with zero infra management.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure serverless
railway Zero-config cloud platform for deploying agent backends, databases, and services from a Git push.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure deploy
vercel Frontend cloud platform with serverless functions and AI SDK integrations for deploying agent-facing UIs.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure deploy
honeycomb Honeycomb's OpenTelemetry-native observability platform: a high-cardinality event store ideal for tracing LLM pipelines and debugging slow agent traces.
75.1 16 stars · 6 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
baserun Baserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and measuring quality.
74.4 16 stars · 5 forks
Harness claude cursor codex opencode gemini
observability tracing testing
humanloop-evals Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.
69.0 12 stars · 3 forks
Harness claude cursor codex opencode gemini
evals human-eval dataset sdk
agentbench Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.
Unranked No signals yet
Harness claude codex
eval benchmark scoring harness
aws-lambda Serverless function-as-a-service platform for event-driven agent compute without provisioning servers.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure serverless
beta9 Open-source serverless GPU container runtime for running AI workloads with fast cold-starts on bare-metal.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure gpu-serverless
blaxel Cloud runtime and control plane for deploying, scaling, and observing production AI agent workloads.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure agent-compute
cloudflare-sandboxes Isolated V8 sandbox environments for multi-tenant agent workloads on top of Cloudflare Workers.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure sandbox
cloudflare-workers Serverless edge-compute platform for deploying agent functions and MCP servers at the network edge.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure edge-compute
coder Self-hosted remote development environment platform for provisioning agent dev workspaces on any cloud.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure dev-environments
Previous Page 4 of 6 Next
Browse · Armory