227 results in Observability, Evals, CLIs & Tools · page 5 of 10
Harness Claude Code Cursor Codex Gemini OpenCode claude-code-usage-monitor A real-time terminal-based tool for monitoring Claude Code token usage. It shows live token consumption, burn rate, and predictions for token depletion. Features include visual progress bars, session-aware analytics, and support for multiple subscription plans.
99.1 8,668 stars · 458 forks
Harness claude
claude-code tooling
swe-bench SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
99.1 5,762 stars · 957 forks · 9 mentions
Harness claude cursor codex opencode gemini
evals code benchmark agents
helicone Open-source LLM observability platform: proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.
99.0 6,123 stars · 662 forks
Harness claude cursor codex opencode gemini
observability proxy logging
alexzhang13-rlm General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.
Contributed by Sentinel
99.0 5,644 stars · 905 forks · 3 mentions
Harness claude codex cursor gemini opencode
clawd-on-desk A desktop pet that reacts to your Claude Code sessions in real time: thinking, typing, juggling, sleeping and more.
99.0 6,093 stars · 636 forks
Harness claude
claude-code tooling
gepa-ai-gepa Optimize prompts, code, and more with AI-powered Reflective Optimization
Contributed by Sentinel
99.0 6,346 stars · 533 forks · 5 mentions
Harness claude codex cursor gemini opencode
giskard Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.0 5,838 stars · 542 forks
Harness claude cursor codex opencode gemini
evals safety vulnerability scan
grafana-tempo Grafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM stacks.
99.0 5,461 stars · 746 forks
Harness claude cursor codex opencode gemini
observability tracing distributed
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
claudable Claudable is an open-source web builder that leverages local CLI agents, such as Claude Code and Cursor Agent, to build and deploy products effortlessly.
98.9 4,054 stars · 625 forks
Harness claude
client cli
big-bench Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.8 3,247 stars · 618 forks
Harness claude cursor codex opencode gemini
evals benchmark google academic
pydantic-logfire Logfire by Pydantic: OpenTelemetry-based structured logging and tracing for Python applications with built-in support for FastAPI, SQLAlchemy, and Anthropic.
98.8 4,450 stars · 283 forks
Harness claude cursor codex opencode gemini
observability opentelemetry logging
langwatch LangWatch provides real-time LLM analytics, guardrails, and evaluation pipelines with a visual studio for monitoring multi-step agent conversations.
98.7 3,522 stars · 362 forks
Harness claude cursor codex opencode gemini
observability tracing guardrails
trulens Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.7 3,530 stars · 335 forks · 1 mention
Harness claude cursor codex opencode gemini
evals rag tracking dashboard
claude-devtools A desktop app that shows your Claude Code sessions by reading their logs: context use per turn across categories, compaction, sub-agent execution trees and custom notification triggers.
98.7 3,893 stars · 298 forks
Harness claude
claude-code tooling
xlang-ai-osworld [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.7 3,117 stars · 530 forks · 7 mentions
Harness claude codex cursor gemini opencode
inspect-ai UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.7 2,683 stars · 685 forks
Harness claude cursor codex opencode gemini
evals safety aisi government
letta-ai-letta-code Stateful agents that are like people, with memory, identity, and the ability to learn and adapt
Contributed by Sentinel
98.7 3,185 stars · 385 forks · 6 mentions
Harness claude codex cursor gemini opencode
helm Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.7 2,898 stars · 412 forks
Harness claude cursor codex opencode gemini
evals benchmark academic stanford
lighteval Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.
98.6 2,533 stars · 553 forks
Harness claude cursor codex opencode gemini
evals huggingface benchmark lightweight
ccometixline-claude-code-statusline A high-performance Claude Code statusline tool written in Rust with Git integration, usage tracking, interactive TUI configuration, and Claude Code enhancement utilities.
98.6 3,456 stars · 215 forks
Harness claude
claude-code status-lines
openlit OpenLIT is an OpenTelemetry-native LLM observability toolkit with GPU monitoring, cost tracking, and a prompt hub (one-line setup for 20+ providers).
98.6 2,736 stars · 367 forks
Harness claude cursor codex opencode gemini
observability opentelemetry gpu
laminar Laminar is an open-source platform for tracing, evaluating, and labeling LLM and agent pipelines with a TypeScript/Python SDK and a self-hostable backend.
98.6 3,218 stars · 229 forks · 1 mention
Harness claude cursor codex opencode gemini
observability tracing evals
Previous Page 5 of 10 Next
Browse · Armory