109 results in Identity, Evals, Observability · page 2 of 5
Harness Claude Code Cursor Codex Gemini OpenCode trulens Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.7 3,530 stars · 335 forks · 1 mention
Harness claude cursor codex opencode gemini
evals rag tracking dashboard
xlang-ai-osworld [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.7 3,117 stars · 530 forks · 7 mentions
Harness claude codex cursor gemini opencode
inspect-ai UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.7 2,683 stars · 685 forks
Harness claude cursor codex opencode gemini
evals safety aisi government
helm Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.7 2,898 stars · 412 forks
Harness claude cursor codex opencode gemini
evals benchmark academic stanford
lighteval Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.
98.6 2,533 stars · 553 forks
Harness claude cursor codex opencode gemini
evals huggingface benchmark lightweight
ccometixline-claude-code-statusline A high-performance Claude Code statusline tool written in Rust with Git integration, usage tracking, interactive TUI configuration, and Claude Code enhancement utilities.
98.6 3,456 stars · 215 forks
Harness claude
claude-code status-lines
openlit OpenLIT is an OpenTelemetry-native LLM observability toolkit with GPU monitoring, cost tracking, and a prompt hub (one-line setup for 20+ providers).
98.6 2,736 stars · 367 forks
Harness claude cursor codex opencode gemini
observability opentelemetry gpu
laminar Laminar is an open-source platform for tracing, evaluating, and labeling LLM and agent pipelines with a TypeScript/Python SDK and a self-hostable backend.
98.6 3,218 stars · 229 forks · 1 mention
Harness claude cursor codex opencode gemini
observability tracing evals
metapriseai-orgkernel Use when every agent action must be traceable to a specific agent instance with its own key, scoped token and tamper-evident log.
98.5 2,699 stars · 247 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
evalplus Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.
98.3 1,819 stars · 208 forks
Harness claude cursor codex opencode gemini
evals code-generation humaneval benchmark
infisical-agent-vault A HTTP credential proxy and vault for AI agents like Claude Code, OpenClaw, Hermes, custom agents + harnesses, and more.
Contributed by Sentinel
98.3 2,171 stars · 141 forks · 3 mentions
Harness claude codex cursor gemini opencode
webarena WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.
98.2 1,592 stars · 249 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents browser benchmark
tau-bench Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
98.1 1,416 stars · 215 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents tool-use benchmark
harbor-framework-terminal-bench Measuring and evolving with the frontier of agent work
Contributed by Sentinel
98.1 588 stars · 434 forks · 13 mentions
Harness claude codex cursor gemini opencode
openinference OpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI.
98.1 1,192 stars · 302 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
langsmith LangChain's platform for tracing, evaluating, and monitoring LLM applications: deep integration with LangChain/LangGraph plus a REST API for any stack.
98.0 1,043 stars · 288 forks
Harness claude cursor codex opencode gemini
observability tracing evals
wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
langtrace Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.8 1,228 stars · 127 forks
Harness claude cursor codex opencode gemini
evals observability opentelemetry tracing
letta-ai-agent-file Use when you need to save, share, version or move a whole agent — its persona, memory and behaviour — as one portable file.
97.8 1,197 stars · 114 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
claude-powerline A vim-style powerline statusline for Claude Code with real-time usage tracking, git integration, custom themes, and more
97.6 1,163 stars · 82 forks
Harness claude
claude-code status-lines
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
unicity-aos-capsule-identity Use when an agent's identity must be persisted as state and assembled into its system prompt at boot, rather than pasted into a prompt by hand.
96.8 8,503 stars · 19 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
Previous Page 2 of 5 Next
Browse · Armory