790 results in Observability, Evals, Workflows · page 3 of 33
Harness Claude Code Cursor Codex Gemini OpenCode openlit OpenLIT is an OpenTelemetry-native LLM observability toolkit with GPU monitoring, cost tracking, and a prompt hub (one-line setup for 20+ providers).
98.6 2,736 stars · 367 forks
Harness claude cursor codex opencode gemini
observability opentelemetry gpu
laminar Laminar is an open-source platform for tracing, evaluating, and labeling LLM and agent pipelines with a TypeScript/Python SDK and a self-hostable backend.
98.6 3,218 stars · 229 forks · 1 mention
Harness claude cursor codex opencode gemini
observability tracing evals
claude-codepro A development environment for Claude Code with a spec-driven workflow, TDD enforcement, cross-session memory, semantic search, quality hooks and modular rules. Large, with wide coverage.
98.3 2,063 stars · 176 forks
Harness claude
claude-code workflows-knowledge-guides
evalplus Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.
98.3 1,819 stars · 208 forks
Harness claude cursor codex opencode gemini
evals code-generation humaneval benchmark
webarena WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.
98.2 1,592 stars · 249 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents browser benchmark
tau-bench Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
98.1 1,416 stars · 215 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents tool-use benchmark
harbor-framework-terminal-bench Measuring and evolving with the frontier of agent work
Contributed by Sentinel
98.1 588 stars · 434 forks · 13 mentions
Harness claude codex cursor gemini opencode
openinference OpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI.
98.1 1,192 stars · 302 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
langsmith LangChain's platform for tracing, evaluating, and monitoring LLM applications: deep integration with LangChain/LangGraph plus a REST API for any stack.
98.0 1,043 stars · 288 forks
Harness claude cursor codex opencode gemini
observability tracing evals
the-ralph-playbook A detailed guide to the Ralph Wiggum technique for autonomous coding loops, with the reasoning behind it and practical guidelines.
97.9 1,029 stars · 266 forks
Harness claude
claude-code workflows-knowledge-guides
wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
langtrace Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.8 1,228 stars · 127 forks
Harness claude cursor codex opencode gemini
evals observability opentelemetry tracing
claude-code-documentation-mirror A mirror of Anthropic's documentation pages for Claude Code, updated every few hours.
97.7 983 stars · 138 forks
Harness claude
claude-code workflows-knowledge-guides
claude-powerline A vim-style powerline statusline for Claude Code with real-time usage tracking, git integration, custom themes, and more
97.6 1,163 stars · 82 forks
Harness claude
claude-code status-lines
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
awesome-ralph A curated list of resources about Ralph, the AI coding technique that runs AI coding agents in automated loops until specifications are fulfilled.
97.3 918 stars · 74 forks
Harness claude
claude-code workflows-knowledge-guides
ralph-wiggum-marketer A Claude Code plugin that provides an autonomous AI copywriter: research agents gather market knowledge into custom knowledge bases, and a Ralph loop writes the copy.
97.2 774 stars · 85 forks
Harness claude
claude-code workflows-knowledge-guides
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
claude-code-repos-index An index of 75+ Claude Code repositories by one author, covering content management, system design, deep research, IoT, agentic workflows, server management and personal health.
96.7 536 stars · 69 forks
Harness claude
claude-code workflows-knowledge-guides
claudopro-directory Well-crafted, wide selection of Claude Code hooks, slash commands, subagent files, and more, covering a range of specialized tasks and workflows. Better resources than your average "Claude-template-for-everything" site.
96.7 300 stars · 144 forks
Harness claude
workflow guide
simone A broader project management workflow for Claude Code that encompasses not just a set of commands, but a system of documents, guidelines, and processes to facilitate project planning and execution.
96.4 558 stars · 46 forks
Harness claude
claude-code workflows-knowledge-guides
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
claude-code-statusline Enhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring
95.9 476 stars · 35 forks
Harness claude
statusline observability
Previous Page 3 of 33 Next
Browse · Armory