1,686 results in Evals, Identity, Observability, Sub-Agents, Hooks, Workflows · page 1 of 71
Harness Claude Code Cursor Codex Gemini OpenCode n8n-io-n8n Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
Contributed by Sentinel
99.9 206,033 stars · 60,896 forks · 26 mentions
Harness claude codex cursor gemini opencode
karpathy-autoresearch AI agents running research on single-GPU nanochat training automatically
Contributed by Sentinel
99.9 95,090 stars · 13,383 forks · 7 mentions
Harness claude codex cursor gemini opencode
learn-claude-code An analysis of how coding agents like Claude Code are designed, which breaks an agent into its basic parts and rebuilds it with minimal code: a rudimentary agent with skills, sub-agents and a to-do list in a few hundred lines of Python.
99.8 75,849 stars · 12,224 forks
Harness claude
claude-code workflows-knowledge-guides
sentry-llm-monitoring Sentry's error and performance monitoring extended to LLM applications. It captures exceptions, latency, and AI token usage with OpenTelemetry integration.
99.7 44,709 stars · 4,832 forks · 15 mentions
Harness claude cursor codex opencode gemini
observability errors apm
wshobson-agents Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity
Contributed by Sentinel
99.7 39,338 stars · 4,195 forks
Harness claude codex cursor gemini opencode
subagent
mlflow-tracing MLflow's LLM tracing module instruments model calls, agent steps, and tool invocations, storing them alongside experiment runs for reproducibility.
99.6 27,768 stars · 6,246 forks
Harness claude cursor codex opencode gemini
observability tracing experiment-tracking
langfuse Open-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications.
99.6 34,067 stars · 3,678 forks · 10 mentions
Harness claude cursor codex opencode gemini
observability tracing evals
openai-symphony Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding agents.
Contributed by Sentinel
99.5 26,991 stars · 2,771 forks · 4 mentions
Harness claude codex cursor gemini opencode
a2a-protocol-github-repository Google's official repository for A2A protocol
99.5 25,941 stars · 2,633 forks
Harness claude cursor codex opencode gemini
a2a agent-to-agent official-resources
voltagent-awesome-claude-code-subagents A collection of 100+ specialized Claude Code subagents covering a wide range of development use cases
Contributed by Sentinel
99.5 24,793 stars · 2,870 forks
Harness claude codex cursor gemini opencode
subagent
promptfoo CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.
99.5 24,737 stars · 2,255 forks
Harness claude cursor codex opencode gemini
evals red-teaming ci cli
claude-hud A status line for Claude Code that shows context usage, tools, agents, to-dos and more. Highly configurable, and maintained when it was listed.
99.5 27,778 stars · 1,286 forks
Harness claude
statusline observability
openai-evals OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.
99.5 19,509 stars · 3,093 forks
Harness claude cursor codex opencode gemini
evals registry benchmark
lm-evaluation-harness EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.
99.5 13,860 stars · 3,533 forks
Harness claude cursor codex opencode gemini
evals academic benchmark harness
anthropic-quickstarts Offers comprehensive development guides for three distinct AI-powered demo projects with standardized workflows, strict code style guidelines, and containerization instructions.
99.4 17,588 stars · 3,031 forks
Harness claude
awesome-claude-code official-documentation
comet-opik Opik by Comet is an open-source LLM evaluation and tracing platform: log traces, run automated evals, create datasets, and track prompt improvements over time.
99.4 22,249 stars · 1,827 forks · 1 mention
Harness claude cursor codex opencode gemini
observability evals tracing
opentelemetry-genai OpenTelemetry semantic conventions and instrumentation for GenAI/LLM spans, traces, and metrics via the GenAI semconv working group.
99.4 4,897 stars · 3,851 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
deepeval Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.
99.4 18,041 stars · 1,890 forks
Harness claude cursor codex opencode gemini
evals metrics rag ci
ragas Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.
99.4 15,853 stars · 1,727 forks · 1 mention · failed install test
Harness claude cursor codex opencode gemini
evals rag metrics
claude-code-system-prompts All parts of Claude Code's system prompt, including builtin tool descriptions, sub agent prompts (Plan/Explore/Task), utility prompts (CLAUDE.md, compact, Bash cmd, security review, agent creation, etc.). Updated for each Claude Code version.
99.3 12,546 stars · 2,043 forks
Harness claude
workflow guide
portkey-ai-gateway Open-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load balancing.
99.3 12,873 stars · 1,280 forks
Harness claude cursor codex opencode gemini
observability gateway proxy
pocketflow Pocket Flow: 100-line LLM framework that lets agents build agents
99.2 11,139 stars · 1,217 forks
Harness claude cursor codex opencode gemini
a2a agent-to-agent frameworks
phoenix Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.
99.2 11,286 stars · 1,086 forks · 2 mentions
Harness claude cursor codex opencode gemini
evals observability rag agents
claude-code-infrastructure-showcase An approach to working with Skills that uses hooks to make Claude select and activate the right Skill for the current context. Documented, and adaptable to other projects and workflows.
99.2 10,016 stars · 1,231 forks
Harness claude
claude-code workflows-knowledge-guides
Page 1 of 71 Next
Browse · Armory