231 results in Evals, Identity, CLIs & Tools · page 4 of 10
Harness Claude Code Cursor Codex Gemini OpenCode deepeval Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.
99.4 18,041 stars · 1,890 forks
Harness claude cursor codex opencode gemini
evals metrics rag ci
theagent-net-webagent A Go framework that stands up a web or business agent from a declarative spec: pick a provider for each slot (model, memory, guardrail, channel, actions over MCP) and get a running agent.
Contributed by Sentinel
99.4 558 stars · 16 forks · 1 mention · passed install test
Harness claude codex cursor gemini opencode
clis-tools
microsoft-skillopt SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
Contributed by Sentinel
99.4 16,613 stars · 1,561 forks · 5 mentions
Harness claude codex cursor gemini opencode
ragas Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.
99.4 15,853 stars · 1,727 forks · 1 mention · failed install test
Harness claude cursor codex opencode gemini
evals rag metrics
auto-claude An autonomous multi-agent coding framework built on the Claude Agent SDK that plans, builds and validates software across the development life cycle, with a kanban-style interface.
99.4 14,545 stars · 1,917 forks
Harness claude
claude-code tooling
fathah-hermes-desktop Desktop Companion for Hermes Agent
Contributed by Sentinel
99.3 14,102 stars · 1,602 forks · 3 mentions
Harness claude codex cursor gemini opencode
ag-ui-protocol-ag-ui AG-UI: the Agent-User Interaction Protocol. Bring Agents into Frontend Applications.
Contributed by Sentinel
99.3 15,679 stars · 1,411 forks · 6 mentions
Harness claude codex cursor gemini opencode
cli
langchain-ai-openwiki OpenWiki is a CLI that writes and maintains agent documentation for your codebase.
Contributed by Sentinel
99.3 15,979 stars · 1,159 forks · 4 mentions
Harness claude codex cursor gemini opencode
cc-usage A CLI tool that reads local Claude Code logs and reports usage: cost, token consumption and more, in a dashboard.
99.3 18,282 stars · 817 forks
Harness claude
claude-code tooling
phoenix Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.
99.2 11,286 stars · 1,086 forks · 2 mentions
Harness claude cursor codex opencode gemini
evals observability rag agents
humanlayer-humanlayer The best way to get AI coding agents to solve hard problems in complex codebases.
Contributed by Sentinel
99.2 11,361 stars · 943 forks · 4 mentions
Harness claude codex cursor gemini opencode
twilio Unleash the power of Twilio from your command prompt
99.2 192 stars · 104 forks · passed install test
comms
artidoro-qlora QLoRA: Efficient Finetuning of Quantized LLMs
Contributed by Sentinel
99.2 11,021 stars · 876 forks · 4 mentions · failed install test
Harness claude codex cursor gemini opencode
clis-tools
harbor-framework-harbor Framework for evaluating and improving agents
Contributed by Sentinel
99.1 4,867 stars · 1,704 forks · 10 mentions
Harness claude codex cursor gemini opencode
claude-squad Claude Squad is a terminal app that manages multiple Claude Code, Codex (and other local agents including Aider) in separate workspaces, allowing you to work on multiple tasks simultaneously.
99.1 8,536 stars · 621 forks · 1 mention
Harness claude
claude-code tooling
claude-code-usage-monitor A real-time terminal-based tool for monitoring Claude Code token usage. It shows live token consumption, burn rate, and predictions for token depletion. Features include visual progress bars, session-aware analytics, and support for multiple subscription plans.
99.1 8,668 stars · 458 forks
Harness claude
claude-code tooling
swe-bench SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
99.1 5,762 stars · 957 forks · 9 mentions
Harness claude cursor codex opencode gemini
evals code benchmark agents
alexzhang13-rlm General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.
Contributed by Sentinel
99.0 5,644 stars · 905 forks · 3 mentions
Harness claude codex cursor gemini opencode
clawd-on-desk A desktop pet that reacts to your Claude Code sessions in real time: thinking, typing, juggling, sleeping and more.
99.0 6,093 stars · 636 forks
Harness claude
claude-code tooling
gepa-ai-gepa Optimize prompts, code, and more with AI-powered Reflective Optimization
Contributed by Sentinel
99.0 6,346 stars · 533 forks · 5 mentions
Harness claude codex cursor gemini opencode
giskard Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.0 5,838 stars · 542 forks
Harness claude cursor codex opencode gemini
evals safety vulnerability scan
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
claudable Claudable is an open-source web builder that leverages local CLI agents, such as Claude Code and Cursor Agent, to build and deploy products effortlessly.
98.9 4,054 stars · 625 forks
Harness claude
client cli
Previous Page 4 of 10 Next
Browse · Armory