231 results in Identity, CLIs & Tools, Evals · page 5 of 10
Harness Claude Code Cursor Codex Gemini OpenCode big-bench Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.8 3,247 stars · 618 forks
Harness claude cursor codex opencode gemini
evals benchmark google academic
trulens Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.7 3,530 stars · 335 forks · 1 mention
Harness claude cursor codex opencode gemini
evals rag tracking dashboard
claude-devtools A desktop app that shows your Claude Code sessions by reading their logs: context use per turn across categories, compaction, sub-agent execution trees and custom notification triggers.
98.7 3,893 stars · 298 forks
Harness claude
claude-code tooling
xlang-ai-osworld [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.7 3,117 stars · 530 forks · 7 mentions
Harness claude codex cursor gemini opencode
inspect-ai UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.7 2,683 stars · 685 forks
Harness claude cursor codex opencode gemini
evals safety aisi government
letta-ai-letta-code Stateful agents that are like people, with memory, identity, and the ability to learn and adapt
Contributed by Sentinel
98.7 3,185 stars · 385 forks · 6 mentions
Harness claude codex cursor gemini opencode
helm Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.7 2,898 stars · 412 forks
Harness claude cursor codex opencode gemini
evals benchmark academic stanford
lighteval Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.
98.6 2,533 stars · 553 forks
Harness claude cursor codex opencode gemini
evals huggingface benchmark lightweight
metapriseai-orgkernel Use when every agent action must be traceable to a specific agent instance with its own key, scoped token and tamper-evident log.
98.5 2,699 stars · 247 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
crystal A full-fledged desktop application for orchestrating, monitoring, and interacting with Claude Code agents.
98.5 3,114 stars · 197 forks
Harness claude
client cli
omnara A command center for AI agents that syncs Claude Code sessions across terminal, web, and mobile. Allows for remote monitoring, human-in-the-loop interaction, and team collaboration.
98.5 2,780 stars · 213 forks · 4 mentions
Harness claude
claude-code alternative-clients
primeintellect-ai-prime-rl Agentic RL Training at Scale
Contributed by Sentinel
98.5 2,007 stars · 421 forks · 3 mentions
Harness claude codex cursor gemini opencode
tweakcc Command-line tool to customize your Claude Code styling.
98.4 2,476 stars · 199 forks
Harness claude
claude-code tooling
evalplus Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.
98.3 1,819 stars · 208 forks
Harness claude cursor codex opencode gemini
evals code-generation humaneval benchmark
infisical-agent-vault A HTTP credential proxy and vault for AI agents like Claude Code, OpenClaw, Hermes, custom agents + harnesses, and more.
Contributed by Sentinel
98.3 2,171 stars · 141 forks · 3 mentions
Harness claude codex cursor gemini opencode
webarena WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.
98.2 1,592 stars · 249 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents browser benchmark
claude-code-tools Well-crafted toolset for session continuity, featuring skills/commands to avoid compaction and recover context across sessions with cross-agent handoff between Claude Code and Codex CLI. Includes a fast Rust/Tantivy-powered full-text session search (TUI for humans, skill/CLI for agents), tmux-cli skill + command for interacting with scripts and CLI agents, and safety hooks to block dangerous commands.
98.2 1,989 stars · 132 forks
Harness claude
claude-code tooling
cc-sessions An opinionated approach to productive development with Claude Code
98.1 1,551 stars · 191 forks
Harness claude
claude-code tooling
tau-bench Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
98.1 1,416 stars · 215 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents tool-use benchmark
harbor-framework-terminal-bench Measuring and evolving with the frontier of agent work
Contributed by Sentinel
98.1 588 stars · 434 forks · 13 mentions
Harness claude codex cursor gemini opencode
claude-code-ide-el claude-code-ide.el integrates Claude Code with Emacs, like Anthropic’s VS Code/IntelliJ extensions. It shows ediff-based code suggestions, pulls LSP/flymake/flycheck diagnostics, and tracks buffer context. It adds an extensible MCP tool support for symbol refs/defs, project metadata, and tree-sitter AST queries.
98.0 1,658 stars · 112 forks
Harness claude
claude-code tooling
rulesync A Node.js CLI tool that automatically generates configs (rules, ignore files, MCP servers, commands, and subagents) for various AI coding agents. Rulesync can convert configs between Claude Code and other AI agents in both directions.
98.0 1,373 stars · 143 forks
Harness claude
tooling cli
claude-code-nvim A seamless integration between Claude Code AI assistant and Neovim.
97.9 2,096 stars · 70 forks
Harness claude
tooling cli ide-integrations
wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
Previous Page 5 of 10 Next
Browse · Armory