264 results in Infrastructure, Identity, Evals, Hooks · page 3 of 11
Harness Claude Code Cursor Codex Gemini OpenCode wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
langtrace Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.8 1,228 stars · 127 forks
Harness claude cursor codex opencode gemini
evals observability opentelemetry tracing
letta-ai-agent-file Use when you need to save, share, version or move a whole agent — its persona, memory and behaviour — as one portable file.
97.8 1,197 stars · 114 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
webvoyager Research browser agent from Zhejiang University and HKU. It uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).
97.8 1,127 stars · 124 forks · 1 mention
Harness claude cursor codex opencode gemini
browser research
surf-computer-use E2B Surf: a Stagehand-powered computer-use interface layer for E2B Firecracker sandboxes; connects the act/extract/observe primitives directly to microVM display output for lightweight headless computer use.
97.6 856 stars · 141 forks
Harness claude cursor codex opencode gemini
browser e2b
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
Infrastructure Experimental modal-labs-modal-client Modal — serverless GPU/CPU containers for running agents, sandboxes and model inference from Python (this is the client SDK).
Contributed by Sentinel
97.0 518 stars · 132 forks · 11 mentions
Harness claude codex cursor gemini opencode
infrastructure
unicity-aos-capsule-identity Use when an agent's identity must be persisted as state and assembled into its system prompt at boot, rather than pasted into a prompt by hand.
96.8 8,503 stars · 19 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
claude-code-hook-comms-hcom A lightweight CLI tool for real-time communication between Claude Code sub-agents through hooks, with @-mention targeting, a live monitoring dashboard and no dependencies. It was described as unstable when it was listed.
96.6 470 stars · 70 forks
Harness claude
claude-code hooks
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
agntcy-oasf Use when agents must describe themselves to other systems in a common schema so they can be catalogued, discovered and verified.
95.7 332 stars · 47 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
claude-hooks A TypeScript-based system for configuring and customizing Claude Code hooks with a powerful and flexible interface.
95.2 389 stars · 26 forks
Harness claude
claude-code hooks
telagod-code-abyss Use when you want a coding agent to have a consistent, composable personality and voice across Claude Code, Codex, Gemini CLI and OpenClaw.
94.6 239 stars · 32 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
agntcy-dir Use when agents and multi-agent systems need to announce themselves and be found across organisations rather than hardcoded.
94.6 184 stars · 55 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
dippy Auto-approve safe bash commands using AST-based parsing, while prompting for destructive operations. Solves permission fatigue without disabling safety. Supports Claude Code, Gemini CLI, and Cursor.
94.0 243 stars · 21 forks
Harness claude
hook
character-card-spec-v2 Use when reading or writing the character-card files that the roleplay-agent ecosystem actually ships, including the PNG-embedded variant.
93.9 188 stars · 29 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
cc-notify CCNotify provides desktop notifications for Claude Code, alerting you to input needs or task completion, with one-click jumps back to VS Code and task duration display.
93.9 216 stars · 23 forks
Harness claude
claude-code hooks
cloudflare-web-bot-auth Use when a website has to be able to tell that a request really came from your agent and not from someone impersonating it.
93.8 157 stars · 40 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
highflame-ai-zeroid Use when a fleet of autonomous agents needs issued identities with a lifecycle — created, rotated and revoked.
92.7 163 stars · 18 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
typescript-quality-hooks A quality-check hook for Node.js TypeScript projects: TypeScript compilation, ESLint auto-fixing and Prettier formatting, with SHA256 config caching that keeps validation under 5 ms during editing.
92.4 178 stars · 14 forks
Harness claude
claude-code hooks
Previous Page 3 of 11 Next
Browse · Armory