Armory
Source

Browse

Search and filter by type across the catalog

226 results in Infrastructure, Hooks, Evals · page 3 of 10

InfrastructurePreview

webvoyager

Research browser agent from Zhejiang University and HKU. It uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).

97.81,127 stars · 124 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserresearch
No one-command install · SourceDetails
InfrastructurePreview

surf-computer-use

E2B Surf: a Stagehand-powered computer-use interface layer for E2B Firecracker sandboxes; connects the act/extract/observe primitives directly to microVM display output for lightweight headless computer use.

97.6856 stars · 141 forks
Harnessclaudecursorcodexopencodegemini
browsere2b
No one-command install · SourceDetails
EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars · 107 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureExperimental

modal-labs-modal-client

Modal — serverless GPU/CPU containers for running agents, sandboxes and model inference from Python (this is the client SDK).

Contributed by Sentinel

97.0518 stars · 132 forks · 11 mentions
Harnessclaudecodexcursorgeminiopencode
infrastructure
No one-command install · SourceDetails
EvalsExperimental

stonybrooknlp-appworld

🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars · 78 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
HooksPreview

claude-code-hook-comms-hcom

A lightweight CLI tool for real-time communication between Claude Code sub-agents through hooks, with @-mention targeting, a live monitoring dashboard and no dependencies. It was described as unstable when it was listed.

96.6470 stars · 70 forks
Harnessclaude
claude-codehooks
No one-command install · SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars · 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install · SourceDetails
HooksPreview

claude-hooks

A TypeScript-based system for configuring and customizing Claude Code hooks with a powerful and flexible interface.

95.2389 stars · 26 forks
Harnessclaude
claude-codehooks
No one-command install · SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars · 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars · 62 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install · SourceDetails
HooksPreview

dippy

Auto-approve safe bash commands using AST-based parsing, while prompting for destructive operations. Solves permission fatigue without disabling safety. Supports Claude Code, Gemini CLI, and Cursor.

94.0243 stars · 21 forks
Harnessclaude
hook
No one-command install · SourceDetails
HooksPreview

cc-notify

CCNotify provides desktop notifications for Claude Code, alerting you to input needs or task completion, with one-click jumps back to VS Code and task duration display.

93.9216 stars · 23 forks
Harnessclaude
claude-codehooks
No one-command install · SourceDetails
HooksPreview

typescript-quality-hooks

A quality-check hook for Node.js TypeScript projects: TypeScript compilation, ESLint auto-fixing and Prettier formatting, with SHA256 config caching that keeps validation under 5 ms during editing.

92.4178 stars · 14 forks
Harnessclaude
claude-codehooks
No one-command install · SourceDetails
HooksPreview

cchooks

A lightweight Python SDK with a clean API and good documentation; simplifies the process of writing hooks and integrating them into your codebase, providing a nice abstraction over the JSON configuration files.

90.7130 stars · 11 forks
Harnessclaude
hook
No one-command install · SourceDetails
EvalsPreview

vellum-evals

Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.

90.582 stars · 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
No one-command install · SourceDetails
HooksPreview

claudio

A small library that plays OS-native sounds for Claude Code events through hooks.

89.2113 stars · 8 forks
Harnessclaude
claude-codehooks
No one-command install · SourceDetails
HooksPreview

claude-code-hooks-sdk

A Laravel-inspired PHP SDK for building Claude Code hook responses with a clean, fluent API. This SDK makes it easy to create structured JSON responses for Claude Code hooks using an expressive, chainable interface.

87.068 stars · 8 forks
Harnessclaude
hook
No one-command install · SourceDetails
EvalsPreview

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

82.327 stars · 12 forks
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingsdk
No one-command install · SourceDetails
EvalsPreview

galileo-evaluate

Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.

80.322 stars · 11 forks
Harnessclaudecursorcodexopencodegemini
evalshallucinationobservabilitysdk
No one-command install · SourceDetails
InfrastructurePreview

e2b

Open-source secure cloud sandboxes (Firecracker microVMs) for running AI-generated code. ~150ms cold start.

80.0passed install test
Harnessclaudecursorcodexopencodegemini
infrastructuresandbox
No one-command install · SourceDetails
InfrastructurePreview

modal

Serverless cloud platform for running Python functions, containers, and AI workloads with zero infra management.

80.0passed install test
Harnessclaudecursorcodexopencodegemini
infrastructureserverless
No one-command install · SourceDetails
InfrastructurePreview

railway

Zero-config cloud platform for deploying agent backends, databases, and services from a Git push.

80.0passed install test
Harnessclaudecursorcodexopencodegemini
infrastructuredeploy
No one-command install · SourceDetails
InfrastructurePreview

vercel

Frontend cloud platform with serverless functions and AI SDK integrations for deploying agent-facing UIs.

80.0passed install test
Harnessclaudecursorcodexopencodegemini
infrastructuredeploy
No one-command install · SourceDetails
Browse · Armory