Armory
Source

Browse

Search and filter by type across the catalog

1,943 results in Workflows, Evals, Skills, Infrastructure · page 4 of 81

InfrastructurePreview

open-operator

Browserbase Open Operator: an open-source Operator-style web agent built on Stagehand; demonstrates full task decomposition, action planning, and evidence collection using the Browserbase cloud.

98.41,953 stars · 325 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserstagehand
No one-command install · SourceDetails
WorkflowsPreview

claude-codepro

A development environment for Claude Code with a spec-driven workflow, TDD enforcement, cross-session memory, semantic search, quality hooks and modular rules. Large, with wide coverage.

98.32,063 stars · 176 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
InfrastructurePreview

notte

Notte open-source web agent environment. It converts browser sessions into a Markov Decision Process with structured observation/action spaces, making browsers first-class RL and LLM agent environments.

98.32,003 stars · 181 forks
Harnessclaudecursorcodexopencodegemini
browserrl
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
EvalsPreview

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructurePreview

e2b-desktop

E2B Desktop Sandbox: a cloud virtual desktop (Ubuntu + VNC) with Python SDK for screenshot, mouse, keyboard, and process control; designed for AI agents that need a full GUI environment in an isolated VM.

98.11,461 stars · 179 forks
Harnessclaudecursorcodexopencodegemini
browsere2b
No one-command install · SourceDetails
SkillsPreview

context-engineering-kit

Hand-crafted collection of advanced context engineering techniques and patterns with minimal token footprint focused on improving agent result quality.

98.01,508 stars · 154 forks
Harnessclaude
skill
No one-command install · SourceDetails
InfrastructurePreview

agent-e

Emergence AI agent-E browser agent: hierarchical LLM-based web automation that uses DOM distillation and action abstraction layers to achieve significantly higher benchmark accuracy than prior browser agents.

98.01,249 stars · 190 forks
Harnessclaudecursorcodexopencodegemini
browseremergence
No one-command install · SourceDetails
WorkflowsPreview

the-ralph-playbook

A detailed guide to the Ralph Wiggum technique for autonomous coding loops, with the reasoning behind it and practical guidelines.

97.91,029 stars · 266 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
EvalsPreview

wandb-weave-evals

Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.

97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
No one-command install · SourceDetails
SkillsPreview

codex-skill

Enables users to prompt codex from claude code. Unlike the raw codex mcp server, this skill infers parameters such as model, reasoning effort, sandboxing from your prompt or asks you to specify them. It also simplifies continuing prior codex sessions so that codex can continue with the prior context.

97.91,424 stars · 109 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
EvalsPreview

langtrace

Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.

97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
No one-command install · SourceDetails
InfrastructurePreview

webvoyager

Research browser agent from Zhejiang University and HKU. It uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).

97.81,127 stars · 124 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserresearch
No one-command install · SourceDetails
WorkflowsPreview

claude-code-documentation-mirror

A mirror of Anthropic's documentation pages for Claude Code, updated every few hours.

97.7983 stars · 138 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
SkillsPreview

claude-codex-settings

A set of plugins for core developer tasks, covering GitHub, Azure, MongoDB, Tavily, Playwright and more. Also works with a few other providers.

97.71,117 stars · 107 forks
Harnessclaude
skill
No one-command install · SourceDetails
SkillsPreview

agentsys

Workflow automation system for Claude with a group of useful plugins, agents, and skills. Automates task-to-production workflows, PR management, code cleanup, performance investigation, drift detection, and multi-agent code review. Includes agnix for linting agent configurations. Built on thousands of lines of code with thousands of tests. Uses deterministic detection (regex, AST) with LLM judgment for efficiency. Used on many production systems.

97.6981 stars · 113 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
InfrastructurePreview

surf-computer-use

E2B Surf: a Stagehand-powered computer-use interface layer for E2B Firecracker sandboxes; connects the act/extract/observe primitives directly to microVM display output for lightweight headless computer use.

97.6856 stars · 141 forks
Harnessclaudecursorcodexopencodegemini
browsere2b
No one-command install · SourceDetails
EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
WorkflowsPreview

awesome-ralph

A curated list of resources about Ralph, the AI coding technique that runs AI coding agents in automated loops until specifications are fulfilled.

97.3918 stars · 74 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
WorkflowsPreview

ralph-wiggum-marketer

A Claude Code plugin that provides an autonomous AI copywriter: research agents gather market knowledge into custom knowledge bases, and a Ralph loop writes the copy.

97.2774 stars · 85 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars · 107 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureExperimental

modal-labs-modal-client

Modal — serverless GPU/CPU containers for running agents, sandboxes and model inference from Python (this is the client SDK).

Contributed by Sentinel

97.0518 stars · 132 forks · 11 mentions
Harnessclaudecodexcursorgeminiopencode
infrastructure
No one-command install · SourceDetails
Browse · Armory