Armory
Source

Browse

Search and filter by type across the catalog

979 results in Workflows, Evals, Hooks, Infrastructure, Observability · page 1 of 41

WorkflowsExperimental

n8n-io-n8n

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

Contributed by Sentinel

99.9206,033 stars · 60,896 forks · 26 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructurePreview

browser-use

Python library that makes web browsers accessible to AI agents; built on Playwright and LangChain. Supports multi-tab, vision + accessibility-tree hybrid mode, custom actions, and a self-correcting agent loop.

99.9111,989 stars · 12,312 forks · 11 mentions · passed install test
Harnessclaudecursorcodexopencodegemini
browserbrowser-use
No one-command install · SourceDetails
InfrastructurePreview

daytona

Secure and elastic sandboxes for running AI-generated code. The public repository stopped updating in June 2026, when development moved to a private codebase.

99.971,846 stars · 5,651 forks · 9 mentions · passed install test
Harnessclaudecursorcodexopencodegemini
infrastructuredev-environments
No one-command install · SourceDetails
InfrastructureExperimental

ggml-org-llama-cpp

LLM inference in C/C++

Contributed by Sentinel

99.9126,728 stars · 22,647 forks · 5 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureExperimental

vllm-project-vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

Contributed by Sentinel

99.990,743 stars · 21,577 forks · 15 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
WorkflowsExperimental

karpathy-autoresearch

AI agents running research on single-GPU nanochat training automatically

Contributed by Sentinel

99.995,090 stars · 13,383 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureStable

e2b-sandbox

Use as the default runtime when an agent must execute untrusted code or commands. Firecracker microVMs with ~150ms cold start give each run an isolated, disposable computer.

99.813,635 stars · 1,015 forks · 10 mentions · passed install test
Harnessclaudecodex
sandboxruntimefirecrackermicrovm
No one-command install · SourceDetails
WorkflowsPreview

learn-claude-code

An analysis of how coding agents like Claude Code are designed, which breaks an agent into its basic parts and rebuilds it with minimal code: a rudimentary agent with skills, sub-agents and a to-do list in a few hundred lines of Python.

99.875,849 stars · 12,224 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
InfrastructureExperimental

diegosouzapw-omniroute

Never stop coding. Free MIT AI gateway: one endpoint, 352 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by 550+ contributors

Contributed by Sentinel

99.860,005 stars · 8,342 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
InfrastructureExperimental

sgl-project-sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

Contributed by Sentinel

99.733,202 stars · 8,472 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
ObservabilityPreview

sentry-llm-monitoring

Sentry's error and performance monitoring extended to LLM applications. It captures exceptions, latency, and AI token usage with OpenTelemetry integration.

99.744,709 stars · 4,832 forks · 15 mentions
Harnessclaudecursorcodexopencodegemini
observabilityerrorsapm
No one-command install · SourceDetails
InfrastructureStable

stripe-agent-toolkit

Use as the payments rail when an agent should earn or spend money in code (create customers, prices, payment links, and usage-based billing): the infrastructure behind an agent that funds its own compute.

99.71,785 stars · 329 forks · passed install test
Harnessclaudecodex
paymentsstripebillingfinancial-rails
No one-command install · SourceDetails
ObservabilityPreview

mlflow-tracing

MLflow's LLM tracing module instruments model calls, agent steps, and tool invocations, storing them alongside experiment runs for reproducibility.

99.627,768 stars · 6,246 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingexperiment-tracking
No one-command install · SourceDetails
ObservabilityPreview

langfuse

Open-source LLM engineering platform with traces, evals, prompt management, and datasets for debugging and improving LLM applications.

99.634,067 stars · 3,678 forks · 10 mentions
Harnessclaudecursorcodexopencodegemini
observabilitytracingevals
No one-command install · SourceDetails
WorkflowsExperimental

openai-symphony

Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding agents.

Contributed by Sentinel

99.526,991 stars · 2,771 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
WorkflowsExperimental

a2a-protocol-github-repository

Google's official repository for A2A protocol

99.525,941 stars · 2,633 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agentofficial-resources
No one-command install · SourceDetails
EvalsPreview

promptfoo

CLI and library for testing, comparing, and red-teaming LLM prompts and agents with assertions and CI integration.

99.524,737 stars · 2,255 forks
Harnessclaudecursorcodexopencodegemini
evalsred-teamingcicli
No one-command install · SourceDetails
ObservabilityPreview

claude-hud

A status line for Claude Code that shows context usage, tools, agents, to-dos and more. Highly configurable, and maintained when it was listed.

99.527,778 stars · 1,286 forks
Harnessclaude
statuslineobservability
No one-command install · SourceDetails
InfrastructurePreview

skyvern

Open-source agent platform that automates browser-based workflows using LLMs and computer vision. It identifies interactive elements via screenshots, handles CAPTCHAs, and supports complex multi-step form flows.

99.522,907 stars · 2,152 forks
Harnessclaudecursorcodexopencodegemini
browserskyvern
No one-command install · SourceDetails
EvalsPreview

openai-evals

OpenAI's official framework for evaluating LLMs and LLM-powered systems, with a registry of community eval sets.

99.519,509 stars · 3,093 forks
Harnessclaudecursorcodexopencodegemini
evalsregistrybenchmark
No one-command install · SourceDetails
EvalsPreview

lm-evaluation-harness

EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.

99.513,860 stars · 3,533 forks
Harnessclaudecursorcodexopencodegemini
evalsacademicbenchmarkharness
No one-command install · SourceDetails
InfrastructureStable

browserbase-bb

Use when an agent must operate the live web (navigate, act, and extract on real pages) via a cloud browser driven by act/extract/observe primitives, with a local-Chromium escape hatch using the same code.

99.424,125 stars · 1,664 forks · 3 mentions
Harnessclaudecodex
browserweb-automationstagehandbrowserbase
No one-command install · SourceDetails
WorkflowsExperimental

anthropic-quickstarts

Offers comprehensive development guides for three distinct AI-powered demo projects with standardized workflows, strict code style guidelines, and containerization instructions.

99.417,588 stars · 3,031 forks
Harnessclaude
awesome-claude-codeofficial-documentation
No one-command install · SourceDetails
ObservabilityPreview

comet-opik

Opik by Comet is an open-source LLM evaluation and tracing platform: log traces, run automated evals, create datasets, and track prompt improvements over time.

99.422,249 stars · 1,827 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
observabilityevalstracing
No one-command install · SourceDetails
Browse · Armory