Armory
Source

Browse

Search and filter by type across the catalog

260 results in Infrastructure, Hooks, Evals, Observability · page 2 of 11

InfrastructurePreview

nanobrowser

Open-source Chrome extension that runs a multi-agent browser automation system locally. Planner, Navigator, and Validator agents collaborate inside the browser with no external API calls for web tasks.

99.313,713 stars · 1,451 forks
Harnessclaudecursorcodexopencodegemini
browsermulti-agent
No one-command install · SourceDetails
ObservabilityPreview

portkey-ai-gateway

Open-source AI gateway providing a unified API across 100+ LLM providers with built-in observability, request logging, fallbacks, caching, and load balancing.

99.312,873 stars · 1,280 forks
Harnessclaudecursorcodexopencodegemini
observabilitygatewayproxy
No one-command install · SourceDetails
InfrastructurePreview

browserless

Browserless.io headless browser service. It provides a Docker-deployable or cloud-hosted Chrome endpoint with REST and WebSocket APIs for screenshot, PDF, scraping, and Puppeteer/Playwright remote sessions.

99.313,654 stars · 1,039 forks
Harnessclaudecursorcodexopencodegemini
browserbrowserless
No one-command install · SourceDetails
InfrastructurePreview

bytebot

Bytebot open-source computer-use agent: a Docker-based Ubuntu desktop with AI-controlled mouse and keyboard; exposes an HTTP API for agents to send click, type, screenshot, and macro commands.

99.311,084 stars · 1,506 forks
Harnessclaudecursorcodexopencodegemini
browsercomputer-use
No one-command install · SourceDetails
EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
ObservabilityPreview

ccstatusline

A highly customizable status line formatter for Claude Code CLI that displays model info, git branch, token usage, and other metrics in your terminal.

99.213,035 stars · 579 forks
Harnessclaude
statuslineobservability
No one-command install · SourceDetails
ObservabilityPreview

openllmetry

OpenTelemetry-based observability for LLM applications. It auto-instruments OpenAI, Anthropic, LangChain, and 20+ providers with zero code changes.

99.17,452 stars · 1,099 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
InfrastructurePreview

steel-browser

Open-source browser API optimised for AI agents. It provides session management, stealth settings, proxy rotation, and a REST/WebSocket interface on top of Chromium for cloud-scale agent browser access.

99.17,696 stars · 982 forks
Harnessclaudecursorcodexopencodegemini
browsersteel
No one-command install · SourceDetails
HooksPreview

plannotator

Interactive plan review UI that intercepts ExitPlanMode via hooks, letting users visually annotate plans with comments, deletions, and replacements before approving or denying with detailed feedback.

99.18,344 stars · 618 forks
Harnessclaude
hook
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
InfrastructurePreview

microsandbox

Use as the OSS self-hosted sandbox when you need to run agent code on your own infra: libkrun-based microVM isolation with no per-sandbox vendor cost, the escape hatch from a managed runtime at high volume.

99.18,436 stars · 454 forks
Harnessclaudecodex
sandboxself-hostedlibkrunmicrovm
No one-command install · SourceDetails
ObservabilityPreview

helicone

Open-source LLM observability platform: proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.

99.06,123 stars · 662 forks
Harnessclaudecursorcodexopencodegemini
observabilityproxylogging
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
ObservabilityPreview

grafana-tempo

Grafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM stacks.

99.05,461 stars · 746 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingdistributed
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
InfrastructurePreview

claude-managed-agents-selfhost

Use for enterprise hosted-control agents (the agent loop runs at the provider while execution happens on your own infrastructure) when you want a managed control plane but data and code must stay on your machines.

98.93,876 stars · 836 forks
Harnessclaude
managed-agentshosted-controlenterpriseself-hosted-execution
No one-command install · SourceDetails
EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
ObservabilityPreview

pydantic-logfire

Logfire by Pydantic: OpenTelemetry-based structured logging and tracing for Python applications with built-in support for FastAPI, SQLAlchemy, and Anthropic.

98.84,450 stars · 283 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrylogging
No one-command install · SourceDetails
ObservabilityPreview

langwatch

LangWatch provides real-time LLM analytics, guardrails, and evaluation pipelines with a visual studio for monitoring multi-step agent conversations.

98.73,522 stars · 362 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingguardrails
No one-command install · SourceDetails
EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
EvalsPreview

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
No one-command install · SourceDetails
Browse · Armory