Armory
Source

Browse

Search and filter by type across the catalog

227 results in CLIs & Tools, Evals, Observability · page 9 of 10

ObservabilityPreview

baserun

Baserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and measuring quality.

74.416 stars · 5 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingtesting
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-rules-doctor

CLI that detects dead `.claude/rules/` files by checking if `paths:` globs actually match files in your repo. Catches silent rule failures where renamed directories or typos in glob patterns cause rules to never apply. Features CI mode (exit 1 on dead rules), JSON output, and verbose mode showing matched files.

70.514 stars · 3 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars · 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install · SourceDetails
CLIs & ToolsExperimental

a2a-ui

UI for Google A2A made using Next.js, TypeScript and Shadcn

67.112 stars · 2 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agentclients
No one-command install · SourceDetails
CLIs & ToolsStable

agentgrid

Use to run many agents in parallel as a visible grid of terminal panes: create an NxM layout, name and monitor panes, broadcast or target prompts, and save/restore whole company configurations.

60.69 stars · 1 fork
Harnessclaudecodexopencode
orchestrationgridtmuxparallelism
No one-command install · SourceDetails
CLIs & ToolsExperimental

ka

AI agent accessible via CLI or network, A2A compatible

32.22 stars
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agentclients
No one-command install · SourceDetails
CLIs & ToolsPreview

armory

Armory: a ranked catalog of 64,000+ open-source agent-harness components with a CLI (`armory search|install|init|rank`), an MCP server and a REST API; every row carries one 0–100 score from public signals.

23.31 fork
Harnessclaudecodexcursorgeminiopencode
registryclimcpcatalog
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-harness

Use to scaffold a complete agent-native harness: the CLAUDE.md, rules, skills, hooks, sub-agents, and memory layout, so a new project starts with the full discipline stack instead of an empty repo.

22.51 star
Harnessclaudecodexcursorgeminiopencode
harnessscaffoldbootstrapclaude-md
No one-command install · SourceDetails
CLIs & ToolsExperimental

anon

Delegated account access for agents: a user grants access without sharing credentials.

3.2failed install test
identity
No one-command install · SourceDetails
CLIs & ToolsPreview

agentswarm

Use to orchestrate a swarm of sub-agents under a CEO pattern when one agent isn't enough but a full visible grid is overkill: break a mission into roles, dispatch them, and coordinate via signal files.

UnrankedNo signals yet
Harnessclaudecodex
orchestrationswarmsub-agentsdispatch
No one-command install · SourceDetails
CLIs & ToolsPreview

agentmoney

Use to track and cap what an agent run costs: meter token/compute spend, set budgets, and surface cost as a first-class signal so an autonomous agent doesn't quietly burn through its limit.

UnrankedNo signals yet
Harnessclaudecodex
costbudgetobservabilitymetering
No one-command install · SourceDetails
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install · SourceDetails
CLIs & ToolsPreview

agent-booster

Use to apply deterministic, zero-LLM code transforms: fast, repeatable edits that don't need a model, so an agent offloads mechanical changes to a cheap tool instead of spending tokens reasoning through them.

UnrankedNo signals yet
Harnessclaudecodexcursor
code-transformdeterministiczero-llmtooling
No one-command install · SourceDetails
CLIs & ToolsPreview

skillsmith

Use to author, test, and share agent skills from the command line: scaffold a SKILL.md, validate its structure, and package it for reuse, turning a one-off procedure into a portable capability.

UnrankedNo signals yet
Harnessclaudecodex
skillsauthoringclisharing
No one-command install · SourceDetails
CLIs & ToolsExperimental

a2a-test

Testing implementation for A2A protocol

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agenttools
No one-command install · SourceDetails
CLIs & ToolsExperimental

a2a-learning

Repository for learning and mastering the A2A protocol

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agenttools
No one-command install · SourceDetails
CLIs & ToolsExperimental

a2a-code-along

Code along examples for A2A

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agenttools
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-chat

An elegant and user-friendly Claude Code chat interface for VS Code.

UnrankedNo signals yet
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-swarm

Launch Claude Code session that is connected to a swarm of Claude Code Agents.

UnrankedNo signals yet
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
ObservabilityPreview

datadog-llm-observability

Datadog's managed LLM Observability product. It traces LLM calls, monitors prompt/completion quality, detects anomalies, and integrates with existing APM.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitymanagedapm
No one-command install · SourceDetails
EvalsPreview

evals-cookbooks

OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalscookbooktemplatesopenai
No one-command install · SourceDetails
ObservabilityPreview

fiddler-ai

Fiddler AI Observability platform monitors LLM applications for hallucinations, toxicity, bias, and drift, with explainability and alerting for production AI.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitymanagedsafety
No one-command install · SourceDetails
EvalsPreview

gaia-benchmark

GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkagentstool-use
No one-command install · SourceDetails
EvalsPreview

honeyhive

LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingtracingdataset
No one-command install · SourceDetails
Browse · Armory