Armory
Source

Browse

Search and filter by type across the catalog

229 results in Memory, CLIs & Tools, Evals · page 9 of 10

CLIs & ToolsPreview

claude-session-restore

Efficiently restore context from previous Claude Code sessions by analyzing session files and git history. Features multi-factor data collection across numerous Claude Code capacities with time-based filtering. Uses tail-based parsing for efficient handling of large session files up to 2GB. Includes both a CLI tool for manual analysis and a Claude Code skill for automatic session restoration.

83.536 stars · 9 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsPreview

cc-tools

High-performance Go implementation of Claude Code hooks and utilities. Provides smart linting, testing, and statusline generation with minimal overhead.

83.350 stars · 5 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
MemoryExperimental

kyros-ai

Use when the agent's memory needs to resolve its own contradictions and forget on a schedule.

82.596 stars · 2 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
CLIs & ToolsExperimental

pr-agent-mentoring

Pull request agent for code reviews and mentoring

82.54 stars · 26 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agenttools
No one-command install · SourceDetails
EvalsPreview

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

82.327 stars · 12 forks
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingsdk
No one-command install · SourceDetails
CLIs & ToolsExperimental

a2a-cli

Command-line client for interacting with A2A servers

81.427 stars · 9 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agentclients
No one-command install · SourceDetails
EvalsPreview

galileo-evaluate

Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.

80.322 stars · 11 forks
Harnessclaudecursorcodexopencodegemini
evalshallucinationobservabilitysdk
No one-command install · SourceDetails
CLIs & ToolsPreview

agentdial

Use to give an agent a universal identity and reachable channels, a stable handle the agent presents across surfaces, when agents need to be addressable and authenticated rather than anonymous processes.

80.0passed install test
Harnessclaudecodex
identitychannelsagent-protocoladdressing
No one-command install · SourceDetails
CLIs & ToolsExperimental

bb

Browse CLI (`bb`): Stagehand-powered browser automation from the terminal.

80.0passed install test
browser
No one-command install · SourceDetails
CLIs & ToolsExperimental

mailtm

Disposable email inboxes with a free REST API: receive mail without an account.

80.0passed install test
identity
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-rules-doctor

CLI that detects dead `.claude/rules/` files by checking if `paths:` globs actually match files in your repo. Catches silent rule failures where renamed directories or typos in glob patterns cause rules to never apply. Features CI mode (exit 1 on dead rules), JSON output, and verbose mode showing matched files.

70.514 stars · 3 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars · 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install · SourceDetails
CLIs & ToolsExperimental

a2a-ui

UI for Google A2A made using Next.js, TypeScript and Shadcn

67.112 stars · 2 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agentclients
No one-command install · SourceDetails
MemoryPreview

wikimem

Use to give an agent a queryable wiki knowledge base when memory should be a navigable knowledge graph, not just a flat log: ingest files, folders, and URLs into a linked vault, then search or ask it in natural language.

63.47 stars · 4 forks
Harnessclaude
memoryknowledge-basewikiingest
No one-command install · SourceDetails
CLIs & ToolsStable

agentgrid

Use to run many agents in parallel as a visible grid of terminal panes: create an NxM layout, name and monitor panes, broadcast or target prompts, and save/restore whole company configurations.

60.69 stars · 1 fork
Harnessclaudecodexopencode
orchestrationgridtmuxparallelism
No one-command install · SourceDetails
CLIs & ToolsExperimental

ka

AI agent accessible via CLI or network, A2A compatible

32.22 stars
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agentclients
No one-command install · SourceDetails
CLIs & ToolsPreview

armory

Armory: a ranked catalog of 64,000+ open-source agent-harness components with a CLI (`armory search|install|init|rank`), an MCP server and a REST API; every row carries one 0–100 score from public signals.

23.31 fork
Harnessclaudecodexcursorgeminiopencode
registryclimcpcatalog
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-harness

Use to scaffold a complete agent-native harness: the CLAUDE.md, rules, skills, hooks, sub-agents, and memory layout, so a new project starts with the full discipline stack instead of an empty repo.

22.51 star
Harnessclaudecodexcursorgeminiopencode
harnessscaffoldbootstrapclaude-md
No one-command install · SourceDetails
CLIs & ToolsExperimental

anon

Delegated account access for agents: a user grants access without sharing credentials.

3.2failed install test
identity
No one-command install · SourceDetails
CLIs & ToolsPreview

agentswarm

Use to orchestrate a swarm of sub-agents under a CEO pattern when one agent isn't enough but a full visible grid is overkill: break a mission into roles, dispatch them, and coordinate via signal files.

UnrankedNo signals yet
Harnessclaudecodex
orchestrationswarmsub-agentsdispatch
No one-command install · SourceDetails
CLIs & ToolsPreview

agentmoney

Use to track and cap what an agent run costs: meter token/compute spend, set budgets, and surface cost as a first-class signal so an autonomous agent doesn't quietly burn through its limit.

UnrankedNo signals yet
Harnessclaudecodex
costbudgetobservabilitymetering
No one-command install · SourceDetails
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install · SourceDetails
CLIs & ToolsPreview

agent-booster

Use to apply deterministic, zero-LLM code transforms: fast, repeatable edits that don't need a model, so an agent offloads mechanical changes to a cheap tool instead of spending tokens reasoning through them.

UnrankedNo signals yet
Harnessclaudecodexcursor
code-transformdeterministiczero-llmtooling
No one-command install · SourceDetails
CLIs & ToolsPreview

skillsmith

Use to author, test, and share agent skills from the command line: scaffold a SKILL.md, validate its structure, and package it for reuse, turning a one-off procedure into a portable capability.

UnrankedNo signals yet
Harnessclaudecodex
skillsauthoringclisharing
No one-command install · SourceDetails
Browse · Armory