Armory
Source

Browse

Search and filter by type across the catalog

231 results in Identity, CLIs & Tools, Evals · page 4 of 10

EvalsPreview

deepeval

Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.

99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
No one-command install · SourceDetails
CLIs & ToolsExperimental

theagent-net-webagent

A Go framework that stands up a web or business agent from a declarative spec: pick a provider for each slot (model, memory, guardrail, channel, actions over MCP) and get a running agent.

Contributed by Sentinel

99.4558 stars · 16 forks · 1 mention · passed install test
Harnessclaudecodexcursorgeminiopencode
clis-tools
No one-command install · SourceDetails
CLIs & ToolsExperimental

microsoft-skillopt

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.

Contributed by Sentinel

99.416,613 stars · 1,561 forks · 5 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

ragas

Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.

99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
No one-command install · SourceDetails
CLIs & ToolsPreview

auto-claude

An autonomous multi-agent coding framework built on the Claude Agent SDK that plans, builds and validates software across the development life cycle, with a kanban-style interface.

99.414,545 stars · 1,917 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsExperimental

fathah-hermes-desktop

Desktop Companion for Hermes Agent

Contributed by Sentinel

99.314,102 stars · 1,602 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsExperimental

ag-ui-protocol-ag-ui

AG-UI: the Agent-User Interaction Protocol. Bring Agents into Frontend Applications.

Contributed by Sentinel

99.315,679 stars · 1,411 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
cli
No one-command install · SourceDetails
CLIs & ToolsExperimental

langchain-ai-openwiki

OpenWiki is a CLI that writes and maintains agent documentation for your codebase.

Contributed by Sentinel

99.315,979 stars · 1,159 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsPreview

cc-usage

A CLI tool that reads local Claude Code logs and reports usage: cost, token consumption and more, in a dashboard.

99.318,282 stars · 817 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
CLIs & ToolsExperimental

humanlayer-humanlayer

The best way to get AI coding agents to solve hard problems in complex codebases.

Contributed by Sentinel

99.211,361 stars · 943 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsExperimental

twilio

Unleash the power of Twilio from your command prompt

99.2192 stars · 104 forks · passed install test
comms
No one-command install · SourceDetails
CLIs & ToolsExperimental

artidoro-qlora

QLoRA: Efficient Finetuning of Quantized LLMs

Contributed by Sentinel

99.211,021 stars · 876 forks · 4 mentions · failed install test
Harnessclaudecodexcursorgeminiopencode
clis-tools
No one-command install · SourceDetails
CLIs & ToolsExperimental

harbor-framework-harbor

Framework for evaluating and improving agents

Contributed by Sentinel

99.14,867 stars · 1,704 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-squad

Claude Squad is a terminal app that manages multiple Claude Code, Codex (and other local agents including Aider) in separate workspaces, allowing you to work on multiple tasks simultaneously.

99.18,536 stars · 621 forks · 1 mention
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-usage-monitor

A real-time terminal-based tool for monitoring Claude Code token usage. It shows live token consumption, burn rate, and predictions for token depletion. Features include visual progress bars, session-aware analytics, and support for multiple subscription plans.

99.18,668 stars · 458 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
CLIs & ToolsExperimental

alexzhang13-rlm

General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.

Contributed by Sentinel

99.05,644 stars · 905 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsPreview

clawd-on-desk

A desktop pet that reacts to your Claude Code sessions in real time: thinking, typing, juggling, sleeping and more.

99.06,093 stars · 636 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsExperimental

gepa-ai-gepa

Optimize prompts, code, and more with AI-powered Reflective Optimization

Contributed by Sentinel

99.06,346 stars · 533 forks · 5 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
CLIs & ToolsPreview

claudable

Claudable is an open-source web builder that leverages local CLI agents, such as Claude Code and Cursor Agent, to build and deploy products effortlessly.

98.94,054 stars · 625 forks
Harnessclaude
clientcli
No one-command install · SourceDetails
Browse · Armory