Armory
Source

Browse

Search and filter by type across the catalog

912 results in Workflows, CLIs & Tools, Evals · page 4 of 38

EvalsPreview

lm-evaluation-harness

EleutherAI's unified framework for evaluating language models on hundreds of academic benchmarks.

99.513,860 stars · 3,533 forks
Harnessclaudecursorcodexopencodegemini
evalsacademicbenchmarkharness
No one-command install · SourceDetails
WorkflowsExperimental

anthropic-quickstarts

Offers comprehensive development guides for three distinct AI-powered demo projects with standardized workflows, strict code style guidelines, and containerization instructions.

99.417,588 stars · 3,031 forks
Harnessclaude
awesome-claude-codeofficial-documentation
No one-command install · SourceDetails
CLIs & ToolsExperimental

swe-agent-swe-agent

SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]

Contributed by Sentinel

99.420,190 stars · 2,209 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsExperimental

primeintellect-ai-prime-agent

A self-improving RLM agent for coding workflows and long-running autonomous tasks.

Contributed by Sentinel

99.419,613 stars · 2,139 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsExperimental

workos-auth-md

An open protocol that lets agents register for services on behalf of users — discoverable through a Markdown file at your domain.

Contributed by Sentinel

99.4613 stars · 52 forks · 2 mentions · passed install test
Harnessclaudecodexcursorgeminiopencode
clis-tools
No one-command install · SourceDetails
CLIs & ToolsExperimental

dzhng-deep-research

An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the simplest implementation of a deep research agent - e.g. an agent that can refine its research direction overtime and deep dive into a topic.

Contributed by Sentinel

99.419,626 stars · 1,995 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

deepeval

Open-source LLM evaluation framework with 14+ metrics (hallucination, faithfulness, answer relevancy) and CI support.

99.418,041 stars · 1,890 forks
Harnessclaudecursorcodexopencodegemini
evalsmetricsragci
No one-command install · SourceDetails
CLIs & ToolsExperimental

theagent-net-webagent

A Go framework that stands up a web or business agent from a declarative spec: pick a provider for each slot (model, memory, guardrail, channel, actions over MCP) and get a running agent.

Contributed by Sentinel

99.4558 stars · 16 forks · 1 mention · passed install test
Harnessclaudecodexcursorgeminiopencode
clis-tools
No one-command install · SourceDetails
CLIs & ToolsExperimental

microsoft-skillopt

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.

Contributed by Sentinel

99.416,613 stars · 1,561 forks · 5 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

ragas

Reference-free evaluation of retrieval-augmented generation pipelines; measures faithfulness, answer relevance, and context precision.

99.415,853 stars · 1,727 forks · 1 mention · failed install test
Harnessclaudecursorcodexopencodegemini
evalsragmetrics
No one-command install · SourceDetails
CLIs & ToolsPreview

auto-claude

An autonomous multi-agent coding framework built on the Claude Agent SDK that plans, builds and validates software across the development life cycle, with a kanban-style interface.

99.414,545 stars · 1,917 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsExperimental

fathah-hermes-desktop

Desktop Companion for Hermes Agent

Contributed by Sentinel

99.314,102 stars · 1,602 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsExperimental

ag-ui-protocol-ag-ui

AG-UI: the Agent-User Interaction Protocol. Bring Agents into Frontend Applications.

Contributed by Sentinel

99.315,679 stars · 1,411 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
cli
No one-command install · SourceDetails
CLIs & ToolsExperimental

langchain-ai-openwiki

OpenWiki is a CLI that writes and maintains agent documentation for your codebase.

Contributed by Sentinel

99.315,979 stars · 1,159 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
WorkflowsPreview

claude-code-system-prompts

All parts of Claude Code's system prompt, including builtin tool descriptions, sub agent prompts (Plan/Explore/Task), utility prompts (CLAUDE.md, compact, Bash cmd, security review, agent creation, etc.). Updated for each Claude Code version.

99.312,546 stars · 2,043 forks
Harnessclaude
workflowguide
No one-command install · SourceDetails
CLIs & ToolsPreview

cc-usage

A CLI tool that reads local Claude Code logs and reports usage: cost, token consumption and more, in a dashboard.

99.318,282 stars · 817 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
WorkflowsExperimental

pocketflow

Pocket Flow: 100-line LLM framework that lets agents build agents

99.211,139 stars · 1,217 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agentframeworks
No one-command install · SourceDetails
EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
WorkflowsPreview

claude-code-infrastructure-showcase

An approach to working with Skills that uses hooks to make Claude select and activate the right Skill for the current context. Documented, and adaptable to other projects and workflows.

99.210,016 stars · 1,231 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
CLIs & ToolsExperimental

humanlayer-humanlayer

The best way to get AI coding agents to solve hard problems in complex codebases.

Contributed by Sentinel

99.211,361 stars · 943 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
WorkflowsPreview

harness

A meta-skill that designs domain-specific agent teams, defines specialized agents, and generates the skills they use. Resources are in Korean but can produce high-quality English-language output.

99.28,876 stars · 1,256 forks
Harnessclaude
workflowguideteams
No one-command install · SourceDetails
CLIs & ToolsExperimental

twilio

Unleash the power of Twilio from your command prompt

99.2192 stars · 104 forks · passed install test
comms
No one-command install · SourceDetails
CLIs & ToolsExperimental

artidoro-qlora

QLoRA: Efficient Finetuning of Quantized LLMs

Contributed by Sentinel

99.211,021 stars · 876 forks · 4 mentions · failed install test
Harnessclaudecodexcursorgeminiopencode
clis-tools
No one-command install · SourceDetails
WorkflowsPreview

claude-code-tips

35+ short Claude Code tips covering voice input, system prompt patching, container workflows for risky tasks, conversation cloning, multi-model orchestration with Gemini CLI and more, with demos, working scripts and a plugin.

99.210,008 stars · 809 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
Browse · Armory