Armory
Source

Browse

Search and filter by type across the catalog

1,325 results in Evals, Skills, CLIs & Tools · page 5 of 56

CLIs & ToolsExperimental

humanlayer-humanlayer

The best way to get AI coding agents to solve hard problems in complex codebases.

Contributed by Sentinel

99.211,361 stars · 943 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
SkillsExperimental

agricidaniel-claude-ads

Claude-first paid-media operations skill for Claude Code across 12 ad platforms (Google, Meta, YouTube, LinkedIn, TikTok, Microsoft, Apple, Amazon, Reddit, Pinterest, Snapchat, X): source-grounded audits, deterministic scoring, versioned JSON reports, and capability-gated account changes.

Contributed by Sentinel

99.28,670 stars · 1,293 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsExperimental

twilio

Unleash the power of Twilio from your command prompt

99.2192 stars · 104 forks · passed install test
comms
No one-command install · SourceDetails
CLIs & ToolsExperimental

artidoro-qlora

QLoRA: Efficient Finetuning of Quantized LLMs

Contributed by Sentinel

99.211,021 stars · 876 forks · 4 mentions · failed install test
Harnessclaudecodexcursorgeminiopencode
clis-tools
No one-command install · SourceDetails
CLIs & ToolsExperimental

harbor-framework-harbor

Framework for evaluating and improving agents

Contributed by Sentinel

99.14,867 stars · 1,704 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-squad

Claude Squad is a terminal app that manages multiple Claude Code, Codex (and other local agents including Aider) in separate workspaces, allowing you to work on multiple tasks simultaneously.

99.18,536 stars · 621 forks · 1 mention
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-usage-monitor

A real-time terminal-based tool for monitoring Claude Code token usage. It shows live token consumption, burn rate, and predictions for token depletion. Features include visual progress bars, session-aware analytics, and support for multiple subscription plans.

99.18,668 stars · 458 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
SkillsPreview

trail-of-bits-security-skills

A very professional collection of over a dozen security-focused skills for code auditing and vulnerability detection. Includes skills for static analysis with CodeQL and Semgrep, variant analysis across codebases, fix verification, and differential code review.

99.06,939 stars · 597 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
CLIs & ToolsExperimental

alexzhang13-rlm

General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.

Contributed by Sentinel

99.05,644 stars · 905 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsPreview

clawd-on-desk

A desktop pet that reacts to your Claude Code sessions in real time: thinking, typing, juggling, sleeping and more.

99.06,093 stars · 636 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsExperimental

gepa-ai-gepa

Optimize prompts, code, and more with AI-powered Reflective Optimization

Contributed by Sentinel

99.06,346 stars · 533 forks · 5 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
SkillsPreview

codebase-to-course

A Claude Code skill that turns any codebase into an interactive single-page HTML course for non-technical readers.

99.05,505 stars · 551 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
CLIs & ToolsPreview

claudable

Claudable is an open-source web builder that leverages local CLI agents, such as Claude Code and Cursor Agent, to build and deploy products effortlessly.

98.94,054 stars · 625 forks
Harnessclaude
clientcli
No one-command install · SourceDetails
EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-devtools

A desktop app that shows your Claude Code sessions by reading their logs: context use per turn across categories, compaction, sub-agent execution trees and custom notification triggers.

98.73,893 stars · 298 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
CLIs & ToolsExperimental

letta-ai-letta-code

Stateful agents that are like people, with memory, identity, and the ability to learn and adapt

Contributed by Sentinel

98.73,185 stars · 385 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
No one-command install · SourceDetails
Browse · Armory