Leaderboard
Scored on public signals; components with none are listed as Unranked · Formula
Component
Domain
Vertical
| Rank | Score | Component | Description | Evidence | Last commit | Install |
|---|---|---|---|---|---|---|
| 1 | Unranked | bullmq-specialistskill · back-end | BullMQ expert for Redis-backed job queues, background processing, and reliable async execution in Node.js/TypeScript applications. Use when: bullmq… | No signals yet | No commit datelisted | armory install bullmq-specialist --cli claude |
| 2 | Unranked | evaluating-llms-harnessskill · back-end | Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models… | No signals yet | No commit datelisted | armory install evaluating-llms-harness --cli claude |
| 3 | Unranked | mcp-builderskill · back-end | Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools… | No signals yet | No commit datelisted | armory install mcp-builder --cli claude |
| 4 | Unranked | nemo-evaluator-sdkskill · back-end | Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable… | No signals yet | No commit datelisted | armory install nemo-evaluator-sdk --cli claude |
| 5 | Unranked | verl-rl-trainingskill · back-end | Provides guidance for training LLMs with reinforcement learning using verl (Volcano Engine RL). Use when implementing RLHF, GRPO, PPO, or other RL… | No signals yet | No commit datelisted | armory install verl-rl-training --cli claude |
Score colour shows how many signals stand behind it, never how good it is: amber, three or more; dimmer amber, two; grey, one. The Evidence column names them.
Stars, forks and last commit are as GitHub reported them when Armory last read each repository: for most, or later. A repository may have changed since.