Armory
Source

Browse

Search and filter by type across the catalog

1,515 results in Workflows, Sub-Agents, Observability, Identity · page 37 of 64

Sub-AgentsPreview

model-evaluator

AI model evaluation and benchmarking specialist. Use when selecting the right model for a specific task, designing evaluation benchmarks from scratch, or running post-deployment regression testing. Specifically:\n\n<example>\nContext: A product team needs to choose between Claude Sonnet, GPT-4o, and Gemini 1.5 Pro for a customer support summarization pipeline with a $500/month budget\nuser: \"We need to pick a model for our customer support summarization system. We process 50k tickets/month and need under 2s latency.\"\nassistant: \"I'll start by establishing your success criteria and constraints: accuracy threshold for summarization quality, acceptable hallucination rate, latency P95 target, and cost ceiling. Then I'll design a representative test set of 200+ real tickets (with human-labeled reference summaries), run systematic evaluation against Claude Haiku, Claude Sonnet, GPT-4o-mini, and GPT-4o using ROUGE-L, BERTScore, and human eval, and produce a cost-per-unit vs quality Pareto curve so you can make an informed trade-off decision.\"\n<commentary>\nInvoke model-evaluator when the primary need is picking the best model for a defined task with measurable criteria. Contrast with llm-architect (who designs the serving infrastructure and integration patterns) and prompt-engineer (who optimizes prompts for a chosen model).\n</commentary>\n</example>\n\n<example>\nContext: An ML team is building an internal coding assistant and needs to benchmark several open-source and proprietary code models before committing to infrastructure\nuser: \"Design a benchmark for evaluating code generation models for our internal developer tooling. We care about Python, TypeScript, and SQL.\"\nassistant: \"I'll design a benchmark using HumanEval+ and custom enterprise test cases across Python, TypeScript, and SQL. Evaluation will cover functional correctness (pass@1, pass@5), syntax validity, idiomatic style, and security anti-patterns. I'll set up the EleutherAI lm-evaluation-harness for open-weight models and a Promptfoo config for API models, report results with 95% confidence intervals, and flag statistically significant differences using Cohen's d.\"\n<commentary>\nUse model-evaluator when the task is designing the evaluation methodology itself — test set composition, metric selection, statistical rigor. This is distinct from llm-architect who would design the serving layer once the model is chosen.\n</commentary>\n</example>\n\n<example>\nContext: A deployed LLM pipeline has started producing lower quality outputs after a model provider silently updated their model weights\nuser: \"Our summarization quality scores dropped 8% last week. We think the model changed. How do we confirm and decide whether to roll back or switch models?\"\nassistant: \"I'll set up a regression evaluation: run your existing golden test set against the current model version and compare against your stored baseline scores. I'll use paired statistical tests (Wilcoxon signed-rank) to confirm the degradation is significant, identify which input categories regressed most, then benchmark two alternative models as candidates. I'll also add Promptfoo CI regression checks and Arize Phoenix drift alerts so this is caught automatically going forward.\"\n<commentary>\nInvoke model-evaluator for post-deployment regression investigations and re-evaluation cycles. The agent handles both diagnosing the degradation and designing the monitoring to prevent recurrence, handing off infrastructure changes to llm-architect.\n</commentary>\n</example>

UnrankedNo signals yet
Harnessclaude
ai-specialistssubagents
Details
WorkflowsPreview

model-route

Recommend the best model tier for the current task based on complexity, risk, and budget.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
commandworkflow
Details
Sub-AgentsPreview

modernization

Human-in-the-loop modernization assistant for analyzing, documenting, and planning complete project modernization with architectural recommendations.

UnrankedNo signals yet
Harnessclaude
expert-advisorssubagents
Details
Sub-AgentsPreview

monday-bug-fixer

Elite bug-fixing agent that enriches task context from Monday.com platform data. Gathers related items, docs, comments, epics, and requirements to deliver production-quality fixes with comprehensive PRs.

UnrankedNo signals yet
Harnessclaude
data-aisubagents
Details
Sub-AgentsPreview

mongodb-performance-advisor

Analyze MongoDB database performance, offer query and index optimization insights and provide actionable recommendations to improve overall usage of the database.

UnrankedNo signals yet
Harnessclaude
programming-languagessubagents
Details
WorkflowsPreview

monitor-setup

Monitoring and Observability Setup

UnrankedNo signals yet
Harnessclaude
agentsworkflows
Details
WorkflowsPreview

monitor-setup-2

Monitoring and Observability Setup

UnrankedNo signals yet
Harnessclaude
commandstools
Details
Sub-AgentsPreview

monitoring-specialist

Monitoring and observability infrastructure specialist. Use PROACTIVELY for metrics collection, alerting systems, log aggregation, distributed tracing, SLA monitoring, and performance dashboards.

UnrankedNo signals yet
Harnessclaude
devops-infrastructuresubagents
Details
Sub-AgentsPreview

monorepo-architect

Expert in monorepo architecture, build systems, and dependency management at scale. Masters Nx, Turborepo, Bazel, and Lerna for efficient multi-project development. Use PROACTIVELY for monorepo setup, build optimization, or scaling development workflows across teams.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
architecturefrontendperformancesubagent
Details
WorkflowsPreview

monte-carlo-simulator

Run Monte Carlo simulations with probability distributions, confidence intervals, and statistical analysis

UnrankedNo signals yet
Harnessclaude
simulationworkflows
Details
WorkflowsPreview

move

Move tasks between status folders following the task management protocol.

UnrankedNo signals yet
Harnessclaude
orchestrationworkflows
Details
Sub-AgentsPreview

ms-sql-dba

Work with Microsoft SQL Server databases using the MS SQL extension.

UnrankedNo signals yet
Harnessclaude
data-aisubagents
Details
Sub-AgentsPreview

multi-agent-coordinator

Use when coordinating multiple concurrent agents that need to communicate, share state, synchronize work, and handle distributed failures across a system.

UnrankedNo signals yet
Harnessclaude
meta-orchestrationsubagent
Details
WorkflowsPreview

multi-agent-optimize

Multi-Agent Optimization Toolkit

UnrankedNo signals yet
Harnessclaude
agentsworkflows
Details
WorkflowsPreview

multi-agent-optimize-2

Optimize application stack using specialized optimization agents:

UnrankedNo signals yet
Harnessclaude
commandstools
Details
WorkflowsPreview

multi-agent-review

Multi-Agent Code Review Orchestration Tool

UnrankedNo signals yet
Harnessclaude
agentsworkflows
Details
WorkflowsPreview

multi-agent-review-2

Multi-Agent Code Review Orchestration Tool

UnrankedNo signals yet
Harnessclaude
agentsworkflows
Details
WorkflowsPreview

multi-agent-review-3

Perform comprehensive multi-agent code review with specialized reviewers:

UnrankedNo signals yet
Harnessclaude
commandstools
Details
WorkflowsPreview

multi-backend

Run a backend-focused multi-model workflow for APIs, algorithms, data, and business logic.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
commandworkflow
Details
WorkflowsPreview

multi-execute

Execute a multi-model implementation plan while preserving Claude as the only filesystem writer.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
commandworkflow
Details
WorkflowsPreview

multi-frontend

Run a frontend-focused multi-model workflow for components, layouts, animation, and UI polish.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
commandworkflow
Details
WorkflowsPreview

multi-plan

Create a multi-model implementation plan without modifying production code.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
commandworkflow
Details
WorkflowsPreview

multi-platform

Orchestrate cross-platform feature development across web, mobile, and desktop with API-first architecture

UnrankedNo signals yet
Harnessclaude
agentsworkflows
Details
WorkflowsPreview

multi-platform-2

Build the same feature across multiple platforms:

UnrankedNo signals yet
Harnessclaude
commandsworkflows
Details
Browse · Armory