Armory
Source

Browse

Search and filter by type across the catalog

1,922 results in Evals, Workflows, Observability, Skills · page 4 of 81

EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
ObservabilityPreview

openinference

OpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI.

98.11,192 stars · 302 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
SkillsPreview

context-engineering-kit

Hand-crafted collection of advanced context engineering techniques and patterns with minimal token footprint focused on improving agent result quality.

98.01,508 stars · 154 forks
Harnessclaude
skill
No one-command install · SourceDetails
ObservabilityPreview

langsmith

LangChain's platform for tracing, evaluating, and monitoring LLM applications: deep integration with LangChain/LangGraph plus a REST API for any stack.

98.01,043 stars · 288 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingevals
No one-command install · SourceDetails
WorkflowsPreview

the-ralph-playbook

A detailed guide to the Ralph Wiggum technique for autonomous coding loops, with the reasoning behind it and practical guidelines.

97.91,029 stars · 266 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
EvalsPreview

wandb-weave-evals

Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.

97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
No one-command install · SourceDetails
SkillsPreview

codex-skill

Enables users to prompt codex from claude code. Unlike the raw codex mcp server, this skill infers parameters such as model, reasoning effort, sandboxing from your prompt or asks you to specify them. It also simplifies continuing prior codex sessions so that codex can continue with the prior context.

97.91,424 stars · 109 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
EvalsPreview

langtrace

Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.

97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
No one-command install · SourceDetails
WorkflowsPreview

claude-code-documentation-mirror

A mirror of Anthropic's documentation pages for Claude Code, updated every few hours.

97.7983 stars · 138 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
SkillsPreview

claude-codex-settings

A set of plugins for core developer tasks, covering GitHub, Azure, MongoDB, Tavily, Playwright and more. Also works with a few other providers.

97.71,117 stars · 107 forks
Harnessclaude
skill
No one-command install · SourceDetails
SkillsPreview

agentsys

Workflow automation system for Claude with a group of useful plugins, agents, and skills. Automates task-to-production workflows, PR management, code cleanup, performance investigation, drift detection, and multi-agent code review. Includes agnix for linting agent configurations. Built on thousands of lines of code with thousands of tests. Uses deterministic detection (regex, AST) with LLM judgment for efficiency. Used on many production systems.

97.6981 stars · 113 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
ObservabilityPreview

claude-powerline

A vim-style powerline statusline for Claude Code with real-time usage tracking, git integration, custom themes, and more

97.61,163 stars · 82 forks
Harnessclaude
claude-codestatus-lines
No one-command install · SourceDetails
EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
WorkflowsPreview

awesome-ralph

A curated list of resources about Ralph, the AI coding technique that runs AI coding agents in automated loops until specifications are fulfilled.

97.3918 stars · 74 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
WorkflowsPreview

ralph-wiggum-marketer

A Claude Code plugin that provides an autonomous AI copywriter: research agents gather market knowledge into custom knowledge bases, and a Ralph loop writes the copy.

97.2774 stars · 85 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars · 107 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsExperimental

stonybrooknlp-appworld

🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars · 78 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
WorkflowsPreview

claude-code-repos-index

An index of 75+ Claude Code repositories by one author, covering content management, system design, deep research, IoT, agentic workflows, server management and personal health.

96.7536 stars · 69 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
WorkflowsPreview

claudopro-directory

Well-crafted, wide selection of Claude Code hooks, slash commands, subagent files, and more, covering a range of specialized tasks and workflows. Better resources than your average "Claude-template-for-everything" site.

96.7300 stars · 144 forks
Harnessclaude
workflowguide
No one-command install · SourceDetails
WorkflowsPreview

simone

A broader project management workflow for Claude Code that encompasses not just a set of commands, but a system of documents, guidelines, and processes to facilitate project planning and execution.

96.4558 stars · 46 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
SkillsPreview

web-assets-generator-skill

Easily generate web assets from Claude Code including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn. Handles image resizing, text-to-image generation, emojis, and provides proper HTML meta tags.

96.4490 stars · 52 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
SkillsExperimental

superdesign-skill

The design skill for Claude Code, Cursor and any coding agent. Stop shipping AI-slop UI: turn it into shippable, tasteful frontend. Install: npx skills add superdesigndev/superdesign-skill. Powered by superdesign.dev

Contributed by Sentinel

96.2520 stars · 38 forks
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars · 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install · SourceDetails
ObservabilityPreview

claude-code-statusline

Enhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring

95.9476 stars · 35 forks
Harnessclaude
statuslineobservability
No one-command install · SourceDetails
Browse · Armory