Armory
Source

Browse

Search and filter by type across the catalog

1,924 results in Evals, Memory, Workflows, Skills · page 5 of 81

EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

codeabra-iai-personal-memory-engine

Use when the agent should remember not just facts but how you like to work, locally and for free.

97.5862 stars · 105 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
WorkflowsPreview

awesome-ralph

A curated list of resources about Ralph, the AI coding technique that runs AI coding agents in automated loops until specifications are fulfilled.

97.3918 stars · 74 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
WorkflowsPreview

ralph-wiggum-marketer

A Claude Code plugin that provides an autonomous AI copywriter: research agents gather market knowledge into custom knowledge bases, and a Ralph loop writes the copy.

97.2774 stars · 85 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars · 107 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsExperimental

stonybrooknlp-appworld

🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars · 78 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
WorkflowsPreview

claude-code-repos-index

An index of 75+ Claude Code repositories by one author, covering content management, system design, deep research, IoT, agentic workflows, server management and personal health.

96.7536 stars · 69 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
WorkflowsPreview

claudopro-directory

Well-crafted, wide selection of Claude Code hooks, slash commands, subagent files, and more, covering a range of specialized tasks and workflows. Better resources than your average "Claude-template-for-everything" site.

96.7300 stars · 144 forks
Harnessclaude
workflowguide
No one-command install · SourceDetails
WorkflowsPreview

simone

A broader project management workflow for Claude Code that encompasses not just a set of commands, but a system of documents, guidelines, and processes to facilitate project planning and execution.

96.4558 stars · 46 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
SkillsPreview

web-assets-generator-skill

Easily generate web assets from Claude Code including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn. Handles image resizing, text-to-image generation, emojis, and provides proper HTML meta tags.

96.4490 stars · 52 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
SkillsExperimental

superdesign-skill

The design skill for Claude Code, Cursor and any coding agent. Stop shipping AI-slop UI: turn it into shippable, tasteful frontend. Install: npx skills add superdesigndev/superdesign-skill. Powered by superdesign.dev

Contributed by Sentinel

96.2520 stars · 38 forks
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars · 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install · SourceDetails
WorkflowsPreview

learn-faster-kit

An educational framework for Claude Code based on the "FASTER" approach to self-teaching, with agents, slash commands and tools for learning at your own pace through active learning and spaced repetition.

95.8380 stars · 41 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
SkillsPreview

cc-devops-skills

Skills for DevOps work that generate and validate infrastructure-as-code with shell scripts and CLI tools, for most deployment platforms. Also useful as documentation.

95.2303 stars · 34 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
WorkflowsPreview

agentic-workflow-patterns

A collection of agentic patterns from Anthropic's docs, each with a Mermaid diagram and a code example: sub-agent orchestration, progressive skills, parallel tool calling, master-clone architecture, wizard workflows and more. Also works with other providers.

95.2302 stars · 33 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
WorkflowsExperimental

deepseek-implementation

DeepSeek application development course companion code

94.8160 stars · 67 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agenttutorials-learning-resources
No one-command install · SourceDetails
WorkflowsExperimental

a2a-samples

Samples for A2A implementation

94.6107 stars · 70 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agenttutorials-learning-resources
No one-command install · SourceDetails
WorkflowsExperimental

mentis

A powerful multi-agent orchestration framework built on LangGraph

94.5296 stars · 23 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agentframeworks
No one-command install · SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars · 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars · 62 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install · SourceDetails
MemoryExperimental

nemori-ai-nemori

Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.

93.5207 stars · 20 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
WorkflowsExperimental

claude-mcp

Claude Unified Model Context Interaction Protocol

92.848 stars · 56 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agentframeworks
No one-command install · SourceDetails
WorkflowsPreview

ab-method

A principled, spec-driven workflow that transforms large problems into focused, incremental missions using Claude Code's specialized sub agents. Includes slash-commands, sub agents, and specialized workflows designed for specific parts of the SDLC.

92.5189 stars · 14 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
WorkflowsExperimental

a2a-mcp-tutorial

A tutorial on how to use Model Context Protocol by Anthropic and Agent2Agent Protocol by Google

92.4114 stars · 29 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agenttutorials-learning-resources
No one-command install · SourceDetails
Browse · Armory