Armory
Source

Browse

Search and filter by type across the catalog

1,207 results in Identity, Skills, Evals · page 3 of 51

EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars · 107 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
IdentityExperimental

unicity-aos-capsule-identity

Use when an agent's identity must be persisted as state and assembled into its system prompt at boot, rather than pasted into a prompt by hand.

96.88,503 stars · 19 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsExperimental

stonybrooknlp-appworld

🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars · 78 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
SkillsPreview

web-assets-generator-skill

Easily generate web assets from Claude Code including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn. Handles image resizing, text-to-image generation, emojis, and provides proper HTML meta tags.

96.4490 stars · 52 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
SkillsExperimental

superdesign-skill

The design skill for Claude Code, Cursor and any coding agent. Stop shipping AI-slop UI: turn it into shippable, tasteful frontend. Install: npx skills add superdesigndev/superdesign-skill. Powered by superdesign.dev

Contributed by Sentinel

96.2520 stars · 38 forks
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars · 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install · SourceDetails
IdentityExperimental

agntcy-oasf

Use when agents must describe themselves to other systems in a common schema so they can be catalogued, discovered and verified.

95.7332 stars · 47 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
SkillsPreview

cc-devops-skills

Skills for DevOps work that generate and validate infrastructure-as-code with shell scripts and CLI tools, for most deployment platforms. Also useful as documentation.

95.2303 stars · 34 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
IdentityExperimental

telagod-code-abyss

Use when you want a coding agent to have a consistent, composable personality and voice across Claude Code, Codex, Gemini CLI and OpenClaw.

94.6239 stars · 32 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

agntcy-dir

Use when agents and multi-agent systems need to announce themselves and be found across organisations rather than hardcoded.

94.6184 stars · 55 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars · 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars · 62 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install · SourceDetails
IdentityExperimental

character-card-spec-v2

Use when reading or writing the character-card files that the roleplay-agent ecosystem actually ships, including the PNG-embedded variant.

93.9188 stars · 29 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

cloudflare-web-bot-auth

Use when a website has to be able to tell that a request really came from your agent and not from someone impersonating it.

93.8157 stars · 40 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

highflame-ai-zeroid

Use when a fleet of autonomous agents needs issued identities with a lifecycle — created, rotated and revoked.

92.7163 stars · 18 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
SkillsPreview

claude-code-agents

Comprehensive E2E development workflow with helpful Claude Code subagent prompts for solo devs. Run multiple auditors in parallel, automate fix cycles with micro-checkpoint protocols, and do browser-based QA. Includes strict protocols to prevent AI going rogue.

91.9147 stars · 15 forks
Harnessclaude
skill
No one-command install · SourceDetails
IdentityExperimental

agentmail-to-agentmail-toolkit

Use when an agent needs its own email address so people and systems can reach it, and it can act on what arrives.

91.698 stars · 26 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
SkillsPreview

book-factory

A comprehensive pipeline of Skills that replicates traditional publishing infrastructure for nonfiction book creation using specialized Claude skills.

91.6108 stars · 20 forks
Harnessclaude
skill
No one-command install · SourceDetails
IdentityExperimental

agntcy-identity

Use when agents, MCP servers and multi-agent systems all need issued identities that another party can verify.

91.4101 stars · 20 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsPreview

vellum-evals

Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.

90.582 stars · 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
No one-command install · SourceDetails
IdentityExperimental

character-card-spec-v3

Use when authoring an agent persona against the current character-card standard, with lorebooks, assets and decorators.

90.2109 stars · 11 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

agentrhq-authsome

Use when agents must stay logged in to third-party services without ever seeing your credentials.

88.687 stars · 9 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

gebruder-wirken

Use when autonomous agents need one gateway that holds their credentials, isolates them per channel, and logs every session.

88.6171 stars · 5 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

opena2a-agent-identity-management

Use when non-human identities need the same lifecycle a workforce IAM gives people — issue, authorize, audit, revoke.

88.559 stars · 18 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
Browse · Armory