Armory
Source

Browse

Search and filter by type across the catalog

1,243 results in Identity, Skills, Memory, Evals · page 4 of 52

SkillsPreview

codex-skill

Enables users to prompt codex from claude code. Unlike the raw codex mcp server, this skill infers parameters such as model, reasoning effort, sandboxing from your prompt or asks you to specify them. It also simplifies continuing prior codex sessions so that codex can continue with the prior context.

97.91,424 stars · 109 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
EvalsPreview

langtrace

Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.

97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
No one-command install · SourceDetails
MemoryExperimental

agiresearch-a-mem

A-MEM: Agentic Memory for LLM Agents

Contributed by Sentinel

97.81,164 stars · 121 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
IdentityExperimental

letta-ai-agent-file

Use when you need to save, share, version or move a whole agent — its persona, memory and behaviour — as one portable file.

97.81,197 stars · 114 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
SkillsPreview

claude-codex-settings

A set of plugins for core developer tasks, covering GitHub, Azure, MongoDB, Tavily, Playwright and more. Also works with a few other providers.

97.71,117 stars · 107 forks
Harnessclaude
skill
No one-command install · SourceDetails
SkillsPreview

agentsys

Workflow automation system for Claude with a group of useful plugins, agents, and skills. Automates task-to-production workflows, PR management, code cleanup, performance investigation, drift detection, and multi-agent code review. Includes agnix for linting agent configurations. Built on thousands of lines of code with thousands of tests. Uses deterministic detection (regex, AST) with LLM judgment for efficiency. Used on many production systems.

97.6981 stars · 113 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
EvalsExperimental

xiaowu0162-longmemeval

Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)

Contributed by Sentinel

97.51,049 stars · 81 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

codeabra-iai-personal-memory-engine

Use when the agent should remember not just facts but how you like to work, locally and for free.

97.5862 stars · 105 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars · 107 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
IdentityExperimental

unicity-aos-capsule-identity

Use when an agent's identity must be persisted as state and assembled into its system prompt at boot, rather than pasted into a prompt by hand.

96.88,503 stars · 19 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsExperimental

stonybrooknlp-appworld

🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars · 78 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
SkillsPreview

web-assets-generator-skill

Easily generate web assets from Claude Code including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn. Handles image resizing, text-to-image generation, emojis, and provides proper HTML meta tags.

96.4490 stars · 52 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
SkillsExperimental

superdesign-skill

The design skill for Claude Code, Cursor and any coding agent. Stop shipping AI-slop UI: turn it into shippable, tasteful frontend. Install: npx skills add superdesigndev/superdesign-skill. Powered by superdesign.dev

Contributed by Sentinel

96.2520 stars · 38 forks
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install · SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars · 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install · SourceDetails
IdentityExperimental

agntcy-oasf

Use when agents must describe themselves to other systems in a common schema so they can be catalogued, discovered and verified.

95.7332 stars · 47 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
SkillsPreview

cc-devops-skills

Skills for DevOps work that generate and validate infrastructure-as-code with shell scripts and CLI tools, for most deployment platforms. Also useful as documentation.

95.2303 stars · 34 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
IdentityExperimental

telagod-code-abyss

Use when you want a coding agent to have a consistent, composable personality and voice across Claude Code, Codex, Gemini CLI and OpenClaw.

94.6239 stars · 32 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

agntcy-dir

Use when agents and multi-agent systems need to announce themselves and be found across organisations rather than hardcoded.

94.6184 stars · 55 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars · 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars · 62 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install · SourceDetails
IdentityExperimental

character-card-spec-v2

Use when reading or writing the character-card files that the roleplay-agent ecosystem actually ships, including the PNG-embedded variant.

93.9188 stars · 29 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
IdentityExperimental

cloudflare-web-bot-auth

Use when a website has to be able to tell that a request really came from your agent and not from someone impersonating it.

93.8157 stars · 40 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
MemoryExperimental

nemori-ai-nemori

Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.

93.5207 stars · 20 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
IdentityExperimental

highflame-ai-zeroid

Use when a fleet of autonomous agents needs issued identities with a lifecycle — created, rotated and revoked.

92.7163 stars · 18 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedidentity
No one-command install · SourceDetails
Browse · Armory