1,207 results in Skills, Identity, Evals · page 3 of 51
Harness Claude Code Cursor Codex Gemini OpenCode princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
unicity-aos-capsule-identity Use when an agent's identity must be persisted as state and assembled into its system prompt at boot, rather than pasted into a prompt by hand.
96.8 8,503 stars · 19 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
web-assets-generator-skill Easily generate web assets from Claude Code including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn. Handles image resizing, text-to-image generation, emojis, and provides proper HTML meta tags.
96.4 490 stars · 52 forks
Harness claude
claude-code agent-skills
superdesign-skill The design skill for Claude Code, Cursor and any coding agent. Stop shipping AI-slop UI: turn it into shippable, tasteful frontend. Install: npx skills add superdesigndev/superdesign-skill. Powered by superdesign.dev
Contributed by Sentinel
96.2 520 stars · 38 forks
Harness claude codex cursor gemini opencode
skills
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
agntcy-oasf Use when agents must describe themselves to other systems in a common schema so they can be catalogued, discovered and verified.
95.7 332 stars · 47 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
cc-devops-skills Skills for DevOps work that generate and validate infrastructure-as-code with shell scripts and CLI tools, for most deployment platforms. Also useful as documentation.
95.2 303 stars · 34 forks
Harness claude
claude-code agent-skills
telagod-code-abyss Use when you want a coding agent to have a consistent, composable personality and voice across Claude Code, Codex, Gemini CLI and OpenClaw.
94.6 239 stars · 32 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
agntcy-dir Use when agents and multi-agent systems need to announce themselves and be found across organisations rather than hardcoded.
94.6 184 stars · 55 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
character-card-spec-v2 Use when reading or writing the character-card files that the roleplay-agent ecosystem actually ships, including the PNG-embedded variant.
93.9 188 stars · 29 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
cloudflare-web-bot-auth Use when a website has to be able to tell that a request really came from your agent and not from someone impersonating it.
93.8 157 stars · 40 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
highflame-ai-zeroid Use when a fleet of autonomous agents needs issued identities with a lifecycle — created, rotated and revoked.
92.7 163 stars · 18 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
claude-code-agents Comprehensive E2E development workflow with helpful Claude Code subagent prompts for solo devs. Run multiple auditors in parallel, automate fix cycles with micro-checkpoint protocols, and do browser-based QA. Includes strict protocols to prevent AI going rogue.
91.9 147 stars · 15 forks
Harness claude
skill
agentmail-to-agentmail-toolkit Use when an agent needs its own email address so people and systems can reach it, and it can act on what arrives.
91.6 98 stars · 26 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
book-factory A comprehensive pipeline of Skills that replicates traditional publishing infrastructure for nonfiction book creation using specialized Claude skills.
91.6 108 stars · 20 forks
Harness claude
skill
agntcy-identity Use when agents, MCP servers and multi-agent systems all need issued identities that another party can verify.
91.4 101 stars · 20 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
character-card-spec-v3 Use when authoring an agent persona against the current character-card standard, with lorebooks, assets and decorators.
90.2 109 stars · 11 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
agentrhq-authsome Use when agents must stay logged in to third-party services without ever seeing your credentials.
88.6 87 stars · 9 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
gebruder-wirken Use when autonomous agents need one gateway that holds their credentials, isolates them per channel, and logs every session.
88.6 171 stars · 5 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
opena2a-agent-identity-management Use when non-human identities need the same lifecycle a workforce IAM gives people — issue, authorize, audit, revoke.
88.5 59 stars · 18 forks
Harness claude codex cursor gemini opencode
cp138-seed identity
Previous Page 3 of 51 Next
Browse · Armory