Armory
Source

Browse

Search and filter by type across the catalog

1,239 results in Observability, Skills, Memory, Evals 路 page 5 of 52

EvalsExperimental

princeton-nlp-webshop

[NeurIPS 2022] 馃洅WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Contributed by Sentinel

97.1589 stars 路 107 forks 路 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
EvalsExperimental

stonybrooknlp-appworld

馃實 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.

Contributed by Sentinel

96.7500 stars 路 78 forks 路 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install 路 SourceDetails
SkillsPreview

web-assets-generator-skill

Easily generate web assets from Claude Code including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn. Handles image resizing, text-to-image generation, emojis, and provides proper HTML meta tags.

96.4490 stars 路 52 forks
Harnessclaude
claude-codeagent-skills
No one-command install 路 SourceDetails
SkillsExperimental

superdesign-skill

The design skill for Claude Code, Cursor and any coding agent. Stop shipping AI-slop UI: turn it into shippable, tasteful frontend. Install: npx skills add superdesigndev/superdesign-skill. Powered by superdesign.dev

Contributed by Sentinel

96.2520 stars 路 38 forks
Harnessclaudecodexcursorgeminiopencode
skills
No one-command install 路 SourceDetails
EvalsPreview

continuous-eval

Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.

96.2517 stars 路 38 forks
Harnessclaudecursorcodexopencodegemini
evalsragagentsmetrics
No one-command install 路 SourceDetails
ObservabilityPreview

claude-code-statusline

Enhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring

95.9476 stars 路 35 forks
Harnessclaude
statuslineobservability
No one-command install 路 SourceDetails
ObservabilityPreview

phospho

Phospho is a text analytics and evaluation platform for LLM apps. It logs sessions, runs clustering, detects failures, and surfaces actionable insights.

95.8439 stars 路 35 forks
Harnessclaudecursorcodexopencodegemini
observabilityanalyticsevals
No one-command install 路 SourceDetails
SkillsPreview

cc-devops-skills

Skills for DevOps work that generate and validate infrastructure-as-code with shell scripts and CLI tools, for most deployment platforms. Also useful as documentation.

95.2303 stars 路 34 forks
Harnessclaude
claude-codeagent-skills
No one-command install 路 SourceDetails
ObservabilityPreview

athina-ai

Athina AI provides developer-focused LLM monitoring and eval framework: real-time inference logging, automated evals, and regression detection in CI.

94.6301 stars 路 23 forks
Harnessclaudecursorcodexopencodegemini
observabilityevalslogging
No one-command install 路 SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars 路 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install 路 SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars 路 62 forks 路 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install 路 SourceDetails
ObservabilityPreview

claude-pace

A lightweight Bash + jq statusline for Claude Code that displays rate limit pace delta (burn rate vs. time remaining), 5h/7d usage percentage, context window usage, git branch and diff stats. Compares current consumption rate against time remaining in each rate limit window to indicate whether quota is being used faster or slower than the window allows. Single file with no external dependencies beyond jq.

93.6229 stars 路 19 forks
Harnessclaude
claude-codestatus-lines
No one-command install 路 SourceDetails
MemoryExperimental

nemori-ai-nemori

Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.

93.5207 stars 路 20 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install 路 SourceDetails
MemoryExperimental

jean-technologies-jean-memory

Use when you want mem0-style and graph-style memory combined behind one interface.

92.1171 stars 路 13 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install 路 SourceDetails
SkillsPreview

claude-code-agents

Comprehensive E2E development workflow with helpful Claude Code subagent prompts for solo devs. Run multiple auditors in parallel, automate fix cycles with micro-checkpoint protocols, and do browser-based QA. Includes strict protocols to prevent AI going rogue.

91.9147 stars 路 15 forks
Harnessclaude
skill
No one-command install 路 SourceDetails
SkillsPreview

book-factory

A comprehensive pipeline of Skills that replicates traditional publishing infrastructure for nonfiction book creation using specialized Claude skills.

91.6108 stars 路 20 forks
Harnessclaude
skill
No one-command install 路 SourceDetails
EvalsPreview

vellum-evals

Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.

90.582 stars 路 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
No one-command install 路 SourceDetails
MemoryExperimental

kyros-ai

Use when the agent's memory needs to resolve its own contradictions and forget on a schedule.

82.596 stars 路 2 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install 路 SourceDetails
EvalsPreview

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

82.327 stars 路 12 forks
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingsdk
No one-command install 路 SourceDetails
ObservabilityPreview

claudia-statusline

High-performance Rust-based statusline for Claude Code with persistent stats tracking, progress bars, and optional cloud sync. Features SQLite-first persistence, git integration, context progress bars, burn rate calculation, XDG-compliant with theme support (dark/light, NO_COLOR).

81.236 stars 路 5 forks
Harnessclaude
claude-codestatus-lines
No one-command install 路 SourceDetails
EvalsPreview

galileo-evaluate

Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.

80.322 stars 路 11 forks
Harnessclaudecursorcodexopencodegemini
evalshallucinationobservabilitysdk
No one-command install 路 SourceDetails
ObservabilityPreview

honeycomb

Honeycomb's OpenTelemetry-native observability platform: a high-cardinality event store ideal for tracing LLM pipelines and debugging slow agent traces.

75.116 stars 路 6 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install 路 SourceDetails
ObservabilityPreview

baserun

Baserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and measuring quality.

74.416 stars 路 5 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingtesting
No one-command install 路 SourceDetails
SkillsPreview

claude-mountaineering-skills

Claude Code skill that automates mountain route research for North American peaks. Aggregates data from 10+ mountaineering sources like Mountaineers.org, PeakBagger.com and SummitPost.com to generate detailed route beta reports with weather, avalanche conditions, and trip reports.

73.333 stars 路 1 fork
Harnessclaude
skill
No one-command install 路 SourceDetails
Browse 路 Armory