Armory
Source

Browse

Search and filter by type across the catalog

532 results in CLAUDE.md / Rules, Evals, Observability · page 3 of 23

ObservabilityPreview

phospho

Phospho is a text analytics and evaluation platform for LLM apps. It logs sessions, runs clustering, detects failures, and surfaces actionable insights.

95.8439 stars · 35 forks
Harnessclaudecursorcodexopencodegemini
observabilityanalyticsevals
No one-command install · SourceDetails
ObservabilityPreview

athina-ai

Athina AI provides developer-focused LLM monitoring and eval framework: real-time inference logging, automated evals, and regression detection in CI.

94.6301 stars · 23 forks
Harnessclaudecursorcodexopencodegemini
observabilityevalslogging
No one-command install · SourceDetails
EvalsPreview

metr-task-standard

METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.

94.3192 stars · 37 forks
Harnessclaudecursorcodexopencodegemini
evalsagentstask-standardsafety
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench-2-1

Terminal-Bench 2.1

Contributed by Sentinel

94.3119 stars · 62 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
evals
No one-command install · SourceDetails
ObservabilityPreview

claude-pace

A lightweight Bash + jq statusline for Claude Code that displays rate limit pace delta (burn rate vs. time remaining), 5h/7d usage percentage, context window usage, git branch and diff stats. Compares current consumption rate against time remaining in each rate limit window to indicate whether quota is being used faster or slower than the window allows. Single file with no external dependencies beyond jq.

93.6229 stars · 19 forks
Harnessclaude
claude-codestatus-lines
No one-command install · SourceDetails
EvalsPreview

vellum-evals

Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.

90.582 stars · 20 forks
Harnessclaudecursorcodexopencodegemini
evalssdkcidataset
No one-command install · SourceDetails
EvalsPreview

braintrust

Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.

82.327 stars · 12 forks
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingsdk
No one-command install · SourceDetails
ObservabilityPreview

claudia-statusline

High-performance Rust-based statusline for Claude Code with persistent stats tracking, progress bars, and optional cloud sync. Features SQLite-first persistence, git integration, context progress bars, burn rate calculation, XDG-compliant with theme support (dark/light, NO_COLOR).

81.236 stars · 5 forks
Harnessclaude
claude-codestatus-lines
No one-command install · SourceDetails
EvalsPreview

galileo-evaluate

Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.

80.322 stars · 11 forks
Harnessclaudecursorcodexopencodegemini
evalshallucinationobservabilitysdk
No one-command install · SourceDetails
ObservabilityPreview

honeycomb

Honeycomb's OpenTelemetry-native observability platform: a high-cardinality event store ideal for tracing LLM pipelines and debugging slow agent traces.

75.116 stars · 6 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
ObservabilityPreview

baserun

Baserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and measuring quality.

74.416 stars · 5 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingtesting
No one-command install · SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars · 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install · SourceDetails
CLAUDE.md / RulesPreview

pre-commit-hooks

A repository of pre-commit hooks whose CLAUDE.md and .claude/ documentation is a thorough, compact example of project instructions for Claude Code.

57.35 stars · 3 forks
Harnessclaude
claude-codeclaude-md-files
No one-command install · SourceDetails
CLAUDE.md / RulesPreview

angular-coding-style

angular rule: apply when working on angular and you need Angular Coding Style.

16.31 mention
Harnessclaudecodexcursorgeminiopencode
rulesangular
Details
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install · SourceDetails
CLAUDE.md / RulesPreview

additional-directories

Grant access to additional directories outside the current project. Useful for monorepo setups, shared libraries, or when working with documentation stored in separate repositories.

UnrankedNo signals yet
Harnessclaude
permissionsclaudemd-rules
Details
CLAUDE.md / RulesPreview

ai-agent-specialist

Cursor rules for TypeScript, React, Node.js, clean architecture, testing, and WHY-oriented engineering guidance.

UnrankedNo signals yet
Harnessclaudecursor
cursor-rulesclaude-md-filesrules
Details
CLAUDE.md / RulesPreview

ai-intellij-plugin

Provides comprehensive Gradle commands for IntelliJ plugin development with platform-specific coding patterns, detailed package structure guidelines, and clear internationalization standards.

UnrankedNo signals yet
Harnessclaude
claude-mdruleslanguage-specific
Details
CLAUDE.md / RulesPreview

allow-git-operations

Allow common git operations for version control workflow. Permits git status, diff, add, commit, and push operations while maintaining security by requiring explicit permission for potentially destructive operations.

UnrankedNo signals yet
Harnessclaude
permissionsclaudemd-rules
Details
CLAUDE.md / RulesPreview

allow-npm-commands

Allow common npm development commands (lint, test, build, start).

UnrankedNo signals yet
Harnessclaude
permissionsclaudemd-rules
Details
CLAUDE.md / RulesPreview

alpha-skills-quant-factor-research

Quantitative factor research skills for Cursor. Evaluate factors, run backtests, mine new alpha through natural language.

UnrankedNo signals yet
Harnessclaudecursor
cursor-rulesclaude-md-filesrules
Details
CLAUDE.md / RulesPreview

android-jetpack-compose-cursorrules-prompt-file

Cursor rules for Android development with Jetpack Compose integration.

UnrankedNo signals yet
Harnessclaudecursor
cursor-rulesclaude-md-filesrules
Details
CLAUDE.md / RulesPreview

angular-hooks

angular rule: apply when working on angular and you need Angular Hooks.

UnrankedNo signals yet
Harnessclaudecodexcursorgeminiopencode
rulesangular
Details
CLAUDE.md / RulesPreview

angular-novo-elements-cursorrules-prompt-file

Cursor rules for Angular development with Novo Elements UI library.

UnrankedNo signals yet
Harnessclaudecursor
cursor-rulesclaude-md-filesrules
Details
Browse · Armory