Armory
Source

Browse

Search and filter by type across the catalog

1,514 results in Workflows, Sub-Agents, Evals, Observability · page 2 of 64

ObservabilityPreview

ccstatusline

A highly customizable status line formatter for Claude Code CLI that displays model info, git branch, token usage, and other metrics in your terminal.

99.213,035 stars · 579 forks
Harnessclaude
statuslineobservability
No one-command install · SourceDetails
WorkflowsPreview

harness

A meta-skill that designs domain-specific agent teams, defines specialized agents, and generates the skills they use. Resources are in Korean but can produce high-quality English-language output.

99.28,876 stars · 1,256 forks
Harnessclaude
workflowguideteams
No one-command install · SourceDetails
WorkflowsPreview

claude-code-tips

35+ short Claude Code tips covering voice input, system prompt patching, container workflows for risky tasks, conversation cloning, multi-model orchestration with Gemini CLI and more, with demos, working scripts and a plugin.

99.210,008 stars · 809 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
ObservabilityPreview

openllmetry

OpenTelemetry-based observability for LLM applications. It auto-instruments OpenAI, Anthropic, LangChain, and 20+ providers with zero code changes.

99.17,452 stars · 1,099 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
WorkflowsPreview

ralph-for-claude-code

An autonomous AI development framework that enables Claude Code to work iteratively on projects until completion. Features intelligent exit detection, rate limiting, circuit breaker patterns, and comprehensive safety guardrails to prevent infinite loops and API overuse. Built with Bash, integrated with tmux for live monitoring, and includes 75+ comprehensive tests.

99.19,613 stars · 721 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
WorkflowsPreview

claude-code-pm

A project-management workflow for Claude Code with specialised agents, slash commands and documentation.

99.18,362 stars · 837 forks
Harnessclaude
workflowguide
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
WorkflowsPreview

claude-code-ultimate-guide

A guide to Claude Code from beginner to power user, with templates for its features, guides on agentic workflows, quizzes and a cheatsheet. Check that it is current before relying on it.

99.05,869 stars · 770 forks
Harnessclaude
claude-codeworkflows-knowledge-guides
No one-command install · SourceDetails
ObservabilityPreview

helicone

Open-source LLM observability platform: proxy-based logging, cost tracking, caching, and rate limiting for OpenAI-compatible APIs.

99.06,123 stars · 662 forks
Harnessclaudecursorcodexopencodegemini
observabilityproxylogging
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
ObservabilityPreview

grafana-tempo

Grafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM stacks.

99.05,461 stars · 746 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingdistributed
No one-command install · SourceDetails
WorkflowsExperimental

learn-agentic-ai

Learn Agentic AI using Dapr Agentic Cloud Ascent (DACA) Design Pattern: OpenAI Agents SDK, Memory, MCP, A2A, Knowledge Graphs, Rancher Desktop, and Kubernetes

99.04,351 stars · 1,008 forks
Harnessclaudecursorcodexopencodegemini
a2aagent-to-agenttutorials-learning-resources
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
WorkflowsExperimental

ysymyth-react

[ICLR 2023] ReAct: Synergizing Reasoning and Acting in Language Models

Contributed by Sentinel

98.84,140 stars · 401 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
ObservabilityPreview

pydantic-logfire

Logfire by Pydantic: OpenTelemetry-based structured logging and tracing for Python applications with built-in support for FastAPI, SQLAlchemy, and Anthropic.

98.84,450 stars · 283 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrylogging
No one-command install · SourceDetails
ObservabilityPreview

langwatch

LangWatch provides real-time LLM analytics, guardrails, and evaluation pipelines with a visual studio for monitoring multi-step agent conversations.

98.73,522 stars · 362 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingguardrails
No one-command install · SourceDetails
EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
WorkflowsExperimental

noahshinn-reflexion

[NeurIPS 2023] Reflexion: Language Agents with Verbal Reinforcement Learning

Contributed by Sentinel

98.73,251 stars · 317 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

helm

Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.

98.72,898 stars · 412 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkacademicstanford
No one-command install · SourceDetails
EvalsPreview

lighteval

Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.

98.62,533 stars · 553 forks
Harnessclaudecursorcodexopencodegemini
evalshuggingfacebenchmarklightweight
No one-command install · SourceDetails
Browse · Armory