Armory
Source

Browse

Search and filter by type across the catalog

1,260 results in Evals, Skills, Memory, Infrastructure · page 4 of 53

EvalsPreview

lighteval

Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.

98.62,533 stars · 553 forks
Harnessclaudecursorcodexopencodegemini
evalshuggingfacebenchmarklightweight
No one-command install · SourceDetails
MemoryExperimental

kayba-ai-agentic-context-engine

Use when an agent should carry forward what it learned from its own successes and failures into later runs.

98.52,564 stars · 307 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

memodb-io-memobase

User Profile-Based Long-Term Memory for AI Chatbot Applications.

Contributed by Sentinel

98.52,875 stars · 232 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
SkillsPreview

t-ches-claude-code-resources

A set of sub-agents, skills and commands with a focus on meta-tools, such as a skill auditor and hook creation, meant to be adapted to your workflow.

98.51,973 stars · 411 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
MemoryExperimental

zilliztech-memsearch

Use when several coding agents should share one memory store instead of each keeping its own notes.

98.52,571 stars · 238 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

kingjulio8238-memary

Use when you want a well-known reference implementation of agent memory over a knowledge graph to read or fork.

98.52,644 stars · 205 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
InfrastructurePreview

open-operator

Browserbase Open Operator: an open-source Operator-style web agent built on Stagehand; demonstrates full task decomposition, action planning, and evidence collection using the Browserbase cloud.

98.41,953 stars · 325 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
browserstagehand
No one-command install · SourceDetails
InfrastructurePreview

notte

Notte open-source web agent environment. It converts browser sessions into a Markov Decision Process with structured observation/action spaces, making browsers first-class RL and LLM agent environments.

98.32,003 stars · 181 forks
Harnessclaudecursorcodexopencodegemini
browserrl
No one-command install · SourceDetails
MemoryExperimental

redplanethq-core

Use when one memory graph should serve Claude Code, Codex and your other assistants at once.

98.31,963 stars · 189 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

evalplus

Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.

98.31,819 stars · 208 forks
Harnessclaudecursorcodexopencodegemini
evalscode-generationhumanevalbenchmark
No one-command install · SourceDetails
EvalsPreview

webarena

WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.

98.21,592 stars · 249 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentsbrowserbenchmark
No one-command install · SourceDetails
MemoryExperimental

langchain-ai-langmem

Long-term memory for agents: tools that extract what matters from conversations, refine prompts from feedback and keep memory across sessions, with LangGraph's store built in.

Contributed by Sentinel

98.21,684 stars · 192 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

tau-bench

Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.

98.11,416 stars · 215 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsagentstool-usebenchmark
No one-command install · SourceDetails
EvalsExperimental

harbor-framework-terminal-bench

Measuring and evolving with the frontier of agent work

Contributed by Sentinel

98.1588 stars · 434 forks · 13 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
MemoryExperimental

bai-lab-memoryos

Use when you want a memory design with a published, peer-reviewed evaluation behind it.

98.11,570 stars · 161 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
MemoryExperimental

cortexkit-magic-context

Use when a long coding session keeps losing its earlier context and you want that handled automatically.

98.12,044 stars · 106 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
InfrastructurePreview

e2b-desktop

E2B Desktop Sandbox: a cloud virtual desktop (Ubuntu + VNC) with Python SDK for screenshot, mouse, keyboard, and process control; designed for AI agents that need a full GUI environment in an isolated VM.

98.11,461 stars · 179 forks
Harnessclaudecursorcodexopencodegemini
browsere2b
No one-command install · SourceDetails
SkillsPreview

context-engineering-kit

Hand-crafted collection of advanced context engineering techniques and patterns with minimal token footprint focused on improving agent result quality.

98.01,508 stars · 154 forks
Harnessclaude
skill
No one-command install · SourceDetails
MemoryExperimental

dataojitori-nocturne-memory

Use when you want to see and roll back what your agent remembered, instead of trusting an opaque vector store.

98.01,373 stars · 170 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
InfrastructurePreview

agent-e

Emergence AI agent-E browser agent: hierarchical LLM-based web automation that uses DOM distillation and action abstraction layers to achieve significantly higher benchmark accuracy than prior browser agents.

98.01,249 stars · 190 forks
Harnessclaudecursorcodexopencodegemini
browseremergence
No one-command install · SourceDetails
MemoryExperimental

claudiodrews-memory-os

Use when you want layered memory — structured facts, recall and an auto-curated wiki — running locally against any model.

97.91,355 stars · 128 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

wandb-weave-evals

Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.

97.91,130 stars · 170 forks · 3 mentions
Harnessclaudecursorcodexopencodegemini
evalswandbexperiment-trackingscoring
No one-command install · SourceDetails
SkillsPreview

codex-skill

Enables users to prompt codex from claude code. Unlike the raw codex mcp server, this skill infers parameters such as model, reasoning effort, sandboxing from your prompt or asks you to specify them. It also simplifies continuing prior codex sessions so that codex can continue with the prior context.

97.91,424 stars · 109 forks
Harnessclaude
claude-codeagent-skills
No one-command install · SourceDetails
EvalsPreview

langtrace

Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.

97.81,228 stars · 127 forks
Harnessclaudecursorcodexopencodegemini
evalsobservabilityopentelemetrytracing
No one-command install · SourceDetails
Browse · Armory