Armory
Source

Browse

Search and filter by type across the catalog

229 results in Memory, CLIs & Tools, Evals · page 5 of 10

EvalsPreview

phoenix

Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.

99.211,286 stars · 1,086 forks · 2 mentions
Harnessclaudecursorcodexopencodegemini
evalsobservabilityragagents
No one-command install · SourceDetails
CLIs & ToolsExperimental

humanlayer-humanlayer

The best way to get AI coding agents to solve hard problems in complex codebases.

Contributed by Sentinel

99.211,361 stars · 943 forks · 4 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsExperimental

twilio

Unleash the power of Twilio from your command prompt

99.2192 stars · 104 forks · passed install test
comms
No one-command install · SourceDetails
CLIs & ToolsExperimental

artidoro-qlora

QLoRA: Efficient Finetuning of Quantized LLMs

Contributed by Sentinel

99.211,021 stars · 876 forks · 4 mentions · failed install test
Harnessclaudecodexcursorgeminiopencode
clis-tools
No one-command install · SourceDetails
CLIs & ToolsExperimental

harbor-framework-harbor

Framework for evaluating and improving agents

Contributed by Sentinel

99.14,867 stars · 1,704 forks · 10 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-squad

Claude Squad is a terminal app that manages multiple Claude Code, Codex (and other local agents including Aider) in separate workspaces, allowing you to work on multiple tasks simultaneously.

99.18,536 stars · 621 forks · 1 mention
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
MemoryExperimental

plastic-labs-honcho

Memory library for building stateful agents

Contributed by Sentinel

99.16,980 stars · 865 forks · 6 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-code-usage-monitor

A real-time terminal-based tool for monitoring Claude Code token usage. It shows live token consumption, burn rate, and predictions for token depletion. Features include visual progress bars, session-aware analytics, and support for multiple subscription plans.

99.18,668 stars · 458 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsPreview

swe-bench

SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.

99.15,762 stars · 957 forks · 9 mentions
Harnessclaudecursorcodexopencodegemini
evalscodebenchmarkagents
No one-command install · SourceDetails
CLIs & ToolsExperimental

alexzhang13-rlm

General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.

Contributed by Sentinel

99.05,644 stars · 905 forks · 3 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
CLIs & ToolsPreview

clawd-on-desk

A desktop pet that reacts to your Claude Code sessions in real time: thinking, typing, juggling, sleeping and more.

99.06,093 stars · 636 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
CLIs & ToolsExperimental

gepa-ai-gepa

Optimize prompts, code, and more with AI-powered Reflective Optimization

Contributed by Sentinel

99.06,346 stars · 533 forks · 5 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

giskard

Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.

99.05,838 stars · 542 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyvulnerabilityscan
No one-command install · SourceDetails
MemoryExperimental

getzep-zep

Examples, framework integrations and tools for Zep Cloud, Zep's hosted agent memory service; the repository says it is not the product itself. The open-source knowledge-graph engine behind Zep is Graphiti (getzep-graphiti).

Contributed by Sentinel

99.04,882 stars · 651 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
EvalsPreview

agenta

Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.

98.94,670 stars · 661 forks
Harnessclaudecursorcodexopencodegemini
evalsplaygroundab-testingci
No one-command install · SourceDetails
EvalsPreview

openai-simple-evals

OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.

98.94,621 stars · 509 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkmmlusimple
No one-command install · SourceDetails
MemoryExperimental

caviraoss-openmemory

Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.

Contributed by Sentinel

98.94,478 stars · 504 forks
Harnessclaudecodexcursorgeminiopencode
memory
No one-command install · SourceDetails
CLIs & ToolsPreview

claudable

Claudable is an open-source web builder that leverages local CLI agents, such as Claude Code and Cursor Agent, to build and deploy products effortlessly.

98.94,054 stars · 625 forks
Harnessclaude
clientcli
No one-command install · SourceDetails
EvalsPreview

big-bench

Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.

98.83,247 stars · 618 forks
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkgoogleacademic
No one-command install · SourceDetails
MemoryExperimental

flowelement-m-flow

Use when recall should follow associations between memories rather than nearest-neighbour similarity alone.

98.84,497 stars · 255 forks
Harnessclaudecodexcursorgeminiopencode
cp138-seedmemory
No one-command install · SourceDetails
EvalsPreview

trulens

Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.

98.73,530 stars · 335 forks · 1 mention
Harnessclaudecursorcodexopencodegemini
evalsragtrackingdashboard
No one-command install · SourceDetails
CLIs & ToolsPreview

claude-devtools

A desktop app that shows your Claude Code sessions by reading their logs: context use per turn across categories, compaction, sub-agent execution trees and custom notification triggers.

98.73,893 stars · 298 forks
Harnessclaude
claude-codetooling
No one-command install · SourceDetails
EvalsExperimental

xlang-ai-osworld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Contributed by Sentinel

98.73,117 stars · 530 forks · 7 mentions
Harnessclaudecodexcursorgeminiopencode
No one-command install · SourceDetails
EvalsPreview

inspect-ai

UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.

98.72,683 stars · 685 forks
Harnessclaudecursorcodexopencodegemini
evalssafetyaisigovernment
No one-command install · SourceDetails
Browse · Armory