229 results in Memory, CLIs & Tools, Evals · page 5 of 10
Harness Claude Code Cursor Codex Gemini OpenCode phoenix Arize Phoenix: open-source LLM observability with built-in evals, span tracing, and dataset curation for RAG and agents.
99.2 11,286 stars · 1,086 forks · 2 mentions
Harness claude cursor codex opencode gemini
evals observability rag agents
humanlayer-humanlayer The best way to get AI coding agents to solve hard problems in complex codebases.
Contributed by Sentinel
99.2 11,361 stars · 943 forks · 4 mentions
Harness claude codex cursor gemini opencode
twilio Unleash the power of Twilio from your command prompt
99.2 192 stars · 104 forks · passed install test
comms
artidoro-qlora QLoRA: Efficient Finetuning of Quantized LLMs
Contributed by Sentinel
99.2 11,021 stars · 876 forks · 4 mentions · failed install test
Harness claude codex cursor gemini opencode
clis-tools
harbor-framework-harbor Framework for evaluating and improving agents
Contributed by Sentinel
99.1 4,867 stars · 1,704 forks · 10 mentions
Harness claude codex cursor gemini opencode
claude-squad Claude Squad is a terminal app that manages multiple Claude Code, Codex (and other local agents including Aider) in separate workspaces, allowing you to work on multiple tasks simultaneously.
99.1 8,536 stars · 621 forks · 1 mention
Harness claude
claude-code tooling
plastic-labs-honcho Memory library for building stateful agents
Contributed by Sentinel
99.1 6,980 stars · 865 forks · 6 mentions
Harness claude codex cursor gemini opencode
memory
claude-code-usage-monitor A real-time terminal-based tool for monitoring Claude Code token usage. It shows live token consumption, burn rate, and predictions for token depletion. Features include visual progress bars, session-aware analytics, and support for multiple subscription plans.
99.1 8,668 stars · 458 forks
Harness claude
claude-code tooling
swe-bench SWE-bench: benchmark for evaluating LLMs on real-world GitHub issue resolution across 12 popular Python repositories.
99.1 5,762 stars · 957 forks · 9 mentions
Harness claude cursor codex opencode gemini
evals code benchmark agents
alexzhang13-rlm General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.
Contributed by Sentinel
99.0 5,644 stars · 905 forks · 3 mentions
Harness claude codex cursor gemini opencode
clawd-on-desk A desktop pet that reacts to your Claude Code sessions in real time: thinking, typing, juggling, sleeping and more.
99.0 6,093 stars · 636 forks
Harness claude
claude-code tooling
gepa-ai-gepa Optimize prompts, code, and more with AI-powered Reflective Optimization
Contributed by Sentinel
99.0 6,346 stars · 533 forks · 5 mentions
Harness claude codex cursor gemini opencode
giskard Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.0 5,838 stars · 542 forks
Harness claude cursor codex opencode gemini
evals safety vulnerability scan
getzep-zep Examples, framework integrations and tools for Zep Cloud, Zep's hosted agent memory service; the repository says it is not the product itself. The open-source knowledge-graph engine behind Zep is Graphiti (getzep-graphiti).
Contributed by Sentinel
99.0 4,882 stars · 651 forks · 7 mentions
Harness claude codex cursor gemini opencode
memory
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
caviraoss-openmemory Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
Contributed by Sentinel
98.9 4,478 stars · 504 forks
Harness claude codex cursor gemini opencode
memory
claudable Claudable is an open-source web builder that leverages local CLI agents, such as Claude Code and Cursor Agent, to build and deploy products effortlessly.
98.9 4,054 stars · 625 forks
Harness claude
client cli
big-bench Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.8 3,247 stars · 618 forks
Harness claude cursor codex opencode gemini
evals benchmark google academic
flowelement-m-flow Use when recall should follow associations between memories rather than nearest-neighbour similarity alone.
98.8 4,497 stars · 255 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
trulens Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.7 3,530 stars · 335 forks · 1 mention
Harness claude cursor codex opencode gemini
evals rag tracking dashboard
claude-devtools A desktop app that shows your Claude Code sessions by reading their logs: context use per turn across categories, compaction, sub-agent execution trees and custom notification triggers.
98.7 3,893 stars · 298 forks
Harness claude
claude-code tooling
xlang-ai-osworld [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.7 3,117 stars · 530 forks · 7 mentions
Harness claude codex cursor gemini opencode
inspect-ai UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.7 2,683 stars · 685 forks
Harness claude cursor codex opencode gemini
evals safety aisi government
Previous Page 5 of 10 Next
Browse · Armory