162 results in Evals, Memory, Observability, Infrastructure · page 5 of 7
Harness Claude Code Cursor Codex Gemini OpenCode surf-computer-use E2B Surf: a Stagehand-powered computer-use interface layer for E2B Firecracker sandboxes; connects the act/extract/observe primitives directly to microVM display output for lightweight headless computer use.
97.6 856 stars · 141 forks
Harness claude cursor codex opencode gemini
browser e2b
claude-powerline A vim-style powerline statusline for Claude Code with real-time usage tracking, git integration, custom themes, and more
97.6 1,163 stars · 82 forks
Harness claude
claude-code status-lines
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
codeabra-iai-personal-memory-engine Use when the agent should remember not just facts but how you like to work, locally and for free.
97.5 862 stars · 105 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
Infrastructure Experimental modal-labs-modal-client Modal — serverless GPU/CPU containers for running agents, sandboxes and model inference from Python (this is the client SDK).
Contributed by Sentinel
97.0 518 stars · 132 forks · 11 mentions
Harness claude codex cursor gemini opencode
infrastructure
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
claude-code-statusline Enhanced 4-line statusline for Claude Code with themes, cost tracking, and MCP server monitoring
95.9 476 stars · 35 forks
Harness claude
statusline observability
phospho Phospho is a text analytics and evaluation platform for LLM apps. It logs sessions, runs clustering, detects failures, and surfaces actionable insights.
95.8 439 stars · 35 forks
Harness claude cursor codex opencode gemini
observability analytics evals
athina-ai Athina AI provides developer-focused LLM monitoring and eval framework: real-time inference logging, automated evals, and regression detection in CI.
94.6 301 stars · 23 forks
Harness claude cursor codex opencode gemini
observability evals logging
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
claude-pace A lightweight Bash + jq statusline for Claude Code that displays rate limit pace delta (burn rate vs. time remaining), 5h/7d usage percentage, context window usage, git branch and diff stats. Compares current consumption rate against time remaining in each rate limit window to indicate whether quota is being used faster or slower than the window allows. Single file with no external dependencies beyond jq.
93.6 229 stars · 19 forks
Harness claude
claude-code status-lines
nemori-ai-nemori Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.
93.5 207 stars · 20 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
jean-technologies-jean-memory Use when you want mem0-style and graph-style memory combined behind one interface.
92.1 171 stars · 13 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
kyros-ai Use when the agent's memory needs to resolve its own contradictions and forget on a schedule.
82.5 96 stars · 2 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
braintrust Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
82.3 27 stars · 12 forks
Harness claude cursor codex opencode gemini
evals experiment-tracking sdk
claudia-statusline High-performance Rust-based statusline for Claude Code with persistent stats tracking, progress bars, and optional cloud sync. Features SQLite-first persistence, git integration, context progress bars, burn rate calculation, XDG-compliant with theme support (dark/light, NO_COLOR).
81.2 36 stars · 5 forks
Harness claude
claude-code status-lines
galileo-evaluate Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.
80.3 22 stars · 11 forks
Harness claude cursor codex opencode gemini
evals hallucination observability sdk
e2b Open-source secure cloud sandboxes (Firecracker microVMs) for running AI-generated code. ~150ms cold start.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure sandbox
modal Serverless cloud platform for running Python functions, containers, and AI workloads with zero infra management.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure serverless
railway Zero-config cloud platform for deploying agent backends, databases, and services from a Git push.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure deploy
Previous Page 5 of 7 Next
Browse · Armory