128 results in Memory, Evals, Infrastructure · page 4 of 6
Harness Claude Code Cursor Codex Gemini OpenCode agent-e Emergence AI agent-E browser agent: hierarchical LLM-based web automation that uses DOM distillation and action abstraction layers to achieve significantly higher benchmark accuracy than prior browser agents.
98.0 1,249 stars · 190 forks
Harness claude cursor codex opencode gemini
browser emergence
claudiodrews-memory-os Use when you want layered memory — structured facts, recall and an auto-curated wiki — running locally against any model.
97.9 1,355 stars · 128 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
langtrace Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.8 1,228 stars · 127 forks
Harness claude cursor codex opencode gemini
evals observability opentelemetry tracing
agiresearch-a-mem A-MEM: Agentic Memory for LLM Agents
Contributed by Sentinel
97.8 1,164 stars · 121 forks
Harness claude codex cursor gemini opencode
memory
webvoyager Research browser agent from Zhejiang University and HKU. It uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).
97.8 1,127 stars · 124 forks · 1 mention
Harness claude cursor codex opencode gemini
browser research
surf-computer-use E2B Surf: a Stagehand-powered computer-use interface layer for E2B Firecracker sandboxes; connects the act/extract/observe primitives directly to microVM display output for lightweight headless computer use.
97.6 856 stars · 141 forks
Harness claude cursor codex opencode gemini
browser e2b
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
codeabra-iai-personal-memory-engine Use when the agent should remember not just facts but how you like to work, locally and for free.
97.5 862 stars · 105 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
Infrastructure Experimental modal-labs-modal-client Modal — serverless GPU/CPU containers for running agents, sandboxes and model inference from Python (this is the client SDK).
Contributed by Sentinel
97.0 518 stars · 132 forks · 11 mentions
Harness claude codex cursor gemini opencode
infrastructure
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
nemori-ai-nemori Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.
93.5 207 stars · 20 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
jean-technologies-jean-memory Use when you want mem0-style and graph-style memory combined behind one interface.
92.1 171 stars · 13 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
kyros-ai Use when the agent's memory needs to resolve its own contradictions and forget on a schedule.
82.5 96 stars · 2 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
braintrust Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
82.3 27 stars · 12 forks
Harness claude cursor codex opencode gemini
evals experiment-tracking sdk
galileo-evaluate Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.
80.3 22 stars · 11 forks
Harness claude cursor codex opencode gemini
evals hallucination observability sdk
e2b Open-source secure cloud sandboxes (Firecracker microVMs) for running AI-generated code. ~150ms cold start.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure sandbox
modal Serverless cloud platform for running Python functions, containers, and AI workloads with zero infra management.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure serverless
railway Zero-config cloud platform for deploying agent backends, databases, and services from a Git push.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure deploy
Previous Page 4 of 6 Next
Browse · Armory