852 results in Evals, Memory, Sub-Agents, Infrastructure · page 4 of 36
Harness Claude Code Cursor Codex Gemini OpenCode e2b-desktop E2B Desktop Sandbox: a cloud virtual desktop (Ubuntu + VNC) with Python SDK for screenshot, mouse, keyboard, and process control; designed for AI agents that need a full GUI environment in an isolated VM.
98.1 1,461 stars · 179 forks
Harness claude cursor codex opencode gemini
browser e2b
dataojitori-nocturne-memory Use when you want to see and roll back what your agent remembered, instead of trusting an opaque vector store.
98.0 1,373 stars · 170 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
agent-e Emergence AI agent-E browser agent: hierarchical LLM-based web automation that uses DOM distillation and action abstraction layers to achieve significantly higher benchmark accuracy than prior browser agents.
98.0 1,249 stars · 190 forks
Harness claude cursor codex opencode gemini
browser emergence
claudiodrews-memory-os Use when you want layered memory — structured facts, recall and an auto-curated wiki — running locally against any model.
97.9 1,355 stars · 128 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
langtrace Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.8 1,228 stars · 127 forks
Harness claude cursor codex opencode gemini
evals observability opentelemetry tracing
agiresearch-a-mem A-MEM: Agentic Memory for LLM Agents
Contributed by Sentinel
97.8 1,164 stars · 121 forks
Harness claude codex cursor gemini opencode
memory
webvoyager Research browser agent from Zhejiang University and HKU. It uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).
97.8 1,127 stars · 124 forks · 1 mention
Harness claude cursor codex opencode gemini
browser research
surf-computer-use E2B Surf: a Stagehand-powered computer-use interface layer for E2B Firecracker sandboxes; connects the act/extract/observe primitives directly to microVM display output for lightweight headless computer use.
97.6 856 stars · 141 forks
Harness claude cursor codex opencode gemini
browser e2b
xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
codeabra-iai-personal-memory-engine Use when the agent should remember not just facts but how you like to work, locally and for free.
97.5 862 stars · 105 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
Infrastructure Experimental modal-labs-modal-client Modal — serverless GPU/CPU containers for running agents, sandboxes and model inference from Python (this is the client SDK).
Contributed by Sentinel
97.0 518 stars · 132 forks · 11 mentions
Harness claude codex cursor gemini opencode
infrastructure
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
nemori-ai-nemori Use when you want to see whether aligning memory to episode-sized chunks beats a heavier memory framework.
93.5 207 stars · 20 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
jean-technologies-jean-memory Use when you want mem0-style and graph-style memory combined behind one interface.
92.1 171 stars · 13 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
kyros-ai Use when the agent's memory needs to resolve its own contradictions and forget on a schedule.
82.5 96 stars · 2 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
braintrust Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
82.3 27 stars · 12 forks
Harness claude cursor codex opencode gemini
evals experiment-tracking sdk
galileo-evaluate Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.
80.3 22 stars · 11 forks
Harness claude cursor codex opencode gemini
evals hallucination observability sdk
e2b Open-source secure cloud sandboxes (Firecracker microVMs) for running AI-generated code. ~150ms cold start.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure sandbox
Previous Page 4 of 36 Next
Browse · Armory