162 results in Observability, Infrastructure, Memory, Evals · page 4 of 7
Harness Claude Code Cursor Codex Gemini OpenCode kayba-ai-agentic-context-engine Use when an agent should carry forward what it learned from its own successes and failures into later runs.
98.5 2,564 stars · 307 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
memodb-io-memobase User Profile-Based Long-Term Memory for AI Chatbot Applications.
Contributed by Sentinel
98.5 2,875 stars · 232 forks
Harness claude codex cursor gemini opencode
memory
zilliztech-memsearch Use when several coding agents should share one memory store instead of each keeping its own notes.
98.5 2,571 stars · 238 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
kingjulio8238-memary Use when you want a well-known reference implementation of agent memory over a knowledge graph to read or fork.
98.5 2,644 stars · 205 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
open-operator Browserbase Open Operator: an open-source Operator-style web agent built on Stagehand; demonstrates full task decomposition, action planning, and evidence collection using the Browserbase cloud.
98.4 1,953 stars · 325 forks · 1 mention
Harness claude cursor codex opencode gemini
browser stagehand
notte Notte open-source web agent environment. It converts browser sessions into a Markov Decision Process with structured observation/action spaces, making browsers first-class RL and LLM agent environments.
98.3 2,003 stars · 181 forks
Harness claude cursor codex opencode gemini
browser rl
redplanethq-core Use when one memory graph should serve Claude Code, Codex and your other assistants at once.
98.3 1,963 stars · 189 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
evalplus Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.
98.3 1,819 stars · 208 forks
Harness claude cursor codex opencode gemini
evals code-generation humaneval benchmark
webarena WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.
98.2 1,592 stars · 249 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents browser benchmark
langchain-ai-langmem Long-term memory for agents: tools that extract what matters from conversations, refine prompts from feedback and keep memory across sessions, with LangGraph's store built in.
Contributed by Sentinel
98.2 1,684 stars · 192 forks
Harness claude codex cursor gemini opencode
memory
tau-bench Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
98.1 1,416 stars · 215 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents tool-use benchmark
harbor-framework-terminal-bench Measuring and evolving with the frontier of agent work
Contributed by Sentinel
98.1 588 stars · 434 forks · 13 mentions
Harness claude codex cursor gemini opencode
bai-lab-memoryos Use when you want a memory design with a published, peer-reviewed evaluation behind it.
98.1 1,570 stars · 161 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
cortexkit-magic-context Use when a long coding session keeps losing its earlier context and you want that handled automatically.
98.1 2,044 stars · 106 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
e2b-desktop E2B Desktop Sandbox: a cloud virtual desktop (Ubuntu + VNC) with Python SDK for screenshot, mouse, keyboard, and process control; designed for AI agents that need a full GUI environment in an isolated VM.
98.1 1,461 stars · 179 forks
Harness claude cursor codex opencode gemini
browser e2b
openinference OpenInference is an open standard and Python/JS instrumentation library for capturing LLM and agent traces in OpenTelemetry format, built by Arize AI.
98.1 1,192 stars · 302 forks
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
dataojitori-nocturne-memory Use when you want to see and roll back what your agent remembered, instead of trusting an opaque vector store.
98.0 1,373 stars · 170 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
agent-e Emergence AI agent-E browser agent: hierarchical LLM-based web automation that uses DOM distillation and action abstraction layers to achieve significantly higher benchmark accuracy than prior browser agents.
98.0 1,249 stars · 190 forks
Harness claude cursor codex opencode gemini
browser emergence
langsmith LangChain's platform for tracing, evaluating, and monitoring LLM applications: deep integration with LangChain/LangGraph plus a REST API for any stack.
98.0 1,043 stars · 288 forks
Harness claude cursor codex opencode gemini
observability tracing evals
claudiodrews-memory-os Use when you want layered memory — structured facts, recall and an auto-curated wiki — running locally against any model.
97.9 1,355 stars · 128 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
langtrace Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.8 1,228 stars · 127 forks
Harness claude cursor codex opencode gemini
evals observability opentelemetry tracing
agiresearch-a-mem A-MEM: Agentic Memory for LLM Agents
Contributed by Sentinel
97.8 1,164 stars · 121 forks
Harness claude codex cursor gemini opencode
memory
webvoyager Research browser agent from Zhejiang University and HKU. It uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).
97.8 1,127 stars · 124 forks · 1 mention
Harness claude cursor codex opencode gemini
browser research
Previous Page 4 of 7 Next
Browse · Armory