92 results in Infrastructure, Evals · page 2 of 4
Harness Claude Code Cursor Codex Gemini OpenCode microsandbox Use as the OSS self-hosted sandbox when you need to run agent code on your own infra — libkrun-based microVM isolation with no per-sandbox vendor cost, the escape hatch from a managed runtime at high volume.
99.1 8,436 stars · 454 forks
Harness claude codex
sandbox self-hosted libkrun microvm
giskard Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.0 5,838 stars · 542 forks
Harness claude cursor codex opencode gemini
evals safety vulnerability scan
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
claude-managed-agents-selfhost Use for enterprise hosted-control agents — the agent loop runs at the provider while execution happens on your own infrastructure — when you want a managed control plane but data and code must stay on your machines.
98.9 3,876 stars · 836 forks
Harness claude
managed-agents hosted-control enterprise self-hosted-execution
big-bench Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.8 3,247 stars · 618 forks
Harness claude cursor codex opencode gemini
evals benchmark google academic
trulens Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.7 3,530 stars · 335 forks · 1 mention
Harness claude cursor codex opencode gemini
evals rag tracking dashboard
xlang-ai-osworld [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.7 3,117 stars · 530 forks · 7 mentions
Harness claude codex cursor gemini opencode
inspect-ai UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.7 2,683 stars · 685 forks
Harness claude cursor codex opencode gemini
evals safety aisi government
helm Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.7 2,898 stars · 412 forks
Harness claude cursor codex opencode gemini
evals benchmark academic stanford
sysbox Next-generation container runtime enabling Docker-in-Docker and VM-like isolation without privileged containers.
98.6 3,847 stars · 228 forks
Harness claude cursor codex opencode gemini
infrastructure containers
lighteval Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.
98.6 2,533 stars · 553 forks
Harness claude cursor codex opencode gemini
evals huggingface benchmark lightweight
open-operator Browserbase Open Operator — open-source Operator-style web agent built on Stagehand; demonstrates full task decomposition, action planning, and evidence collection using the Browserbase cloud.
98.4 1,953 stars · 325 forks · 1 mention
Harness claude cursor codex opencode gemini
browser stagehand
notte Notte open-source web agent environment — converts browser sessions into a Markov Decision Process with structured observation/action spaces, making browsers first-class RL and LLM agent environments.
98.3 2,003 stars · 181 forks
Harness claude cursor codex opencode gemini
browser rl
evalplus Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.
98.3 1,819 stars · 208 forks
Harness claude cursor codex opencode gemini
evals code-generation humaneval benchmark
webarena WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.
98.2 1,592 stars · 249 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents browser benchmark
tau-bench Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
98.1 1,416 stars · 215 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents tool-use benchmark
harbor-framework-terminal-bench Measuring and evolving with the frontier of agent work
Contributed by Sentinel
98.1 588 stars · 434 forks · 13 mentions
Harness claude codex cursor gemini opencode
e2b-desktop E2B Desktop Sandbox — cloud virtual desktop (Ubuntu + VNC) with Python SDK for screenshot, mouse, keyboard, and process control; designed for AI agents that need a full GUI environment in an isolated VM.
98.1 1,461 stars · 179 forks
Harness claude cursor codex opencode gemini
browser e2b
agent-e Emergence AI agent-E browser agent — hierarchical LLM-based web automation that uses DOM distillation and action abstraction layers to achieve significantly higher benchmark accuracy than prior browser agents.
98.0 1,249 stars · 190 forks
Harness claude cursor codex opencode gemini
browser emergence
wandb-weave-evals Weights & Biases Weave evaluation framework for tracking LLM experiments, scoring model outputs, and comparing runs.
97.9 1,130 stars · 170 forks · 3 mentions
Harness claude cursor codex opencode gemini
evals wandb experiment-tracking scoring
langtrace Open-source observability tool for LLMs with OpenTelemetry-based tracing, automated evals, and annotation workflows.
97.8 1,228 stars · 127 forks
Harness claude cursor codex opencode gemini
evals observability opentelemetry tracing
webvoyager Research browser agent from Zhejiang University and HKU — uses GPT-4V interleaved screenshot + HTML observations to complete open-ended web tasks; established an early web-agent benchmark (WebVoyager).
97.8 1,127 stars · 124 forks · 1 mention
Harness claude cursor codex opencode gemini
browser research
surf-computer-use E2B Surf — a Stagehand-powered computer-use interface layer for E2B Firecracker sandboxes; connects the act/extract/observe primitives directly to microVM display output for lightweight headless computer use.
97.6 856 stars · 141 forks
Harness claude cursor codex opencode gemini
browser e2b
Previous Page 2 of 4 Next
Browse · Armory