92 results in Infrastructure, Evals · page 3 of 4
Harness Claude Code Cursor Codex Gemini OpenCode xiaowu0162-longmemeval Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
Contributed by Sentinel
97.5 1,049 stars · 81 forks · 10 mentions
Harness claude codex cursor gemini opencode
princeton-nlp-webshop [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Contributed by Sentinel
97.1 589 stars · 107 forks · 3 mentions
Harness claude codex cursor gemini opencode
Infrastructure Experimental modal-labs-modal-client Modal — serverless GPU/CPU containers for running agents, sandboxes and model inference from Python (this is the client SDK).
Contributed by Sentinel
97.0 518 stars · 132 forks · 11 mentions
Harness claude codex cursor gemini opencode
infrastructure
stonybrooknlp-appworld 🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
Contributed by Sentinel
96.7 500 stars · 78 forks · 4 mentions
Harness claude codex cursor gemini opencode
continuous-eval Relari's modular evaluation library for LLM pipelines with deterministic + LLM-based metrics for RAG and agent workflows.
96.2 517 stars · 38 forks
Harness claude cursor codex opencode gemini
evals rag agents metrics
metr-task-standard METR's Task Standard: a specification and scaffold for creating agentic tasks used in autonomous agent capability evaluations.
94.3 192 stars · 37 forks
Harness claude cursor codex opencode gemini
evals agents task-standard safety
harbor-framework-terminal-bench-2-1 Terminal-Bench 2.1
Contributed by Sentinel
94.3 119 stars · 62 forks · 3 mentions
Harness claude codex cursor gemini opencode
evals
vellum-evals Vellum evaluation SDK for running LLM test suites with custom metrics, dataset pinning, and CI workflow integration.
90.5 82 stars · 20 forks
Harness claude cursor codex opencode gemini
evals sdk ci dataset
braintrust Developer platform for logging, evaluating, and comparing LLM experiments with dataset versioning and scoring functions.
82.3 27 stars · 12 forks
Harness claude cursor codex opencode gemini
evals experiment-tracking sdk
galileo-evaluate Galileo evaluation and observability SDK for detecting hallucinations, data errors, and model weaknesses in LLM pipelines.
80.3 22 stars · 11 forks
Harness claude cursor codex opencode gemini
evals hallucination observability sdk
e2b Open-source secure cloud sandboxes (Firecracker microVMs) for running AI-generated code. ~150ms cold start.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure sandbox
modal Serverless cloud platform for running Python functions, containers, and AI workloads with zero infra management.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure serverless
railway Zero-config cloud platform for deploying agent backends, databases, and services from a Git push.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure deploy
vercel Frontend cloud platform with serverless functions and AI SDK integrations for deploying agent-facing UIs.
80.0 passed install test
Harness claude cursor codex opencode gemini
infrastructure deploy
humanloop-evals Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.
69.0 12 stars · 3 forks
Harness claude cursor codex opencode gemini
evals human-eval dataset sdk
agentbench Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.
Unranked No signals yet
Harness claude codex
eval benchmark scoring harness
aws-lambda Serverless function-as-a-service platform for event-driven agent compute without provisioning servers.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure serverless
beta9 Open-source serverless GPU container runtime for running AI workloads with fast cold-starts on bare-metal.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure gpu-serverless
blaxel Cloud runtime and control plane for deploying, scaling, and observing production AI agent workloads.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure agent-compute
cloudflare-sandboxes Isolated V8 sandbox environments for multi-tenant agent workloads on top of Cloudflare Workers.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure sandbox
cloudflare-workers Serverless edge-compute platform for deploying agent functions and MCP servers at the network edge.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure edge-compute
coder Self-hosted remote development environment platform for provisioning agent dev workspaces on any cloud.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure dev-environments
deno-deploy Edge serverless runtime for deploying TypeScript agent workers globally with zero config and V8 isolation.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure edge-compute
evals-cookbooks OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals cookbook templates openai
Previous Page 3 of 4 Next
Browse · Armory