126 results in Evals, Infrastructure, Observability · page 5 of 6
Harness Claude Code Cursor Codex Gemini OpenCode datadog-llm-observability Datadog's managed LLM Observability product. It traces LLM calls, monitors prompt/completion quality, detects anomalies, and integrates with existing APM.
Unranked No signals yet
Harness claude cursor codex opencode gemini
observability managed apm
deno-deploy Edge serverless runtime for deploying TypeScript agent workers globally with zero config and V8 isolation.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure edge-compute
evals-cookbooks OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals cookbook templates openai
fermyon-spin WebAssembly serverless framework for building and deploying fast, portable agent microservices.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure wasm-serverless
fiddler-ai Fiddler AI Observability platform monitors LLM applications for hallucinations, toxicity, bias, and drift, with explainability and alerting for production AI.
Unranked No signals yet
Harness claude cursor codex opencode gemini
observability managed safety
firecracker AWS-open-sourced microVM technology powering secure, fast (<125ms) isolation for serverless and container workloads.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure microvm
fly-io Deploy full-stack apps and long-running agent processes globally via Firecracker microVMs close to users.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure deploy
gaia-benchmark GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals benchmark agents tool-use
gvisor Google-open-sourced application kernel providing a sandboxed container runtime for untrusted agent code.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure sandbox
honeyhive LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals experiment-tracking tracing dataset
honeyhive HoneyHive is an AI evaluation and observability platform for tracing agent pipelines, running evaluations, and debugging regressions in production.
Unranked No signals yet
Harness claude cursor codex opencode gemini
observability evals tracing
kata-containers Lightweight VMs that combine container speed with VM-level security isolation for agent runtime workloads.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure microvm
kubernetes Industry-standard container orchestration system for deploying, scaling, and managing agent workload fleets.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure orchestration
langtrace Open-source, OpenTelemetry-compliant LLM observability tool by Scale3Labs. It traces calls to all major LLM providers and frameworks with a self-hostable UI.
Unranked No signals yet
Harness claude cursor codex opencode gemini
observability opentelemetry tracing
lightpanda Ultra-fast headless browser written in Zig specifically for AI and automation workloads. It runs JavaScript natively, uses 9x less memory than Chrome headless, and targets sub-100ms page execution.
Unranked No signals yet
Harness claude cursor codex opencode gemini
browser lightpanda
literal-ai Literal AI is an observability and evaluation platform for conversational AI. It captures multi-step threads, scores responses, and integrates with Chainlit.
Unranked No signals yet
Harness claude cursor codex opencode gemini
observability tracing evals
lunary Open-source LLM observability and prompt management platform. It tracks conversations, errors, costs, and user feedback for production AI applications.
Unranked No signals yet
Harness claude cursor codex opencode gemini
observability logging evals
maxim-ai Maxim AI is an evaluation and observability platform for AI agents. It supports multi-step trace analysis, prompt testing, and production quality monitoring.
Unranked No signals yet
Harness claude cursor codex opencode gemini
observability evals agents
morph-cloud Hypervisor-level snapshotting cloud for agent inference: fork and snapshot VM state for fast agent branching.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure agent-compute
multion MultiOn AI browser agent API: a cloud service that lets developers invoke an autonomous web agent via REST to complete tasks like form filling, data extraction, and multi-step workflows on any site.
Unranked No signals yet
Harness claude cursor codex opencode gemini
browser multion
new-relic-ai-monitoring New Relic AI Monitoring instruments LLM calls end-to-end. It traces model invocations, measures token costs, and surfaces anomalies via the New Relic platform.
Unranked No signals yet
Harness claude cursor codex opencode gemini
observability managed apm
northflank Developer platform for containerized microservices, cron jobs, and persistent agent workloads with GitOps support.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure containers
patronus-ai Automated LLM evaluation and hallucination detection platform with a Python SDK and judge-model scoring.
Unranked No signals yet
Harness claude cursor codex opencode gemini
evals hallucination judge sdk
render Unified cloud for deploying web services, background workers, cron jobs, and databases for agent backends.
Unranked No signals yet
Harness claude cursor codex opencode gemini
infrastructure deploy
Previous Page 5 of 6 Next
Browse · Armory