Armory
Source

Browse

Search and filter by type across the catalog

162 results in Infrastructure, Evals, Memory, Observability · page 6 of 7

InfrastructurePreview

vercel

Frontend cloud platform with serverless functions and AI SDK integrations for deploying agent-facing UIs.

80.0passed install test
Harnessclaudecursorcodexopencodegemini
infrastructuredeploy
No one-command install · SourceDetails
ObservabilityPreview

honeycomb

Honeycomb's OpenTelemetry-native observability platform: a high-cardinality event store ideal for tracing LLM pipelines and debugging slow agent traces.

75.116 stars · 6 forks
Harnessclaudecursorcodexopencodegemini
observabilityopentelemetrytracing
No one-command install · SourceDetails
ObservabilityPreview

baserun

Baserun captures LLM traces via a lightweight decorator-based SDK and provides a dashboard for debugging prompt chains, testing variants, and measuring quality.

74.416 stars · 5 forks
Harnessclaudecursorcodexopencodegemini
observabilitytracingtesting
No one-command install · SourceDetails
EvalsPreview

humanloop-evals

Humanloop Python SDK with integrated evals, dataset versioning, and human + LLM judge scoring for production pipelines.

69.012 stars · 3 forks
Harnessclaudecursorcodexopencodegemini
evalshuman-evaldatasetsdk
No one-command install · SourceDetails
MemoryPreview

wikimem

Use to give an agent a queryable wiki knowledge base when memory should be a navigable knowledge graph, not just a flat log: ingest files, folders, and URLs into a linked vault, then search or ask it in natural language.

63.47 stars · 4 forks
Harnessclaude
memoryknowledge-basewikiingest
No one-command install · SourceDetails
EvalsPreview

agentbench

Use to put a number on harness quality: run an agent harness against a task set and get a score, so harness changes are validated by evidence. It is the eval backbone of a self-improving loop.

UnrankedNo signals yet
Harnessclaudecodex
evalbenchmarkscoringharness
No one-command install · SourceDetails
InfrastructurePreview

aws-lambda

Serverless function-as-a-service platform for event-driven agent compute without provisioning servers.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructureserverless
No one-command install · SourceDetails
InfrastructurePreview

beta9

Open-source serverless GPU container runtime for running AI workloads with fast cold-starts on bare-metal.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructuregpu-serverless
No one-command install · SourceDetails
InfrastructurePreview

blaxel

Cloud runtime and control plane for deploying, scaling, and observing production AI agent workloads.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructureagent-compute
No one-command install · SourceDetails
InfrastructurePreview

cloudflare-sandboxes

Isolated V8 sandbox environments for multi-tenant agent workloads on top of Cloudflare Workers.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructuresandbox
No one-command install · SourceDetails
InfrastructurePreview

cloudflare-workers

Serverless edge-compute platform for deploying agent functions and MCP servers at the network edge.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructureedge-compute
No one-command install · SourceDetails
InfrastructurePreview

coder

Self-hosted remote development environment platform for provisioning agent dev workspaces on any cloud.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructuredev-environments
No one-command install · SourceDetails
ObservabilityPreview

datadog-llm-observability

Datadog's managed LLM Observability product. It traces LLM calls, monitors prompt/completion quality, detects anomalies, and integrates with existing APM.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitymanagedapm
No one-command install · SourceDetails
InfrastructurePreview

deno-deploy

Edge serverless runtime for deploying TypeScript agent workers globally with zero config and V8 isolation.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructureedge-compute
No one-command install · SourceDetails
EvalsPreview

evals-cookbooks

OpenAI Cookbook eval recipes: task-specific templates for summarization, QA, and classification evaluation using the Evals framework.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalscookbooktemplatesopenai
No one-command install · SourceDetails
InfrastructurePreview

fermyon-spin

WebAssembly serverless framework for building and deploying fast, portable agent microservices.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructurewasm-serverless
No one-command install · SourceDetails
ObservabilityPreview

fiddler-ai

Fiddler AI Observability platform monitors LLM applications for hallucinations, toxicity, bias, and drift, with explainability and alerting for production AI.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilitymanagedsafety
No one-command install · SourceDetails
InfrastructurePreview

firecracker

AWS-open-sourced microVM technology powering secure, fast (<125ms) isolation for serverless and container workloads.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructuremicrovm
No one-command install · SourceDetails
InfrastructurePreview

fly-io

Deploy full-stack apps and long-running agent processes globally via Firecracker microVMs close to users.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructuredeploy
No one-command install · SourceDetails
EvalsPreview

gaia-benchmark

GAIA: benchmark of 466 real-world questions requiring multi-step reasoning, web browsing, and tool use for general AI assistants.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsbenchmarkagentstool-use
No one-command install · SourceDetails
InfrastructurePreview

gvisor

Google-open-sourced application kernel providing a sandboxed container runtime for untrusted agent code.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructuresandbox
No one-command install · SourceDetails
EvalsPreview

honeyhive

LLM evaluation and experimentation platform with session tracing, dataset management, and metric-based run comparison.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
evalsexperiment-trackingtracingdataset
No one-command install · SourceDetails
ObservabilityPreview

honeyhive

HoneyHive is an AI evaluation and observability platform for tracing agent pipelines, running evaluations, and debugging regressions in production.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
observabilityevalstracing
No one-command install · SourceDetails
InfrastructurePreview

kata-containers

Lightweight VMs that combine container speed with VM-level security isolation for agent runtime workloads.

UnrankedNo signals yet
Harnessclaudecursorcodexopencodegemini
infrastructuremicrovm
No one-command install · SourceDetails
Browse · Armory