826 results in Observability, Memory, Workflows, Evals · page 3 of 35
Harness Claude Code Cursor Codex Gemini OpenCode grafana-tempo Grafana Tempo is a cost-efficient distributed tracing backend (OpenTelemetry-native) that pairs with Loki for logs and Prometheus for metrics in LLM stacks.
99.0 5,461 stars · 746 forks
Harness claude cursor codex opencode gemini
observability tracing distributed
learn-agentic-ai Learn Agentic AI using Dapr Agentic Cloud Ascent (DACA) Design Pattern: OpenAI Agents SDK, Memory, MCP, A2A, Knowledge Graphs, Rancher Desktop, and Kubernetes
99.0 4,351 stars · 1,008 forks
Harness claude cursor codex opencode gemini
a2a agent-to-agent tutorials-learning-resources
getzep-zep Examples, framework integrations and tools for Zep Cloud, Zep's hosted agent memory service; the repository says it is not the product itself. The open-source knowledge-graph engine behind Zep is Graphiti (getzep-graphiti).
Contributed by Sentinel
99.0 4,882 stars · 651 forks · 7 mentions
Harness claude codex cursor gemini opencode
memory
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
caviraoss-openmemory Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
Contributed by Sentinel
98.9 4,478 stars · 504 forks
Harness claude codex cursor gemini opencode
memory
ysymyth-react [ICLR 2023] ReAct: Synergizing Reasoning and Acting in Language Models
Contributed by Sentinel
98.8 4,140 stars · 401 forks · 13 mentions
Harness claude codex cursor gemini opencode
big-bench Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.8 3,247 stars · 618 forks
Harness claude cursor codex opencode gemini
evals benchmark google academic
pydantic-logfire Logfire by Pydantic: OpenTelemetry-based structured logging and tracing for Python applications with built-in support for FastAPI, SQLAlchemy, and Anthropic.
98.8 4,450 stars · 283 forks
Harness claude cursor codex opencode gemini
observability opentelemetry logging
flowelement-m-flow Use when recall should follow associations between memories rather than nearest-neighbour similarity alone.
98.8 4,497 stars · 255 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
langwatch LangWatch provides real-time LLM analytics, guardrails, and evaluation pipelines with a visual studio for monitoring multi-step agent conversations.
98.7 3,522 stars · 362 forks
Harness claude cursor codex opencode gemini
observability tracing guardrails
trulens Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.7 3,530 stars · 335 forks · 1 mention
Harness claude cursor codex opencode gemini
evals rag tracking dashboard
xlang-ai-osworld [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.7 3,117 stars · 530 forks · 7 mentions
Harness claude codex cursor gemini opencode
inspect-ai UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.7 2,683 stars · 685 forks
Harness claude cursor codex opencode gemini
evals safety aisi government
noahshinn-reflexion [NeurIPS 2023] Reflexion: Language Agents with Verbal Reinforcement Learning
Contributed by Sentinel
98.7 3,251 stars · 317 forks · 3 mentions
Harness claude codex cursor gemini opencode
mirix-ai-mirix Use when what the agent should remember is what actually happened on screen, consolidated into structured memories.
98.7 3,440 stars · 270 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
helm Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.7 2,898 stars · 412 forks
Harness claude cursor codex opencode gemini
evals benchmark academic stanford
lighteval Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.
98.6 2,533 stars · 553 forks
Harness claude cursor codex opencode gemini
evals huggingface benchmark lightweight
ralph-orchestrator Ralph Orchestrator implements the simple but effective "Ralph Wiggum" technique for autonomous task completion, continuously running an AI agent against a prompt file until the task is marked as complete or limits are reached. This implementation provides a robust, well-tested, and feature-complete orchestration system for AI-driven development. Also cited in the Anthropic Ralph plugin documentation.
98.6 3,120 stars · 292 forks
Harness claude
claude-code workflows-knowledge-guides
ccometixline-claude-code-statusline A high-performance Claude Code statusline tool written in Rust with Git integration, usage tracking, interactive TUI configuration, and Claude Code enhancement utilities.
98.6 3,456 stars · 215 forks
Harness claude
claude-code status-lines
openlit OpenLIT is an OpenTelemetry-native LLM observability toolkit with GPU monitoring, cost tracking, and a prompt hub (one-line setup for 20+ providers).
98.6 2,736 stars · 367 forks
Harness claude cursor codex opencode gemini
observability opentelemetry gpu
laminar Laminar is an open-source platform for tracing, evaluating, and labeling LLM and agent pipelines with a TypeScript/Python SDK and a self-hostable backend.
98.6 3,218 stars · 229 forks · 1 mention
Harness claude cursor codex opencode gemini
observability tracing evals
kayba-ai-agentic-context-engine Use when an agent should carry forward what it learned from its own successes and failures into later runs.
98.5 2,564 stars · 307 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
memodb-io-memobase User Profile-Based Long-Term Memory for AI Chatbot Applications.
Contributed by Sentinel
98.5 2,875 stars · 232 forks
Harness claude codex cursor gemini opencode
memory
Previous Page 3 of 35 Next
Browse · Armory