207 results in Memory, Evals, Hooks · page 2 of 9
Harness Claude Code Cursor Codex Gemini OpenCode giskard Open-source LLM testing framework for detecting vulnerabilities (prompt injection, hallucinations, bias) via automated scan.
99.0 5,838 stars · 542 forks
Harness claude cursor codex opencode gemini
evals safety vulnerability scan
getzep-zep Examples, framework integrations and tools for Zep Cloud, Zep's hosted agent memory service; the repository says it is not the product itself. The open-source knowledge-graph engine behind Zep is Graphiti (getzep-graphiti).
Contributed by Sentinel
99.0 4,882 stars · 651 forks · 7 mentions
Harness claude codex cursor gemini opencode
memory
agenta Open-source LLM developer platform with prompt playground, evaluation pipelines, and A/B testing for iterating on LLM apps.
98.9 4,670 stars · 661 forks
Harness claude cursor codex opencode gemini
evals playground ab-testing ci
openai-simple-evals OpenAI's lightweight benchmark suite (MMLU, HumanEval, MATH, GPQA, MGSM) for fast model capability comparisons.
98.9 4,621 stars · 509 forks
Harness claude cursor codex opencode gemini
evals benchmark mmlu simple
caviraoss-openmemory Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
Contributed by Sentinel
98.9 4,478 stars · 504 forks
Harness claude codex cursor gemini opencode
memory
big-bench Google's Beyond the Imitation Game benchmark: 200+ diverse tasks designed to probe capabilities beyond standard NLP benchmarks.
98.8 3,247 stars · 618 forks
Harness claude cursor codex opencode gemini
evals benchmark google academic
flowelement-m-flow Use when recall should follow associations between memories rather than nearest-neighbour similarity alone.
98.8 4,497 stars · 255 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
trulens Evaluation and tracking for LLM and RAG applications with a feedback-function API and experiment dashboard.
98.7 3,530 stars · 335 forks · 1 mention
Harness claude cursor codex opencode gemini
evals rag tracking dashboard
xlang-ai-osworld [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Contributed by Sentinel
98.7 3,117 stars · 530 forks · 7 mentions
Harness claude codex cursor gemini opencode
inspect-ai UK AISI's framework for safety evaluations of large language models, with task-based scaffolding and solver pipelines.
98.7 2,683 stars · 685 forks
Harness claude cursor codex opencode gemini
evals safety aisi government
mirix-ai-mirix Use when what the agent should remember is what actually happened on screen, consolidated into structured memories.
98.7 3,440 stars · 270 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
helm Stanford CRFM Holistic Evaluation of Language Models: standardized benchmark suite covering accuracy, calibration, robustness, and fairness.
98.7 2,898 stars · 412 forks
Harness claude cursor codex opencode gemini
evals benchmark academic stanford
lighteval Hugging Face lightweight evaluation library for LLMs across academic benchmarks, with fast local and remote inference support.
98.6 2,533 stars · 553 forks
Harness claude cursor codex opencode gemini
evals huggingface benchmark lightweight
kayba-ai-agentic-context-engine Use when an agent should carry forward what it learned from its own successes and failures into later runs.
98.5 2,564 stars · 307 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
memodb-io-memobase User Profile-Based Long-Term Memory for AI Chatbot Applications.
Contributed by Sentinel
98.5 2,875 stars · 232 forks
Harness claude codex cursor gemini opencode
memory
zilliztech-memsearch Use when several coding agents should share one memory store instead of each keeping its own notes.
98.5 2,571 stars · 238 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
kingjulio8238-memary Use when you want a well-known reference implementation of agent memory over a knowledge graph to read or fork.
98.5 2,644 stars · 205 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
tdd-guard A hooks-driven system that monitors file operations in real-time and blocks changes that violate TDD principles.
98.4 2,324 stars · 185 forks
Harness claude
hook
redplanethq-core Use when one memory graph should serve Claude Code, Codex and your other assistants at once.
98.3 1,963 stars · 189 forks
Harness claude codex cursor gemini opencode
cp138-seed memory
evalplus Rigorous code generation evaluation framework built on top of HumanEval and MBPP with 80x more test cases.
98.3 1,819 stars · 208 forks
Harness claude cursor codex opencode gemini
evals code-generation humaneval benchmark
webarena WebArena: realistic web-based environment for evaluating autonomous agents on long-horizon browser interaction tasks.
98.2 1,592 stars · 249 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents browser benchmark
langchain-ai-langmem Long-term memory for agents: tools that extract what matters from conversations, refine prompts from feedback and keep memory across sessions, with LangGraph's store built in.
Contributed by Sentinel
98.2 1,684 stars · 192 forks
Harness claude codex cursor gemini opencode
memory
tau-bench Tau-bench: agent benchmark for tool-agent-user interactions in retail and airline domains with policy-grounded evaluation.
98.1 1,416 stars · 215 forks · 1 mention
Harness claude cursor codex opencode gemini
evals agents tool-use benchmark
harbor-framework-terminal-bench Measuring and evolving with the frontier of agent work
Contributed by Sentinel
98.1 588 stars · 434 forks · 13 mentions
Harness claude codex cursor gemini opencode
Previous Page 2 of 9 Next
Browse · Armory