Eval framework. Define correct, test against it, get results.
-
Updated
Feb 17, 2026 - Go
Eval framework. Define correct, test against it, get results.
Find what your AI agent gets wrong — before you have a rubric. Qualitative eval for PMs.
A web-based interactive demo for the GuessArena evaluation framework
Curated AI agent evaluation skills from Microsoft's Eval Guide — plan, generate, run, and interpret eval suites for Copilot Studio agents
One-stop CLI for running structured skill evals across all agent harnesses
4-model parallel planning workflow with eval framework — Claude, Gemini, Codex, GLM-5 · OpenClaw ecosystem
Evaluation framework for testing LLM outputs locally. Define prompt templates and custom scorers, run evals against OpenAI or Anthropic models, store results in SQLite, and browse via Rich CLI dashboard or FastAPI UI. Lightweight, self-contained, extensible Python library.
🚀 基于Java的开源AI自动化评测框架 / An open source AI automation evaluation framework based on Java
Python tool for extracting validated, schema-defined JSON from documents (PDF, HTML, Markdown, plaintext) using LLMs, with source grounding and a built-in eval harness. Works with Anthropic Claude or OpenAI.
Open-source evaluation framework for MCP servers powered by Claude. Auto-discovers tools, generates test scenarios, runs LLM-as-judge scoring to assess correctness and safety, audits for security issues, and visualizes traces in a web dashboard.
Observability layer for multi-step AI pipelines — traces execution, auto-diagnoses root causes, and builds a growing eval dataset from human feedback
Binary safety verdicts (SAFE/HELD/LEAK/MISS/BROKE) + persona fan-out for LLM pipeline evals
Autonomous Code Agent Tooling, Strict Eval Frameworks & Local-First CLI Utilities
Open-source evaluation framework for AI agents. Define test suites with rubrics, run your agent, get LLM-as-judge scores against criteria, inspect full execution traces, and diff runs to catch behavioral regressions.
Define YAML rubrics, run agents through test scenarios, get LLM-judged per-criterion scores with full trajectory traces, and analyze results in an interactive web dashboard.
Lightweight CLI for versioning prompts and running eval suites. Score outputs with deterministic matching or LLM-as-judge, compare prompt versions with rich terminal diffs. No infra, git-friendly, local-first.
Self-hosted evaluation framework for MCP servers. Define YAML test suites, run agents against MCP tools, score with LLM-as-judge rubrics, and monitor results in a FastAPI dashboard with tool-call traces and regression detection.
A FastAPI WebSocket service and CI-friendly CLI for running LangSmith evaluations on demand — bundles LLM-as-judge (Claude) and heuristic evaluators, syncs datasets idempotently, and gates deploys via threshold checks.
Self-hostable LLM evaluation framework for measuring model performance across configurable skills. Run YAML-defined benchmarks against OpenAI/Anthropic models, score with LLM-as-judge, compare results in a CLI and web dashboard.
Multi-model LLM evaluation using Promptfoo — benchmarks Claude, Gemini, and GPT-4.1-mini on a review classification task, combining deterministic JSON checks with LLM-as-a-judge assertions for field names, valid values, and correctness.
To associate your repository with the eval-framework topic, visit your repo's landing page and select "manage topics."