Ship evals before you ship features.
-
Updated
Jun 22, 2026 - Nunjucks
Ship evals before you ship features.
Eval framework. Define correct, test against it, get results.
A guard-railed, closed-loop workflow for AI coding agents: live state bus + execution-level hard intercepts for Claude Code and Codex (GitHub PR / GitLab MR). From step-level to requirement-level; eval-driven, spec-driven, human-in-the-loop.
AI-augmented QA platform for spec-driven development and testing, RAG-grounded analysis, eval-driven development and contract validation across Python, Go, Rust and Solidity.
Autonomous skill improvement loop for Claude Code plugins — inspired by Karpathy's autoresearch. Modify → evaluate → keep/discard → repeat until convergence. Zero-touch quality iteration at scale.
Production harness for a multi-agent BI system — eval-gated, guardrailed, cross-source-validated. LangGraph + hybrid RAG + FastAPI, live on AWS.
正解表なしで5つの性質からソートを採点するメタモルフィック・オラクル | Sort graded by metamorphic relations
非決定的なシャッフルをカイ二乗検定で採点する統計オラクル | Shuffle graded by Monte-Carlo statistical tests
LLM査読の検出力をラベル付き見本で採点するメタ評価ハーネス | Meta-eval: grading an LLM reviewer with labeled specimens
Claude Codeエージェント定義の査読と、その検出力を採点するオラクル | Agent-spec review graded by a detection-power oracle
Multi-agent inspection pipeline for solar cell EL images: EfficientNet-B0 severity classifier + Qwen3-VL (Ollama) reasoning, served via FastAPI. 75.3% on a 20-criteria eval suite.
Companion code for the talk "Managing Production Agents at Scale — from Chaos to Reliability". One Google ADK agent, three production failure modes: eval-driven development, resilience, and zero-trust on Vertex AI.
Fractal design docs: one frame, zoom into any element, same shape at every depth — with a document-structure oracle | 枠1枚から掘れるフラクタル設計書の生成エージェント+文書構造オラクル
A hands-on learning repository exploring Spec-Driven Development (SDD) for building deterministic AI systems. Covers specs, evaluation loops, patterns, experiments, and failures to bridge theory with real-world AI engineering practices.
ランダム入力の集中砲火で「壊れない」を採点するファジング・オラクル | Robust parser graded by a fuzzing (implicit) oracle
Eval 驱动的 LLM Agent 应用开发团队(Claude Code Subagents)——没有 eval 不许改 prompt
答えが一意でない出力を仕様アサーションで採点するEDD実証 | Divisor finder graded by a spec-assertion oracle
JS→TS移行を元JSとの差分テストで採点するEDD実証 | JS→TS migration graded by differential testing
Python 2→3移植をgolden比較で採点するEDD実証 | Py2→3 migration graded by golden-output comparison
Behavioral regression eval: measure whether your AI rules are still needed on current models | AIへのルールが現行モデルにまだ必要かを実測でふるい分ける行動回帰テスト+統計オラクル
To associate your repository with the eval-driven-development topic, visit your repo's landing page and select "manage topics."