BenchCAD evaluates whether a model can understand and write parametric CAD code — the CadQuery programs that generate real mechanical parts. It is built on 17,900 execution-verified CadQuery programs across 106 industrial part families (bevel gears, compression springs, twist drills, threaded adapters, …) drawn from 47 engineering standards (ISO / DIN / EN / ASME / IEC). The benchmark decomposes model ability into four tasks (across three task dirs — QA splits into a vision and a code variant) spanning perception, parametric abstraction, and executable synthesis.
Scoring is execution-grounded and deterministic — there is no LLM judge. Generation tasks grade by voxel IoU between the model's executed STEP solid and the ground-truth STEP; the QA task grades by symmetric ratio accuracy on numeric answers. Every score is reproducible by re-running the scorer, never model-judged.
| Task | Input → Output | Metric |
|---|---|---|
Vision2Code (Vision2Code/) |
rendered views of a part → a CadQuery program | voxel IoU vs ground-truth STEP |
CodeEdit (CodeEdit/) |
instruction (+ part) → edit an existing CadQuery program | normalized IoU (improvement over baseline) |
Vision-QA (QA/, mode: img) |
rendered views + a numeric question → a number | symmetric ratio accuracy |
Code-QA (QA/, mode: code) |
CadQuery code + a numeric question → a number | symmetric ratio accuracy |
- Objective, execution-grounded labels. Ground truth is real geometry; scores come from a CAD kernel, not an LLM judge or human vote — they can't be gamed by fluent-but-wrong outputs.
- Multimodal. Tasks probe code-only, vision-only, and combined reasoning.
- Industry-standard parts. Real mechanical families and standards, not toy primitives.
- Reproducible. Pinned environment, one-command runs, deterministic scoring.
# Python 3.11, managed by uv (https://docs.astral.sh/uv/)
uv sync
# LLM API keys only — benchmark data is public
cp .env.example .env # then paste OPENAI / ANTHROPIC / GEMINI / XAI / OPENROUTER keysAfter uv sync and pasting a key into .env, one command runs everything —
data is pulled from HuggingFace on demand, nothing to download by hand:
# Quick smoke: all three tasks, 5 records each
uv run python benchcad.py --model gpt-4o
# The full benchmark: all three tasks, full split
uv run python benchcad.py --task all --num all --model gpt-4o
# One task, 100 records, reproducible random sample
uv run python benchcad.py --task vision2code --num 100 --model claude-opus-4-7 --seed 42
# Several models at once; composite (fused) scoring for Vision2Code
uv run python benchcad.py --num all --model gpt-4o gemini-3-pro-preview --score composite| Flag | Meaning |
|---|---|
--task |
all (default) / vision2code / codeedit / qa |
--num |
all, or a number of records per task (default 5, a smoke) |
--model |
one or more model names (required) |
--seed |
reproducible random --num N sample |
--score |
Vision2Code scoring: iou (default) or composite |
Each task is also runnable on its own with finer control
(cd Vision2Code && uv run python main.py --config configs/prod.yaml --model gpt-4o);
see the per-task READMEs.
Hosted on HuggingFace at BenchCAD/BenchCAD
and pulled into the gitignored data/ folder on first prod run. One config per task:
| Task | Config | Size | Contents |
|---|---|---|---|
| Vision2Code | code_gen |
17,900 | GT CadQuery code + 4 rendered views per part (106 families) |
| CodeEdit | edit-bench |
748 | instruction-guided edit benchmark |
| Vision-QA / Code-QA | QA |
2,400 | numeric questions over 200 parts (dimensions, counts, ratios); asked from the rendered image (mode: img) or the CadQuery code (mode: code) |
Vision2Code input. The model sees one 524×524 px image: four diagonal views of
the part, each 256×256 px, in a 2×2 grid with a 4 px white gutter. The runner renders
it from the ground-truth STEP with benchcad_core/scoring/views.py, as it has since
the first public release. This is the canonical input, and every Vision2Code score this
repository produces is measured on it. The view_*_png / composite_png columns of
code_gen are 128 px previews from a different renderer. They are not the evaluation
input, and scores measured on them are not comparable with scores from this runner
(#54).
A tiny test_data/ (≈4 records) is committed per task for smoke tests without any
download. Dataset schema and column details are documented on the dataset card.
BenchCAD/
├── benchcad.py one-click runner (--task / --num / --model)
├── pyproject.toml shared, pinned environment
├── Vision2Code/ CodeEdit/ QA/ three task dirs — QA serves Vision-QA + Code-QA (mode: img / code)
├── tools/ regrade / validate_task / validate_family / ingest_to_hf
└── contributions/ community-submitted parts (see CONTRIBUTING.md)
Each task subdir shares the same shape: main.py, configs/{test,prod}.yaml,
pipeline/, scoring/, models/, test_data/, tools/download_*.py, README.md.
See CONTRIBUTING.md. You can contribute a new part/QA item, a model's results (we re-grade raw predictions — numbers are never self-reported), or code / errata fixes.
@misc{benchcad2026,
title = {BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD},
author = {{BenchCAD Authors}},
year = {2026}
}Code is released under the MIT License; the dataset is released under CC BY 4.0.