Skip to content

Repository files navigation

BenchCAD

BenchCAD

A benchmark for evaluating LLMs and multimodal models on programmatic CAD.

Website Leaderboard Dataset Code License: MIT Data License: CC BY 4.0 Python 3.11

Website · Leaderboard · Dataset · Contributing


BenchCAD evaluates whether a model can understand and write parametric CAD code — the CadQuery programs that generate real mechanical parts. It is built on 17,900 execution-verified CadQuery programs across 106 industrial part families (bevel gears, compression springs, twist drills, threaded adapters, …) drawn from 47 engineering standards (ISO / DIN / EN / ASME / IEC). The benchmark decomposes model ability into four tasks (across three task dirs — QA splits into a vision and a code variant) spanning perception, parametric abstraction, and executable synthesis.

Scoring is execution-grounded and deterministic — there is no LLM judge. Generation tasks grade by voxel IoU between the model's executed STEP solid and the ground-truth STEP; the QA task grades by symmetric ratio accuracy on numeric answers. Every score is reproducible by re-running the scorer, never model-judged.

Tasks

Task Input → Output Metric
Vision2Code (Vision2Code/) rendered views of a part → a CadQuery program voxel IoU vs ground-truth STEP
CodeEdit (CodeEdit/) instruction (+ part) → edit an existing CadQuery program normalized IoU (improvement over baseline)
Vision-QA (QA/, mode: img) rendered views + a numeric question → a number symmetric ratio accuracy
Code-QA (QA/, mode: code) CadQuery code + a numeric question → a number symmetric ratio accuracy

Why BenchCAD

  • Objective, execution-grounded labels. Ground truth is real geometry; scores come from a CAD kernel, not an LLM judge or human vote — they can't be gamed by fluent-but-wrong outputs.
  • Multimodal. Tasks probe code-only, vision-only, and combined reasoning.
  • Industry-standard parts. Real mechanical families and standards, not toy primitives.
  • Reproducible. Pinned environment, one-command runs, deterministic scoring.

Installation

# Python 3.11, managed by uv (https://docs.astral.sh/uv/)
uv sync

# LLM API keys only — benchmark data is public
cp .env.example .env   # then paste OPENAI / ANTHROPIC / GEMINI / XAI / OPENROUTER keys

Quick start

After uv sync and pasting a key into .env, one command runs everything — data is pulled from HuggingFace on demand, nothing to download by hand:

# Quick smoke: all three tasks, 5 records each
uv run python benchcad.py --model gpt-4o

# The full benchmark: all three tasks, full split
uv run python benchcad.py --task all --num all --model gpt-4o

# One task, 100 records, reproducible random sample
uv run python benchcad.py --task vision2code --num 100 --model claude-opus-4-7 --seed 42

# Several models at once; composite (fused) scoring for Vision2Code
uv run python benchcad.py --num all --model gpt-4o gemini-3-pro-preview --score composite
Flag Meaning
--task all (default) / vision2code / codeedit / qa
--num all, or a number of records per task (default 5, a smoke)
--model one or more model names (required)
--seed reproducible random --num N sample
--score Vision2Code scoring: iou (default) or composite

Each task is also runnable on its own with finer control (cd Vision2Code && uv run python main.py --config configs/prod.yaml --model gpt-4o); see the per-task READMEs.

Dataset

Hosted on HuggingFace at BenchCAD/BenchCAD and pulled into the gitignored data/ folder on first prod run. One config per task:

Task Config Size Contents
Vision2Code code_gen 17,900 GT CadQuery code + 4 rendered views per part (106 families)
CodeEdit edit-bench 748 instruction-guided edit benchmark
Vision-QA / Code-QA QA 2,400 numeric questions over 200 parts (dimensions, counts, ratios); asked from the rendered image (mode: img) or the CadQuery code (mode: code)

Vision2Code input. The model sees one 524×524 px image: four diagonal views of the part, each 256×256 px, in a 2×2 grid with a 4 px white gutter. The runner renders it from the ground-truth STEP with benchcad_core/scoring/views.py, as it has since the first public release. This is the canonical input, and every Vision2Code score this repository produces is measured on it. The view_*_png / composite_png columns of code_gen are 128 px previews from a different renderer. They are not the evaluation input, and scores measured on them are not comparable with scores from this runner (#54).

A tiny test_data/ (≈4 records) is committed per task for smoke tests without any download. Dataset schema and column details are documented on the dataset card.

Repository layout

BenchCAD/
├── benchcad.py             one-click runner (--task / --num / --model)
├── pyproject.toml          shared, pinned environment
├── Vision2Code/  CodeEdit/  QA/    three task dirs — QA serves Vision-QA + Code-QA (mode: img / code)
├── tools/                  regrade / validate_task / validate_family / ingest_to_hf
└── contributions/          community-submitted parts (see CONTRIBUTING.md)

Each task subdir shares the same shape: main.py, configs/{test,prod}.yaml, pipeline/, scoring/, models/, test_data/, tools/download_*.py, README.md.

Contributing

See CONTRIBUTING.md. You can contribute a new part/QA item, a model's results (we re-grade raw predictions — numbers are never self-reported), or code / errata fixes.

Citation

@misc{benchcad2026,
  title  = {BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD},
  author = {{BenchCAD Authors}},
  year   = {2026}
}

License

Code is released under the MIT License; the dataset is released under CC BY 4.0.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Stars

76 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages