Skip to content

experiment(training): evaluate Jev-Omni as Studio pre-review sidecar - #30

Merged
Teakowa merged 1 commit into
mainfrom
experiment/jev-pre-review
Sep 29, 2026
Merged

Teakowa merged 1 commit into
mainfrom
experiment/jev-pre-review

Conversation

@e54-bot

@e54-bot e54-bot commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Summary

Implements the offline, replayable Jev-Omni evaluation required by #25:

  • training/jev/ materializes residual Studio review rows (rows the deterministic auto-accept/auto-reject rules do not decide) into digest-verified PreReviewRecords: crop SHA-256, de-duplicated OCR candidates, bounded options (+ "none correct" / "not a valid target"), engine evidence, and human truth — ground truth excluded from the input digest.
  • OmniMlxRunner runs Ruiruiz30/Jev-Omni-MLX-4bit locally in an isolated stdlib worker subprocess inside its own venv — no mlx/mlx-vlm in project dependencies, no external service, and production/training paths untouched.
  • Evaluation fits a temperature on a source-grouped fit split, picks the lowest threshold meeting the gate (≥25% review reduction, ≤2% false-confident, ≥97% accuracy), and reports per-ROI review reduction, false-confident accepts/rejects, manual-transcription share, latency/memory on held-out rows.
  • Every decision is persisted to decisions.jsonl; --replay reproduces reports without model calls.

Measured result (929 residual rows, 35/70/140 image-token budgets, two prompt variants): auto-decisions are 100% accurate at safe thresholds but reduce review by at most 9.3% — far under the gate — and auto-reject never engages because most reject reasons are contextual rather than visible in the crop. Report recommendation: remove; final lifecycle call stays with the maintainer per the issue.

training/README.md documents reproduction and the measured outcome.

Fixes #25

Test plan

  • uv run pytest — 291 passed (12 new tests covering records/digest integrity, routing safety invariants, threshold/temperature fitting, replay equivalence, worker protocol)
  • Full MLX experiment over all 5 local Studio batches: 2,787 decisions, 0 runtime errors
  • git diff --check clean; model weights/venv/reports stay under ignored training/.work/

Generated with Devin

Adds an offline, replayable experiment (training/jev + run_jev_experiment.py) that materializes residual Studio review rows into digest-verified records, routes them through the local Jev-Omni-MLX-4bit worker, and fits calibrated routing thresholds on a source-grouped holdout.

Measured on 929 residual rows across 35/70/140 image-token budgets and two prompt variants: auto-decisions reach 100% accuracy at safe thresholds but reduce review by at most 9.3%, far below the 25% gate, and auto-reject never engages because most reject reasons are contextual rather than visible in the crop. Recommendation: remove.

Fixes #25

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@Teakowa
Teakowa merged commit e34eee4 into main Sep 29, 2026
1 check passed
@Teakowa
Teakowa deleted the experiment/jev-pre-review branch September 29, 2026 17:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

experiment(dataset): evaluate Jev-Omni as a dataset and evaluation sidecar

2 participants