Skip to content

feat(training): add Colab GPU recognition runner - #21

Merged
Teakowa merged 15 commits into
mainfrom
feat/ocrkit-20-colab-runner
Sep 27, 2026
Merged

Teakowa merged 15 commits into
mainfrom
feat/ocrkit-20-colab-runner

Conversation

@e54-bot

@e54-bot e54-bot commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Add an OCRKit-owned Colab CLI runner around the existing recognition training and fixture evaluation scripts.
  • Stage the selected reviewed train/holdout labels and referenced crops, provenance, required source and evaluator inputs, and pretrained checkpoint; verify staged and returned checksums.
  • Add CUDA-aware PaddlePaddle setup with a selectable GPU preference, explicit compatibility checks, and runtime teardown on provisioning and execution failures. Keep local CPU training available.

Verification

  • bash -n training/bootstrap.sh training/setup_rec_environment.sh training/run_rec_smoke.sh
  • python3 -m py_compile training/run_rec_colab.py training/colab_remote.py
  • Mocked Colab CLI runs covered success, training failure, evaluation-gate failure, and GPU provisioning failure; failure paths stopped the runtime and did not create top-level success artifacts.
  • Mocked CUDA setup covered a supported GPU and rejection of compute capability 7.5.
  • No live Colab GPU allocation or training run was performed.

Closes #20

@Teakowa Teakowa left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Teardown failure handling needs correction.

Comment thread training/run_rec_colab.py
except Exception as exc:
stop_status = -1
stop_error = f"Colab runtime teardown command failed: {type(exc).__name__}: {exc}"
error = f"{error}; {stop_error}" if error else stop_error

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If colab stop returns nonzero after an earlier provisioning/training error, error is already set, so the manual teardown command is never recorded. The runtime may remain allocated, while status.json only contains the original failure. Always include colab stop -s <session> when stop fails, even when another error occurred.

@Teakowa Teakowa left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review of the Colab GPU runner. The design fits #20: a thin adapter over the existing smoke/evaluation scripts, private inputs staged as a checksummed archive, no publish/promote path, teardown in finally. Three items need fixing before merge.

Must fix

  1. Colab CLI semantics are unverified. colab new -s … --gpu, upload, exec -f, download, stop -s, and the assumption that exec propagates the script exit code were only exercised against mocks. The failure/teardown contract depends on them; one real run or a check against the CLI is needed.
  2. A failed run can leave success-looking output. copy_remote_result (success branch) copies checkpoint/, evaluation/ and run.json into run_dir before the run is final. A later failure (partial copy, teardown failure, status check) leaves them in place. Stage into a temp dir and rename only after teardown succeeds; otherwise move under partial/.
  3. No committed tests. The mocked scenarios (success, training failure, gate failure, GPU provisioning failure, unsupported compute capability) exist only in the PR description. Commit them as a focused test and confirm they fail when teardown or the gate check is removed.

Should fix

  • run_request["training"] hardcodes batch size, first_bs, eval_batch_step and thresholds duplicated from the shell scripts, so the recorded "effective configuration" can drift from what ran. The 0.9604221635883905 gate now lives in three places.
  • setup_rec_environment.sh: unrelated BASH_SOURCE[0] → $0 change and ${var} → $var churn; CUDA path installs, uninstalls and reinstalls Paddle (double download).
  • ocrkit_worktree_dirty is recorded but a dirty tree's revision does not identify the staged code.

Minor

  • No timeout/idle bound on colab exec.
  • If exec fails before the result archive exists, the download error masks the real failure.
  • Confirm paddlepaddle-gpu==3.3.1 exists on the cu129 index.

Stage returned artifacts and promote them to the run directory only after
runtime teardown succeeds; otherwise keep them under partial/. Drop the
duplicated accuracy gates and hardcoded training values so the shared
evaluation script stays the single source of truth. Restore the unrelated
setup script churn, install Paddle once for CUDA, and surface the exec
failure when result retrieval also fails. Add mocked runner tests.

Refs #20
colab exec defaults to a 30 second timeout, which would abort training.
Pass an explicit --timeout-seconds bound and report the remote error when
the run metadata is not successful, since colab exec does not map a remote
SystemExit to its exit status.

Refs #20
PaddlePaddle 3.3.1 cu126 and cu129 run convolution forward/backward and
matmul on a Colab T4 (compute capability 7.5), so drop the capability gate
and rely on the runtime device check. Colab ships without python3-venv, so
bootstrap creates the venv with uv when OCRKIT_TRAINING_PYTHON is set.
Default the GPU preference to T4.

Refs #20
colab upload and download send each file as a single base64 JSON request,
and a 145 MB input archive was rejected with HTTP 400. Send the input in
32 MiB parts that the runtime joins, and return the result as checksummed
parts listed in an index.

Refs #20
paddle2onnx 2.1.0 has no Python 3.13 wheel, and Colab now ships 3.13.
Create the training venv with uv on Python 3.12, matching the local
environment, and drop the OCRKIT_TRAINING_PYTHON override.

Refs #20
Colab provides python3.12 without ensurepip, so the plain venv step fails
there. Keep python3.12 -m venv as the first choice and fall back to a
uv-managed Python 3.12 venv.

Refs #20
The official cu129 PaddlePaddle wheel is 2.9 GB and downloads at about
0.4 MB/s from a Colab runtime. Install a mirror of the identical file from
the OWBastion CDN, verify its sha256, and keep the official index as the
fallback for other CUDA versions.

Refs #20
The remote runner read request.run.epochs, but the request stores it under
run.training. Assert the request keys the remote runner reads.

Refs #20
A checkpoint trained on a CUDA device was exported with a GPU-specific
inference graph that paddle2onnx aborts on. Exporting the same checkpoint
on CPU converts and passes the fixture gate, and matches the existing
CPU-only local flow, so make the export device independent.

Refs #20
CUDA builds of PaddlePaddle export nn.Linear as linear_v2, which
paddle2onnx 2.1.0 aborts on, so the evaluation export cannot run on the
Colab GPU runtime. Train on Colab with --train-only, stop the runtime after
retrieval, and run the unchanged evaluate_rec_checkpoint.sh locally. The run
only succeeds when that evaluation passes; otherwise the checkpoint is kept
under partial/. Stop uploading the evaluation-only inputs.

Refs #20
Return best_accuracy.pdparams, config and the training log instead of the
full checkpoint directory. The optimizer state and latest checkpoint are
not needed until resume is supported and made the result about five times
larger. Retry each part transfer up to three times.

Refs #20
The Colab CLI transfers files as single base64-encoded JSON requests, which
measured well under 1 MB/s for large binaries. Retrieving a 119 MB checkpoint
through it took several minutes and repeatedly lost the connection when the
Colab runtime was recycled mid-transfer, corrupting or losing the run.

Have Colab upload the trained checkpoint directly to a private R2 bucket
using a short-lived, single-object presigned PUT URL the runner generates
locally with existing OCRKIT_R2_* credentials; Colab never receives R2
credentials, and the URL is redacted before it is logged or persisted. The
runner downloads the checkpoint from R2 after stopping the Colab runtime,
verifies its checksum, and deletes the object whether or not the run
succeeded. Only the small run.json and remote.log now travel through the
Colab CLI.

The 125 MB official base checkpoint no longer needs to be uploaded from the
local machine either: Colab fetches it directly from PaddlePaddle's public
model CDN and verifies it against a pinned sha256. A checkpoint passed via
--pretrained-checkpoint continues to upload through the CLI, since it is not
publicly hosted.

Add per-stage timing to the remote run for future capacity planning, and
require the R2 environment variables up front with a clear error instead of
failing partway through a run.

Verified live: presigned PUT from a Colab T4 session to R2 uploaded 120 MB in
about 4 seconds. Retrieval and checksum verification back through boto3 are
covered by new tests, falsified by disabling the delete-on-finally guarantee,
the early-record-before-later-failure guarantee, and the R2 error-code
mapping, each of which failed the corresponding test before the fix.

Refs #20
…ectly

Running `python training/run_rec_colab.py` puts training/, not the repo
root, on sys.path, so the new `from app...` imports failed with
ModuleNotFoundError. Insert the repo root before importing, matching the
sys.path shim already used by training/scripts/*.py.

Refs #20
colab stop can fail with Not Found simply because Colab already reclaimed
an idle or finished runtime on its own; a completed, checksum-verified
training run was being reported as failed and its checkpoint demoted to
partial/ solely because of this race. Verify with `colab sessions` before
treating a failed stop as a real teardown failure.

Verified live: a full T4 run (base checkpoint fetched by Colab, trained,
checksum-verified checkpoint retrieved through R2) hit exactly this race;
colab sessions confirmed no session was actually left running.

Refs #20
@Teakowa
Teakowa merged commit 66fe0e5 into main Sep 27, 2026
2 checks passed
@Teakowa
Teakowa deleted the feat/ocrkit-20-colab-runner branch September 27, 2026 08:23
@Teakowa Teakowa mentioned this pull request Sep 27, 2026
11 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(training): add Colab CLI GPU training runner

2 participants