feat(training): add Colab GPU recognition runner - #21
Conversation
Teakowa
left a comment
There was a problem hiding this comment.
Teardown failure handling needs correction.
| except Exception as exc: | ||
| stop_status = -1 | ||
| stop_error = f"Colab runtime teardown command failed: {type(exc).__name__}: {exc}" | ||
| error = f"{error}; {stop_error}" if error else stop_error |
There was a problem hiding this comment.
If colab stop returns nonzero after an earlier provisioning/training error, error is already set, so the manual teardown command is never recorded. The runtime may remain allocated, while status.json only contains the original failure. Always include colab stop -s <session> when stop fails, even when another error occurred.
Teakowa
left a comment
There was a problem hiding this comment.
Review of the Colab GPU runner. The design fits #20: a thin adapter over the existing smoke/evaluation scripts, private inputs staged as a checksummed archive, no publish/promote path, teardown in finally. Three items need fixing before merge.
Must fix
- Colab CLI semantics are unverified.
colab new -s … --gpu,upload,exec -f,download,stop -s, and the assumption thatexecpropagates the script exit code were only exercised against mocks. The failure/teardown contract depends on them; one real run or a check against the CLI is needed. - A failed run can leave success-looking output.
copy_remote_result(success branch) copiescheckpoint/,evaluation/andrun.jsonintorun_dirbefore the run is final. A later failure (partial copy, teardown failure, status check) leaves them in place. Stage into a temp dir and rename only after teardown succeeds; otherwise move underpartial/. - No committed tests. The mocked scenarios (success, training failure, gate failure, GPU provisioning failure, unsupported compute capability) exist only in the PR description. Commit them as a focused test and confirm they fail when teardown or the gate check is removed.
Should fix
run_request["training"]hardcodes batch size,first_bs,eval_batch_stepand thresholds duplicated from the shell scripts, so the recorded "effective configuration" can drift from what ran. The0.9604221635883905gate now lives in three places.setup_rec_environment.sh: unrelatedBASH_SOURCE[0]→$0change and${var}→$varchurn; CUDA path installs, uninstalls and reinstalls Paddle (double download).ocrkit_worktree_dirtyis recorded but a dirty tree's revision does not identify the staged code.
Minor
- No timeout/idle bound on
colab exec. - If exec fails before the result archive exists, the download error masks the real failure.
- Confirm
paddlepaddle-gpu==3.3.1exists on thecu129index.
Stage returned artifacts and promote them to the run directory only after runtime teardown succeeds; otherwise keep them under partial/. Drop the duplicated accuracy gates and hardcoded training values so the shared evaluation script stays the single source of truth. Restore the unrelated setup script churn, install Paddle once for CUDA, and surface the exec failure when result retrieval also fails. Add mocked runner tests. Refs #20
colab exec defaults to a 30 second timeout, which would abort training. Pass an explicit --timeout-seconds bound and report the remote error when the run metadata is not successful, since colab exec does not map a remote SystemExit to its exit status. Refs #20
PaddlePaddle 3.3.1 cu126 and cu129 run convolution forward/backward and matmul on a Colab T4 (compute capability 7.5), so drop the capability gate and rely on the runtime device check. Colab ships without python3-venv, so bootstrap creates the venv with uv when OCRKIT_TRAINING_PYTHON is set. Default the GPU preference to T4. Refs #20
colab upload and download send each file as a single base64 JSON request, and a 145 MB input archive was rejected with HTTP 400. Send the input in 32 MiB parts that the runtime joins, and return the result as checksummed parts listed in an index. Refs #20
paddle2onnx 2.1.0 has no Python 3.13 wheel, and Colab now ships 3.13. Create the training venv with uv on Python 3.12, matching the local environment, and drop the OCRKIT_TRAINING_PYTHON override. Refs #20
Colab provides python3.12 without ensurepip, so the plain venv step fails there. Keep python3.12 -m venv as the first choice and fall back to a uv-managed Python 3.12 venv. Refs #20
The official cu129 PaddlePaddle wheel is 2.9 GB and downloads at about 0.4 MB/s from a Colab runtime. Install a mirror of the identical file from the OWBastion CDN, verify its sha256, and keep the official index as the fallback for other CUDA versions. Refs #20
The remote runner read request.run.epochs, but the request stores it under run.training. Assert the request keys the remote runner reads. Refs #20
A checkpoint trained on a CUDA device was exported with a GPU-specific inference graph that paddle2onnx aborts on. Exporting the same checkpoint on CPU converts and passes the fixture gate, and matches the existing CPU-only local flow, so make the export device independent. Refs #20
CUDA builds of PaddlePaddle export nn.Linear as linear_v2, which paddle2onnx 2.1.0 aborts on, so the evaluation export cannot run on the Colab GPU runtime. Train on Colab with --train-only, stop the runtime after retrieval, and run the unchanged evaluate_rec_checkpoint.sh locally. The run only succeeds when that evaluation passes; otherwise the checkpoint is kept under partial/. Stop uploading the evaluation-only inputs. Refs #20
Return best_accuracy.pdparams, config and the training log instead of the full checkpoint directory. The optimizer state and latest checkpoint are not needed until resume is supported and made the result about five times larger. Retry each part transfer up to three times. Refs #20
The Colab CLI transfers files as single base64-encoded JSON requests, which measured well under 1 MB/s for large binaries. Retrieving a 119 MB checkpoint through it took several minutes and repeatedly lost the connection when the Colab runtime was recycled mid-transfer, corrupting or losing the run. Have Colab upload the trained checkpoint directly to a private R2 bucket using a short-lived, single-object presigned PUT URL the runner generates locally with existing OCRKIT_R2_* credentials; Colab never receives R2 credentials, and the URL is redacted before it is logged or persisted. The runner downloads the checkpoint from R2 after stopping the Colab runtime, verifies its checksum, and deletes the object whether or not the run succeeded. Only the small run.json and remote.log now travel through the Colab CLI. The 125 MB official base checkpoint no longer needs to be uploaded from the local machine either: Colab fetches it directly from PaddlePaddle's public model CDN and verifies it against a pinned sha256. A checkpoint passed via --pretrained-checkpoint continues to upload through the CLI, since it is not publicly hosted. Add per-stage timing to the remote run for future capacity planning, and require the R2 environment variables up front with a clear error instead of failing partway through a run. Verified live: presigned PUT from a Colab T4 session to R2 uploaded 120 MB in about 4 seconds. Retrieval and checksum verification back through boto3 are covered by new tests, falsified by disabling the delete-on-finally guarantee, the early-record-before-later-failure guarantee, and the R2 error-code mapping, each of which failed the corresponding test before the fix. Refs #20
…ectly Running `python training/run_rec_colab.py` puts training/, not the repo root, on sys.path, so the new `from app...` imports failed with ModuleNotFoundError. Insert the repo root before importing, matching the sys.path shim already used by training/scripts/*.py. Refs #20
colab stop can fail with Not Found simply because Colab already reclaimed an idle or finished runtime on its own; a completed, checksum-verified training run was being reported as failed and its checkpoint demoted to partial/ solely because of this race. Verify with `colab sessions` before treating a failed stop as a real teardown failure. Verified live: a full T4 run (base checkpoint fetched by Colab, trained, checksum-verified checkpoint retrieved through R2) hit exactly this race; colab sessions confirmed no session was actually left running. Refs #20
Summary
Verification
bash -n training/bootstrap.sh training/setup_rec_environment.sh training/run_rec_smoke.shpython3 -m py_compile training/run_rec_colab.py training/colab_remote.pyCloses #20