Skip to content

feat(training): add Kaggle GPU runner - #24

Merged
Teakowa merged 12 commits into
mainfrom
feat/kaggle-gpu-runner
Sep 27, 2026
Merged

Teakowa merged 12 commits into
mainfrom
feat/kaggle-gpu-runner

Conversation

@e54-bot

@e54-bot e54-bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Adds training/run_rec_kaggle.py as a second remote GPU training backend alongside Colab, sharing the local staging/R2-retrieval/evaluation contract via training/remote_gpu_common.py.
  • Input archive and trained checkpoint both travel through short-lived, single-object presigned R2 URLs — Kaggle's own dataset-attachment mechanism has no stable, documented contract for staging a directory tree into a kernel, so it isn't used for input transport at all (see training/kaggle_remote.py's docstring and the commit history below for what was tried and ruled out).
  • training/kaggle_remote.py runs inside the Kaggle kernel; kept self-contained like colab_remote.py since kernels push reads only the code file's own text.
  • Verified end to end against a real Kaggle account: private kernel, real Tesla T4 GPU, PaddlePaddle 3.3.1 CUDA training, checkpoint round-tripped through R2, local evaluation gate passed (field_accuracy 0.9736 vs the 0.9604 release threshold).
  • Several real bugs found during that verification are fixed with individual commits + regression tests: OAuth (kaggle auth login) username resolution, a dataset/kernel slug collision, Kaggle's archive auto-extraction behavior (content-sniffed, not extension-based — two iterations to land on R2-based transport instead), a transient kernels status network error misread as a kernel failure, kernels output needlessly syncing the entire build environment (venv/PaddleOCR clone) back over the network, and a queued-kernel's input archive being deleted out from under it if the local script's --timeout-seconds elapses before Kaggle schedules it.
  • Documents backend selection, Kaggle auth (including the required phone verification for kernel internet access), privacy behavior, and failure handling in training/README.md.

Follow-up filed as #23: integrating this into Studio needs a genuinely different async-job model than Studio's current local-process tracking, which is out of scope here.

Test plan

  • uv run pytest -q — 279 passed
  • Real end-to-end run against a live Kaggle account (private kernel + GPU + R2 + local evaluation), including a real queued-timeout scenario confirming the input archive survives for the kernel to consume later

Fixes #22

Adds a second remote GPU backend alongside Colab: run_rec_kaggle.py
stages the same reviewed-dataset input archive as a private, run-scoped
Kaggle Dataset, pushes a private script kernel (kaggle_remote.py) that
runs the unchanged CUDA training path, retrieves the checkpoint through
the same bounded R2 transfer, and runs the same local evaluation gate.

Extracts the local-side staging/R2-retrieval/evaluation contract shared
by both backends into training/remote_gpu_common.py. colab_remote.py and
kaggle_remote.py stay separate self-contained scripts because both
Colab's exec -f and a Kaggle script kernel's code_file only ever
transfer a single file to the remote runtime.

Cleanup deletes the private input dataset (the only Kaggle-side storage
holding reviewed crops) after every run; the Kaggle CLI has no
kernel-delete command, so the private kernel id is recorded in
status.json for manual removal instead of imitating `colab stop`.

Fixes #22
kaggle auth login (the current CLI's OAuth flow) stores the session in
~/.kaggle/credentials.json instead of the legacy kaggle.json API-key
file. run_rec_kaggle.py only checked KAGGLE_USERNAME and kaggle.json, so
an OAuth-authenticated CLI would fail to name the private dataset/kernel
even though `kaggle` itself was authenticated. Found while verifying the
runner end to end against a real Kaggle account.

Refs #22
kaggle datasets create and kaggle kernels push share one per-account
slug namespace. Reusing the same run-scoped slug for both made
`kernels push` fail with 409 Conflict (SaveKernel) right after the
input dataset finished uploading. Found by running the Kaggle runner
end to end against a real account.

Refs #22
Kaggle auto-extracts recognized archive extensions (.zip, .gz, .tar.gz,
.tgz, ...) when a dataset is mounted into a kernel, so the uploaded
ocrkit-input.tar.gz never survived as a file for kaggle_remote.py to
find and extract itself; only its already-unpacked contents were
mounted, one directory level off from what locate_input_archive()
expected. Renaming the staged file to ocrkit-input.pkg keeps the exact
same gzip+tar bytes (tarfile.open still reads it with mode "r:gz") but
avoids the extension-based auto-extraction. Found by running the Kaggle
runner end to end against a real account and inspecting the retrieved
kernel output.

Refs #22
Renaming the input archive's extension did not stop Kaggle from
auto-extracting it: Kaggle sniffs archive formats (zip/gz/tar.gz/tgz) by
content, not by filename, when mounting a dataset into a kernel. The
uploaded blob never survived as a file for kaggle_remote.py to find.

Restructure staging around that: training/remote_gpu_common.py now
splits stage_inputs (Colab's single tar.gz, unchanged) and a new
stage_input_directory (Kaggle: the same files as plain copies, since
Kaggle datasets keep a directory tree natively). kaggle_remote.py reads
directly from the mounted /kaggle/input/<slug>/ tree via shutil.copytree
instead of extracting a tarfile. Found by running the Kaggle runner end
to end against a real account and inspecting the retrieved kernel log.

Refs #22
-r skip (the alternative to zip/tar) silently drops subdirectories from
the upload entirely rather than uploading them as-is, so dataset/ and
repo/ never reached Kaggle at all with the previous -r skip choice.
-r zip locally zips each subdirectory instead, and Kaggle reliably
auto-unzips .zip files back into their folder when a dataset is
mounted, reconstructing the same tree kaggle_remote.py expects. The
top-level request.json is never a folder, so it is never zipped and
always survives as a plain file. Found by running the Kaggle runner end
to end against a real account and reading the CLI's own upload log.

Refs #22
…lure

Two consecutive real-account runs still failed to find request.json
after two different fixes for Kaggle's dataset auto-extraction, so the
next failure needs ground truth instead of another guess: list what
Kaggle actually mounted under /kaggle/input and include it in the
raised error, which flows into run.json and is retrieved locally.

Refs #22
Reading the installed kaggle-api source directly settled the repeated
guessing: kernels push reads only the code_file's own text as the
kernel source and ignores every other file in the push folder, and
directory zip/skip handling on datasets create is a plain client-side
upload transform with no server-side "please re-expand this" signal —
Kaggle's dataset-attachment mechanism has no documented, stable
contract for staging a directory tree into a kernel.

Drop Kaggle Datasets entirely. The input archive now travels through R2
exactly like the checkpoint already travels back: the runner uploads it
to a run-scoped key, generates a bounded presigned GET URL, and
substitutes that URL into kaggle_remote.py's own source text before
pushing it (the only channel available to hand a script kernel any
per-run data). kaggle_remote.py downloads and extracts it with the same
safe_extract Colab's worker already uses.

- app/storage/r2_client.py: add generate_presigned_get_url, symmetric
  to the existing presigned PUT.
- training/remote_gpu_common.py: drop stage_input_directory; add
  upload_via_presigned_url. stage_inputs (the shared tar.gz builder) is
  unchanged and now used by both backends again.
- training/kaggle_remote.py: replace locate_input_directory/copytree
  with download_input_archive + safe_extract, driven by an
  INPUT_ARCHIVE_URL placeholder substituted at push time.
- training/run_rec_kaggle.py: remove all dataset create/status/delete
  orchestration and the dataset/kernel slug-collision workaround it
  needed; the kernel is the only Kaggle resource created.

Refs #22
…ernet access

A kernel's enable_internet metadata has no effect until the Kaggle
account itself is phone-verified; without it, network requests inside
the kernel fail with a DNS resolution error. Found while running the
Kaggle runner end to end against a real (not yet phone-verified)
account.

Refs #22
… kernel failure

poll_kernel_status matched the substring "error" anywhere in the status
command's combined stdout/stderr, so a dropped-connection SSLEOFError
while merely asking Kaggle for status (the kernel was still queued and
had not even started) was misread as the kernel itself having failed.
That short-circuited straight to fetch_kernel_output, which correctly
found nothing yet and reported a confusing "did not include run
metadata" error instead of the real, transient cause.

Only inspect the status text when the CLI call itself succeeded
(returncode 0); a failed status check is retried like any other
not-yet-settled poll. Found by running the Kaggle runner end to end
against a real account and reading its local kaggle.log.

Refs #22
kaggle kernels output syncs back the entire /kaggle/working tree, not
just the results/ directory kaggle_remote.py writes run.json/remote.log
to. With the venv, PaddleOCR clone, wheel downloads, extracted dataset,
and trained checkpoint binary all placed under /kaggle/working, output
retrieval was pulling multiple gigabytes of build artifacts back over
the network for no reason (the checkpoint itself already travels
through R2, and config.yml/train.log ride along as text in run.json).

Move all of that scratch to /tmp/ocrkit-run, which is never part of
kernel output; only results/ stays under /kaggle/working. Found by
running the Kaggle runner end to end against a real account: output
retrieval stalled on a slow connection well before ever reaching the
build directories, at a rate that made the full sync impractical.

Refs #22
… timeout

Kaggle's free-tier GPU queue can run far longer than any single run's
own training time (observed both near-zero and 35+ minute waits in the
same session), and a kernel that never left the queue never ran
kaggle_remote.py, so it never touched its input archive. The previous
unconditional cleanup on timeout would delete that archive out from
under a kernel Kaggle might still schedule and run later, guaranteeing
its eventual download would 404.

poll_kernel_status now returns "timeout_queued" or "timeout_running"
instead of raising on timeout, distinguishing a kernel that never
started from one that did; only the latter (or a definitive
complete/error) triggers cleanup. The presigned URLs a queued kernel
depends on are also decoupled from --timeout-seconds (which only bounds
this script's own patience) onto a fixed, generous RESOURCE_URL_TTL_SECONDS,
so they don't expire out from under a kernel still waiting for capacity.

Confirmed for real: a run timed out locally after 1h still queued, and
the kernel went on to complete successfully on Kaggle's side ~10
minutes later with its checkpoint retrievable exactly as designed.

Refs #22
@Teakowa
Teakowa merged commit ff7c245 into main Sep 27, 2026
2 checks passed
@Teakowa
Teakowa deleted the feat/kaggle-gpu-runner branch September 27, 2026 17:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(training): add Kaggle GPU runner

2 participants