feat(training): add Kaggle GPU runner - #24
Merged
Merged
Conversation
Adds a second remote GPU backend alongside Colab: run_rec_kaggle.py stages the same reviewed-dataset input archive as a private, run-scoped Kaggle Dataset, pushes a private script kernel (kaggle_remote.py) that runs the unchanged CUDA training path, retrieves the checkpoint through the same bounded R2 transfer, and runs the same local evaluation gate. Extracts the local-side staging/R2-retrieval/evaluation contract shared by both backends into training/remote_gpu_common.py. colab_remote.py and kaggle_remote.py stay separate self-contained scripts because both Colab's exec -f and a Kaggle script kernel's code_file only ever transfer a single file to the remote runtime. Cleanup deletes the private input dataset (the only Kaggle-side storage holding reviewed crops) after every run; the Kaggle CLI has no kernel-delete command, so the private kernel id is recorded in status.json for manual removal instead of imitating `colab stop`. Fixes #22
kaggle auth login (the current CLI's OAuth flow) stores the session in ~/.kaggle/credentials.json instead of the legacy kaggle.json API-key file. run_rec_kaggle.py only checked KAGGLE_USERNAME and kaggle.json, so an OAuth-authenticated CLI would fail to name the private dataset/kernel even though `kaggle` itself was authenticated. Found while verifying the runner end to end against a real Kaggle account. Refs #22
kaggle datasets create and kaggle kernels push share one per-account slug namespace. Reusing the same run-scoped slug for both made `kernels push` fail with 409 Conflict (SaveKernel) right after the input dataset finished uploading. Found by running the Kaggle runner end to end against a real account. Refs #22
Kaggle auto-extracts recognized archive extensions (.zip, .gz, .tar.gz, .tgz, ...) when a dataset is mounted into a kernel, so the uploaded ocrkit-input.tar.gz never survived as a file for kaggle_remote.py to find and extract itself; only its already-unpacked contents were mounted, one directory level off from what locate_input_archive() expected. Renaming the staged file to ocrkit-input.pkg keeps the exact same gzip+tar bytes (tarfile.open still reads it with mode "r:gz") but avoids the extension-based auto-extraction. Found by running the Kaggle runner end to end against a real account and inspecting the retrieved kernel output. Refs #22
Renaming the input archive's extension did not stop Kaggle from auto-extracting it: Kaggle sniffs archive formats (zip/gz/tar.gz/tgz) by content, not by filename, when mounting a dataset into a kernel. The uploaded blob never survived as a file for kaggle_remote.py to find. Restructure staging around that: training/remote_gpu_common.py now splits stage_inputs (Colab's single tar.gz, unchanged) and a new stage_input_directory (Kaggle: the same files as plain copies, since Kaggle datasets keep a directory tree natively). kaggle_remote.py reads directly from the mounted /kaggle/input/<slug>/ tree via shutil.copytree instead of extracting a tarfile. Found by running the Kaggle runner end to end against a real account and inspecting the retrieved kernel log. Refs #22
-r skip (the alternative to zip/tar) silently drops subdirectories from the upload entirely rather than uploading them as-is, so dataset/ and repo/ never reached Kaggle at all with the previous -r skip choice. -r zip locally zips each subdirectory instead, and Kaggle reliably auto-unzips .zip files back into their folder when a dataset is mounted, reconstructing the same tree kaggle_remote.py expects. The top-level request.json is never a folder, so it is never zipped and always survives as a plain file. Found by running the Kaggle runner end to end against a real account and reading the CLI's own upload log. Refs #22
…lure Two consecutive real-account runs still failed to find request.json after two different fixes for Kaggle's dataset auto-extraction, so the next failure needs ground truth instead of another guess: list what Kaggle actually mounted under /kaggle/input and include it in the raised error, which flows into run.json and is retrieved locally. Refs #22
Reading the installed kaggle-api source directly settled the repeated guessing: kernels push reads only the code_file's own text as the kernel source and ignores every other file in the push folder, and directory zip/skip handling on datasets create is a plain client-side upload transform with no server-side "please re-expand this" signal — Kaggle's dataset-attachment mechanism has no documented, stable contract for staging a directory tree into a kernel. Drop Kaggle Datasets entirely. The input archive now travels through R2 exactly like the checkpoint already travels back: the runner uploads it to a run-scoped key, generates a bounded presigned GET URL, and substitutes that URL into kaggle_remote.py's own source text before pushing it (the only channel available to hand a script kernel any per-run data). kaggle_remote.py downloads and extracts it with the same safe_extract Colab's worker already uses. - app/storage/r2_client.py: add generate_presigned_get_url, symmetric to the existing presigned PUT. - training/remote_gpu_common.py: drop stage_input_directory; add upload_via_presigned_url. stage_inputs (the shared tar.gz builder) is unchanged and now used by both backends again. - training/kaggle_remote.py: replace locate_input_directory/copytree with download_input_archive + safe_extract, driven by an INPUT_ARCHIVE_URL placeholder substituted at push time. - training/run_rec_kaggle.py: remove all dataset create/status/delete orchestration and the dataset/kernel slug-collision workaround it needed; the kernel is the only Kaggle resource created. Refs #22
…ernet access A kernel's enable_internet metadata has no effect until the Kaggle account itself is phone-verified; without it, network requests inside the kernel fail with a DNS resolution error. Found while running the Kaggle runner end to end against a real (not yet phone-verified) account. Refs #22
… kernel failure poll_kernel_status matched the substring "error" anywhere in the status command's combined stdout/stderr, so a dropped-connection SSLEOFError while merely asking Kaggle for status (the kernel was still queued and had not even started) was misread as the kernel itself having failed. That short-circuited straight to fetch_kernel_output, which correctly found nothing yet and reported a confusing "did not include run metadata" error instead of the real, transient cause. Only inspect the status text when the CLI call itself succeeded (returncode 0); a failed status check is retried like any other not-yet-settled poll. Found by running the Kaggle runner end to end against a real account and reading its local kaggle.log. Refs #22
kaggle kernels output syncs back the entire /kaggle/working tree, not just the results/ directory kaggle_remote.py writes run.json/remote.log to. With the venv, PaddleOCR clone, wheel downloads, extracted dataset, and trained checkpoint binary all placed under /kaggle/working, output retrieval was pulling multiple gigabytes of build artifacts back over the network for no reason (the checkpoint itself already travels through R2, and config.yml/train.log ride along as text in run.json). Move all of that scratch to /tmp/ocrkit-run, which is never part of kernel output; only results/ stays under /kaggle/working. Found by running the Kaggle runner end to end against a real account: output retrieval stalled on a slow connection well before ever reaching the build directories, at a rate that made the full sync impractical. Refs #22
… timeout Kaggle's free-tier GPU queue can run far longer than any single run's own training time (observed both near-zero and 35+ minute waits in the same session), and a kernel that never left the queue never ran kaggle_remote.py, so it never touched its input archive. The previous unconditional cleanup on timeout would delete that archive out from under a kernel Kaggle might still schedule and run later, guaranteeing its eventual download would 404. poll_kernel_status now returns "timeout_queued" or "timeout_running" instead of raising on timeout, distinguishing a kernel that never started from one that did; only the latter (or a definitive complete/error) triggers cleanup. The presigned URLs a queued kernel depends on are also decoupled from --timeout-seconds (which only bounds this script's own patience) onto a fixed, generous RESOURCE_URL_TTL_SECONDS, so they don't expire out from under a kernel still waiting for capacity. Confirmed for real: a run timed out locally after 1h still queued, and the kernel went on to complete successfully on Kaggle's side ~10 minutes later with its checkpoint retrievable exactly as designed. Refs #22
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
training/run_rec_kaggle.pyas a second remote GPU training backend alongside Colab, sharing the local staging/R2-retrieval/evaluation contract viatraining/remote_gpu_common.py.training/kaggle_remote.py's docstring and the commit history below for what was tried and ruled out).training/kaggle_remote.pyruns inside the Kaggle kernel; kept self-contained likecolab_remote.pysincekernels pushreads only the code file's own text.field_accuracy 0.9736vs the0.9604release threshold).kaggle auth login) username resolution, a dataset/kernel slug collision, Kaggle's archive auto-extraction behavior (content-sniffed, not extension-based — two iterations to land on R2-based transport instead), a transientkernels statusnetwork error misread as a kernel failure,kernels outputneedlessly syncing the entire build environment (venv/PaddleOCR clone) back over the network, and a queued-kernel's input archive being deleted out from under it if the local script's--timeout-secondselapses before Kaggle schedules it.training/README.md.Follow-up filed as #23: integrating this into Studio needs a genuinely different async-job model than Studio's current local-process tracking, which is out of scope here.
Test plan
uv run pytest -q— 279 passedFixes #22