diff --git a/docs/capture-storage-design.md b/docs/capture-storage-design.md index 8fdc5c4ed..f63fc389b 100644 --- a/docs/capture-storage-design.md +++ b/docs/capture-storage-design.md @@ -1827,6 +1827,9 @@ bytes, and retained failure details all have explicit caps. writer will be native. See *Phase 6 -- Decision: the production writer is native*. - One process owns a spool directory; cross-process locking is not implemented. + (The native C++ spool does lock it: an owner lock on `/.owner.lock`, + B6, documented in `native/csrc/store/spool.h`. The Python reference + deliberately stays unlocked.) - Durable mode stages synchronously and uploads through a separate explicit uploader, so remote backpressure is isolated from local commit. - The pipeline remains opt-in and is not connected to Ring². diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index b34179302..b952dae58 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -185,6 +185,71 @@ refuses the persistent backend with `ConfigurationError` rather than generating with nothing stored. The catalog takes one publisher per `(database, table_prefix)`, so a second engine on the same catalog is refused at `create_record_runtime`. +The engine spools into a directory of its own under +`capture_sink_config.spool_root`, +`//r-/` (the key is the first 12 +hex digits of a sha256 of where the packs go -- the ClickHouse host and port, +`database`, `table_prefix`, the S3 endpoint and bucket, and `store_id`, as +spelled in the config -- the rank torchrun's +`RANK`, 0 when unset, and the incarnation fresh on every +`create_record_runtime`), and owns it: an flock on its `.owner.lock` and on the +directory itself, taken before the service starts and let go after the sink and +the service are done, when a drained directory is removed. The directory's own +lock keeps it owned should its `.owner.lock` be removed from under it, and +systemd-tmpfiles skips a flocked directory when it ages `/tmp`; other cleaners +may not, so keep `spool_root` out of what they age. If the sink did not seal -- +the release backstop's flush, when the stopping ring lets go of it, did not go +through -- it may still be staging, so the directory stays owned by the +process until it exits (a warning names it) and the next process on the node +adopts it; so does the directory of an engine dropped without `close()`. +A second process on a directory is +refused, naming the holder's pid and host. Once started, the service's +background loop adopts the directories under the same catalog key whose owners +have died: their stale `.open` files are swept, their ready packs uploaded and +indexed, a round at a time and a slice of each cycle, and the directory +removed, so a crashed process's packs reach the catalog through the next one on +the node, whatever run it belongs to. Neither `create_record_runtime` nor +`flush_and_wait` waits for that: a flush covers this process's records (an +adopted pack uploaded and not yet indexed is waited for like its own), waits +for no more of an adoption in progress than the step it is in (one round of +its uploads, or one of its packs validated -- a dead directory's packs are +hashed one per step, so a large backlog's listing does not hold a flush up), +and `close()` cuts an adoption as it cuts the service's own uploads, leaving +what it did not upload in the dead directory for the next process. The +storage part of `capture_status()` reports the adoption (`adopted_spools`, +`adopted_packs`, `adoption_owed`, `live_siblings`). A dead directory the +service can never adopt -- one holding a pack it can never upload, such as one +larger than its `uploader_max_in_flight_bytes` or one whose key already holds a +different object, or one it cannot lock -- is left in place once the rest of +its packs are up, listed in `blocked_siblings`, reported once in `last_error` +and in its `.owner.lock` (a `blocked: ` line), and neither retried nor +owed. `spool_max_bytes` bounds the directory together with what the dead +directories the service can adopt still hold, so restarts while uploads are +blocked cannot each add a whole budget; the room comes back as they are +adopted. A live process's directory is its own budget, and neither a blocked +directory nor one this process keeps owned itself (an earlier engine's whose +sink did not seal) is charged: no adoption here drains them. The +spool root must be +node-local: NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, GPFS, 9p, AFS and OrangeFS +are refused (by statfs `f_type`) unless `NativeSinkConfig.spool_allow_shared_filesystem`, which a FUSE +filesystem that is itself local, such as fuse-overlayfs, needs too. Without +`capture_storage_config` the sink owns `spool_root` itself. With an explicit +`record_sink`, the service drains `spool_root` as that sink writes it, +unswept and adopting nothing. Both of those modes pass over the rank +directories under `spool_root` -- what a crashed or undrained default-mode run +left there is the next default-mode start's to adopt -- so switching to them +(the rollback to an explicit `record_sink` included) works after such a run. +They are refused, naming the holder, while a default-mode process on the node +holds a rank directory under that `spool_root`, and a default-mode start is +refused while one of them holds `spool_root`. Upgrading from an engine without this layout: it +spooled into `/v1/...` and its next start uploaded what a crashed +run left there, but nothing adopts packs outside the layout now -- nor those a +sink-only or explicit-`record_sink` run leaves in `spool_root`. Each +`create_record_runtime` logs a warning while any are there, with their count, +one of them, and a directory of the layout nobody owns +(`//r0-00000000/`); moved into it with their paths +below `spool_root` kept (`v1/...`), packs bound for this catalog and store are +adopted by the next start on the node. To reach a secured catalog, set `clickhouse_scheme="https"` (and the server's TLS HTTP port, usually 8443) on `NativeCaptureStorageConfig`. The client always verifies the server's certificate and name, against libcurl's built-in CA diff --git a/native/csrc/catalog/bindings_store.cpp b/native/csrc/catalog/bindings_store.cpp index dcaf4cdc6..10a8bb1c6 100644 --- a/native/csrc/catalog/bindings_store.cpp +++ b/native/csrc/catalog/bindings_store.cpp @@ -2,6 +2,10 @@ // // StorageService spool -> object store -> catalog, on a background thread // CaptureReader search / select / hydrate against the catalog + store +// SpoolOwnerLock the owner lock of one spool directory (store/spool.h), +// which the engine holds around its sink and service +// spool_rank_directory, spool_catalog_key, spool_owner +// the section 2.3 spool layout, and who owns a directory // // Pure C++ plus libcurl and libcrypto. It uses pybind11's headers but links // nothing from torch, and registers no ring types, so it loads beside @@ -10,12 +14,14 @@ #include #include +#include #include #include #include "catalog/hydration.h" #include "catalog/reader.h" #include "catalog/storage_service.h" +#include "store/spool.h" namespace py = pybind11; namespace dc = dmi_catalog; @@ -75,10 +81,50 @@ dc::ClickHouseConnection clickhouse_connection(const py::dict& d) { return c; } +dmi_store::OwnerLock owner_lock(const std::string& text) { + dmi_store::OwnerLock mode = dmi_store::OwnerLock::kTake; + if (!dmi_store::ParseOwnerLock(text, &mode)) { + throw py::value_error("spool_owner_lock must be 'take' or " + "'held_by_caller', got '" + text + "'"); + } + return mode; +} + +// Where a spool's packs go (store/spool.h); every field is required. +dmi_store::SpoolDestination spool_destination(const py::dict& d) { + for (const char* key : {"clickhouse_host", "clickhouse_port", "database", + "table_prefix", "s3_endpoint", "s3_bucket", + "store_id"}) { + if (!d.contains(key)) { + throw py::key_error(std::string("spool destination needs '") + key + + "'"); + } + } + dmi_store::SpoolDestination out; + out.clickhouse_host = d["clickhouse_host"].cast(); + out.clickhouse_port = d["clickhouse_port"].cast(); + out.database = d["database"].cast(); + out.table_prefix = d["table_prefix"].cast(); + out.s3_endpoint = d["s3_endpoint"].cast(); + out.s3_bucket = d["s3_bucket"].cast(); + out.store_id = d["store_id"].cast(); + return out; +} + dc::StorageServiceConfig service_config(const py::dict& d) { dc::StorageServiceConfig c; c.spool_root = get(d, "spool_root", ""); c.spool_max_bytes = get(d, "spool_max_bytes", c.spool_max_bytes); + c.spool_owner_lock = + owner_lock(get(d, "spool_owner_lock", "take")); + c.spool_allow_shared_filesystem = get( + d, "spool_allow_shared_filesystem", c.spool_allow_shared_filesystem); + c.adopt_sibling_spools = + get(d, "adopt_sibling_spools", c.adopt_sibling_spools); + c.adoption_recheck_interval_ns = get( + d, "adoption_recheck_interval_ns", c.adoption_recheck_interval_ns); + c.adoption_slice_ns = + get(d, "adoption_slice_ns", c.adoption_slice_ns); c.s3 = s3_config(d); c.uploader.store_id = get(d, "store_id", c.uploader.store_id); c.uploader.max_workers = get(d, "uploader_max_workers", c.uploader.max_workers); @@ -129,6 +175,11 @@ py::dict snapshot_dict(const dc::StorageServiceSnapshot& s) { out["swept_on_start"] = s.swept_on_start; out["pending_index"] = s.pending_index; out["rejected_packs"] = s.rejected_packs; + out["adopted_spools"] = s.adopted_spools; + out["adopted_packs"] = s.adopted_packs; + out["adoption_owed"] = s.adoption_owed; + out["live_siblings"] = s.live_siblings; + out["blocked_siblings"] = s.blocked_siblings; out["failed"] = s.failed; out["lease_state"] = s.lease_state; // Seconds on the monotonic clock, comparable with time.monotonic() (both @@ -275,6 +326,23 @@ class CaptureReader { dc::NativeCaptureReader reader_; }; +// A refused SpoolStatus as the Python exception that says what it is: a +// directory another process owns, a refused configuration, or an I/O error. +PyObject* g_spool_owned_error = nullptr; // SpoolOwnedError, set at import + +[[noreturn]] void raise_spool_status(dmi_store::SpoolStatus status, + const std::string& error) { + if (status == dmi_store::SpoolStatus::kOwned) { + PyErr_SetString(g_spool_owned_error, error.c_str()); + throw py::error_already_set(); + } + if (status == dmi_store::SpoolStatus::kBadArgument) { + throw py::value_error(error); + } + PyErr_SetString(PyExc_OSError, error.c_str()); + throw py::error_already_set(); +} + } // namespace PYBIND11_MODULE(_dmi_native_store, m) { @@ -296,6 +364,91 @@ PYBIND11_MODULE(_dmi_native_store, m) { [](const dc::CaptureStorageService& s) { return snapshot_dict(s.snapshot()); }) .def("rethrow_if_failed", &dc::CaptureStorageService::rethrow_if_failed); + // Raised when another owner holds a spool directory; a RuntimeError, so + // callers catching the service's refusals keep catching it. + static py::exception spool_owned_error( + m, "SpoolOwnedError", PyExc_RuntimeError); + g_spool_owned_error = spool_owned_error.ptr(); + + py::class_(m, "SpoolOwnerLock") + .def(py::init([](const std::string& directory, + bool allow_shared_filesystem) { + auto lock = std::make_unique(); + std::string error; + const dmi_store::SpoolStatus status = + dmi_store::SpoolOwnerLock::Acquire( + directory, allow_shared_filesystem, lock.get(), &error); + if (status != dmi_store::SpoolStatus::kOk) { + raise_spool_status(status, error); + } + return lock; + }), + py::arg("directory"), py::arg("allow_shared_filesystem") = false) + .def_property_readonly("directory", &dmi_store::SpoolOwnerLock::directory) + .def_property_readonly("held", &dmi_store::SpoolOwnerLock::held) + .def("release", &dmi_store::SpoolOwnerLock::Release) + .def("release_and_remove_if_empty", + [](dmi_store::SpoolOwnerLock& self) { + std::string error; + return self.ReleaseAndRemoveIfEmpty(&error); + }) + .def("__enter__", + [](dmi_store::SpoolOwnerLock& self) -> dmi_store::SpoolOwnerLock& { + return self; + }, py::return_value_policy::reference) + .def("__exit__", [](dmi_store::SpoolOwnerLock& self, const py::args&) { + self.Release(); + return false; + }); + + m.def("spool_owner", + [](const std::string& directory) -> py::object { + dmi_store::SpoolOwner owner; + if (!dmi_store::ReadSpoolOwner(directory, &owner)) return py::none(); + py::dict out; + out["host"] = owner.host; + out["pid"] = owner.pid; + return out; + }, + py::arg("directory"), + "Who holds a spool directory's owner lock, or None when nothing " + "does."); + m.def("spool_catalog_key", + [](const py::dict& destination) { + return dmi_store::SpoolCatalogKey(spool_destination(destination)); + }, + py::arg("destination"), + "The catalog key of a destination dict (clickhouse_host, " + "clickhouse_port, database, table_prefix, s3_endpoint, s3_bucket, " + "store_id)."); + m.def("spool_rank_directory", + [](const std::string& base, const py::dict& destination, + uint64_t producer_rank, std::optional incarnation) { + const std::string fresh = + incarnation ? *incarnation : dmi_store::NewSpoolIncarnation(); + uint64_t parsed_rank = 0; + std::string parsed; + if (!dmi_store::ParseSpoolRankDirectoryName( + dmi_store::SpoolRankDirectoryName(producer_rank, fresh), + &parsed_rank, &parsed)) { + throw py::value_error("incarnation must be 8 lowercase hex " + "digits"); + } + return dmi_store::SpoolRankDirectory( + base, spool_destination(destination), producer_rank, fresh); + }, + py::arg("base"), py::arg("destination"), py::arg("producer_rank"), + py::arg("incarnation") = py::none(), + "//r-, the section " + "2.3 spool layout; a fresh incarnation when none is given."); + m.def("_set_spool_filesystem_type_for_testing", + [](std::optional f_type) { + dmi_store::SetFilesystemTypeForTesting(f_type ? *f_type : -1); + }, + py::arg("f_type"), + "Test seam: node-local checks in this module read f_type instead " + "of statfs(2); None restores statfs."); + py::class_(m, "CaptureReader") .def(py::init(), py::arg("config")) .def("search", &CaptureReader::search, py::arg("filters")) diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index c1490eec4..16c747e88 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -4,6 +4,9 @@ #include #include #include +#include +#include +#include #include #include #include @@ -152,6 +155,24 @@ class CaptureStorageService::LeaseScope { uint64_t counted_on_entry_ = 0; }; +// The dead sibling a cycle is adopting, kept across cycles: its owner lock, +// a Spool on it (held_by_caller, the lock being this service's), its +// recovery -- the packs listed and those validated so far, a pack a step +// (Spool::BeginRecovery) -- and, once every pack is validated, those not +// uploaded yet. So a backlog is hashed once, however many cycles its +// listing and its upload take. +struct CaptureStorageService::Adoption { + std::string directory; + dmi_store::SpoolOwnerLock lock; + dmi_store::Spool spool; + dmi_store::SpoolRecovery recovery; + bool validated = false; // every listed pack; `remaining` holds the valid + std::deque remaining; + // The first pack this service can never upload (UploadFailure:: + // retryable false); the directory is left, with it, once the rest are up. + std::string blocked; +}; + CaptureStorageService::CaptureStorageService(StorageServiceConfig config) : config_(std::move(config)), s3_(config_.s3), @@ -190,10 +211,41 @@ CaptureStorageService::CaptureStorageService(StorageServiceConfig config) " ms, or raise lease_ttl_ns"); } std::string error; - if (dmi_store::Spool::Open({config_.spool_root, config_.spool_max_bytes}, - &spool_, &error) != dmi_store::SpoolStatus::kOk) { + dmi_store::SpoolConfig spool_config{config_.spool_root, + config_.spool_max_bytes}; + spool_config.owner_lock = config_.spool_owner_lock; + spool_config.allow_shared_filesystem = config_.spool_allow_shared_filesystem; + if (dmi_store::Spool::Open(spool_config, &spool_, &error) != + dmi_store::SpoolStatus::kOk) { throw std::runtime_error("storage service: cannot open spool: " + error); } + if (config_.adopt_sibling_spools) { + // Siblings are adopted INTO this catalog, so this directory must sit + // under this catalog's key: a directory under another catalog's key + // would index that catalog's packs here. + const std::filesystem::path own(spool_.root()); + dmi_store::SpoolDestination destination; + destination.clickhouse_host = config_.clickhouse.host; + destination.clickhouse_port = config_.clickhouse.port; + destination.database = config_.writer.database; + destination.table_prefix = config_.writer.table_prefix; + destination.s3_endpoint = config_.s3.endpoint; + destination.s3_bucket = config_.s3.bucket; + destination.store_id = config_.uploader.store_id; + const std::string key = dmi_store::SpoolCatalogKey(destination); + uint64_t rank = 0; + std::string incarnation; + if (!dmi_store::ParseSpoolRankDirectoryName(own.filename().string(), + &rank, &incarnation) || + own.parent_path().filename().string() != key) { + throw std::invalid_argument( + "storage service: adopt_sibling_spools needs spool_root to be a " + "rank directory /" + key + "/r- (this " + "catalog's key for its ClickHouse host and port, database, " + "table_prefix, S3 endpoint, bucket and store_id), got " + + spool_.root()); + } + } s3_.set_cancellation(&read_cancel_); upload_s3_.set_cancellation(&upload_cancel_); uploader_ = std::make_unique(&spool_, &upload_s3_, @@ -250,8 +302,10 @@ void CaptureStorageService::start() { last_reconcile_ns_ = steady_ns(); { + // With siblings to look at, the loop's first cycle runs at once: it, + // not start(), adopts them. std::lock_guard lock(wake_mutex_); - kick_ = false; + kick_ = adoption_scan_owed_; } { LeaseScope lease(this); // publishes the lease state @@ -265,12 +319,10 @@ void CaptureStorageService::start() { } void CaptureStorageService::sweep_and_reconcile_at_start() { - // After the lease, never before, so a second process pointed at this spool - // usually learns that the catalog is held before it can delete a live - // sink's .open files. Only usually: a holder that is quarantined has let - // its row lapse, and a second process can take the lease in that gap. The - // spool itself is not locked; one process per spool is the caller's job - // until the spool gets an owner lock. + // After the lease: a start refused the catalog never touches the spool. + // The spool's owner lock (taken at construction, by this service or its + // caller) is what keeps another process's writer out of the directory + // this deletes .open files in. if (config_.sweep_spool_on_start) { std::vector recovered; std::string error; @@ -282,6 +334,16 @@ void CaptureStorageService::sweep_and_reconcile_at_start() { state_.swept_on_start = recovered.size(); } + // Dead siblings are the loop's, from its first cycle on: after the lease + // and this directory's own sweep, as the plan orders them, but not here, + // where a backlog the object store refuses would hold start() -- and + // create_record_runtime -- through every pack's retry chain. + if (config_.adopt_sibling_spools) { + adoption_scan_owed_ = true; + std::lock_guard lock(state_mutex_); + state_.adoption_owed = true; + } + // A failed pass is not fatal -- the bucket is still there next time. Nor // is a lease lost while it runs, to a quarantine or to another holder: // that is the running service's case, and the loop handles it as it does @@ -326,6 +388,7 @@ void CaptureStorageService::stop() { // lease has to keep renewing until that is done. stop_lease_thread(); std::lock_guard cycle(cycle_mutex_); + let_go_of_adoption(); if (started_) { started_ = false; // A quarantined writer holds no lease, so it writes no tombstone: the @@ -359,6 +422,15 @@ void CaptureStorageService::stop_lease_thread() { } bool CaptureStorageService::flush(double timeout_s) { + // The loop's adoption gives way to this call at its next step, so the + // cycle lock is not held for the rest of an adoption slice. + struct InProgress { + explicit InProgress(std::atomic* count) : count_(count) { + count_->fetch_add(1, std::memory_order_acq_rel); + } + ~InProgress() { count_->fetch_sub(1, std::memory_order_acq_rel); } + std::atomic* count_; + } in_progress(&flushes_in_progress_); const auto deadline = std::chrono::steady_clock::now() + std::chrono::duration_cast( std::chrono::duration(timeout_s)); @@ -383,8 +455,9 @@ bool CaptureStorageService::flush(double timeout_s) { if (upload_cancel_.cancelled_for_good()) return false; // Its own cycle, bounded by the deadline, and without the reconcile: // the reconcile lists the whole bucket and asks the catalog about - // every page, which no deadline bounds. The loop runs it. - const bool drained = run_cycle(deadline_ns, false).drained; + // every page, which no deadline bounds. The loop runs it. Nor does it + // adopt: a dead backlog is not this process's records. + const bool drained = run_cycle(deadline_ns, false, false).drained; if (!rejected_unreported_.empty()) { std::string message = "storage service: " + std::to_string(rejected_unreported_.size()) + @@ -427,11 +500,22 @@ void CaptureStorageService::loop() { kicked = kick_; kick_ = false; } + bool failed = false; { std::lock_guard lock(state_mutex_); - if (failure_) return; // another publisher holds the catalog + failed = failure_ != nullptr; } std::lock_guard cycle(cycle_mutex_); + if (failed) { + // Another publisher holds the catalog, and no cycle runs again. A + // sibling half adopted would stay locked by this process, with + // nobody working on it, until stop() -- the engine's close(), maybe + // hours of capture later -- and every other process on the node + // would read it as live meanwhile. Nor does it look again. + adoption_scan_owed_ = false; + let_go_of_adoption(); + return; + } { std::lock_guard lock(wake_mutex_); if (stop_requested_) return; @@ -442,8 +526,9 @@ void CaptureStorageService::loop() { // flush -- close()'s order -- by that cycle's catalog work, which stop() // cannot cut. So wait out the rest of the interval first, unless a // fresh lease asked for a cycle now -- once: flushes that keep coming - // do every cycle's work but the reconcile, which only this loop runs, - // so the next wake runs a cycle whatever they did. + // do every cycle's work but the reconcile and the adoption, which only + // this loop runs, so the next wake runs a cycle whatever they did. + // start() kicks the first cycle when there are siblings to look at. const uint64_t since_ns = steady_ns() - last_cycle_end_ns_; if (!kicked && !waited_again && last_cycle_end_ns_ != 0 && since_ns < wait_ns) { @@ -452,7 +537,7 @@ void CaptureStorageService::loop() { continue; } waited_again = false; - run_cycle(0, true); + run_cycle(0, true, true); // poll_interval * 2^streak, capped: flush() shares the streak, so an // outage it saw also slows the loop, and a success from either resets it. wait_ns = config_.poll_interval_ns; @@ -465,7 +550,7 @@ void CaptureStorageService::loop() { } CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle( - uint64_t deadline_ns, bool allow_reconcile) { + uint64_t deadline_ns, bool allow_reconcile, bool adopt) { CycleOutcome outcome; // stop() has begun: it cancelled the uploads and the reads for good, so a // cycle could only make catalog requests -- a lease claim, a replay guard @@ -494,30 +579,10 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle( LeaseScope lease(this); catalog = ensure_publisher_lease(); } - // Indexes refs, keeping whatever does not index owed: it is already gone - // from the spool, so pending_index_ is the only record of it in-process. - // Those a cancel or the deadline left owed are counted in `deferred` - // (index_bounded). + // What a cycle uploads and cannot index stays owed in pending_index_ + // (index_or_owe). Those a cancel or the deadline left owed are counted + // here: they make the cycle cut short, not failed. size_t deferred = 0; - const auto index_or_owe = [this, catalog, deadline_ns, - &deferred](std::vector refs) { - if (!catalog) { - pending_index_.insert(pending_index_.end(), refs.begin(), refs.end()); - return; - } - std::vector unindexed; - try { - if (!refs.empty()) { - deferred += index_bounded(std::move(refs), &unindexed, deadline_ns); - } - } catch (...) { - pending_index_.insert(pending_index_.end(), unindexed.begin(), - unindexed.end()); - throw; - } - pending_index_.insert(pending_index_.end(), unindexed.begin(), - unindexed.end()); - }; try { // 1. Retry what earlier cycles uploaded but could not index. While any // of it is still owed, the catalog is down or refusing: upload @@ -527,7 +592,7 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle( if (catalog && !pending_index_.empty()) { std::vector owed; owed.swap(pending_index_); - index_or_owe(std::move(owed)); + deferred += index_or_owe(std::move(owed), catalog, deadline_ns); } // 2. Upload what the sink has staged -- but only with the lease and @@ -589,52 +654,43 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle( } } const size_t end = std::min(staged.size(), next + chunk); - const dmi_store::UploadBatchResult batch = uploader_->UploadStaged( + ChunkOutcome sent = upload_chunk( + uploader_.get(), std::vector(staged.begin() + next, staged.begin() + end)); next = end; - std::vector to_index; - uint64_t uploaded_packs = 0; - uint64_t uploaded_bytes = 0; - size_t failed_uploads = 0; - size_t cancelled_uploads = 0; - for (size_t i = 0; i < batch.refs.size(); ++i) { - const dmi_store::PackRef& ref = batch.refs[i]; - if (!ref.pack_id.empty()) { - to_index.push_back({ref.pack_id, ref.store_id, ref.object_key, - ref.object_bytes, ref.checksum, - ref.record_count}); - ++uploaded_packs; - uploaded_bytes += ref.object_bytes; - } else if (i < batch.failures.size() && - batch.failures[i].cancelled) { - ++cancelled_uploads; // still staged; not the pack's fault - } else { - // A failed upload stays in the spool, so a later cycle - // retries it. - ++failed_uploads; - if (i < batch.failures.size()) { - record_error("upload failed for " + - batch.failures[i].object_key + ": " + - batch.failures[i].error); - } + for (size_t i = 0; i < sent.batch.failures.size(); ++i) { + const dmi_store::UploadFailure& failure = sent.batch.failures[i]; + // A failed upload stays in the spool, so a later cycle retries + // it; a cancelled one is still staged, and not the pack's fault. + if (!sent.batch.refs[i].pack_id.empty() || failure.cancelled) { + continue; } + record_error("upload failed for " + failure.object_key + ": " + + failure.error); } - if (cancelled_uploads != 0) outcome.cut_short = true; - upload_failures += failed_uploads; - { - std::lock_guard lock(state_mutex_); - state_.uploaded_packs += uploaded_packs; - state_.uploaded_bytes += uploaded_bytes; - state_.upload_failures += failed_uploads; - state_.cancelled_uploads += cancelled_uploads; - } - index_or_owe(std::move(to_index)); + if (sent.cancelled != 0) outcome.cut_short = true; + upload_failures += sent.failed; + deferred += index_or_owe(std::move(sent.to_index), true, deadline_ns); } uploaded_all = next == staged.size() && upload_failures == 0; } } + // 3a. The loop's cycles adopt dead siblings, a slice at a time, under + // the same rule as the uploads above: only with the lease, every + // staged pack of this process's up, nothing owed, and no cancel. + // flush()'s cycles do not: a dead backlog is not this process's + // records. Adoption uploads and indexes through the same chunks, + // and the same Cancellations, as the service's own spool. + bool adoption_failed = false; + if (adopt && config_.adopt_sibling_spools && uploaded_all && + pending_index_.empty() && !upload_cancel_.cancelled()) { + bool adoption_cut = false; + adoption_failed = !adopt_step(deadline_ns, &deferred, &adoption_cut); + if (adoption_cut) outcome.cut_short = true; + } + // 4. Reconcile on its interval, or when the pass at start() lost the // lease before it finished -- in the loop's cycles only, and not // once a cancel came. The lease thread keeps the lease alive. @@ -652,10 +708,14 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle( // catalog, and nothing pending. Without the lease nothing can be // confirmed in the catalog, so the cycle is not drained, and it counts // towards the backoff. Packs a cancel or the deadline left owed make it - // cut short, not failed. + // cut short, not failed. A dead sibling still to adopt is not a + // failure, nor undrained: only an adoption upload that failed backs the + // loop off, and adopted packs uploaded and not indexed are owed like + // the service's own. if (deferred != 0) outcome.cut_short = true; outcome.failed = !catalog || listing_failed || lost_lease || - upload_failures != 0 || pending_index_.size() > deferred; + upload_failures != 0 || adoption_failed || + pending_index_.size() > deferred; outcome.drained = uploaded_all && !outcome.failed && !outcome.cut_short; std::lock_guard lock(state_mutex_); ++state_.cycles; @@ -672,6 +732,7 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle( { std::lock_guard lock(state_mutex_); state_.pending_index = pending_index_.size(); + state_.adoption_owed = adoption_owed(); } // A cycle a cancel cut short proved nothing about the store or the // catalog: it resets the backoff only by succeeding, and grows it only by @@ -685,6 +746,384 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle( return outcome; } +size_t CaptureStorageService::index_or_owe(std::vector refs, + bool catalog, + uint64_t deadline_ns) { + if (!catalog) { + pending_index_.insert(pending_index_.end(), refs.begin(), refs.end()); + return 0; + } + std::vector unindexed; + size_t deferred = 0; + try { + if (!refs.empty()) { + deferred = index_bounded(std::move(refs), &unindexed, deadline_ns); + } + } catch (...) { + pending_index_.insert(pending_index_.end(), unindexed.begin(), + unindexed.end()); + throw; + } + pending_index_.insert(pending_index_.end(), unindexed.begin(), + unindexed.end()); + return deferred; +} + +CaptureStorageService::ChunkOutcome CaptureStorageService::upload_chunk( + dmi_store::SpoolUploader* uploader, + std::vector chunk) { + ChunkOutcome sent; + sent.batch = uploader->UploadStaged(std::move(chunk)); + uint64_t uploaded_bytes = 0; + for (size_t i = 0; i < sent.batch.refs.size(); ++i) { + const dmi_store::PackRef& ref = sent.batch.refs[i]; + if (!ref.pack_id.empty()) { + sent.to_index.push_back({ref.pack_id, ref.store_id, ref.object_key, + ref.object_bytes, ref.checksum, + ref.record_count}); + uploaded_bytes += ref.object_bytes; + } else if (i < sent.batch.failures.size() && + sent.batch.failures[i].cancelled) { + ++sent.cancelled; // still staged; not the pack's fault + } else { + ++sent.failed; // still staged too + } + } + std::lock_guard lock(state_mutex_); + state_.uploaded_packs += sent.to_index.size(); + state_.uploaded_bytes += uploaded_bytes; + state_.upload_failures += sent.failed; + state_.cancelled_uploads += sent.cancelled; + return sent; +} + +void CaptureStorageService::let_go_of_adoption() { + // A sibling half adopted keeps what is left of it, for the next process + // on the node; what was uploaded from it was indexed, or is owed. Its + // directory is not removed: only finish_adoption() removes one, once no + // pack of it is left. + adopting_.reset(); + adoption_queue_.clear(); + std::lock_guard lock(state_mutex_); + state_.adoption_owed = adoption_owed(); +} + +bool CaptureStorageService::adoption_owed() const { + return adoption_scan_owed_ || adopting_ != nullptr || + !adoption_queue_.empty(); +} + +bool CaptureStorageService::stop_requested() { + std::lock_guard lock(wake_mutex_); + return stop_requested_; +} + +bool CaptureStorageService::adopt_step(uint64_t deadline_ns, size_t* deferred, + bool* cut_short) { + *cut_short = false; + const uint64_t started = steady_ns(); + if (adopting_ == nullptr && adoption_queue_.empty()) { + // Look at the siblings when that is owed (from start() on) or, while + // the last look found a live one, again on the recheck interval. + const bool recheck_due = + live_siblings_ && config_.adoption_recheck_interval_ns > 0 && + started - last_adoption_scan_ns_ >= + config_.adoption_recheck_interval_ns; + if (!adoption_scan_owed_ && !recheck_due) return true; + if (!scan_siblings()) return false; + } + bool ok = true; + // Whether this call has taken a step yet: a lock taken and a sweep, one + // pack of a listing validated, a round, or a finish. Each call takes one + // at least, so flushes that keep coming slow adoption down but never stop + // it. + bool stepped = false; + while (!stop_requested()) { + // The service's uploads' Cancellation is adoption's too: stop() cuts + // it between steps -- between the packs a listing validates too -- and + // in a round (its uploads, and, through the reads' Cancellation, its + // index reads). + if (upload_cancel_.cancelled()) { + *cut_short = true; + break; + } + // A flush() is waiting for the cycle: the rest of the slice is the + // next cycle's, which the loop runs once the flush is done. + if (stepped && + flushes_in_progress_.load(std::memory_order_acquire) > 0) { + break; + } + stepped = true; + if (adopting_ == nullptr) { + if (adoption_queue_.empty()) break; + const std::string next = adoption_queue_.front(); + adoption_queue_.pop_front(); + begin_adoption(next); + continue; + } + if (!adopting_->validated) { + // One pack of the listing: validating hashes every byte of it, so + // over a dead backlog a whole listing takes as long as the backlog + // is big, and neither a flush nor stop() waits for more than the + // pack in flight. The slice bounds a listing as it does the rounds. + Adoption& adoption = *adopting_; + if (adoption.spool.ContinueRecovery(&adoption.recovery)) { + // Each pack's identity and object key come from the pack and its + // path in the dead directory, exactly as its owner would have + // uploaded it. + adoption.remaining.assign( + std::make_move_iterator(adoption.recovery.valid.begin()), + std::make_move_iterator(adoption.recovery.valid.end())); + adoption.recovery = dmi_store::SpoolRecovery{}; + adoption.validated = true; + } + if (steady_ns() - started >= config_.adoption_slice_ns) break; + continue; + } + if (adopting_->remaining.empty()) { + finish_adoption(); + continue; + } + // The rules of the cycle's uploads, before every round: none without + // the lease, and none while an uploaded pack is still owed to the + // catalog -- the dead spool is where the rest are durable. + bool catalog = false; + { + LeaseScope lease(this); + catalog = writer_.held_lease() != nullptr; + } + if (!catalog || !pending_index_.empty()) break; + bool cut = false; + if (!upload_adopted_round(deadline_ns, deferred, &cut)) { + ok = false; // the failed packs stay in the dead spool + break; + } + if (cut) { + *cut_short = true; // the cut packs stay in the dead spool + break; + } + if (steady_ns() - started >= config_.adoption_slice_ns) break; + } + return ok; +} + +bool CaptureStorageService::scan_siblings() { + namespace fs = std::filesystem; + const fs::path own(spool_.root()); + std::vector siblings; + std::vector claim_staging; + std::error_code ec; + for (fs::directory_iterator it(own.parent_path(), ec), end; + !ec && it != end; it.increment(ec)) { + uint64_t rank = 0; + std::string incarnation; + std::error_code type_ec; + if (it->path() == own || it->is_symlink(type_ec) || + !it->is_directory(type_ec)) { + continue; + } + const std::string name = it->path().filename().string(); + if (dmi_store::IsSpoolClaimStagingName(name)) { + claim_staging.push_back(it->path()); + continue; + } + if (!dmi_store::ParseSpoolRankDirectoryName(name, &rank, &incarnation) || + blocked_siblings_.count(it->path().string()) != 0) { + continue; + } + siblings.push_back(it->path().string()); + } + if (ec) { + record_error("adoption: cannot list " + own.parent_path().string() + + ": " + ec.message()); + return false; + } + clear_dead_claim_staging(claim_staging); + std::sort(siblings.begin(), siblings.end()); + uint64_t live = 0; + for (const std::string& sibling : siblings) { + // A live owner answers a non-blocking probe at once; TryAdopt would + // retry for a few milliseconds first, on every recheck. + if (dmi_store::ReadSpoolOwner(sibling, nullptr)) { + ++live; + } else { + adoption_queue_.push_back(sibling); + } + } + adoption_scan_owed_ = false; + live_siblings_ = live != 0; + last_adoption_scan_ns_ = steady_ns(); + std::lock_guard state(state_mutex_); + state_.live_siblings = live; + return true; +} + +void CaptureStorageService::begin_adoption(const std::string& directory) { + auto adoption = std::make_unique(); + adoption->directory = directory; + std::string error; + const dmi_store::SpoolStatus locked = + dmi_store::SpoolOwnerLock::TryAdopt(directory, &adoption->lock, &error); + if (locked == dmi_store::SpoolStatus::kOwned) { + // Alive after all (taken since the look): its owner's, and not owed. + live_siblings_ = true; + std::lock_guard state(state_mutex_); + ++state_.live_siblings; + return; + } + if (locked != dmi_store::SpoolStatus::kOk) { + // Another adopter drained and removed it meanwhile: nothing is owed. + if (!std::filesystem::exists(directory)) return; + // Its lock cannot be taken at all -- another user's lock file, say. + block_sibling(directory, "cannot lock it: " + error); + return; + } + dmi_store::SpoolConfig config{directory, config_.spool_max_bytes}; + config.owner_lock = dmi_store::OwnerLock::kHeldByCaller; // adoption->lock + config.allow_shared_filesystem = config_.spool_allow_shared_filesystem; + // Its .open files swept and its ready packs listed, none hashed yet: + // adopt_step validates them, a pack a step. + if (dmi_store::Spool::Open(config, &adoption->spool, &error) != + dmi_store::SpoolStatus::kOk || + adoption->spool.BeginRecovery(&adoption->recovery, &error) != + dmi_store::SpoolStatus::kOk) { + block_sibling(directory, "cannot open it: " + error, &adoption->lock); + return; + } + adopting_ = std::move(adoption); +} + +bool CaptureStorageService::upload_adopted_round(uint64_t deadline_ns, + size_t* deferred, + bool* cut) { + *cut = false; + Adoption& adoption = *adopting_; + // A round is a chunk of the service's own upload path (upload_chunk, then + // index_or_owe): uploaded, then indexed before the next is uploaded, so + // at most one chunk (indexer.max_packs) of it is out of the dead spool + // and not yet in the catalog. It is at most uploader.max_workers packs as + // well, since the adoption slice is checked between rounds. + const size_t round = static_cast(std::min( + std::max(1, config_.uploader.max_workers), + std::max(1, config_.indexer.max_packs))); + std::vector entries; + while (!adoption.remaining.empty() && entries.size() < round) { + entries.push_back(std::move(adoption.remaining.front())); + adoption.remaining.pop_front(); + } + // Through the service's upload client and its Cancellation, so stop() + // cuts an adoption's transfers, retries and backoff as it does the + // service's own. + dmi_store::SpoolUploader uploader(&adoption.spool, &upload_s3_, + config_.uploader); + uploader.set_cancellation(&upload_cancel_); + ChunkOutcome sent = upload_chunk(&uploader, entries); + size_t retryable = 0; + std::vector cancelled; + for (size_t i = 0; i < sent.batch.refs.size(); ++i) { + if (!sent.batch.refs[i].pack_id.empty()) continue; + // Still in the dead spool either way. One a cancel cut short goes + // back to the front, as it was; one a later try could upload is + // retried by a later cycle; one no try by this service can is not, + // and blocks the directory once the rest are up. + const dmi_store::UploadFailure failure = + i < sent.batch.failures.size() ? sent.batch.failures[i] + : dmi_store::UploadFailure{}; + const std::string what = failure.object_key + ": " + failure.error; + if (failure.cancelled) { + cancelled.push_back(entries[i]); + } else if (failure.retryable) { + ++retryable; + adoption.remaining.push_back(entries[i]); + record_error("adopting dead spool " + adoption.directory + + ": upload failed for " + what); + } else if (adoption.blocked.empty()) { + adoption.blocked = "it holds a pack this service can never upload, " + + what; + } + } + adoption.remaining.insert(adoption.remaining.begin(), cancelled.begin(), + cancelled.end()); + { + std::lock_guard state(state_mutex_); + state_.adopted_packs += sent.to_index.size(); + } + *cut = !cancelled.empty(); + // Uploaded, so gone from the dead spool: indexed now, a cancel of the + // uploads or not, or owed in pending_index_ like any pack of this + // service's own. After the requeue above, since a lost lease throws out + // of here and the packs still in the dead spool must stay in remaining. + *deferred += index_or_owe(std::move(sent.to_index), true, deadline_ns); + return retryable == 0; +} + +void CaptureStorageService::finish_adoption() { + std::unique_ptr adoption = std::move(adopting_); + const std::string directory = adoption->directory; + if (!adoption->blocked.empty()) { + // The directory stays, and its lock goes. + block_sibling(directory, adoption->blocked, &adoption->lock); + return; + } + { + std::lock_guard state(state_mutex_); + ++state_.adopted_spools; + } + std::string error; + if (!adoption->lock.ReleaseAndRemoveIfEmpty(&error)) { + // Nothing to upload is left, only files that are not its packs (a + // quarantined one, a spool directory nested in it): the directory + // stays for someone to look at. + block_sibling(directory, "it was drained, but still holds files that " + "are not its packs (a quarantined pack, or a " + "spool directory nested in it)"); + } +} + +void CaptureStorageService::block_sibling(const std::string& directory, + const std::string& reason, + dmi_store::SpoolOwnerLock* lock) { + if (lock != nullptr) { + // Said in its lock file, while this service still holds it, then let + // go of: for a person, and for the sinks on the node, which charge a + // dead directory against their budget only while an adoption can + // drain it (SpoolConfig::charge_dead_siblings). A later take -- by a + // process that can adopt it -- rewrites the mark. + lock->MarkBlocked(reason); + lock->Release(); + } + blocked_siblings_.insert(directory); + record_error("dead spool " + directory + " is left in place, not to be " + "adopted by this service: " + reason); + std::lock_guard state(state_mutex_); + state_.blocked_siblings.assign(blocked_siblings_.begin(), + blocked_siblings_.end()); +} + +void CaptureStorageService::clear_dead_claim_staging( + const std::vector& staging) { + namespace fs = std::filesystem; + // A claim builds its directory's staging copy and renames it into place + // within milliseconds, so one this old whose lock nobody holds belongs + // to a claim that died before its rename. Younger ones are left alone: a + // claim between its mkdir and its flock holds no lock yet, and clearing + // its copy would fail it. Nothing is owed either way. + constexpr auto kDeadAfter = std::chrono::seconds(60); + for (const fs::path& path : staging) { + std::error_code ec; + const auto written = fs::last_write_time(path, ec); + if (ec || fs::file_time_type::clock::now() - written < kDeadAfter) { + continue; + } + dmi_store::SpoolOwnerLock lock; + std::string error; + if (dmi_store::SpoolOwnerLock::TryAdopt(path.string(), &lock, &error) == + dmi_store::SpoolStatus::kOk) { + lock.ReleaseAndRemoveIfEmpty(&error); + } + } +} + size_t CaptureStorageService::index_bounded(std::vector refs, std::vector* unindexed, uint64_t deadline_ns) { @@ -1312,6 +1751,13 @@ void CaptureStorageService::latch_failure(std::exception_ptr failure, std::fprintf(stderr, "dmi capture storage: indexing stopped: %s\n", line.c_str()); std::fflush(stderr); + // Woken now, not after its wait (up to max_backoff_ns in an outage), so + // the loop lets go of a sibling it was adopting at once. + { + std::lock_guard lock(wake_mutex_); + kick_ = true; + } + wake_.notify_all(); } } // namespace dmi_catalog diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index 54c3fa1f0..c1b95a28d 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -23,6 +23,14 @@ // already committed, and indexes the rest. It runs at start(), and // periodically when reconcile_interval_ns is non-zero. // +// One owner per spool directory. The service's spool is opened under the +// directory's owner lock (store/spool.h), so a second process on it is +// refused at construction, naming the holder. With adopt_sibling_spools +// the service's directory is one rank directory of the plan's section 2.3 +// layout, and once started its loop adopts the siblings whose owner has +// died: a crashed process's spool is recovered by the next process on the +// node for the same catalog, whatever run it belongs to. +// // Deployment shape: the service holds the catalog's single publisher lease, so // run ONE service per (database, table_prefix). A second one waits up to // start_lease_wait_ns for the lease at start(), then fails naming the holder. @@ -85,10 +93,13 @@ #include #include #include +#include #include +#include #include #include #include +#include #include #include #include @@ -109,6 +120,74 @@ struct StorageServiceConfig { // sink writes. std::string spool_root; uint64_t spool_max_bytes = 1ull << 40; + // The spool directory's owner lock (store/spool.h). kTake owns it for the + // service's life, and refuses a directory another process owns. A process + // that also runs the sink on it -- the engine -- holds one SpoolOwnerLock + // and opens both with kHeldByCaller: two takes in one process refuse each + // other. + dmi_store::OwnerLock spool_owner_lock = dmi_store::OwnerLock::kTake; + bool spool_allow_shared_filesystem = false; + // Adopt the spools of dead processes bound for this catalog. spool_root + // must then be a rank directory of the plan's section 2.3 layout, + // //r-/ + // under THIS catalog's key (SpoolCatalogKey of clickhouse.host and + // .port, writer.database and .table_prefix, s3.endpoint and .bucket, and + // uploader.store_id), or construction throws. + // The loop's cycles adopt, starting with the first, which start() kicks + // off once it holds the lease and has swept its own directory; start() + // itself uploads nothing of a dead backlog, and neither does flush(). A + // cycle probes the owner lock of every sibling rank directory, then works + // through those whose owner is gone: it takes one's lock and sweeps its + // .open files, validates its .ready packs once, a pack a step -- which + // hashes every byte of it -- and uploads and indexes them a round at a + // time -- a chunk of the service's own upload path, at most + // uploader.max_workers and indexer.max_packs packs, each indexed before + // the next is uploaded -- under the cycle's upload rules, the lease and + // nothing owed to the catalog checked before every round, until + // adoption_slice_ns has passed, a stop is requested, or a flush() is + // running -- a flush waits for the step the adoption is in (one round, + // one pack's validation, or one sibling's lock and sweep), not for the + // rest of a slice or of a listing. Its uploads and index reads go through + // the service's own clients and Cancellations, and its listing stops + // between packs on the uploads' cancel, so stop() cuts an adoption as it + // cuts the service's own work: what it did not upload stays in the dead + // spool, which is let go of and never removed while a pack of it is + // left, for the next process on the node. The sibling's lock, its + // listing and its remaining packs are kept between cycles, so a large + // backlog is hashed once, not on every cycle, and neither a flush() nor + // the service's own uploads wait behind all of it. A drained sibling is + // removed once nothing but its lock file is left. An upload that failed + // stays in the dead spool and is retried by + // a later cycle, after the backoff -- unless no retry by this service can + // ever succeed (dmi_store::UploadFailure::retryable: a pack over + // uploader.max_in_flight_bytes, a different object at its key, bytes that + // no longer match), or the sibling cannot be locked or opened at all. + // Such a sibling is blocked: once the rest of its packs are up it is left + // in place, with its lock let go, for a process that can adopt it (or a + // person); it is reported once (last_error, snapshot blocked_siblings), + // and in its lock file's record (dmi_store::SpoolOwner::blocked), which + // keeps the sinks on the node from charging it against their budgets; + // never retried or re-hashed by this service, and not owed. So is one + // drained of packs that still holds other files. flush() covers this + // process's records: a sibling still to adopt does not keep it from + // reporting drained, though adopted packs uploaded and not yet indexed + // are owed like the service's own. Live siblings -- another rank or job on this + // node, a predecessor still closing -- are left alone, and are not owed. + bool adopt_sibling_spools = false; + // While a pass found a live sibling, the loop passes over the siblings + // again this often, so one whose owner dies later -- a predecessor that + // was still inside close() when this service started, a rank that + // crashes while this one runs -- is adopted then, not at the next + // restart on the node. A live sibling costs one non-blocking lock probe + // per pass. 0 never looks again after the first pass. + uint64_t adoption_recheck_interval_ns = 30'000'000'000ull; + // How long one cycle may spend adopting before it lets go of the cycle -- + // to the service's own uploads, stop() -- and carries on in the next. + // Checked between rounds and between the packs a listing validates, so a + // cycle can outrun it by one round, or one pack's hash. A flush() does + // not wait for it: a cycle that has made one step of adoption lets go of + // the cycle at the next one while a flush is running. + uint64_t adoption_slice_ns = 1'000'000'000ull; dmi_store::S3Config s3; dmi_store::UploaderConfig uploader; // uploader.store_id names the store @@ -168,8 +247,9 @@ struct StorageServiceConfig { // past its row: until it resumes and next checks, or -- after a system // suspend, which the steady clock the deadline runs on does not count -- // until a renewal or publish is refused. Two on different (database, - // table_prefix) pairs each hold a lease and upload freely. The spool - // itself is not locked; one process per spool is the caller's job. + // table_prefix) pairs each hold a lease and upload freely. The spool's + // owner lock (spool_owner_lock) is what keeps a second process off the + // directory itself: it is refused at construction, before any of this. bool sweep_spool_on_start = true; bool reconcile_on_start = true; }; @@ -195,6 +275,17 @@ struct StorageServiceSnapshot { uint64_t swept_on_start = 0; // ready packs Recover() found at start uint64_t pending_index = 0; // uploaded packs awaiting a retried index uint64_t rejected_packs = 0; // set aside: cannot be indexed (see flush) + // adopt_sibling_spools: dead siblings drained, the ready packs of theirs + // that were uploaded, and whether one is still to adopt (or a look at + // the siblings is due). + uint64_t adopted_spools = 0; + uint64_t adopted_packs = 0; + bool adoption_owed = false; + // Siblings whose owner was alive at the last adoption pass. + uint64_t live_siblings = 0; + // Dead siblings this service will not adopt, left in place (see + // adopt_sibling_spools), by directory. + std::vector blocked_siblings; // A foreign lease outlived 2 x TTL: the service stopped for good. bool failed = false; // "none" before start, "held", "quarantined" (an unknown outcome set the @@ -227,7 +318,9 @@ class CaptureStorageService { // start_lease_wait_ns for another holder's to expire, or for a claim that // timed out to go through -- past it, once, to wait out the quarantine // such a claim left), sweep the spool, reconcile once, then start the - // background cycle. The lease renews from the moment it is taken. + // background cycle -- whose first cycle, at once, starts adopting dead + // siblings (adopt_sibling_spools). The lease renews from the moment it is + // taken. // Throws if the lease is still held by another publisher when the wait // ends, or its claim still times out. A lease lost while the reconcile // runs does not fail start(): the loop takes a fresh one, as it would @@ -258,6 +351,11 @@ class CaptureStorageService { // request of at most 5 s that nothing cuts -- past one request timeout // when that is under about 6 s. At zero it still runs one cycle, which // indexes one batch of what earlier cycles owe and uploads nothing. + // Its cycles adopt nothing, and a dead sibling still to adopt does not + // keep it from returning true (adopt_sibling_spools); a loop cycle that + // is adopting when it is called lets go of the cycle after the step it + // is in -- one round of a dead spool's uploads, one of its packs + // validated, or one sibling's lock and sweep. // Throws, once, if packs were set aside since the last flush: they are in // the object store but can never reach the catalog. bool flush(double timeout_s); @@ -274,7 +372,8 @@ class CaptureStorageService { // multipart upload it cut (one attempt, 5 s at most), and the lease // release, each catalog request bounded by the client's request timeout // (under the lease, by the lease deadline); the lease renews until the - // loop is done. + // loop is done. An adoption in flight is cut the same way, its dead + // sibling let go of with what it still holds (adopt_sibling_spools). void stop(); StorageServiceSnapshot snapshot() const; @@ -292,6 +391,16 @@ class CaptureStorageService { bool cut_short = false; }; + // One chunk of the upload path, the service's own spool's or a dead + // sibling's (upload_chunk): the batch, positional as UploadStaged returns + // it, what went up and is to be indexed, and how many were not. + struct ChunkOutcome { + dmi_store::UploadBatchResult batch; + std::vector to_index; + size_t failed = 0; // still staged; counted in upload_failures + size_t cancelled = 0; // still staged: a cancel cut them short + }; + // Holds lease_mutex_ for a stretch of catalog work, and bounds every // request the thread makes meanwhile by the held lease's deadline -- read // afresh per request, so a renewal inside the stretch extends it at once. @@ -308,16 +417,77 @@ class CaptureStorageService { // start() waits out like any holder's. class LeaseScope; + // The dead sibling being adopted (storage_service.cpp). + struct Adoption; + void loop(); // start()'s spool sweep and reconcile, with the lease held and the lease // thread renewing it. Requires cycle_mutex_. void sweep_and_reconcile_at_start(); + // One cycle's share of adoption (see adopt_sibling_spools): looks at the + // siblings when that is due, then adopts a step at a time -- a sibling + // taken and swept, one of its packs validated, a round, a finish -- until + // the slice ends, a flush() is waiting (after one step at least), a stop + // is requested or the uploads' Cancellation is cancelled -- which also + // cuts a round in flight, and *cut_short then says so. Adds what its index + // passes left owed through a cancel to *deferred (index_bounded). False + // when an upload failed, so the cycle backs off. Requires cycle_mutex_. + // Only a lost lease propagates. + bool adopt_step(uint64_t deadline_ns, size_t* deferred, bool* cut_short); + // Probes every sibling rank directory's owner lock and queues the dead + // ones. False when the directory cannot be listed. Requires cycle_mutex_. + bool scan_siblings(); + // Takes a queued sibling's lock, sweeps its .open files and lists its + // ready packs into adopting_, hashing none of them (adopt_step validates + // them, a pack a step); leaves adopting_ empty for a live or vanished + // one, and for one it cannot lock or open, which it blocks. + void begin_adoption(const std::string& directory); + // Uploads and indexes one round of adopting_'s packs, one chunk of the + // service's upload path; *cut when a cancel left some in the dead spool, + // back at the front of adopting_. False when an upload failed. + bool upload_adopted_round(uint64_t deadline_ns, size_t* deferred, + bool* cut); + // adopting_ holds no pack any more: removes the directory, or leaves a + // blocked one. + void finish_adoption(); + // Leaves a dead sibling in place for good, reporting why -- and, given + // its held lock, marking why in it (SpoolOwnerLock::MarkBlocked) before + // letting go of it. + void block_sibling(const std::string& directory, const std::string& reason, + dmi_store::SpoolOwnerLock* lock = nullptr); + bool adoption_owed() const; // requires cycle_mutex_ + // Lets go of the sibling being adopted, and of those queued, leaving what + // is left of them for the next process on the node: at stop(), and once + // the service has latched. Requires cycle_mutex_. + void let_go_of_adoption(); + bool stop_requested(); + // Removes the staging copies (dmi_store::IsSpoolClaimStagingName) that + // claims killed before their rename left under the catalog key. + void clear_dead_claim_staging( + const std::vector& staging); + // Uploads one chunk -- packs a listing returned, in its order -- through + // `uploader`, the service's own or an adoption's over a dead spool, and + // books the outcome. The caller indexes to_index (index_or_owe) before it + // uploads another chunk, so at most one chunk is ever out of a spool and + // not yet in the catalog. Requires cycle_mutex_. + ChunkOutcome upload_chunk(dmi_store::SpoolUploader* uploader, + std::vector chunk); + // Indexes refs that are gone from their spool, keeping whatever does not + // index in pending_index_ -- the only record of it in-process. With no + // catalog, keeps them all. Returns how many of those a cancel or the + // deadline left owed (index_bounded). Only a lost lease propagates. + // Requires cycle_mutex_. + size_t index_or_owe(std::vector refs, bool catalog, + uint64_t deadline_ns); // Stops the lease thread and waits for it. void stop_lease_thread(); // One cycle. A non-zero deadline_ns (steady ns) cancels its uploads at // that moment -- flush()'s -- and allow_reconcile false skips the - // periodic and the owed reconcile. Requires cycle_mutex_. - CycleOutcome run_cycle(uint64_t deadline_ns, bool allow_reconcile); + // periodic and the owed reconcile. `adopt`: the loop's cycles adopt dead + // siblings (adopt_sibling_spools), flush()'s do not. Requires + // cycle_mutex_. + CycleOutcome run_cycle(uint64_t deadline_ns, bool allow_reconcile, + bool adopt); // Indexes refs in bounded batches, appending every ref that did not index // to *unindexed. Only a lost lease propagates; other failures are // recorded. A non-zero deadline_ns (steady ns) starts no batch past it @@ -407,6 +577,19 @@ class CaptureStorageService { // The reconcile at start() lost the lease before it finished; the loop // runs one once it holds a lease again. Guarded by cycle_mutex_. bool reconcile_owed_ = false; + // Adoption's state, guarded by cycle_mutex_: whether a look at the + // siblings is due (from start() on), the dead ones the last look found, + // the one being adopted, whether the last look found a live one and when + // it ran (the loop looks again every adoption_recheck_interval_ns). + bool adoption_scan_owed_ = false; + std::deque adoption_queue_; + std::unique_ptr adopting_; + std::set blocked_siblings_; + bool live_siblings_ = false; + uint64_t last_adoption_scan_ns_ = 0; + // flush() calls in progress. A cycle adopting lets go of the cycle at + // its next step while one is, so a flush never waits out a slice. + std::atomic flushes_in_progress_{0}; int failure_streak_ = 0; // consecutive failed cycles, for the backoff // steady ns at which the last cycle -- the loop's or a flush's -- ended; // 0 before the first. The loop waits its interval from it. Guarded by diff --git a/native/csrc/sink/bindings_sink.cpp b/native/csrc/sink/bindings_sink.cpp index e9bd394aa..bb10b213e 100644 --- a/native/csrc/sink/bindings_sink.cpp +++ b/native/csrc/sink/bindings_sink.cpp @@ -182,8 +182,19 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { uint64_t max_linger_ns, uint64_t spool_max_bytes, const std::string& overload, std::optional admission_timeout_s, - double release_flush_timeout_s) { + double release_flush_timeout_s, + const std::string& owner_lock, + bool allow_shared_filesystem, + bool charge_dead_siblings) { dmi_sink::SinkConfig config; + if (!dmi_store::ParseOwnerLock(owner_lock, + &config.spool_owner_lock)) { + throw py::value_error( + "owner_lock must be 'take' or 'held_by_caller', got '" + + owner_lock + "'"); + } + config.spool_allow_shared_filesystem = allow_shared_filesystem; + config.spool_charge_dead_siblings = charge_dead_siblings; config.overload = ParseOverload(overload); config.admission_timeout_s = ParseAdmissionTimeout(admission_timeout_s); @@ -224,7 +235,17 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { py::arg("release_flush_timeout_s") = std::chrono::duration( dmi_sink::NativePackSink::kDefaultReleaseFlushTimeout) - .count()) + .count(), + // The spool directory's owner lock (store/spool.h): "take" owns + // it for the sink's life; "held_by_caller" when the caller holds + // a SpoolOwnerLock on it, as the engine does around its sink and + // storage service. + py::arg("owner_lock") = "take", + py::arg("allow_shared_filesystem") = false, + // spool_root is a rank directory of the spool layout, and the + // dead incarnations' packs beside it count against + // spool_max_bytes, as the engine's claimed directory does. + py::arg("charge_dead_siblings") = false) .def("attach", [](std::shared_ptr self) { // Simulates engine ownership for tests (the real engine takes @@ -278,6 +299,10 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { [](const dmi_sink::NativePackSink& self) { return std::string(OverloadName(self.sink().config().overload)); }) + // Released by its engine, and the release backstop's flush went + // through: nothing the sink holds can still reach the spool. + .def_property_readonly("sealed_on_release", + &dmi_sink::NativePackSink::sealed_on_release) .def_property_readonly( "release_flush_timeout_s", [](const dmi_sink::NativePackSink& self) { diff --git a/native/csrc/sink/conformance_sink.cpp b/native/csrc/sink/conformance_sink.cpp index 47196a0ff..042f6b9e3 100644 --- a/native/csrc/sink/conformance_sink.cpp +++ b/native/csrc/sink/conformance_sink.cpp @@ -4,7 +4,7 @@ // {"op":"open","root":"...","max_bytes":N,"max_queue_records":N, // "max_queue_bytes":N,"max_pack_bytes":N,"max_pack_records":N, // "max_linger_ns":N,"overload":"block"|"drop_newest", -// "admission_timeout":-1} +// "admission_timeout":-1,"owner_lock":"take"|"held_by_caller" (optional)} // -> {"ok":true} // {"op":"submit","metadata":{...canonical field names...},"payload_b64":"..."} // -> {"ok":true,"admission":"accepted"|...} @@ -209,6 +209,14 @@ int main() { jc::FindString(line, "overload") == "block" ? dmi_sink::Overload::kBlock : dmi_sink::Overload::kDropNewest; + const std::string owner_lock = jc::FindString(line, "owner_lock"); + if (!owner_lock.empty() && + !dmi_store::ParseOwnerLock(owner_lock, &config.spool_owner_lock)) { + std::string out = "{\"ok\":false,\"what\":"; + jc::EscapeJson("unknown owner_lock: " + owner_lock, &out); + std::cout << out << "}\n"; + continue; + } const int64_t workers = Integer(line, "num_workers"); config.num_workers = static_cast(workers > 0 ? workers : 1); // Before the sink exists: a limit that cannot be represented must not diff --git a/native/csrc/sink/native_pack_sink.cpp b/native/csrc/sink/native_pack_sink.cpp index 5f8d9bcbf..afcc79cb1 100644 --- a/native/csrc/sink/native_pack_sink.cpp +++ b/native/csrc/sink/native_pack_sink.cpp @@ -258,12 +258,16 @@ bool NativePackSink::flush_and_wait(Duration timeout) { } void NativePackSink::on_engine_release() noexcept { + sealed_on_release_.store(false, std::memory_order_release); if (release_flush_timeout_ == Duration::zero()) return; try { std::string error; const bool flushed = sink_->Flush( std::chrono::duration(release_flush_timeout_).count(), &error); - if (flushed) return; + if (flushed) { + sealed_on_release_.store(true, std::memory_order_release); + return; + } // Once released, rethrow_if_failed refuses the sink as not attached, // and a timeout latches nothing: this line is what says what became of // the open pack (a failure also counts in the snapshot's failures). diff --git a/native/csrc/sink/native_pack_sink.h b/native/csrc/sink/native_pack_sink.h index 97b2dfeaa..10b2d19bb 100644 --- a/native/csrc/sink/native_pack_sink.h +++ b/native/csrc/sink/native_pack_sink.h @@ -10,6 +10,7 @@ #ifndef DMI_SINK_NATIVE_PACK_SINK_H_ #define DMI_SINK_NATIVE_PACK_SINK_H_ +#include #include #include #include @@ -66,9 +67,22 @@ class NativePackSink final : public ring::RecordSink { PackSink& sink_for_testing() { return *sink_; } const std::string& layout() const { return layout_; } Duration release_flush_timeout() const { return release_flush_timeout_; } + // Whether the sink can no longer write its spool: its engine released it + // and the release backstop's flush went through. A released sink admits + // nothing, so once that flush has persisted everything admitted, no stage + // is left to come -- not from its stagers, its linger, or its destructor. + // False while attached (and from a new engine's acquire on), after a + // release whose flush failed or timed out, and when the backstop is off + // (release_flush_timeout zero): a stage may then still be on its way. + // The engine lets go of the spool directory's owner lock only on true. + bool sealed_on_release() const { + return sealed_on_release_.load(std::memory_order_acquire); + } protected: - void on_engine_acquire() override {} + void on_engine_acquire() override { + sealed_on_release_.store(false, std::memory_order_release); + } // The backstop for a ring that stops without a flush: RingEngine::stop // drains its record worker into submit() and then releases the sink, and // until a flush seals it the open pack is only in memory -- for up to @@ -80,7 +94,8 @@ class NativePackSink final : public ring::RecordSink { // times out writes one line to stderr, and that line is the only report // of a timeout. A pipeline failure also counts in snapshot()["failures"]; // rethrow_if_failed reports it only once the sink is attached again, - // since a released sink refuses the call as not attached. + // since a released sink refuses the call as not attached. Whether the + // flush went through is sealed_on_release(). void on_engine_release() noexcept override; private: @@ -91,6 +106,7 @@ class NativePackSink final : public ring::RecordSink { std::unique_ptr sink_; const std::string layout_; const Duration release_flush_timeout_; + std::atomic sealed_on_release_{false}; // Counters at construction. Losses are judged against it, as the // reference adapter judges them against its pipeline's baseline. SinkSnapshot baseline_; diff --git a/native/csrc/sink/pack_sink.cpp b/native/csrc/sink/pack_sink.cpp index f51e129c6..112f4a5f3 100644 --- a/native/csrc/sink/pack_sink.cpp +++ b/native/csrc/sink/pack_sink.cpp @@ -96,6 +96,9 @@ std::string PackSink::Start(std::string* spool_error) { dmi_store::SpoolConfig spool_config; spool_config.root = config_.spool_root; spool_config.max_bytes = config_.spool_max_bytes; + spool_config.owner_lock = config_.spool_owner_lock; + spool_config.allow_shared_filesystem = config_.spool_allow_shared_filesystem; + spool_config.charge_dead_siblings = config_.spool_charge_dead_siblings; std::string error; const dmi_store::SpoolStatus st = dmi_store::Spool::Open(spool_config, &spool_, &error); diff --git a/native/csrc/sink/pack_sink.h b/native/csrc/sink/pack_sink.h index 214f7478c..9399754ba 100644 --- a/native/csrc/sink/pack_sink.h +++ b/native/csrc/sink/pack_sink.h @@ -68,6 +68,16 @@ struct SinkConfig { double admission_timeout_s = -1.0; std::string spool_root; uint64_t spool_max_bytes = 1ull << 40; + // The spool directory's owner lock (spool.h). kTake owns the directory + // for the sink's life; a process that also runs a storage service on it + // holds one SpoolOwnerLock and opens both with kHeldByCaller. + dmi_store::OwnerLock spool_owner_lock = dmi_store::OwnerLock::kTake; + bool spool_allow_shared_filesystem = false; + // spool_root is a rank directory of the section 2.3 layout, and what the + // dead incarnations beside it still hold counts against spool_max_bytes + // (SpoolConfig::charge_dead_siblings). The engine sets it for the + // directory it claims. + bool spool_charge_dead_siblings = false; // Pack assembler workers. Records route by scope hash // (tenant, session, producer_rank), so one scope always lands on one // worker: per-scope ordering and single-scope packs are preserved at any diff --git a/native/csrc/store/conformance_spool.cpp b/native/csrc/store/conformance_spool.cpp index 9112ce919..b8883b847 100644 --- a/native/csrc/store/conformance_spool.cpp +++ b/native/csrc/store/conformance_spool.cpp @@ -11,6 +11,9 @@ // -> {"ok":true} // {"op":"snapshot","root":"...","max_bytes":N} // -> {"ok":true,"snapshot":{"entries":N,"bytes":N,"peak_bytes":N,"max_bytes":N}} +// Every op re-opens the spool, and takes "owner_lock":"take" (the default) +// or "held_by_caller"; an open another owner refuses answers +// {"ok":false,"status":"open","what":"...owned by pid N on host H..."}. // Errors: {"ok":false,"status":"...","what":"..."}. #include "spool.h" @@ -103,8 +106,18 @@ int main() { continue; } if (config.max_bytes == 0) config.max_bytes = 1ull << 40; + // Optional: "take" (the default) or "held_by_caller". + const std::string owner_lock = jc::FindString(line, "owner_lock"); dmi_store::Spool spool; std::string error; + if (!owner_lock.empty() && + !dmi_store::ParseOwnerLock(owner_lock, &config.owner_lock)) { + std::string out = "{\"ok\":false,\"status\":\"open\",\"what\":"; + jc::EscapeJson("unknown owner_lock: " + owner_lock, &out); + out += "}\n"; + std::cout << out; + continue; + } if (dmi_store::Spool::Open(config, &spool, &error) != dmi_store::SpoolStatus::kOk) { std::string out = "{\"ok\":false,\"status\":\"open\",\"what\":"; diff --git a/native/csrc/store/conformance_store.cpp b/native/csrc/store/conformance_store.cpp index 51fae7ee2..fc264f58f 100644 --- a/native/csrc/store/conformance_store.cpp +++ b/native/csrc/store/conformance_store.cpp @@ -15,6 +15,11 @@ // {"op":"list",...,"prefix":"...","delimiter":"...","max_keys":N, // "continuation":"..."} -> {"ok":true,"truncated":bool,"next_token":"...", // "objects":[{"key":"...","size":N,"etag":"..."}...],"attempts":N} +// {"op":"upload_one"|"upload_pending",...,"root":"...", +// "owner_lock":"take"|"held_by_caller" (optional, take by default)} +// open the spool at root, owner lock included, for the one op; +// upload_pending's failures carry pack_id, object_key, attempts, error, +// cancelled and retryable. // Errors: {"ok":false,"what":"..."}. // // Any op may carry "cancel_after_ms":N: the client (and the uploader) get a @@ -354,7 +359,12 @@ int main() { if (spool_config.max_bytes == 0) spool_config.max_bytes = 1ull << 40; dmi_store::Spool spool; std::string spool_error; - if (dmi_store::Spool::Open(spool_config, &spool, &spool_error) != + const std::string owner_lock = jc::FindString(line, "owner_lock"); + if (!owner_lock.empty() && + !dmi_store::ParseOwnerLock(owner_lock, &spool_config.owner_lock)) { + out += "false,\"what\":"; + jc::EscapeJson("spool open: unknown owner_lock: " + owner_lock, &out); + } else if (dmi_store::Spool::Open(spool_config, &spool, &spool_error) != dmi_store::SpoolStatus::kOk) { out += "false,\"what\":"; jc::EscapeJson("spool open: " + spool_error, &out); @@ -400,7 +410,12 @@ int main() { if (spool_config.max_bytes == 0) spool_config.max_bytes = 1ull << 40; dmi_store::Spool spool; std::string spool_error; - if (dmi_store::Spool::Open(spool_config, &spool, &spool_error) != + const std::string owner_lock = jc::FindString(line, "owner_lock"); + if (!owner_lock.empty() && + !dmi_store::ParseOwnerLock(owner_lock, &spool_config.owner_lock)) { + out += "false,\"what\":"; + jc::EscapeJson("spool open: unknown owner_lock: " + owner_lock, &out); + } else if (dmi_store::Spool::Open(spool_config, &spool, &spool_error) != dmi_store::SpoolStatus::kOk) { out += "false,\"what\":"; jc::EscapeJson("spool open: " + spool_error, &out); @@ -439,6 +454,8 @@ int main() { jc::EscapeJson(failure.error, &out); out += std::string(",\"cancelled\":") + (failure.cancelled ? "true" : "false"); + out += std::string(",\"retryable\":") + + (failure.retryable ? "true" : "false"); out += "}"; first = false; } diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 5be18b7ca..ac392c013 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -5,12 +5,29 @@ #include #include +#include +#include +#include +#include #include +#include #include #include #include +#include +#include #include +#include +#include +#include +#if defined(__linux__) +#include +#else +// statfs(2) with f_fstypename: macOS and the BSDs. +#include +#include +#endif #include namespace dmi_store { @@ -20,6 +37,96 @@ namespace { constexpr const char* kReadySuffix = ".dmi-pack.ready"; constexpr const char* kOpenSuffix = ".open"; +constexpr const char* kOwnerLockFile = ".owner.lock"; +// SpoolOwnerLock::Acquire builds a new directory as +// /..<8 hex>.creating and renames it into place. +constexpr const char* kClaimStagingSuffix = ".creating"; +// /_refs/: the upload handoff's ref files (plan section 2.4). Not a +// legal object-key component (those start with an alphanumeric), so no pack +// is ever staged under it. +constexpr const char* kRefsDirectory = "_refs"; + +// statfs(2) f_type values of network filesystems, whose flock does not +// keep out a process on another node -- or is not a place a node-local +// spool can be (linux/magic.h has NFS, SMB2, CIFS, FUSE, 9p and both AFS +// values; the others are their own). BeeGFS keeps flock client-local +// unless tuneUseGlobalFileLocks is set, and GPFS (IBM Storage Scale) keeps +// it node-local. FUSE covers network filesystems (sshfs, s3fs, gcsfuse, +// GlusterFS) and local ones alike, and f_type cannot tell them apart, so a +// local one needs the override. AFS is OpenAFS's and kAFS's. +constexpr uint32_t kNfsSuperMagic = 0x6969; +constexpr uint32_t kLustreSuperMagic = 0x0BD00BD0; +constexpr uint32_t kBeeGfsSuperMagic = 0x19830326; +constexpr uint32_t kCifsSuperMagic = 0xFF534D42; +constexpr uint32_t kSmb2SuperMagic = 0xFE534D42; +constexpr uint32_t kFuseSuperMagic = 0x65735546; +constexpr uint32_t kGpfsSuperMagic = 0x47504653; +constexpr uint32_t kV9fsMagic = 0x01021997; +constexpr uint32_t kAfsSuperMagic = 0x5346414F; +constexpr uint32_t kAfsFsMagic = 0x6B414653; +constexpr uint32_t kOrangeFsSuperMagic = 0x20030528; + +std::atomic g_filesystem_type_for_testing{-1}; + +#if !defined(__linux__) +// Where statfs names the filesystem rather than giving Linux's magic: the +// same refusal, by f_fstypename (FreeBSD spells a FUSE mount +// "fusefs."). +const char* SharedFilesystemTypeName(const char* name) { + static const char* const kShared[][2] = { + {"nfs", "NFS"}, {"smbfs", "SMB"}, {"afpfs", "AFP"}, + {"webdav", "WebDAV"}, {"lustre", "Lustre"}, {"macfuse", "FUSE"}, + {"osxfuse", "FUSE"}, {"fusefs", "FUSE"}, {"afs", "AFS"}}; + for (const auto& entry : kShared) { + const size_t n = std::strlen(entry[0]); + if (std::strncmp(name, entry[0], n) == 0 && + (name[n] == '\0' || name[n] == '.')) { + return entry[1]; + } + } + return nullptr; +} +#endif +std::function& LockOpenHookForTesting() { + static auto* hook = new std::function; + return *hook; +} + +// Whether a recursive walk of a spool root is at /_refs, which no scan +// enters. +bool AtRefsDirectory(const fs::recursive_directory_iterator& it) { + std::error_code ec; + return it.depth() == 0 && it->path().filename() == kRefsDirectory && + it->is_directory(ec); +} + +// Whether a recursive walk of a spool is at a subdirectory with a lock +// file of its own: another spool directory nested in this one -- a rank +// directory of the layout under a flat spool_root, a claim's staging copy, +// a root someone put inside a dead rank directory -- live or dead. No walk +// of this spool enters it, for counting, sweeping, listing or uploading: a +// live one's owner is writing it, and a dead one's packs are its +// successor's to adopt, under its own keys, not this spool's. +bool AtNestedSpool(const fs::recursive_directory_iterator& it) { + std::error_code ec; + return it->is_directory(ec) && !it->is_symlink(ec) && + fs::exists(it->path() / kOwnerLockFile, ec); +} + +// Every walk of a spool's own files skips these two. +bool AtSkippedDirectory(const fs::recursive_directory_iterator& it) { + return AtRefsDirectory(it) || AtNestedSpool(it); +} + +std::string Hostname() { + char host[256] = {0}; + if (::gethostname(host, sizeof(host) - 1) != 0 || host[0] == '\0') { + return "unknown-host"; + } + return host; +} + +std::string Errno(int error) { return std::strerror(error); } bool HasSuffix(const std::string& name, const char* suffix) { const size_t n = std::strlen(suffix); @@ -238,14 +345,951 @@ std::string ReadyName(const std::string& pack_id, uint64_t created, std::to_string(records) + "." + checksum + ".dmi-pack.ready"; } +void FsyncParent(const std::string& path) { + FsyncDir(fs::path(path).parent_path().string(), nullptr); +} + +// The record's roles, on the line after " ". +constexpr const char* kAdoptingRole = "adopting"; +constexpr const char* kBlockedRole = "blocked: "; +constexpr size_t kOwnerRecordBytes = 2048; + +// " \n", then a role line when there is one ("adopting", or +// "blocked: "): whoever holds the lock records itself, so a refused +// process can say who holds the directory, and a sink beside it whether +// an adoption can drain it. +void WriteOwnerRecord(int fd, const std::string& role = "") { + std::string record = Hostname() + " " + std::to_string(::getpid()) + "\n"; + if (!role.empty()) { + std::string line = role.substr(0, kOwnerRecordBytes - record.size() - 2); + std::replace(line.begin(), line.end(), '\n', ' '); + record += line + "\n"; + } + if (::ftruncate(fd, 0) == 0) { + (void)!::pwrite(fd, record.data(), record.size(), 0); + } +} + +void ReadOwnerRecord(int fd, SpoolOwner* owner) { + char buffer[kOwnerRecordBytes]; + const ssize_t n = ::pread(fd, buffer, sizeof(buffer), 0); + *owner = SpoolOwner{}; + if (n <= 0) return; + std::string record(buffer, static_cast(n)); + const auto trim = [](std::string* text) { + while (!text->empty() && std::isspace(static_cast( + text->back()))) { + text->pop_back(); + } + }; + const size_t newline = record.find('\n'); + if (newline != std::string::npos) { + std::string role = record.substr(newline + 1); + record.resize(newline); + trim(&role); + const size_t blocked = std::strlen(kBlockedRole); + if (role == kAdoptingRole) { + owner->adopting = true; + } else if (role.compare(0, blocked, kBlockedRole) == 0) { + owner->blocked = role.substr(blocked); + if (owner->blocked.empty()) owner->blocked = "blocked"; + } + } + trim(&record); + const size_t space = record.rfind(' '); + if (space == std::string::npos) { + owner->host = record; + return; + } + owner->host = record.substr(0, space); + const std::string pid = record.substr(space + 1); + if (!pid.empty() && pid.size() <= 18 && + std::all_of(pid.begin(), pid.end(), + [](char c) { return c >= '0' && c <= '9'; })) { + owner->pid = std::strtoll(pid.c_str(), nullptr, 10); + } +} + +std::string OwnedMessage(const std::string& dir, const SpoolOwner& owner) { + const std::string file = dir + "/" + kOwnerLockFile; + if (owner.pid <= 0) { + return "spool directory " + dir + " is owned by another process, which " + "holds " + file + " and has not recorded itself yet; a spool " + "directory has one owner process"; + } + return "spool directory " + dir + " is owned by pid " + + std::to_string(owner.pid) + " on host " + owner.host + + " (it holds " + file + "); a spool directory has one owner process"; +} + +// Whether `fd` is still the file at `path`: a remover unlinks the lock file +// before it removes a drained directory, and a lock taken on the unlinked +// file guards nothing. +bool IsFileAt(int fd, const std::string& path) { + struct stat by_fd{}, by_path{}; + return ::fstat(fd, &by_fd) == 0 && ::stat(path.c_str(), &by_path) == 0 && + by_fd.st_dev == by_path.st_dev && by_fd.st_ino == by_path.st_ino; +} + +// A path that may not exist yet, absolute, with its existing prefix's +// symlinks resolved. +std::string CanonicalPath(const std::string& path, std::string* error) { + std::error_code ec; + const fs::path absolute = fs::absolute(path, ec); + if (ec) { + if (error) *error = "cannot resolve " + path + ": " + ec.message(); + return ""; + } + const fs::path canonical = fs::weakly_canonical(absolute, ec); + if (ec) { + if (error) *error = "cannot resolve " + path + ": " + ec.message(); + return ""; + } + std::string out = canonical.string(); + while (out.size() > 1 && out.back() == '/') out.pop_back(); + return out; +} + +// A spool directory must not be nested under, or contain, another one that +// is owned: Scan walks recursively, and although every walk passes over a +// subdirectory with a lock file of its own (AtNestedSpool), the outer +// spool's walk can reach one before its owner's lock file is there -- a +// take of an existing directory creates it -- and sweep the .open files +// its owner then writes. Run AFTER `dir`'s own lock is taken, so that of +// two takes racing on an outer directory and one inside it, at least one +// sees the other: each publishes its lock before it looks. Only a HELD +// lock refuses, in either direction. One nobody holds is a spool that was: +// every take leaves its file behind, a spool_root a sink-only run once +// owned holds one, and a crashed default-mode run leaves its rank +// directory's under spool_root. Such a directory's packs are left alone by +// every walk of the other (AtNestedSpool) -- a dead one inside is for its +// successor to adopt -- so it refuses nothing: the rank directories under +// a flat spool_root and that spool_root's own modes (sink-only, explicit +// record_sink) take turns, and never run at once. +SpoolStatus CheckNotNested(const std::string& dir, std::string* error) { + const auto holder = [](const SpoolOwner& owner) { + return owner.pid > 0 ? "pid " + std::to_string(owner.pid) + " on host " + + owner.host + : std::string("another owner"); + }; + fs::path ancestor(dir); + while (ancestor.has_parent_path() && ancestor.parent_path() != ancestor) { + ancestor = ancestor.parent_path(); + std::error_code ec; + SpoolOwner owner; + if (fs::exists(ancestor / kOwnerLockFile, ec) && + ReadSpoolOwner(ancestor.string(), &owner)) { + if (error) { + *error = "spool directory " + dir + " is nested under the spool " + "directory " + ancestor.string() + ", which " + + holder(owner) + " holds (" + kOwnerLockFile + "), and whose " + "recovery would sweep this one; use a directory outside " + "it, or wait for that owner to end"; + } + return SpoolStatus::kBadArgument; + } + } + std::error_code ec; + for (auto it = fs::recursive_directory_iterator( + dir, fs::directory_options::skip_permission_denied, ec); + !ec && it != fs::recursive_directory_iterator(); it.increment(ec)) { + if (it->path().filename() != kOwnerLockFile) continue; + const fs::path owned = it->path().parent_path(); + if (owned == fs::path(dir)) continue; + SpoolOwner owner; + if (!ReadSpoolOwner(owned.string(), &owner)) continue; // dead: left be + if (error) { + *error = "spool directory " + dir + " contains the spool directory " + + owned.string() + ", which " + holder(owner) + " holds (" + + kOwnerLockFile + "), and which this one's recovery could " + "sweep while its owner writes it; use a directory that does " + "not contain it, or wait for that owner to end"; + } + return SpoolStatus::kBadArgument; + } + return SpoolStatus::kOk; +} + +} // namespace + +const char* OwnerLockName(OwnerLock mode) { + return mode == OwnerLock::kHeldByCaller ? "held_by_caller" : "take"; +} + +bool ParseOwnerLock(const std::string& text, OwnerLock* mode) { + if (text == "take") { + *mode = OwnerLock::kTake; + return true; + } + if (text == "held_by_caller") { + *mode = OwnerLock::kHeldByCaller; + return true; + } + return false; +} + +const char* SharedFilesystemName(int64_t f_type) { + switch (static_cast(f_type)) { + case kNfsSuperMagic: return "NFS"; + case kLustreSuperMagic: return "Lustre"; + case kBeeGfsSuperMagic: return "BeeGFS"; + case kCifsSuperMagic: return "CIFS"; + case kSmb2SuperMagic: return "SMB2"; + case kFuseSuperMagic: return "FUSE"; + case kGpfsSuperMagic: return "GPFS"; + case kV9fsMagic: return "9p"; + case kAfsSuperMagic: return "AFS"; + case kAfsFsMagic: return "AFS"; + case kOrangeFsSuperMagic: return "OrangeFS"; + default: return nullptr; + } +} + +void SetFilesystemTypeForTesting(int64_t f_type) { + g_filesystem_type_for_testing.store(f_type); +} + +void SetLockOpenHookForTesting(std::function hook) { + LockOpenHookForTesting() = std::move(hook); +} + +SpoolStatus CheckNodeLocal(const std::string& dir, + bool allow_shared_filesystem, std::string* error) { + const int64_t f_type = g_filesystem_type_for_testing.load(); + const char* shared = nullptr; + std::string seen; // what statfs said, for the refusal + const auto magic = [](int64_t value) { + char text[48]; + std::snprintf(text, sizeof(text), "statfs f_type 0x%llx", + static_cast(value)); + return std::string(text); + }; + if (f_type >= 0) { + shared = SharedFilesystemName(f_type); + seen = magic(f_type); + } else { + struct statfs info{}; + if (::statfs(dir.c_str(), &info) != 0) { + if (error) *error = "cannot statfs " + dir + ": " + Errno(errno); + return SpoolStatus::kIo; + } +#if defined(__linux__) + const int64_t type = + static_cast(static_cast(info.f_type)); + shared = SharedFilesystemName(type); + seen = magic(type); +#else + shared = SharedFilesystemTypeName(info.f_fstypename); + seen = std::string("statfs f_fstypename ") + info.f_fstypename; +#endif + } + if (shared == nullptr || allow_shared_filesystem) return SpoolStatus::kOk; + if (error) { + *error = "spool directory " + dir + " is on " + shared + " (" + seen + + "): a spool must be node-local, " + "since its owner lock (flock) does not keep out a process on " + "another node there. Use a local disk, or set " + "allow_shared_filesystem if no process on another node can " + "reach this directory" + + (std::string(shared) == "FUSE" + ? " (a FUSE filesystem that is itself local, such as " + "fuse-overlayfs or ntfs-3g, is one)" + : std::string()); + } + return SpoolStatus::kBadArgument; +} + +namespace { +// A probe of one flock: if it can be taken nobody holds it, and it is let +// go at once. (A take racing the probe retries, see LockInPlace.) An +// unlock on the probe's own description, so a child forked meanwhile +// keeps nothing either. +bool FlockHeld(int fd) { + if (::flock(fd, LOCK_EX | LOCK_NB) == 0) { + ::flock(fd, LOCK_UN); + return false; + } + return errno == EWOULDBLOCK; +} } // namespace +bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner) { + const std::string file = dir + "/" + kOwnerLockFile; + const int fd = ::open(file.c_str(), O_RDONLY | O_CLOEXEC); + bool held = fd >= 0 && FlockHeld(fd); + if (!held) { + // The directory's own lock, which its owner keeps however its lock + // file is replaced (SpoolOwnerLock). + const int dir_fd = ::open(dir.c_str(), O_RDONLY | O_DIRECTORY | O_CLOEXEC); + if (dir_fd >= 0) { + held = FlockHeld(dir_fd); + ::close(dir_fd); + } + } + if (owner != nullptr) { + if (fd >= 0) { + ReadOwnerRecord(fd, owner); + } else { + *owner = SpoolOwner{}; + } + } + if (fd >= 0) ::close(fd); + return held; +} + +bool IsSpoolClaimStagingName(const std::string& name) { + // "." + + "." + 8 hex + ".creating", not empty. + const size_t suffix = std::strlen(kClaimStagingSuffix); + if (name.size() < 1 + 1 + 1 + 8 + suffix || name[0] != '.' || + !HasSuffix(name, kClaimStagingSuffix)) { + return false; + } + const size_t dot = name.size() - suffix - 9; + if (name[dot] != '.') return false; + return std::all_of(name.begin() + dot + 1, name.end() - suffix, [](char c) { + return (c >= '0' && c <= '9') || (c >= 'a' && c <= 'f'); + }); +} + +namespace { +std::atomic g_fdinfo_hides_locks_for_testing{false}; + +// Whether /proc/self/fdinfo/ lists a write flock on the descriptor's +// open file description ("lock: 1: FLOCK ADVISORY WRITE ..."). +bool FdinfoShowsWriteFlock(const std::string& fd) { + if (g_fdinfo_hides_locks_for_testing.load()) return false; + std::FILE* in = std::fopen(("/proc/self/fdinfo/" + fd).c_str(), "re"); + if (in == nullptr) return false; + bool shown = false; + char line[512]; + while (!shown && std::fgets(line, sizeof(line), in) != nullptr) { + shown = std::strncmp(line, "lock:", 5) == 0 && + std::strstr(line, " FLOCK ") != nullptr && + std::strstr(line, " WRITE ") != nullptr; + } + std::fclose(in); + return shown; +} +// Whether this kernel lists flocks in /proc/self/fdinfo at all. Linux does +// (since 3.8); gVisor's procfs prints only pos, flags and mnt_id, and +// WSL1's lists no locks either. Probed once, on a flock taken on a +// temporary file; no temporary file, no telling, and the record decides. +bool FdinfoListsFlocks() { + if (g_fdinfo_hides_locks_for_testing.load()) return false; + static const bool lists = [] { + std::FILE* temp = std::tmpfile(); + if (temp == nullptr) return false; + const int fd = ::fileno(temp); + const bool shown = ::flock(fd, LOCK_EX | LOCK_NB) == 0 && + FdinfoShowsWriteFlock(std::to_string(fd)); + std::fclose(temp); + return shown; + }(); + return lists; +} +} // namespace + +void SetFdinfoHidesLocksForTesting(bool hide) { + g_fdinfo_hides_locks_for_testing.store(hide); +} + +bool SpoolOwnedByThisProcess(const std::string& dir) { + const std::string file = dir + "/" + kOwnerLockFile; + // The lock file, and the directory itself, which its owner locks too: a + // lock file replaced behind the owner's back is no longer the one it + // holds, but the directory is. + struct stat targets[2]{}; + size_t n_targets = 0; + if (::stat(file.c_str(), &targets[n_targets]) == 0) ++n_targets; + if (::stat(dir.c_str(), &targets[n_targets]) == 0 && + S_ISDIR(targets[n_targets].st_mode)) { + ++n_targets; + } + if (n_targets == 0) return false; + // /proc/self/fdinfo/ lists the flocks each open file description + // holds ("lock: 1: FLOCK ADVISORY WRITE ..."), so the kernel says + // whether one of this process's descriptors on the file holds the lock -- + // a SpoolOwnerLock's, or one a Spool took with kTake. The record in the + // file is only a fallback, where /proc cannot be read or lists no flocks + // (FdinfoListsFlocks): it is written after the lock is taken, and a pid + // says nothing across pid namespaces. + DIR* fds = FdinfoListsFlocks() ? ::opendir("/proc/self/fd") : nullptr; + if (fds == nullptr) { + SpoolOwner owner; + return ReadSpoolOwner(dir, &owner) && owner.pid == ::getpid() && + owner.host == Hostname(); + } + const int listing = ::dirfd(fds); + bool held = false; + while (!held) { + const dirent* entry = ::readdir(fds); + if (entry == nullptr) break; + char* end = nullptr; + const long fd = std::strtol(entry->d_name, &end, 10); + if (end == entry->d_name || *end != '\0' || fd == listing) continue; + struct stat by_fd{}; + if (::fstat(static_cast(fd), &by_fd) != 0) continue; + bool on_target = false; + for (size_t i = 0; i < n_targets; ++i) { + on_target = on_target || (by_fd.st_dev == targets[i].st_dev && + by_fd.st_ino == targets[i].st_ino); + } + if (!on_target) continue; + held = FdinfoShowsWriteFlock(entry->d_name); + } + ::closedir(fds); + return held; +} + +namespace { + +// Every descriptor this binary has open on a spool owner lock file, or on +// the directory it locks with it: held, or between its open() and its +// flock, or on its way to close(). The fork +// handlers close the child's copies of all of them. Tracking starts at the +// open() and ends at the close(), each under the mutex that BeforeFork +// takes, so no fork -- from any thread, at any point of a take or a +// release -- hands a child a copy the handler does not know of: a copy +// made before the flock shares the description the flock then locks. +// Leaked on purpose, so no static destructor runs while one is open. Each +// binary that compiles spool.cpp (the store and sink extensions, the +// drivers) keeps its own set and its own handlers, for its own descriptors. +std::mutex& LockDescriptorsMutex() { + static std::mutex* mutex = new std::mutex; + return *mutex; +} +std::unordered_set& LockDescriptors() { + static auto* descriptors = new std::unordered_set; + return *descriptors; +} +// Bumped in each forked child, where every SpoolOwnerLock taken before the +// fork then reads as released (SpoolOwnerLock::held). +std::atomic g_fork_generation{0}; + +void BeforeForkLockDescriptors() { LockDescriptorsMutex().lock(); } +void AfterForkLockDescriptorsInParent() { LockDescriptorsMutex().unlock(); } +void AfterForkLockDescriptorsInChild() { + // Close, never LOCK_UN: an unlock on the shared description would drop + // the parent's hold too, and closing one of its descriptors does not. + for (const int fd : LockDescriptors()) ::close(fd); + LockDescriptors().clear(); + g_fork_generation.fetch_add(1, std::memory_order_relaxed); + LockDescriptorsMutex().unlock(); +} + +int OpenLockDescriptor(const char* path, int flags, mode_t mode) { + static std::once_flag handlers; + std::call_once(handlers, [] { + ::pthread_atfork(&BeforeForkLockDescriptors, + &AfterForkLockDescriptorsInParent, + &AfterForkLockDescriptorsInChild); + }); + std::lock_guard guard(LockDescriptorsMutex()); + const int fd = ::open(path, flags, mode); + const int failure = errno; + if (fd >= 0) LockDescriptors().insert(fd); + errno = failure; + return fd; +} + +void CloseLockDescriptor(int fd) { + if (fd < 0) return; + std::lock_guard guard(LockDescriptorsMutex()); + LockDescriptors().erase(fd); + ::close(fd); // closing the last descriptor unlocks +} + +// The two locks of a held spool directory: its lock file's, which records +// the holder, and the directory's own. The directory cannot be unlinked +// while it holds anything, so a lock file removed behind a live owner's +// back -- by an age-based cleaner such as systemd-tmpfiles, which also +// skips a directory that is flocked, or by a person -- leaves the +// directory owned: the next take meets its lock and is refused, where by +// the new lock file alone it took the live directory for a dead one. +struct HeldLock { + int file_fd = -1; + int dir_fd = -1; +}; + +void CloseHeldLock(HeldLock* lock) { + CloseLockDescriptor(lock->dir_fd); + CloseLockDescriptor(lock->file_fd); + *lock = HeldLock{}; +} + +// Takes `dir`'s own lock beside its lock file's, which `lock` holds. kOwned +// while another holder has it: an owner whose lock file was replaced (or +// is being probed, which a retry outlasts). +SpoolStatus LockDirectory(const std::string& dir, HeldLock* lock, + std::string* error) { + lock->dir_fd = OpenLockDescriptor(dir.c_str(), + O_RDONLY | O_DIRECTORY | O_CLOEXEC, 0); + if (lock->dir_fd < 0) { + if (error) *error = "cannot open spool directory " + dir + ": " + + Errno(errno); + return SpoolStatus::kIo; + } + if (::flock(lock->dir_fd, LOCK_EX | LOCK_NB) == 0) return SpoolStatus::kOk; + const int failure = errno; + CloseLockDescriptor(lock->dir_fd); + lock->dir_fd = -1; + if (failure != EWOULDBLOCK) { + if (error) *error = "cannot lock spool directory " + dir + ": " + + Errno(failure); + return SpoolStatus::kIo; + } + if (error) { + *error = "spool directory " + dir + " is owned by another process, " + "which holds the directory's own lock: its " + kOwnerLockFile + + " was replaced since, so that file does not name it; a spool " + "directory has one owner process"; + } + return SpoolStatus::kOwned; +} + +// Locks an existing directory -- its lock file, creating the file if it +// has none, and the directory itself. Retries a lock lost to a remover's +// unlink, and a refusal as brief as another process's ReadSpoolOwner +// probe. +SpoolStatus LockInPlace(const std::string& dir, const char* role, + HeldLock* out, std::string* error) { + const std::string file = dir + "/" + kOwnerLockFile; + for (int attempt = 0; attempt < 8; ++attempt) { + HeldLock lock; + lock.file_fd = + OpenLockDescriptor(file.c_str(), O_RDWR | O_CREAT | O_CLOEXEC, 0644); + const int fd = lock.file_fd; + if (fd < 0) { + if (error) *error = "cannot open " + file + ": " + Errno(errno); + return SpoolStatus::kIo; + } + if (LockOpenHookForTesting()) LockOpenHookForTesting()(file); + if (::flock(fd, LOCK_EX | LOCK_NB) != 0) { + const int failure = errno; + SpoolOwner owner; + ReadOwnerRecord(fd, &owner); + CloseHeldLock(&lock); + if (failure != EWOULDBLOCK) { + if (error) *error = "cannot lock " + file + ": " + Errno(failure); + return SpoolStatus::kIo; + } + if (attempt < 2) { + std::this_thread::sleep_for(std::chrono::milliseconds(2)); + continue; + } + if (error) *error = OwnedMessage(dir, owner); + return SpoolStatus::kOwned; + } + if (!IsFileAt(fd, file)) { + CloseHeldLock(&lock); + if (!fs::is_directory(dir)) { + if (error) *error = "spool directory " + dir + " was removed while " + "it was being locked"; + return SpoolStatus::kIo; + } + continue; + } + const SpoolStatus directory = LockDirectory(dir, &lock, error); + if (directory != SpoolStatus::kOk) { + CloseHeldLock(&lock); + if (directory == SpoolStatus::kOwned && attempt < 2) { + std::this_thread::sleep_for(std::chrono::milliseconds(2)); + continue; + } + return directory; + } + WriteOwnerRecord(fd, role); + *out = lock; + return SpoolStatus::kOk; + } + if (error) *error = "cannot lock " + file + ": it keeps being replaced"; + return SpoolStatus::kIo; +} + +// Creates `dir` with its locks already held: built under a hidden name +// beside it, then renamed into place, so a scan of the parent never meets +// the directory unowned (an adopter would otherwise take a brand-new +// sibling for a dead one). Falls back to LockInPlace if `dir` appears +// meanwhile. +SpoolStatus CreateLocked(const std::string& dir, HeldLock* out, + std::string* error) { + const fs::path target(dir); + const std::string parent = target.parent_path().string(); + const std::string name = target.filename().string(); + std::random_device random; + for (int attempt = 0; attempt < 8; ++attempt) { + char suffix[16]; + std::snprintf(suffix, sizeof(suffix), "%08x", + static_cast(random())); + // IsSpoolClaimStagingName's pattern. + const std::string staging = + parent + "/." + name + "." + suffix + kClaimStagingSuffix; + if (::mkdir(staging.c_str(), 0755) != 0) { + if (errno == EEXIST) continue; + if (error) *error = "cannot create " + staging + ": " + Errno(errno); + return SpoolStatus::kIo; + } + const std::string file = staging + "/" + kOwnerLockFile; + HeldLock lock; + lock.file_fd = OpenLockDescriptor( + file.c_str(), O_RDWR | O_CREAT | O_EXCL | O_CLOEXEC, 0644); + const int fd = lock.file_fd; + if (fd < 0 || ::flock(fd, LOCK_EX | LOCK_NB) != 0) { + const int failure = errno; + CloseHeldLock(&lock); + ::unlink(file.c_str()); + ::rmdir(staging.c_str()); + if (error) *error = "cannot lock " + file + ": " + Errno(failure); + return SpoolStatus::kIo; + } + // Nobody else knows the staging copy yet, so its own lock is free. + if (LockDirectory(staging, &lock, error) != SpoolStatus::kOk) { + CloseHeldLock(&lock); + ::unlink(file.c_str()); + ::rmdir(staging.c_str()); + return SpoolStatus::kIo; + } + WriteOwnerRecord(fd); + ::fsync(fd); + FsyncDir(staging, nullptr); + // A rename that refuses an existing target: plain rename() silently + // replaces an empty directory. +#if defined(RENAME_NOREPLACE) + const int renamed = ::renameat2(AT_FDCWD, staging.c_str(), AT_FDCWD, + dir.c_str(), RENAME_NOREPLACE); +#elif defined(__APPLE__) && defined(RENAME_EXCL) + const int renamed = + ::renamex_np(staging.c_str(), dir.c_str(), RENAME_EXCL); +#else + // Neither: refuse a target that exists before renaming. The window + // left is between the check and the rename, and only another claim of + // the same fresh incarnation could fall into it. + int renamed = -1; + if (::access(dir.c_str(), F_OK) == 0) { + errno = EEXIST; + } else { + renamed = ::rename(staging.c_str(), dir.c_str()); + } +#endif + if (renamed != 0) { + const int failure = errno; + CloseHeldLock(&lock); + ::unlink(file.c_str()); + ::rmdir(staging.c_str()); + if (failure == EEXIST || failure == ENOTEMPTY) { + return LockInPlace(dir, "", out, error); + } + if (error) { + *error = "cannot create spool directory " + dir + ": " + + Errno(failure); + } + return SpoolStatus::kIo; + } + FsyncDir(parent, nullptr); + *out = lock; // the directory's lock went with the rename + return SpoolStatus::kOk; + } + if (error) *error = "cannot create spool directory " + dir; + return SpoolStatus::kIo; +} + +} // namespace + +void SpoolOwnerLock::Hold(int fd, int dir_fd, std::string dir) { + fd_ = fd; + dir_fd_ = dir_fd; + dir_ = std::move(dir); + generation_ = g_fork_generation.load(std::memory_order_relaxed); +} + +bool SpoolOwnerLock::held() const { + return fd_ >= 0 && + generation_ == g_fork_generation.load(std::memory_order_relaxed); +} + +SpoolOwnerLock::~SpoolOwnerLock() { Release(); } + +SpoolOwnerLock::SpoolOwnerLock(SpoolOwnerLock&& other) noexcept { + if (other.held()) { + fd_ = other.fd_; + dir_fd_ = other.dir_fd_; + dir_ = std::move(other.dir_); + generation_ = other.generation_; + } + other.fd_ = -1; + other.dir_fd_ = -1; + other.dir_.clear(); +} + +SpoolOwnerLock& SpoolOwnerLock::operator=(SpoolOwnerLock&& other) noexcept { + if (this != &other) { + Release(); + if (other.held()) { + fd_ = other.fd_; + dir_fd_ = other.dir_fd_; + dir_ = std::move(other.dir_); + generation_ = other.generation_; + } + other.fd_ = -1; + other.dir_fd_ = -1; + other.dir_.clear(); + } + return *this; +} + +void SpoolOwnerLock::Release() { + // Each untracked and closed in one step (CloseLockDescriptor). In a + // forked child the fork handler closed them already, and their numbers + // may be other files' by now: nothing to close. + if (held()) { + CloseLockDescriptor(dir_fd_); + CloseLockDescriptor(fd_); + } + fd_ = -1; + dir_fd_ = -1; + dir_.clear(); +} + +SpoolStatus SpoolOwnerLock::Acquire(const std::string& dir, + bool allow_shared_filesystem, + SpoolOwnerLock* out, std::string* error) { + out->Release(); + if (dir.empty()) { + if (error) *error = "spool directory must not be empty"; + return SpoolStatus::kBadArgument; + } + const std::string canonical = CanonicalPath(dir, error); + if (canonical.empty()) return SpoolStatus::kIo; + std::error_code ec; + const bool exists = fs::is_directory(canonical, ec); + if (!exists && fs::exists(canonical, ec)) { + if (error) *error = "spool directory " + canonical + " is not a directory"; + return SpoolStatus::kBadArgument; + } + const std::string parent = fs::path(canonical).parent_path().string(); + if (!exists) { + fs::create_directories(parent, ec); + if (ec) { + if (error) *error = "cannot create " + parent + ": " + ec.message(); + return SpoolStatus::kIo; + } + } + SpoolStatus status = + CheckNodeLocal(exists ? canonical : parent, allow_shared_filesystem, + error); + if (status != SpoolStatus::kOk) return status; + const std::string lock_file = canonical + "/" + kOwnerLockFile; + const bool had_lock_file = exists && fs::exists(lock_file, ec); + HeldLock lock; + status = exists ? LockInPlace(canonical, "", &lock, error) + : CreateLocked(canonical, &lock, error); + if (status != SpoolStatus::kOk) return status; + out->Hold(lock.file_fd, lock.dir_fd, canonical); + // Only now, with this lock published: see CheckNotNested. + status = CheckNotNested(canonical, error); + if (status != SpoolStatus::kOk) { + // Leave nothing of this take behind: the directory it created (while + // it is still empty), or the lock file it added to one that existed. + if (!exists) { + std::string ignored; + out->ReleaseAndRemoveIfEmpty(&ignored); + } else { + if (!had_lock_file) ::unlink(lock_file.c_str()); + out->Release(); + } + return status; + } + return SpoolStatus::kOk; +} + +SpoolStatus SpoolOwnerLock::TryAdopt(const std::string& dir, + SpoolOwnerLock* out, + std::string* error) { + out->Release(); + char resolved[4096]; + if (::realpath(dir.c_str(), resolved) == nullptr || + !fs::is_directory(resolved)) { + if (error) *error = "no spool directory to adopt at " + dir; + return SpoolStatus::kBadArgument; + } + HeldLock lock; + const SpoolStatus status = + LockInPlace(resolved, kAdoptingRole, &lock, error); + if (status != SpoolStatus::kOk) return status; + out->Hold(lock.file_fd, lock.dir_fd, resolved); + return SpoolStatus::kOk; +} + +bool SpoolOwnerLock::MarkBlocked(const std::string& reason) { + if (!held()) return false; + WriteOwnerRecord(fd_, kBlockedRole + (reason.empty() ? "blocked" : reason)); + ::fsync(fd_); + return true; +} + +bool SpoolOwnerLock::ReleaseAndRemoveIfEmpty(std::string* error) { + if (!held()) return false; + const std::string dir = dir_; + const fs::path lock_file = fs::path(dir) / kOwnerLockFile; + std::vector subdirectories; + std::error_code ec; + for (auto it = fs::recursive_directory_iterator(dir, ec); + !ec && it != fs::recursive_directory_iterator(); it.increment(ec)) { + std::error_code type_ec; + if (it->is_directory(type_ec) && !it->is_symlink(type_ec)) { + subdirectories.push_back(it->path()); + } else if (it->path() != lock_file) { + Release(); // something is left: the directory stays as it is + return false; + } + } + if (ec) { + if (error) *error = "cannot list " + dir + ": " + ec.message(); + Release(); + return false; + } + // Deepest first, so each is empty when its turn comes. + std::sort(subdirectories.begin(), subdirectories.end(), + [](const fs::path& a, const fs::path& b) { + return a.string().size() > b.string().size(); + }); + for (const fs::path& subdirectory : subdirectories) { + ::rmdir(subdirectory.c_str()); + } + // Unlinked while held: a process that opens the file from here on creates + // a new one (and the rmdir below then fails, leaving it the directory); one + // that opened the old file first finds, once it locks it, that the file + // is no longer at the path (IsFileAt), and lets it go. + ::unlink(lock_file.c_str()); + const bool removed = ::rmdir(dir.c_str()) == 0; + if (!removed && error) { + *error = "cannot remove " + dir + ": " + Errno(errno); + } + if (removed) FsyncParent(dir); + Release(); + return removed; +} + +std::string SpoolCatalogKey(const SpoolDestination& destination) { + const std::string text = + destination.database + "/" + destination.table_prefix + "/" + + destination.store_id + "\nclickhouse " + destination.clickhouse_host + + ":" + std::to_string(destination.clickhouse_port) + "\ns3 " + + destination.s3_endpoint + "/" + destination.s3_bucket; + unsigned char digest[SHA256_DIGEST_LENGTH]; + SHA256(reinterpret_cast(text.data()), text.size(), + digest); + static const char* kHex = "0123456789abcdef"; + std::string out; + for (int i = 0; i < 6; ++i) { + out.push_back(kHex[digest[i] >> 4]); + out.push_back(kHex[digest[i] & 0xF]); + } + return out; +} + +namespace { +bool IsLowerHex(const std::string& text, size_t size) { + return text.size() == size && + std::all_of(text.begin(), text.end(), [](char c) { + return (c >= '0' && c <= '9') || (c >= 'a' && c <= 'f'); + }); +} +} // namespace + +bool IsSpoolCatalogKey(const std::string& name) { return IsLowerHex(name, 12); } + +std::string SpoolRankDirectoryName(uint64_t producer_rank, + const std::string& incarnation) { + return "r" + std::to_string(producer_rank) + "-" + incarnation; +} + +bool ParseSpoolRankDirectoryName(const std::string& name, + uint64_t* producer_rank, + std::string* incarnation) { + if (name.size() < 4 || name[0] != 'r') return false; + const size_t dash = name.find('-'); + if (dash == std::string::npos || dash < 2) return false; + const std::string digits = name.substr(1, dash - 1); + if (digits.size() > 19 || (digits.size() > 1 && digits[0] == '0') || + !std::all_of(digits.begin(), digits.end(), + [](char c) { return c >= '0' && c <= '9'; })) { + return false; + } + const std::string tail = name.substr(dash + 1); + if (!IsLowerHex(tail, 8)) return false; + *producer_rank = std::strtoull(digits.c_str(), nullptr, 10); + *incarnation = tail; + return true; +} + +std::string NewSpoolIncarnation() { + std::random_device random; + char out[16]; + std::snprintf(out, sizeof(out), "%08x", static_cast(random())); + return out; +} + +std::string SpoolRankDirectory(const std::string& base, + const SpoolDestination& destination, + uint64_t producer_rank, + const std::string& incarnation) { + std::string root = base; + while (root.size() > 1 && root.back() == '/') root.pop_back(); + return root + "/" + SpoolCatalogKey(destination) + "/" + + SpoolRankDirectoryName(producer_rank, incarnation); +} + SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { if (config.max_bytes == 0) { if (error) *error = "max_bytes must be positive"; return SpoolStatus::kBadArgument; } + if (config.root.empty()) { + if (error) *error = "spool root must not be empty"; + return SpoolStatus::kBadArgument; + } + // A re-opened object gives up the directory it owned first. + out->owner_lock_.Release(); std::error_code ec; + if (config.owner_lock == OwnerLock::kTake) { + // Before anything reads the directory: the accounting walk below, and + // above all Recover(), belong to its one owner. Creates the root. + const SpoolStatus locked = SpoolOwnerLock::Acquire( + config.root, config.allow_shared_filesystem, &out->owner_lock_, + error); + if (locked != SpoolStatus::kOk) return locked; + } else { + // The caller took the lock, so the directory and its lock file exist. + char held[4096]; + SpoolOwner owner; + if (::realpath(config.root.c_str(), held) == nullptr || + !ReadSpoolOwner(held, &owner)) { + if (error) { + *error = "spool owner_lock=held_by_caller, but nothing holds " + + config.root + "/" + kOwnerLockFile + + ": take a SpoolOwnerLock on the directory first, or open " + "it with owner_lock=take"; + } + return SpoolStatus::kBadArgument; + } + // Held, but by THIS process? "Someone holds it" passes exactly when + // another live process owns the directory, and this Spool's Recover + // would then delete that owner's in-flight .open files. + if (!SpoolOwnedByThisProcess(held)) { + if (error) { + *error = "spool owner_lock=held_by_caller, but this process does " + "not hold the owner lock of " + std::string(held) + ": " + + OwnedMessage(held, owner) + ". held_by_caller is for a " + "second Spool in the process that holds the directory's " + "SpoolOwnerLock"; + } + return SpoolStatus::kOwned; + } + const SpoolStatus local = + CheckNodeLocal(held, config.allow_shared_filesystem, error); + if (local != SpoolStatus::kOk) return local; + } fs::create_directories(config.root, ec); if (ec) { if (error) *error = "cannot create spool root: " + ec.message(); @@ -260,6 +1304,8 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { } out->root_ = resolved; out->max_bytes_ = config.max_bytes; + out->charge_dead_siblings_ = config.charge_dead_siblings; + out->sibling_bytes_ = 0; out->committed_bytes_ = 0; out->committed_entries_ = 0; out->reserved_bytes_ = 0; @@ -272,9 +1318,13 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { // files plus stale .open files both count until Recover() runs, and only // the ready PATHS are remembered (_accounted_ready = ready_bytes), so a // later retry of one of them is recognised as already counted. - for (const auto& entry : - fs::recursive_directory_iterator(out->root_, ec)) { - if (ec) break; + for (auto it = fs::recursive_directory_iterator(out->root_, ec); + it != fs::recursive_directory_iterator(); ++it) { + if (AtSkippedDirectory(it)) { + it.disable_recursion_pending(); + continue; + } + const fs::directory_entry& entry = *it; if (!entry.is_regular_file()) continue; const std::string name = entry.path().filename().string(); const bool is_ready = HasSuffix(name, kReadySuffix); @@ -287,9 +1337,75 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { } } out->peak_bytes_ = out->committed_bytes_; + if (out->charge_dead_siblings_) { + out->sibling_bytes_ = out->ChargedSiblingBytes(); + } return SpoolStatus::kOk; } +namespace { +// Whether this process could take `dir`'s lock and empty it: write its lock +// file (or create one), and unlink in it. +bool CouldAdopt(const std::string& dir) { + const std::string file = dir + "/" + kOwnerLockFile; + if (::access(dir.c_str(), W_OK | X_OK) != 0) return false; + return ::access(file.c_str(), F_OK) != 0 || + ::access(file.c_str(), R_OK | W_OK) == 0; +} +} // namespace + +uint64_t Spool::ChargedSiblingBytes() const { + const fs::path own(root_); + uint64_t bytes = 0; + std::error_code ec; + for (fs::directory_iterator it(own.parent_path(), ec), end; + !ec && it != end; it.increment(ec)) { + uint64_t rank = 0; + std::string incarnation; + std::error_code type_ec; + if (it->path() == own || it->is_symlink(type_ec) || + !it->is_directory(type_ec) || + !ParseSpoolRankDirectoryName(it->path().filename().string(), &rank, + &incarnation)) { + continue; + } + const std::string sibling = it->path().string(); + // Only what adoption can drain: a dead directory, or one this process's + // adoption holds. + SpoolOwner owner; + if (ReadSpoolOwner(sibling, &owner)) { + // Another live process's directory is its own budget. One this + // process holds for its own writing -- an earlier engine's claim, + // kept owned while its unsealed sink may still stage -- no adoption + // here drains (its service reads it as live), until the process + // exits and the next one on the node adopts it. + if (!owner.adopting || !SpoolOwnedByThisProcess(sibling)) continue; + } else if (!owner.blocked.empty() || !CouldAdopt(sibling)) { + // Dead, but left for good by an adopter that could never drain it + // (its lock file says why), or not one this process could take and + // empty at all -- another user's, say. + continue; + } + std::error_code walk_ec; + for (fs::recursive_directory_iterator walk(sibling, walk_ec), last; + !walk_ec && walk != last; walk.increment(walk_ec)) { + if (AtSkippedDirectory(walk)) { + walk.disable_recursion_pending(); + continue; + } + std::error_code entry_ec; + if (!walk->is_regular_file(entry_ec)) continue; + const std::string name = walk->path().filename().string(); + if (!HasSuffix(name, kReadySuffix) && !HasSuffix(name, kOpenSuffix)) { + continue; + } + const uint64_t size = walk->file_size(entry_ec); + if (!entry_ec) bytes += size; + } + } + return bytes; +} + bool Spool::AccountReadyLocked(const std::string& path, uint64_t object_bytes) { if (accounted_ready_.count(path) != 0) return false; @@ -330,9 +1446,13 @@ void Spool::ReconcileCommittedLocked() { uint64_t ready_count = 0; std::unordered_map seen_ready; std::error_code walk_ec; - for (const auto& entry : - fs::recursive_directory_iterator(root_, walk_ec)) { - if (walk_ec) break; + for (auto it = fs::recursive_directory_iterator(root_, walk_ec); + it != fs::recursive_directory_iterator(); ++it) { + if (AtSkippedDirectory(it)) { + it.disable_recursion_pending(); + continue; + } + const fs::directory_entry& entry = *it; if (!entry.is_regular_file()) continue; const std::string name = entry.path().filename().string(); if (HasSuffix(name, kReadySuffix)) { @@ -352,11 +1472,28 @@ void Spool::ReconcileCommittedLocked() { // The path ledger is rebuilt with the aggregate it describes, so the two // never disagree about which files the committed account holds. accounted_ready_ = std::move(seen_ready); + // And the dead siblings' charge with it, so what adoption has drained + // since is capacity again. + if (charge_dead_siblings_) sibling_bytes_ = ChargedSiblingBytes(); // The scan can raise the committed total (files another object wrote), and // peak_bytes_ must never read below what the account holds right now. peak_bytes_ = std::max(peak_bytes_, committed_bytes_ + reserved_bytes_); } +std::string Spool::FullMessage(uint64_t n) const { + std::string message = + "spool byte limit exceeded: " + + std::to_string(committed_bytes_ + reserved_bytes_ + sibling_bytes_ + n) + + " > " + std::to_string(max_bytes_); + if (sibling_bytes_ > 0) { + message += " (" + std::to_string(sibling_bytes_) + + " bytes of it in dead spool directories beside this one, " + "which this process's storage service adopts: the room " + "comes back as it drains them)"; + } + return message; +} + void Spool::SetStageHookForTesting(std::function hook) { std::lock_guard lock(mutex_); stage_hook_for_testing_ = std::move(hook); @@ -435,14 +1572,12 @@ SpoolStatus Spool::Stage(const std::string& pack_id, uint64_t created_at_ns, return SpoolStatus::kConflict; } } - if (committed_bytes_ + reserved_bytes_ + n > max_bytes_) { + if (committed_bytes_ + reserved_bytes_ + sibling_bytes_ + n > + max_bytes_) { ReconcileCommittedLocked(); - if (committed_bytes_ + reserved_bytes_ + n > max_bytes_) { - if (error) { - *error = "spool byte limit exceeded: " + - std::to_string(committed_bytes_ + reserved_bytes_ + n) + - " > " + std::to_string(max_bytes_); - } + if (committed_bytes_ + reserved_bytes_ + sibling_bytes_ + n > + max_bytes_) { + if (error) *error = FullMessage(n); return SpoolStatus::kFull; } } @@ -505,12 +1640,9 @@ SpoolStatus Spool::Stage(const std::string& pack_id, uint64_t created_at_ns, // ready file -- a state Python cannot reach at all, since it holds its // lock across the whole of stage()), while a serial retry under a lowered // cap is admitted the way the reference admits it. - if (reserved_bytes_ > 0 && committed_bytes_ + reserved_bytes_ > max_bytes_) { - if (error) { - *error = "spool byte limit exceeded: " + - std::to_string(committed_bytes_ + reserved_bytes_) + " > " + - std::to_string(max_bytes_); - } + if (reserved_bytes_ > 0 && + committed_bytes_ + reserved_bytes_ + sibling_bytes_ > max_bytes_) { + if (error) *error = FullMessage(0); return SpoolStatus::kFull; } if (AccountReadyLocked(ready, n)) ++generation_; @@ -625,6 +1757,29 @@ SpoolStatus Spool::Recover(std::vector* out, std::string* error) { return Scan(out, true, error); } +SpoolStatus Spool::BeginRecovery(SpoolRecovery* recovery, std::string* error) { + *recovery = SpoolRecovery{}; + std::lock_guard lock(mutex_); + uint64_t open_bytes = 0; // stays 0: every .open file is swept + ListReadyLocked(true, &recovery->listed, &open_bytes); + (void)error; + return SpoolStatus::kOk; +} + +bool Spool::ContinueRecovery(SpoolRecovery* recovery) { + std::lock_guard lock(mutex_); + if (recovery->next < recovery->listed.size()) { + StagedPack staged; + if (ValidateReadyLocked(recovery->listed[recovery->next], &staged)) { + recovery->valid.push_back(std::move(staged)); + } + ++recovery->next; + } + if (recovery->next < recovery->listed.size()) return false; + CommitListingLocked(recovery->valid, 0); + return true; +} + SpoolStatus Spool::ListPending(std::vector* out, std::string* error) { return Scan(out, false, error); } @@ -640,13 +1795,36 @@ SpoolStatus Spool::Scan(std::vector* out, bool discard_open_files, out->clear(); if (cut) *cut = false; std::lock_guard lock(mutex_); - std::error_code ec; std::vector readies; - uint64_t bytes = 0; - std::unordered_map seen_ready; - for (const auto& entry : - fs::recursive_directory_iterator(root_, ec)) { - if (ec) break; + uint64_t open_bytes = 0; + ListReadyLocked(discard_open_files, &readies, &open_bytes); + for (const std::string& path : readies) { + if (cancel != nullptr && cancel->cancelled()) { + // Before the next pack's hash. The account below is rebuilt from a + // whole listing only; quarantines already made stand. + out->clear(); + if (cut) *cut = true; + return SpoolStatus::kOk; + } + StagedPack staged; + if (ValidateReadyLocked(path, &staged)) out->push_back(std::move(staged)); + } + CommitListingLocked(*out, open_bytes); + (void)error; + return SpoolStatus::kOk; +} + +void Spool::ListReadyLocked(bool discard_open_files, + std::vector* readies, + uint64_t* open_bytes) { + std::error_code ec; + for (auto it = fs::recursive_directory_iterator(root_, ec); + it != fs::recursive_directory_iterator(); ++it) { + if (AtSkippedDirectory(it)) { + it.disable_recursion_pending(); + continue; + } + const fs::directory_entry& entry = *it; if (!entry.is_regular_file()) continue; const std::string path = entry.path().string(); const std::string name = entry.path().filename().string(); @@ -658,65 +1836,65 @@ SpoolStatus Spool::Scan(std::vector* out, bool discard_open_files, FsyncDir(entry.path().parent_path().string(), nullptr); } else { const uint64_t size = entry.file_size(ec); - if (!ec) bytes += size; + if (!ec) *open_bytes += size; ec.clear(); // Another writer may have just committed its temp. } continue; } if (HasSuffix(name, kReadySuffix)) { - readies.push_back(path); + readies->push_back(path); } } - std::sort(readies.begin(), readies.end()); - for (const std::string& path : readies) { - if (cancel != nullptr && cancel->cancelled()) { - // Before the next pack's hash. The account below is rebuilt from a - // whole listing only; quarantines already made stand. - out->clear(); - if (cut) *cut = true; - return SpoolStatus::kOk; - } - const std::string name = fs::path(path).filename().string(); - std::string id, sum; - uint64_t created = 0, records = 0; - const uint64_t size = fs::file_size(path, ec); - if (!ParseReadyName(name, &id, &created, &records, &sum) || ec || - Sha256HexFile(path, nullptr) != sum) { - // Quarantine: keep the bytes, drop the .ready suffix. - const std::string target = path.substr(0, path.size() - 6) + - ".quarantined"; - ::rename(path.c_str(), target.c_str()); - FsyncDir(fs::path(path).parent_path().string(), nullptr); - ++generation_; - continue; - } - const std::string rel = fs::relative(path, root_, ec).string(); - const size_t slash = rel.rfind('/'); - const std::string parent = (slash == std::string::npos) ? "" : rel.substr(0, slash); - StagedPack staged; - staged.pack_id = id; - staged.created_at_ns = created; - staged.record_count = records; - staged.checksum = sum; - staged.object_key = (parent.empty() ? "" : parent + "/") + id + ".dmi-pack"; - staged.path = path; - staged.object_bytes = size; - out->push_back(std::move(staged)); - bytes += size; - seen_ready.emplace(path, size); + std::sort(readies->begin(), readies->end()); +} + +bool Spool::ValidateReadyLocked(const std::string& path, StagedPack* out) { + std::error_code ec; + const std::string name = fs::path(path).filename().string(); + std::string id, sum; + uint64_t created = 0, records = 0; + const uint64_t size = fs::file_size(path, ec); + if (!ParseReadyName(name, &id, &created, &records, &sum) || ec || + Sha256HexFile(path, nullptr) != sum) { + // Quarantine: keep the bytes, drop the .ready suffix. + const std::string target = path.substr(0, path.size() - 6) + + ".quarantined"; + ::rename(path.c_str(), target.c_str()); + FsyncDir(fs::path(path).parent_path().string(), nullptr); + ++generation_; + return false; } + const std::string rel = fs::relative(path, root_, ec).string(); + const size_t slash = rel.rfind('/'); + const std::string parent = (slash == std::string::npos) ? "" : rel.substr(0, slash); + out->pack_id = id; + out->created_at_ns = created; + out->record_count = records; + out->checksum = sum; + out->object_key = (parent.empty() ? "" : parent + "/") + id + ".dmi-pack"; + out->path = path; + out->object_bytes = size; + return true; +} + +void Spool::CommitListingLocked(const std::vector& valid, + uint64_t open_bytes) { // Recovery rebuilds the committed account only; a stage in flight on // another thread keeps its reservation. The path ledger is rebuilt with // it (_commit_recovery_locked does the same), so the surviving entries are // exactly the ones a later retry will recognise as already counted, and // the quarantined ones are simply absent. + uint64_t bytes = open_bytes; + std::unordered_map seen_ready; + for (const StagedPack& staged : valid) { + bytes += staged.object_bytes; + seen_ready.emplace(staged.path, staged.object_bytes); + } committed_bytes_ = bytes; - committed_entries_ = out->size(); + committed_entries_ = valid.size(); accounted_ready_ = std::move(seen_ready); peak_bytes_ = std::max(peak_bytes_, committed_bytes_ + reserved_bytes_); ++generation_; - (void)error; - return SpoolStatus::kOk; } SpoolStatus Spool::Remove(const StagedPack& staged, std::string* error) { @@ -792,6 +1970,7 @@ SpoolSnapshot Spool::Snapshot() const { snapshot.bytes = committed_bytes_ + reserved_bytes_; snapshot.peak_bytes = peak_bytes_; snapshot.max_bytes = max_bytes_; + snapshot.sibling_bytes = sibling_bytes_; return snapshot; } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 60599977f..31b4a10db 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -11,6 +11,50 @@ // the directory chain to the root. Idempotent: re-staging validates the // existing ready file (name, size, sha256) and returns it; a different pack // under the same pack_id is a conflict. +// +// One owner per directory (B6). Recover() deletes every .open file this +// object is not writing, so it is only safe while no other PROCESS writes +// there. A spool directory therefore has an owner lock -- flock(LOCK_EX) on +// /.owner.lock, whose content records the holder's host and pid, and +// on itself -- and a second process that tries to take it is +// refused, told who holds it. The lock goes with its holder, even one +// killed with SIGKILL. +// - owner_lock=kTake (the default) takes it in Open(), before anything +// reads the directory, and holds it for the Spool object's life. +// - owner_lock=kHeldByCaller takes none: the calling process holds a +// SpoolOwnerLock on the directory already. Two Spools in ONE process +// that both take refuse each other (flock binds to an open file +// description, not to the process), so a process running a sink and a +// storage service on one directory holds one SpoolOwnerLock and opens +// both Spools with kHeldByCaller. Open() refuses it unless one of THIS +// process's descriptors holds the lock (SpoolOwnedByThisProcess): kOwned, +// naming the holder, beside another process's lock, and kBadArgument +// when nothing holds it. Standalone callers -- the drivers, adoption -- +// take. +// A spool never walks into a subdirectory with a .owner.lock of its own -- +// another spool directory nested in it, live or dead -- to count, sweep, +// list or upload what it holds: a dead one's packs are for its successor to +// adopt, under its own keys. So a flat spool_root (the sink-only mode, an +// explicit record_sink) passes over the rank directories a crashed +// default-mode run left under it, and an adopter over a spool someone put +// inside the dead directory it drains. A directory nested under, or +// containing, a HELD spool directory is refused at the take, since the +// outer spool's walk could meet the inner one before its lock file is +// there; one nobody holds refuses nothing (every take leaves its file +// behind). The check runs after the lock is taken, so of two processes +// taking an outer and a nested directory at once, at least one is refused. +// /_refs/ is never scanned: the upload handoff's ref files live there +// (plan section 2.4). The spool must be node-local: NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, +// GPFS, 9p, AFS and OrangeFS are refused by statfs f_type unless +// allow_shared_filesystem is set, since none guarantees a flock that +// excludes a process on another node. +// +// The Python DurablePackSpool (spool.py) takes no lock, and its recover() +// deletes every .open file under its root; the C++ spool is deliberately +// stricter, and that is not ported to the reference. Nor is skipping +// /_refs/, or a nested spool directory: the reference counts, sweeps +// and quarantines .open and .ready files there like any others, which the +// C++ spool leaves alone -- C++-only divergences, deliberately not ported. #ifndef DMI_STORE_SPOOL_H_ #define DMI_STORE_SPOOL_H_ @@ -27,9 +71,41 @@ namespace dmi_store { class Cancellation; // cancel.h +enum class OwnerLock { + kTake = 0, // Open() takes /.owner.lock for the Spool's life + kHeldByCaller, // the caller holds a SpoolOwnerLock on +}; + +// "take" / "held_by_caller". +const char* OwnerLockName(OwnerLock mode); +bool ParseOwnerLock(const std::string& text, OwnerLock* mode); + struct SpoolConfig { std::string root; uint64_t max_bytes = 0; + OwnerLock owner_lock = OwnerLock::kTake; + // Admit a root on a shared filesystem (SharedFilesystemName), or on a + // FUSE filesystem that is local after all. Only safe when every process + // that could open the directory runs on this node. + bool allow_shared_filesystem = false; + // The root is a rank directory of the section 2.3 layout, and what its + // SIBLING rank directories hold (ready packs and temp files) counts + // against max_bytes as well, while an adoption by this process's storage + // service can drain it: a sibling nobody holds (a dead incarnation's, + // waiting to be adopted), and one this process holds through an adoption + // (SpoolOwner::adopting). Not charged: a sibling another live process + // holds, which is that process's own budget; one this process holds for + // its own writing -- an earlier engine's claim kept owned while its + // unsealed sink may still stage -- which no adoption here drains until + // the process exits; and one an adopter left blocked (SpoolOwner:: + // blocked), or that this process could not take and empty at all. Every + // process start gets a fresh rank directory, so without this each + // crash-restart while uploads are blocked would add a whole max_bytes to + // the node's spool; before the layout every restart reused one directory + // and one budget. The charge is refreshed wherever the committed account + // is (Open, and before a stage is refused), so the capacity comes back as + // adoption drains them. + bool charge_dead_siblings = false; }; struct StagedPack { @@ -42,11 +118,22 @@ struct StagedPack { uint64_t object_bytes = 0; }; +// A recovery in steps (Spool::BeginRecovery): the ready packs it listed, +// sorted, how many of them it has validated, and those that were valid. +struct SpoolRecovery { + std::vector listed; + size_t next = 0; + std::vector valid; +}; + struct SpoolSnapshot { uint64_t entries = 0; uint64_t bytes = 0; uint64_t peak_bytes = 0; uint64_t max_bytes = 0; + // charge_dead_siblings: what the sibling directories were charged, as of + // the last refresh; capacity is judged against bytes plus this. + uint64_t sibling_bytes = 0; }; enum class SpoolStatus { @@ -56,6 +143,7 @@ enum class SpoolStatus { kIntegrity, // ready file fails validation kIo, // filesystem error kBadArgument, + kOwned, // another owner holds the directory's owner lock }; inline const char* SpoolStatusName(SpoolStatus s) { @@ -66,14 +154,196 @@ inline const char* SpoolStatusName(SpoolStatus s) { case SpoolStatus::kIntegrity: return "ready pack failed validation"; case SpoolStatus::kIo: return "spool filesystem error"; case SpoolStatus::kBadArgument: return "invalid argument"; + case SpoolStatus::kOwned: return "spool directory is owned by another process"; } return "unknown"; } +// The statfs f_type names of the shared filesystems a spool refuses: "NFS", +// "Lustre", "BeeGFS", "CIFS", "SMB2", "FUSE", "GPFS", "9p", "AFS", +// "OrangeFS", or nullptr for any other. The list needs maintenance as +// deployments meet new ones. +const char* SharedFilesystemName(int64_t f_type); + +// Refuses `dir` (which must exist) on a shared filesystem unless +// `allow_shared_filesystem`. +SpoolStatus CheckNodeLocal(const std::string& dir, + bool allow_shared_filesystem, std::string* error); + +// Test seam: every node-local check in this binary reads `f_type` instead of +// calling statfs(2). A negative value restores statfs. +void SetFilesystemTypeForTesting(int64_t f_type); + +// Test seam: every /proc/self/fdinfo read in this binary sees no "lock:" +// lines, as under gVisor or WSL1, whose procfs lists none. +void SetFdinfoHidesLocksForTesting(bool hide); + +// Test seam: taking an existing directory's lock calls `hook` with the lock +// file's path after opening the file and before locking it -- the window in +// which a remover can unlink it. An empty function removes the hook. +void SetLockOpenHookForTesting(std::function hook); + +// The record in a directory's owner lock file: its last holder, which +// wrote it on taking the lock, and why it held the directory. Read with +// the lock free, it is the last holder's, whom the kernel has let go of. +struct SpoolOwner { + std::string host; + int64_t pid = 0; + // Taken by an adopter (SpoolOwnerLock::TryAdopt), to drain a dead + // directory, rather than by the process writing it. + bool adopting = false; + // Why an adopter left the directory for good (MarkBlocked), or empty. + std::string blocked; +}; + +// Whether 's owner lock is held right now (by any process, this one +// included) -- its lock file's, or the directory's own. Fills *owner with +// the lock file's record either way (empty without one). A holder that has +// locked but not yet written its record, or whose lock file was replaced, +// reads as an empty host and pid 0. +bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner); + +// Whether `name` is the staging copy of a directory SpoolOwnerLock::Acquire +// is creating, "..<8 hex>.creating": built beside its target with its +// lock file held, then renamed into place. One nobody holds was left by a +// claim killed before its rename; it holds nothing but its lock file, is +// ignored by the nesting check, and an adopter clears it. +bool IsSpoolClaimStagingName(const std::string& name); + +// Whether one of THIS process's descriptors holds 's owner lock (its +// lock file's, or the directory's own), as the kernel reports it in +// /proc/self/fdinfo (falling back to the recorded host and pid where /proc +// cannot be read, or its fdinfo lists no flocks at all, as under gVisor or +// WSL1). What kHeldByCaller requires. +bool SpoolOwnedByThisProcess(const std::string& dir); + +// The owner lock of one spool directory: flock(LOCK_EX) on /.owner.lock +// and on itself, released with the object (or Release()), and by the +// kernel when the process dies. The directory's lock keeps a live owner's +// directory owned should its lock file be removed from under it -- by an +// age-based cleaner (systemd-tmpfiles, which also leaves a flocked +// directory and everything below it alone) or by a person: the next take +// meets the directory's lock, not only a new lock file nobody holds. flock binds to an open file description, which a child +// shares after fork(): the descriptor is close-on-exec, and a child forked +// WITHOUT exec (a fork-started worker) closes its copy of every lock +// descriptor at once (pthread_atfork) -- held, or opened and not yet locked +// or not yet closed, since a fork from another thread can land anywhere in +// a take or a release -- so the lock never outlives its owner in a child. +// The owner's own hold is untouched, and the child's objects read as not +// held. A child the owner spawns through posix_spawn or vfork runs no +// atfork handler, and loses the descriptor at exec. +class SpoolOwnerLock { + public: + static constexpr const char* kFileName = ".owner.lock"; + + SpoolOwnerLock() = default; + ~SpoolOwnerLock(); + SpoolOwnerLock(SpoolOwnerLock&& other) noexcept; + SpoolOwnerLock& operator=(SpoolOwnerLock&& other) noexcept; + SpoolOwnerLock(const SpoolOwnerLock&) = delete; + SpoolOwnerLock& operator=(const SpoolOwnerLock&) = delete; + + // Takes the lock on `dir`, creating it when it does not exist -- beside + // its lock file, already held, and renamed into place, so no scan of the + // parent ever meets the directory before its owner holds it -- and records + // this host and pid in it. kOwned, naming the holder, when another holder + // has it; kBadArgument for a shared filesystem or a directory nested + // under, or containing, a held one (see above), after letting go of the + // lock and of whatever this call created. + static SpoolStatus Acquire(const std::string& dir, + bool allow_shared_filesystem, + SpoolOwnerLock* out, std::string* error); + + // Adoption's try-lock: takes the lock of an EXISTING directory whose owner + // is gone, creating its lock file if it has none, and records the take as + // an adoption (SpoolOwner::adopting). kOwned while its owner lives; never + // creates the directory. + static SpoolStatus TryAdopt(const std::string& dir, SpoolOwnerLock* out, + std::string* error); + + // Records in the lock file, while it is held, that the directory is left + // for good and why (SpoolOwner::blocked): an adopter that can never drain + // it lets go of it after this. The next take rewrites the record. Returns + // whether it was written. + bool MarkBlocked(const std::string& reason); + + bool held() const; + // The canonical path of the directory, while held. + const std::string& directory() const { return dir_; } + + void Release(); + + // Releases the lock, first removing the directory if it holds nothing but + // its lock file and empty subdirectories. Anything else -- a pack, a + // quarantined file, a ref -- keeps the directory, which the next owner or + // adopter meets as it was left. Returns whether the directory was removed. + bool ReleaseAndRemoveIfEmpty(std::string* error); + + private: + // Takes over the two descriptors, which the fork handler has tracked + // since their open(). + void Hold(int fd, int dir_fd, std::string dir); + + int fd_ = -1; // on /.owner.lock + int dir_fd_ = -1; // on itself + std::string dir_; + // The fork generation the lock was taken in (spool.cpp): in a forked + // child, whose copies of the descriptors the fork handler closed, the + // object reads as not held, and never closes a descriptor number the + // child may have reused since. + uint64_t generation_ = 0; +}; + +// Where a spool's packs go: the catalog -- its ClickHouse server, database +// and table prefix -- and the store -- its S3 endpoint, bucket and store id. +struct SpoolDestination { + std::string clickhouse_host; + uint64_t clickhouse_port = 0; + std::string database; + std::string table_prefix; + std::string s3_endpoint; + std::string s3_bucket; + std::string store_id; +}; + +// The spool layout of the plan's section 2.3. A capture process spools into +// //r-/ +// where catalog_key is the first 12 hex digits of the sha256 of +// "//\n" +// "clickhouse :\n" +// "s3 /" +// and incarnation is 8 hex digits fresh for every process start, so no two +// processes -- two jobs on one node, or a restart of the same rank -- ever +// share a directory. The directories under one catalog key are siblings: +// packs bound for one catalog and store, which a successor on the node +// adopts once their owner has died (CaptureStorageService, +// adopt_sibling_spools). The servers are in the key, not only the names: +// with the plan's sha256(database/table_prefix/store_id), two deployments +// sharing a spool_root and the default names but not a server adopted each +// other's dead directories into the wrong catalog and bucket. The servers +// are hashed as spelled, so spell them alike on every process of a +// deployment, or a dead directory waits for a process that does. +std::string SpoolCatalogKey(const SpoolDestination& destination); +bool IsSpoolCatalogKey(const std::string& name); +std::string SpoolRankDirectoryName(uint64_t producer_rank, + const std::string& incarnation); +// "r-<8 lowercase hex>", the rank in canonical decimal. +bool ParseSpoolRankDirectoryName(const std::string& name, + uint64_t* producer_rank, + std::string* incarnation); +std::string NewSpoolIncarnation(); +std::string SpoolRankDirectory(const std::string& base, + const SpoolDestination& destination, + uint64_t producer_rank, + const std::string& incarnation); + class Spool { public: - // Opens (creating) the root. Recovery of pre-existing files is explicit - // via Recover(), matching the Python constructor + recover() split. + // Opens (creating) the root, after the node-local check and, with kTake, + // after taking its owner lock (kOwned when another holder has it); + // kHeldByCaller is refused unless this process holds it. Recovery of + // pre-existing files is explicit via Recover(), matching the Python + // constructor + recover() split. static SpoolStatus Open(SpoolConfig config, Spool* out, std::string* error); Spool() = default; @@ -88,8 +358,23 @@ class Spool { // Startup cleanup: delete abandoned "*.open" files, validate ready packs, // quarantine failures, and rebuild accounting. Other writers on this root - // must be stopped; use ListPending() while they are running. + // must be stopped; use ListPending() while they are running. The owner + // lock keeps other processes out; writers in this process sharing it + // (kHeldByCaller) are the caller's to order. SpoolStatus Recover(std::vector* out, std::string* error); + // Recover() a pack at a time, for an adopter's recovery of a dead spool: + // Recover() hashes every byte of a dead backlog in one call, and whatever + // waits for the adopter -- a stop(), a flush() behind its cycle -- waits + // for all of it. BeginRecovery() is Recover()'s sweep of the .open files + // and its listing of the ready packs, and hashes none of them. Each + // ContinueRecovery() then validates the next pack listed, quarantining + // one that fails as Recover() does; once none is left it rebuilds the + // account as Recover() does and returns true, recovery->valid then the + // packs Recover() would have listed. The account is left as it was + // until then. Nothing else may validate or remove this spool's packs + // meanwhile. + SpoolStatus BeginRecovery(SpoolRecovery* recovery, std::string* error); + bool ContinueRecovery(SpoolRecovery* recovery); // Validate and list ready packs without deleting in-progress writes. SpoolStatus ListPending(std::vector* out, std::string* error); @@ -106,6 +391,9 @@ class Spool { SpoolSnapshot Snapshot() const; + // The canonical root, once opened. + const std::string& root() const { return root_; } + // Test seam: called by Stage() after its capacity reservation is taken and // before the temp file is written, outside the lock. Lets a test hold one // stager at exactly the point where its reservation exists but nothing is @@ -116,6 +404,18 @@ class Spool { SpoolStatus Scan(std::vector* out, bool discard_open_files, std::string* error, const Cancellation* cancel = nullptr, bool* cut = nullptr); + // A scan's parts. The walk: every ready path, sorted, into *readies; + // the .open files this object is not writing deleted (discard_open_files) + // or their bytes added to *open_bytes. The validation of one listed ready + // path: true with *out filled when its name parses and its size and + // sha256 match, and otherwise it is quarantined. And the account rebuilt + // from a whole listing's valid packs. `mutex_` must be held for each. + void ListReadyLocked(bool discard_open_files, + std::vector* readies, + uint64_t* open_bytes); + bool ValidateReadyLocked(const std::string& path, StagedPack* out); + void CommitListingLocked(const std::vector& valid, + uint64_t open_bytes); // Count/uncount one ready path in the committed account, at most once each // -- Python's _account_ready_locked / _unaccount_ready_locked. `mutex_` // must be held. Both return whether they actually changed the account. @@ -135,9 +435,19 @@ class Spool { // stale-high. The retry/EEXIST-loser paths run it too, so a charge that // would exceed the cap is judged against the same durable truth. void ReconcileCommittedLocked(); + // charge_dead_siblings: the bytes of ready and temp files in the sibling + // rank directories an adoption can drain (SpoolConfig). + uint64_t ChargedSiblingBytes() const; + // The kFull refusal of a stage of `n` bytes. `mutex_` must be held. + std::string FullMessage(uint64_t n) const; std::string root_; uint64_t max_bytes_ = 0; + bool charge_dead_siblings_ = false; + // What ChargedSiblingBytes() found last. Under `mutex_`. + uint64_t sibling_bytes_ = 0; + // Held for the object's life under OwnerLock::kTake; empty otherwise. + SpoolOwnerLock owner_lock_; mutable std::mutex mutex_; // Two accounts, kept apart on purpose: // committed_bytes_/committed_entries_ -- ready files this object knows diff --git a/native/csrc/store/uploader.cpp b/native/csrc/store/uploader.cpp index bdd756acf..e40b20082 100644 --- a/native/csrc/store/uploader.cpp +++ b/native/csrc/store/uploader.cpp @@ -103,7 +103,8 @@ SpoolUploader::SpoolUploader(Spool* spool, S3Client* client, bool SpoolUploader::UploadOne(const StagedPack& staged, PackRef* ref, int* attempts_out, std::string* error, - bool* cancelled_out) { + bool* cancelled_out, bool* retryable_out) { + if (retryable_out) *retryable_out = true; const std::string& key = staged.object_key; std::mt19937_64 rng( static_cast(std::hash{}(staged.pack_id))); @@ -190,6 +191,7 @@ bool SpoolUploader::UploadOne(const StagedPack& staged, PackRef* ref, // the retry-exhausted exit carries one. if (attempts_out) *attempts_out = attempts; if (error) *error = last_error; + if (retryable_out) *retryable_out = false; return false; // NOT retryable } } @@ -207,6 +209,7 @@ bool SpoolUploader::UploadOne(const StagedPack& staged, PackRef* ref, // Corrupt staged bytes: no retry can fix local corruption, but report // it as the failure rather than uploading garbage. last_error = "staged bytes do not match the staged checksum"; + if (retryable_out) *retryable_out = false; break; } std::string etag; @@ -313,7 +316,8 @@ UploadBatchResult SpoolUploader::UploadStaged(std::vector pending) { // test_mixed_batch_reports_oversized_pack_at_its_position), and // docs/benchmarks.md records the accounting decision behind it. result.failures[i] = {pending[i].pack_id, pending[i].object_key, 0, - "pack exceeds the in-flight byte limit"}; + "pack exceeds the in-flight byte limit", + /*cancelled=*/false, /*retryable=*/false}; ++result.snapshot.attempted_packs; ++result.snapshot.failed_packs; } @@ -376,7 +380,7 @@ UploadBatchResult SpoolUploader::UploadStaged(std::vector pending) { pending[left.index].object_key, 0, "upload cancelled before it started; the pack stays " "staged", - true}; + /*cancelled=*/true, /*retryable=*/true}; ++result.snapshot.cancelled_packs; finished[left.index] = true; } @@ -406,7 +410,9 @@ UploadBatchResult SpoolUploader::UploadStaged(std::vector pending) { std::string error; int attempts = 0; bool cancelled = false; - const bool ok = UploadOne(*staged, &ref, &attempts, &error, &cancelled); + bool retryable = true; + const bool ok = UploadOne(*staged, &ref, &attempts, &error, &cancelled, + &retryable); const int64_t elapsed = NowNs() - started; { std::lock_guard lock(mutex); @@ -428,7 +434,8 @@ UploadBatchResult SpoolUploader::UploadStaged(std::vector pending) { } else { result.refs[slot.index] = PackRef{}; result.failures[slot.index] = {staged->pack_id, staged->object_key, - attempts, error, cancelled}; + attempts, error, cancelled, + retryable}; if (cancelled) { ++result.snapshot.cancelled_packs; } else { diff --git a/native/csrc/store/uploader.h b/native/csrc/store/uploader.h index 6219ad050..843965e26 100644 --- a/native/csrc/store/uploader.h +++ b/native/csrc/store/uploader.h @@ -65,6 +65,11 @@ struct UploadFailure { // failed_packs. Attempts that ran out on failures of their own stay // failures, even when a cancel came in meanwhile. bool cancelled = false; + // False when no retry by this uploader can ever succeed: a pack over + // max_in_flight_bytes, a different object already at its key (a + // conflict), staged bytes that no longer match their checksum. The pack + // stays staged either way. A cancelled pack is retryable. + bool retryable = true; }; struct UploadSnapshot { @@ -116,13 +121,18 @@ class SpoolUploader { // Upload `pending`, packs a ListPending already returned, as UploadPending // does after its listing: in their order, refs and failures positional. - // A caller that uploads a listing in parts lists it once. + // A caller that uploads a listing in parts lists it once (the service's + // own spool a chunk at a time, and adoption, which lists a dead spool + // once, a pack a step through Spool::BeginRecovery, and uploads it a + // round at a time). UploadBatchResult UploadStaged(std::vector pending); // Upload one staged entry with retry. Public for tests. *cancelled_out - // (when given) says whether a cancel ended it. + // (when given) says whether a cancel ended it; *retryable_out says, on + // failure, whether a later try could succeed (UploadFailure). bool UploadOne(const StagedPack& staged, PackRef* ref, int* attempts_out, - std::string* error, bool* cancelled_out = nullptr); + std::string* error, bool* cancelled_out = nullptr, + bool* retryable_out = nullptr); private: Spool* spool_; diff --git a/src/dmi/engine.py b/src/dmi/engine.py index a155b3161..6c487e8bc 100644 --- a/src/dmi/engine.py +++ b/src/dmi/engine.py @@ -162,6 +162,9 @@ def __init__( "NativeCaptureStorageConfig") # The running storage service, while a record runtime is attached. self._capture_storage: Optional[Any] = None + # This process's own spool directory and its owner lock, held from + # before the service starts until the sink and the service are done. + self._spool_claim: Optional[Any] = None host_configured = host_engine is not None or db_config is not None if self._storage_backend == "in-memory" and not host_configured: raise ValueError( @@ -422,7 +425,7 @@ def create_record_runtime( # spool: its start sweeps a crashed sink's stale .open files, which # is safe only while nothing writes there. An explicit record_sink # may already hold the spool open, so it is left unswept. - storage = self._start_capture_storage(sweep_spool=record_sink is None) + storage = self._start_capture_storage(record_sink) try: runtime = self._attach_record_runtime( record_format, record_schema, record_sink, @@ -431,7 +434,10 @@ def create_record_runtime( except BaseException: if storage is not None: self._capture_storage = None - storage.stop() + try: + storage.stop() + finally: + self._release_spool_claim() raise return runtime @@ -486,23 +492,77 @@ def _refuse_an_unbounded_sink_admission( "NativePackSink (storage_backend='persistent', or record_sink=) " "to use a budget") - def _start_capture_storage(self, *, sweep_spool: bool) -> Optional[Any]: + def _start_capture_storage(self, record_sink: Optional[Any]) -> Optional[Any]: + """Start the storage service, and claim the spool it drains. + + With the default sink, this process spools into a directory of its + own under ``capture_sink_config.spool_root`` -- + ``//r-/``, fresh on every + start -- and owns it: its owner lock is taken here, before the + service opens it, and held until the sink and the service are done + (``_release_spool_claim``). The service and the sink both open it + ``held_by_caller``; two takes in one process refuse each other. The + service adopts the directories of dead processes beside it. + + An explicit ``record_sink`` writes where it was built to: the + service drains ``spool_root`` itself, unswept and with no siblings + to adopt, beside the sink's lock if this process holds one. It + passes over the rank directories under ``spool_root``, which a + default-mode start adopts. + """ config = self._capture_storage_config if config is None or self._storage_backend != "persistent": return None - from .storage.native_capture import NativeCaptureStorage + from .storage.native_capture import ( + NativeCaptureStorage, + claim_spool_directory, + spool_owner_lock_beside, + ) sink_config = self._capture_sink_config - storage = NativeCaptureStorage( - config, - spool_root=sink_config.spool_root, - spool_max_bytes=sink_config.spool_max_bytes, - sweep_spool=sweep_spool, - ) - storage.start() + shared = sink_config.spool_allow_shared_filesystem + if record_sink is not None: + storage = NativeCaptureStorage( + config, + spool_root=sink_config.spool_root, + spool_max_bytes=sink_config.spool_max_bytes, + sweep_spool=False, + spool_owner_lock=spool_owner_lock_beside( + sink_config.spool_root), + spool_allow_shared_filesystem=shared, + ) + storage.start() + self._capture_storage = storage + return storage + claim = claim_spool_directory(sink_config, config) + try: + storage = NativeCaptureStorage( + config, + spool_root=claim.directory, + spool_max_bytes=sink_config.spool_max_bytes, + sweep_spool=True, + spool_owner_lock="held_by_caller", + adopt_sibling_spools=True, + spool_allow_shared_filesystem=shared, + ) + storage.start() + except BaseException: + claim.release() + raise + self._spool_claim = claim self._capture_storage = storage return storage + def _release_spool_claim(self) -> None: + """Let go of this process's spool directory, once nothing writes it. + + A drained directory is removed; one still holding packs stays, for + the next process on the node for this catalog to adopt. + """ + claim, self._spool_claim = self._spool_claim, None + if claim is not None: + claim.release() + def _attach_record_runtime( self, record_format: "RecordFormat[MetadataT]", @@ -526,8 +586,16 @@ def _attach_record_runtime( ): from .storage.capture.native_sink import create_native_pack_sink + # Into the directory the engine claimed for the service, under + # its lock, with what dead incarnations left beside it charged + # against its budget; without a service the sink owns the spool + # root. + claim = self._spool_claim record_sink = create_native_pack_sink( - self._capture_sink_config + self._capture_sink_config, + spool_root=None if claim is None else claim.directory, + owner_lock="take" if claim is None else "held_by_caller", + charge_dead_siblings=claim is not None, ).native_sink _native_engine = _native_module() @@ -769,8 +837,9 @@ def enable_ring_transport( drain_deadline = None if storage is None else ( time.monotonic() + self._capture_storage_config.close_flush_timeout_s) - if storage is not None: - self._seal_capture_sink(drain_deadline) + old_record_sink = self._record_sink + sealed = storage is not None and self._seal_capture_sink( + drain_deadline) if old_record_mode: self._report_capture_failure() try: @@ -782,6 +851,9 @@ def enable_ring_transport( # A record sink remains leased while its worker may still # call it. Preserve the transport so shutdown can retry. raise + if storage is not None: + sealed = self._sink_sealed_after_release(old_record_sink, + sealed) try: _rt.deactivate() except Exception: @@ -791,7 +863,8 @@ def enable_ring_transport( self._record_mode = False self._record_sink = None if storage is not None: - self._retire_capture_storage(storage, drain_deadline) + self._retire_capture_storage(storage, drain_deadline, + sink_sealed=sealed) # Pass the DMXHostEngine C++ object directly; RingEngine builds a # SubmitFn that calls submit_direct without touching Python/GIL. @@ -831,8 +904,9 @@ def next_auto_group_id(self) -> int: self._auto_batch_group_id += 1 return gid - def _seal_capture_sink(self, deadline: float) -> None: - """Flush the record sink before its ring stops, within ``deadline``. + def _seal_capture_sink(self, deadline: float) -> bool: + """Flush the record sink before its ring stops, within ``deadline``; + whether it sealed. Releasing the sink from the stopping ring flushes it too (the native pack sink's release backstop), but only as a last resort: bounded @@ -845,8 +919,30 @@ def _seal_capture_sink(self, deadline: float) -> None: max(0.0, deadline - time.monotonic())) except Exception as exc: _LOG.warning("capture sink did not flush: %s", exc) + return False + return True + + @staticmethod + def _sink_sealed_after_release(sink: Any, sealed_before_stop: bool) -> bool: + """Whether nothing can still stage into the spool, the ring stopped. + + Stopping the ring released the sink, and the native pack sink's + release backstop flushed it then; ``sealed_on_release`` says whether + that went through, and a released sink admits nothing more. It + decides either way: it seals a sink whose flush before the stop ran + out of close()'s budget, and a sink it did not get through -- one + wedged since that flush, holding what the stopping ring drained into + it, or with the backstop off -- may still stage, whatever that flush + said. A sink that does not say is judged by the flush before the + stop. + """ + released = getattr(sink, "sealed_on_release", None) + if released is None: + return sealed_before_stop + return bool(released) - def _retire_capture_storage(self, storage: Any, deadline: float) -> None: + def _retire_capture_storage(self, storage: Any, deadline: float, *, + sink_sealed: bool) -> None: """Drain the storage service until ``deadline``, then stop it.""" if self._capture_storage is not storage: return @@ -861,7 +957,37 @@ def _retire_capture_storage(self, storage: Any, deadline: float) -> None: except Exception as exc: _LOG.warning("capture storage did not drain: %s", exc) finally: - storage.stop() + try: + storage.stop() + finally: + # Last: the ring is stopped by now, and the service has + # stopped touching the directory. The directory is let go of + # only if the sink can stage no more + # (_sink_sealed_after_release). + if sink_sealed: + self._release_spool_claim() + else: + self._keep_spool_claim_held() + + def _keep_spool_claim_held(self) -> None: + """Leave the spool directory owned by this process until it exits. + + A sink that did not seal may still be staging: it outlives the ring + (the user's RecordRuntime keeps it), its stagers carry on, and its + destructor stages what it still holds. Let go of now, the directory + is another process's to adopt (or this process's next engine's), + and that adoption would sweep a stage in flight; removed, it would + be recreated by the next stage with no lock at all. The claim stays + in ``native_capture``'s registry, so nothing lets go of it until the + kernel does, at exit; the next process on the node adopts it then. + """ + claim, self._spool_claim = self._spool_claim, None + if claim is not None: + _LOG.warning( + "capture sink did not seal, so its spool directory %s stays " + "owned by this process until it exits; the next process on " + "the node for this catalog adopts what it holds", + claim.directory) def close(self) -> None: """Tear down backend resources. @@ -877,7 +1003,9 @@ def close(self) -> None: ``NativeCaptureStorageConfig.close_flush_timeout_s`` says by how much. What does not drain in time is logged, not raised, and stays where the next start recovers it; ``flush_and_wait`` is the call - that raises. + that raises. The spool directory's owner lock goes last, and only + once the sink can stage no more -- its release backstop went + through; otherwise this process keeps the directory until it exits. """ storage = self._capture_storage @@ -885,8 +1013,12 @@ def close(self) -> None: # pack, then getting everything staged into the catalog. drain_deadline = None if storage is None else ( time.monotonic() + self._capture_storage_config.close_flush_timeout_s) + # Whether nothing can still stage into the spool: no record sink to + # seal, or one sealed by its flush or by its release from the ring. + sealed = True if self._ring_transport is not None: record_mode = self._record_mode + record_sink = self._record_sink stopped = False # Best-effort reset of the device-global native null flag. This is # needed only after callers explicitly disabled capture; the normal @@ -897,7 +1029,7 @@ def close(self) -> None: except Exception: pass if record_mode and storage is not None: - self._seal_capture_sink(drain_deadline) + sealed = self._seal_capture_sink(drain_deadline) if record_mode: self._report_capture_failure() try: @@ -911,6 +1043,8 @@ def close(self) -> None: # alive. Leave the state intact so close can be retried. if record_mode and not stopped: return + if record_mode and storage is not None: + sealed = self._sink_sealed_after_release(record_sink, sealed) try: _rt = _ring_module() _rt.deactivate() @@ -922,7 +1056,8 @@ def close(self) -> None: self._record_sink = None if storage is not None: - self._retire_capture_storage(storage, drain_deadline) + self._retire_capture_storage(storage, drain_deadline, + sink_sealed=sealed) if self._host_engine is not None: try: diff --git a/src/dmi/storage/capture/native_sink.py b/src/dmi/storage/capture/native_sink.py index bc34b942c..78d4d6889 100644 --- a/src/dmi/storage/capture/native_sink.py +++ b/src/dmi/storage/capture/native_sink.py @@ -86,12 +86,33 @@ def _load_native_sink_extension() -> Any: class NativePackSinkHandle: """Owns a native pack sink; mirrors CapturePackReferenceSink's surface - (``record_format`` + ``native_sink``) so call sites switch by factory.""" - - def __init__(self, config: NativeSinkConfig) -> None: + (``record_format`` + ``native_sink``) so call sites switch by factory. + + ``spool_root`` overrides ``config.spool_root`` (the engine passes its + own rank directory under it), and ``owner_lock`` is the spool + directory's owner-lock mode: ``"take"`` owns the directory for the + sink's life, ``"held_by_caller"`` when the caller holds a + ``SpoolOwnerLock`` on it -- as the engine does around its sink and its + storage service, which would otherwise refuse each other. + ``charge_dead_siblings`` says the spool root is a rank directory of the + spool layout, and what the dead incarnations beside it still hold + counts against ``spool_max_bytes`` (the engine's claimed directory). + """ + + def __init__( + self, + config: NativeSinkConfig, + *, + spool_root: str | None = None, + owner_lock: str = "take", + charge_dead_siblings: bool = False, + ) -> None: module = _load_native_sink_extension() self._native_sink = module.NativePackSink( - spool_root=config.spool_root, + spool_root=config.spool_root if spool_root is None else spool_root, + owner_lock=owner_lock, + allow_shared_filesystem=config.spool_allow_shared_filesystem, + charge_dead_siblings=charge_dead_siblings, layout=LAYOUT_NAME, num_workers=config.num_workers, max_queue_records=config.max_queue_records, @@ -120,12 +141,20 @@ def native_sink(self) -> Any: return self._native_sink -def create_native_pack_sink(config: NativeSinkConfig) -> NativePackSinkHandle: +def create_native_pack_sink( + config: NativeSinkConfig, + *, + spool_root: str | None = None, + owner_lock: str = "take", + charge_dead_siblings: bool = False, +) -> NativePackSinkHandle: """Select the native capture writer for one record runtime.""" if not isinstance(config, NativeSinkConfig): raise TypeError("config must be a NativeSinkConfig") - return NativePackSinkHandle(config) + return NativePackSinkHandle(config, spool_root=spool_root, + owner_lock=owner_lock, + charge_dead_siblings=charge_dead_siblings) __all__ = [ diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index 585015ace..d27bec6e7 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -21,11 +21,18 @@ ``table_prefix``), so run one capture process per catalog. A second engine on the same catalog waits ``start_lease_wait_s`` for the lease, then fails at ``create_record_runtime`` with the lease held, naming the holder. + +Each spool directory has one owner process (an flock on its +``.owner.lock``). The engine claims a fresh directory of its own under +``spool_root`` (:func:`claim_spool_directory`) and its service adopts the +spools of dead processes beside it, so a crashed run's packs reach the +catalog through the next process on the node for the same catalog. """ from __future__ import annotations import json +import logging import math import os import re @@ -124,6 +131,21 @@ class NativeSinkConfig: A record larger than ``max_queue_bytes`` or ``max_pack_bytes`` can never be admitted; ``validate_capture_bounds`` refuses such a bound at attach, before any forward runs. + + ``spool_root`` must be node-local, and each spool directory has one + owner process, held by an flock on its ``.owner.lock``. With + ``capture_storage_config`` the engine spools into a directory of its own + under it, ``//r-/`` (see + :func:`claim_spool_directory`), and its storage service adopts the + directories of dead processes beside it. ``spool_max_bytes`` then + bounds this directory together with what the dead incarnations beside + it still hold, so crash-restarts while uploads are blocked cannot each + add a whole budget. Without one, the sink owns ``spool_root`` itself. + A root on NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, + GPFS, 9p, AFS or OrangeFS is refused unless + ``spool_allow_shared_filesystem``: flock there does not keep out a + process on another node (a FUSE filesystem that is local, such as + fuse-overlayfs, needs the override too). """ spool_root: str @@ -136,10 +158,13 @@ class NativeSinkConfig: max_linger_ns: int = 1_000_000_000 overload: str = "block" admission_timeout_s: Optional[float] = 2.0 + spool_allow_shared_filesystem: bool = False def __post_init__(self) -> None: if not self.spool_root: raise ValueError("spool_root is required") + if type(self.spool_allow_shared_filesystem) is not bool: + raise TypeError("spool_allow_shared_filesystem must be bool") for name in ( "spool_max_bytes", "num_workers", @@ -288,7 +313,13 @@ class NativeCaptureStorageConfig: # budget cuts a multipart upload; against a slow catalog that still # answers, by one batch of statements and the release. When the sink # itself is stuck, the flush its release from the ring makes adds up to - # 30 s. close() logs what did not drain; flush_and_wait is what raises. + # 30 s. The budget also pays for the adoption step the background loop + # is in when the drain starts (adopt_sibling_spools, the engine's + # default): the drain's cycle waits for it -- one of a dead process's + # packs validated, which hashes it, or one round of their uploads -- + # but never for a dead backlog's whole listing, and past the budget + # stopping the service cuts it. close() logs what did not drain; + # flush_and_wait is what raises. close_flush_timeout_s: float = 60.0 # Bytes of packs the uploader holds in flight at once. A staged pack # larger than this is never uploaded, so the sink's max_pack_bytes must @@ -532,6 +563,20 @@ def _native_dict(self) -> dict[str, Any]: "uploader_max_in_flight_bytes": self.uploader_max_in_flight_bytes, } + def _spool_destination(self) -> dict[str, Any]: + """Where the packs go, as the spool's catalog key hashes it: the + catalog's server and names, and the store's endpoint, bucket and + id (native/csrc/store/spool.h, SpoolDestination).""" + return { + "clickhouse_host": self.clickhouse_host, + "clickhouse_port": self.clickhouse_port, + "database": self.database, + "table_prefix": self.table_prefix, + "s3_endpoint": self.s3_endpoint, + "s3_bucket": self.s3_bucket, + "store_id": self.store_id, + } + def _native_reader_dict(self) -> dict[str, Any]: """The reader's native config: the reader account, when one is set.""" native = self._native_dict() @@ -597,8 +642,160 @@ def validate_capture_bounds( "uploader_max_in_flight_bytes") +_LOG = logging.getLogger(__name__) + +# The spool directory's owner-lock modes (native/csrc/store/spool.h). +SPOOL_OWNER_LOCKS = ("take", "held_by_caller") + +# The spool layout's directory names (native/csrc/store/spool.h). +_CATALOG_KEY = re.compile(r"[0-9a-f]{12}") +_RANK_DIRECTORY = re.compile(r"r(?:0|[1-9][0-9]{0,18})-[0-9a-f]{8}") + + +def _packs_outside_the_layout(spool_root: str) -> tuple[int, Optional[str]]: + """Ready packs under ``spool_root`` that no rank directory of the layout + holds -- left by an engine from before the layout, or by a sink-only or + explicit-record_sink run, which write into ``spool_root`` itself -- and + the first one met. Only names are read, and the rank directories, which + adoption drains, are not walked.""" + count, example = 0, None + for directory, subdirectories, files in os.walk(spool_root): + relative = os.path.relpath(directory, spool_root) + depth = 0 if relative == "." else relative.count(os.sep) + 1 + if depth == 1 and _CATALOG_KEY.fullmatch(os.path.basename(directory)): + subdirectories[:] = [name for name in subdirectories + if not _RANK_DIRECTORY.fullmatch(name)] + for name in files: + if name.endswith(".dmi-pack.ready"): + count += 1 + if example is None: + example = os.path.join(directory, name) + return count, example + + +def _spool_producer_rank() -> int: + """The rank a spool directory is named for: torchrun's global ``RANK``, + 0 for a single process or anything that is not a rank. It only labels + the directory; the incarnation is what keeps two processes apart.""" + text = os.environ.get("RANK", "") + return int(text) if text.isdigit() else 0 + + +# Every SpoolClaim not yet released. Dropping a claim must not let go of its +# directory: an engine dropped without close() drops its claim while the +# ring and the sink it activated may still be capturing into the directory, +# and another process's adoption would then sweep it from under them. So a +# claim lives until release(), or until the process exits and the kernel +# drops its lock. +_HELD_SPOOL_CLAIMS: set["SpoolClaim"] = set() + + +class SpoolClaim: + """This process's own spool directory, owned through its lock. + + From :func:`claim_spool_directory`. Hold it for as long as anything in + the process writes or reads the directory -- the engine holds it from + before its storage service starts until the sink and the service are + done -- then :meth:`release` it. Only ``release()`` lets go: a claim + that is merely dropped stays held until the process exits. + """ + + def __init__(self, lock: Any) -> None: + self._lock = lock + self.directory: str = lock.directory + _HELD_SPOOL_CLAIMS.add(self) + + @property + def held(self) -> bool: + return bool(self._lock.held) + + def release(self) -> bool: + """Let go of the directory, removing it if nothing but its lock + file is left. Whatever did not drain stays, and the next process on + the node for this catalog adopts it. Returns whether it was + removed. Call it only once nothing in the process can still write + the directory.""" + _HELD_SPOOL_CLAIMS.discard(self) + return bool(self._lock.release_and_remove_if_empty()) + + +def claim_spool_directory( + sink_config: NativeSinkConfig, + storage_config: NativeCaptureStorageConfig, +) -> SpoolClaim: + """Create and lock this process's spool directory. + + ``//r-/``: the catalog key is + the first 12 hex digits of a sha256 of where the packs go -- the + ClickHouse host and port, ``database``, ``table_prefix``, the S3 endpoint + and bucket, and ``store_id`` -- so every directory under it holds packs + for this catalog and store, and two deployments that share the default + names but not a server never adopt each other's directories; the + incarnation is fresh for every call, so no two processes -- two jobs on + one node, or a restart -- share a directory; the rank is torchrun's + ``RANK`` (0 when unset). The directory is created with its owner lock + already held. Raises ``SpoolOwnedError`` (a ``RuntimeError``) if another + process holds it, and ``ValueError`` for a shared filesystem or a + directory nested in a spool that another owner holds (a sink-only or + explicit-``record_sink`` run on ``spool_root``). + + Ready packs under ``spool_root`` outside the layout -- an engine from + before it spooled into ``/v1/...``, and a sink-only or + explicit-``record_sink`` run still does -- are adopted by nothing, so a + claim logs a warning naming how many there are, one of them, and a + directory of the layout to move them into for adoption. + """ + module = _load_native_store_extension() + directory = module.spool_rank_directory( + sink_config.spool_root, storage_config._spool_destination(), + _spool_producer_rank()) + claim = SpoolClaim(module.SpoolOwnerLock( + directory, + allow_shared_filesystem=sink_config.spool_allow_shared_filesystem)) + count, example = _packs_outside_the_layout(sink_config.spool_root) + if count: + orphanage = os.path.join(sink_config.spool_root, + os.path.basename(os.path.dirname(directory)), + "r0-00000000") + _LOG.warning( + "spool_root %s holds %d ready pack(s) outside the per-process " + "layout (for example %s), which no engine adopts: left by an " + "engine from before the layout, or by a sink-only or " + "explicit-record_sink run. If they are bound for this catalog " + "and store, move them, keeping their paths below spool_root " + "(v1/...), into a directory of the layout nobody owns, such as " + "%s, and the next start on this node adopts them", + sink_config.spool_root, count, example, orphanage) + return claim + + +def spool_owner_lock_beside(spool_root: str) -> str: + """The owner-lock mode for a second Spool on a directory: ``held_by_caller`` + when this process already holds its lock (a sink the caller built took + it), else ``take``, which another process's lock refuses by name.""" + owner = _load_native_store_extension().spool_owner(spool_root) + if (owner is not None and owner["pid"] == os.getpid() + and owner["host"] == socket.gethostname()): + return "held_by_caller" + return "take" + + class NativeCaptureStorage: - """The in-process storage service: spool -> object store -> catalog.""" + """The in-process storage service: spool -> object store -> catalog. + + The spool directory has one owner process, held by an flock on + ``/.owner.lock``. ``spool_owner_lock="take"`` makes this + service its owner for the service's life, and refuses a directory + another process owns, naming it. A process that also runs the sink on + the directory -- the engine -- holds one ``SpoolOwnerLock`` and passes + ``"held_by_caller"`` here and to the sink: two takes in one process + refuse each other. + + ``adopt_sibling_spools`` needs ``spool_root`` to be a rank directory of + the spool layout (``spool_rank_directory``); once started, the service's + background loop drains the sibling directories whose owners have died + into this catalog, a slice per cycle. + """ def __init__( self, @@ -607,9 +804,22 @@ def __init__( spool_root: str, spool_max_bytes: int, sweep_spool: bool, + spool_owner_lock: str = "take", + adopt_sibling_spools: bool = False, + spool_allow_shared_filesystem: bool = False, ) -> None: if not isinstance(config, NativeCaptureStorageConfig): raise TypeError("config must be a NativeCaptureStorageConfig") + if spool_owner_lock not in SPOOL_OWNER_LOCKS: + raise ValueError( + f"spool_owner_lock must be one of {SPOOL_OWNER_LOCKS}, got " + f"{spool_owner_lock!r}") + for name, value in ( + ("adopt_sibling_spools", adopt_sibling_spools), + ("spool_allow_shared_filesystem", + spool_allow_shared_filesystem)): + if type(value) is not bool: + raise TypeError(f"{name} must be bool") module = _load_native_store_extension() native = config._native_dict() native.update( @@ -622,6 +832,9 @@ def __init__( reconcile_prefix=config.reconcile_prefix, reconcile_interval_ns=int(config.reconcile_interval_s * 1e9), sweep_spool_on_start=sweep_spool, + spool_owner_lock=spool_owner_lock, + adopt_sibling_spools=adopt_sibling_spools, + spool_allow_shared_filesystem=spool_allow_shared_filesystem, **config._lease_native(), ) self._config = config @@ -651,6 +864,16 @@ def flush(self, timeout_s: float) -> None: 6 s late when the deadline cuts a multipart upload, whose abort nothing cuts (see ``NativeCaptureStorageConfig.close_flush_timeout_s``). + The dead processes' spools the service adopts + (``adopt_sibling_spools``) are not part of it -- they are the + background loop's, and ``snapshot()`` reports them + (``adopted_spools``, ``adoption_owed``) -- though an adopted pack + uploaded and not yet indexed is waited for like this process's own. + Nor does a flush adopt, but it does wait, within ``timeout_s``, for + the adoption step the loop is in when it is called: one of a dead + spool's packs validated, which hashes it, or one round of their + uploads (at most four packs, one per upload worker) -- not a dead + backlog's whole listing, which goes a pack a step. """ if not self._service.flush(float(timeout_s)): snapshot = self._service.snapshot() @@ -862,6 +1085,7 @@ def read( __all__ = [ "PACK_FRAMING_RESERVE_BYTES", "SINK_OVERLOAD_POLICIES", + "SPOOL_OWNER_LOCKS", "NativeSinkConfig", "NativeCapture", "NativeCapturePage", @@ -869,5 +1093,8 @@ def read( "NativeCaptureSelection", "NativeCaptureStorage", "NativeCaptureStorageConfig", + "SpoolClaim", + "claim_spool_directory", + "spool_owner_lock_beside", "validate_capture_bounds", ] diff --git a/tests/native/live_spool_stage.cpp b/tests/native/live_spool_stage.cpp index 53047cc94..ab69b629b 100644 --- a/tests/native/live_spool_stage.cpp +++ b/tests/native/live_spool_stage.cpp @@ -1,15 +1,29 @@ -// Hold a real Stage after opening its temp file, while an uploader scans. +// Hold a real Stage after opening its temp file, while another process tries +// the same spool. +// +// live_spool_stage [packs] +// +// Opens with the default owner lock (take) and stages `packs` packs +// (default 1). When the LAST one has created its temp file it prints OPEN +// and waits for a line on stdin, so the test can act -- or SIGKILL it -- +// with a live writer and an in-flight .open file on disk. #include "store/spool.h" #include #include #include +#include #include #include #include #include #include +namespace { +int g_pause_at = 1; +int g_opened = 0; +} // namespace + extern "C" int __real_open(const char*, int, ...); extern "C" int __wrap_open(const char* path, int flags, ...) { mode_t mode = 0; @@ -20,7 +34,8 @@ extern "C" int __wrap_open(const char* path, int flags, ...) { va_end(args); } const int fd = __real_open(path, flags, mode); - if (fd >= 0 && (flags & O_CREAT) && std::strstr(path, ".open")) { + if (fd >= 0 && (flags & O_CREAT) && std::strstr(path, ".open") && + ++g_opened == g_pause_at) { std::cout << "OPEN" << std::endl; std::string release; std::getline(std::cin, release); @@ -29,23 +44,34 @@ extern "C" int __wrap_open(const char* path, int flags, ...) { } int main(int argc, char** argv) { - if (argc != 2) return 2; + if (argc != 2 && argc != 3) return 2; + const int packs = argc == 3 ? std::atoi(argv[2]) : 1; + if (packs < 1) return 2; + g_pause_at = packs; dmi_store::Spool spool; std::string error; - if (dmi_store::Spool::Open({argv[1], 5000}, &spool, &error) != - dmi_store::SpoolStatus::kOk) return 3; - const std::vector data(1000, 42); - unsigned char digest[SHA256_DIGEST_LENGTH]; - SHA256(data.data(), data.size(), digest); - char checksum[65]; - for (int i = 0; i < SHA256_DIGEST_LENGTH; ++i) { - std::snprintf(checksum + 2 * i, 3, "%02x", digest[i]); + if (dmi_store::Spool::Open({argv[1], 50000}, &spool, &error) != + dmi_store::SpoolStatus::kOk) { + std::cout << "open failed: " << error << std::endl; + return 3; + } + for (int n = 1; n <= packs; ++n) { + const std::vector data(1000, static_cast(41 + n)); + unsigned char digest[SHA256_DIGEST_LENGTH]; + SHA256(data.data(), data.size(), digest); + char checksum[65]; + for (int i = 0; i < SHA256_DIGEST_LENGTH; ++i) { + std::snprintf(checksum + 2 * i, 3, "%02x", digest[i]); + } + char id[40]; + std::snprintf(id, sizeof(id), "018f0000-0000-7000-8000-%012d", n); + dmi_store::StagedPack staged; + const auto status = spool.Stage( + id, 1700000000000000000ull + n, 1, checksum, + std::string(id) + ".dmi-pack", data.data(), data.size(), &staged, + &error); + std::cout << dmi_store::SpoolStatusName(status) << ": " << error << '\n'; + if (status != dmi_store::SpoolStatus::kOk) return 1; } - const std::string id = "018f0000-0000-7000-8000-000000000001"; - dmi_store::StagedPack staged; - const auto status = spool.Stage( - id, 1700000000000000000ull, 1, checksum, id + ".dmi-pack", - data.data(), data.size(), &staged, &error); - std::cout << dmi_store::SpoolStatusName(status) << ": " << error << '\n'; - return status == dmi_store::SpoolStatus::kOk ? 0 : 1; + return 0; } diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp new file mode 100644 index 000000000..cab47b6c8 --- /dev/null +++ b/tests/native/test_spool_owner_lock.cpp @@ -0,0 +1,1344 @@ +// B6: the spool owner lock, in process and across a fork. +// +// 1. Two Spool objects that both TAKE one directory refuse each other, +// even in one process: flock binds to an open file description, not to +// the process. That is why the engine holds one SpoolOwnerLock and both +// of its Spools (sink and service) open with held_by_caller. +// 2. held_by_caller opens beside a holder in this process, and is +// refused beside another process's holder or when nothing holds the +// lock. +// 3. A second process is refused, told the holder's pid and host; the +// lock goes with its holder, even one killed with SIGKILL, and a child +// it forked without exec does not keep it. +// 4. Nesting: a directory under, or containing, a HELD one is refused -- +// also when two processes take the pair at once -- and a dead one +// nested in a spool is left alone by all of its walks. +// 5. The node-local check refuses NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, +// GPFS, 9p, AFS and OrangeFS by statfs f_type, unless explicitly +// allowed (a test seam stands in for statfs). +// 6. Adoption's try-lock never creates a directory, and a released +// directory that holds nothing but its lock file can be removed; an +// adopter recovers a dead spool a pack at a time. +// 7. The directory layout of the plan's section 2.3. +// 8. The spool budget charges what dead sibling directories hold. +// +// Built and run by tests/test_native_spool_owner_lock_unit.py. + +#include +#include +#include +#include +#include + +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "store/spool.h" + +namespace fs = std::filesystem; +using dmi_store::OwnerLock; +using dmi_store::Spool; +using dmi_store::SpoolConfig; +using dmi_store::SpoolOwnerLock; +using dmi_store::SpoolStatus; + +namespace { + +int g_failures = 0; + +#define CHECK(cond) \ + do { \ + if (!(cond)) { \ + std::cerr << __FILE__ << ":" << __LINE__ << ": CHECK failed: " #cond \ + << "\n"; \ + ++g_failures; \ + } \ + } while (0) + +bool Contains(const std::string& text, const std::string& part) { + return text.find(part) != std::string::npos; +} + +std::string FreshRoot(const char* tag) { + const char* base = std::getenv("SPOOL_TEST_ROOT"); + const std::string root = + std::string(base != nullptr ? base : "/tmp") + "/owner-" + tag; + fs::remove_all(root); + fs::create_directories(root); + return fs::canonical(root).string(); +} + +std::string Hostname() { + char host[256] = {0}; + ::gethostname(host, sizeof(host) - 1); + return host; +} + +std::string Sha256Hex(const std::string& data) { + unsigned char digest[SHA256_DIGEST_LENGTH]; + SHA256(reinterpret_cast(data.data()), data.size(), + digest); + static const char* kHex = "0123456789abcdef"; + std::string out(64, '0'); + for (int i = 0; i < 32; ++i) { + out[2 * i] = kHex[digest[i] >> 4]; + out[2 * i + 1] = kHex[digest[i] & 0xF]; + } + return out; +} + +SpoolStatus StageOne(Spool& spool, int n, std::string* error) { + char id[40]; + std::snprintf(id, sizeof(id), "018f0000-0000-7000-8000-%012d", n); + const std::string data(100, static_cast('a' + n)); + dmi_store::StagedPack staged; + return spool.Stage(id, 1700000000000000000ull + n, 1, Sha256Hex(data), + std::string("v1/") + id + ".dmi-pack", + reinterpret_cast(data.data()), + data.size(), &staged, error); +} + +// (1) The regression a per-Spool lock would cause in the engine. +void TestTwoTakesInOneProcessRefuseEachOther() { + const std::string root = FreshRoot("two-takes") + "/spool"; + Spool service, sink; + std::string error; + CHECK(Spool::Open({root, 1 << 20}, &service, &error) == SpoolStatus::kOk); + error.clear(); + CHECK(Spool::Open({root, 1 << 20}, &sink, &error) == SpoolStatus::kOwned); + CHECK(Contains(error, "pid " + std::to_string(::getpid()))); + CHECK(Contains(error, Hostname())); +} + +// (2) held_by_caller beside a holder -- the engine's shape: one lock, two +// Spools, both staging and listing. +void TestHeldByCallerOpensBesideTheHolder() { + const std::string root = FreshRoot("held") + "/spool"; + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk); + CHECK(lock.held()); + SpoolConfig config{root, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool service, sink; + CHECK(Spool::Open(config, &service, &error) == SpoolStatus::kOk); + CHECK(Spool::Open(config, &sink, &error) == SpoolStatus::kOk); + CHECK(StageOne(sink, 1, &error) == SpoolStatus::kOk); + std::vector pending; + CHECK(service.ListPending(&pending, &error) == SpoolStatus::kOk); + CHECK(pending.size() == 1); + // A take is still refused while the engine's lock is held. + Spool rival; + CHECK(Spool::Open({root, 1 << 20}, &rival, &error) == SpoolStatus::kOwned); +} + +// (2b) held_by_caller is for a Spool in the process that holds the lock. +// Beside ANOTHER process's lock it is refused, naming that holder: were +// "something holds it" enough, any process could open a live writer's +// directory that way, and its Recover would delete the writer's .open file. +void TestHeldByCallerBesideAnotherProcessIsRefused() { + const std::string root = FreshRoot("held-elsewhere") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t child = ::fork(); + if (child == 0) { + ::close(ready[0]); + SpoolOwnerLock lock; + std::string error; + const bool ok = SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk; + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + + SpoolConfig config{root, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool spool; + std::string error; + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kOwned); + CHECK(Contains(error, "held_by_caller")); + CHECK(Contains(error, "pid " + std::to_string(child))); + CHECK(Contains(error, Hostname())); + + ::kill(child, SIGKILL); + int status = 0; + ::waitpid(child, &status, 0); +} + +// (2c) Where the kernel lists no flocks in /proc/self/fdinfo -- gVisor's +// procfs prints only pos/flags/mnt_id, WSL1's none either -- the check +// fell back to the owner record only when /proc/self/fd could not be +// opened, so the process that really held the lock was refused its own +// held_by_caller Spools, naming its own pid, and no default-mode capture +// could start. There the record decides: this host and pid, or not. +void TestHeldByCallerWhereTheKernelListsNoFlocks() { + const std::string root = FreshRoot("no-fdinfo-locks") + "/spool"; + dmi_store::SetFdinfoHidesLocksForTesting(true); + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk); + CHECK(dmi_store::SpoolOwnedByThisProcess(root)); + SpoolConfig config{root, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool spool; + error.clear(); + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kOk); + CHECK(error.empty()); + + // Beside another process's lock it is still refused, naming that holder. + const std::string other = FreshRoot("no-fdinfo-locks-other") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t child = ::fork(); + if (child == 0) { + ::close(ready[0]); + SpoolOwnerLock held; + std::string child_error; + const bool ok = SpoolOwnerLock::Acquire(other, false, &held, + &child_error) == SpoolStatus::kOk; + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + CHECK(!dmi_store::SpoolOwnedByThisProcess(other)); + SpoolConfig beside{other, 1 << 20}; + beside.owner_lock = OwnerLock::kHeldByCaller; + Spool refused; + error.clear(); + CHECK(Spool::Open(beside, &refused, &error) == SpoolStatus::kOwned); + CHECK(Contains(error, "pid " + std::to_string(child))); + ::kill(child, SIGKILL); + int status = 0; + ::waitpid(child, &status, 0); + dmi_store::SetFdinfoHidesLocksForTesting(false); +} + +void TestHeldByCallerWithoutAHolderIsRefused() { + const std::string root = FreshRoot("unheld") + "/spool"; + SpoolConfig config{root, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool spool; + std::string error; + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kBadArgument); + CHECK(Contains(error, "held_by_caller")); + // A lock file nobody holds is refused too. + { + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk); + } + error.clear(); + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kBadArgument); + CHECK(Contains(error, "held_by_caller")); +} + +void TestTheLockGoesWithItsSpool() { + const std::string root = FreshRoot("scope") + "/spool"; + std::string error; + { + Spool first; + CHECK(Spool::Open({root, 1 << 20}, &first, &error) == SpoolStatus::kOk); + } + Spool second; + CHECK(Spool::Open({root, 1 << 20}, &second, &error) == SpoolStatus::kOk); +} + +// (3) Across a fork: the child takes the lock, the parent is refused and +// told who holds it; the lock goes with the child, even on SIGKILL. +void TestASecondProcessIsRefusedUntilTheHolderDies() { + const std::string root = FreshRoot("fork") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t child = ::fork(); + if (child == 0) { + ::close(ready[0]); + SpoolOwnerLock lock; + std::string error; + const bool ok = SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk; + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOwned); + CHECK(Contains(error, "pid " + std::to_string(child))); + CHECK(Contains(error, Hostname())); + CHECK(Contains(error, root)); + dmi_store::SpoolOwner owner; + CHECK(dmi_store::ReadSpoolOwner(root, &owner)); + CHECK(owner.pid == child); + CHECK(owner.host == Hostname()); + Spool spool; + CHECK(Spool::Open({root, 1 << 20}, &spool, &error) == SpoolStatus::kOwned); + + ::kill(child, SIGKILL); + int status = 0; + ::waitpid(child, &status, 0); + CHECK(!dmi_store::ReadSpoolOwner(root, &owner)); + error.clear(); + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk); + CHECK(dmi_store::ReadSpoolOwner(root, &owner)); + CHECK(owner.pid == ::getpid()); +} + +// (3b) flock binds to an open file description, which a child forked +// without exec shares. Such a child -- a fork-started worker -- used to keep +// the lock after its parent was SIGKILLed, so the dead parent's directory +// read as owned (naming the dead pid) and was never adopted. A forked child +// now lets go of its copies of every held lock at once, and the parent's +// hold is untouched. +void TestAForkedChildDoesNotKeepTheLockPastItsParent() { + const std::string root = FreshRoot("fork-child") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t owner = ::fork(); + if (owner == 0) { + ::close(ready[0]); + SpoolOwnerLock lock; + std::string error; + if (SpoolOwnerLock::Acquire(root, false, &lock, &error) != + SpoolStatus::kOk) { + ::_exit(3); + } + const pid_t worker = ::fork(); // no exec + if (worker == 0) { + // The child's own view: it holds nothing, and the lock is still held + // (by the parent). + const char mine = lock.held() ? 'H' : 'h'; + const char held = dmi_store::ReadSpoolOwner(root, nullptr) ? 'P' : 'p'; + if (::write(ready[1], &mine, 1) != 1 || ::write(ready[1], &held, 1) != 1) { + ::_exit(3); + } + ::pause(); // outlives its parent until killed + ::_exit(0); + } + char pid_text[32]; + const int n = std::snprintf(pid_text, sizeof(pid_text), "%d\n", worker); + if (::write(ready[1], pid_text, n) != n) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + // The worker's two bytes and the owner's "\n", in any order. + std::string seen; + char byte = 0; + int flags = 0; + while (seen.find('\n') == std::string::npos || flags < 2) { + if (::read(ready[0], &byte, 1) != 1) break; + seen.push_back(byte); + if (byte == 'h' || byte == 'H' || byte == 'p' || byte == 'P') ++flags; + } + ::close(ready[0]); + std::string digits; + for (const char c : seen) { + if (c >= '0' && c <= '9') digits.push_back(c); + } + const pid_t worker = static_cast(std::atoi(digits.c_str())); + CHECK(worker > 0); + CHECK(seen.find('h') != std::string::npos); // the child holds nothing + CHECK(seen.find('P') != std::string::npos); // the parent still does + dmi_store::SpoolOwner owner_record; + CHECK(dmi_store::ReadSpoolOwner(root, &owner_record)); + CHECK(owner_record.pid == owner); + + ::kill(owner, SIGKILL); + int status = 0; + ::waitpid(owner, &status, 0); + // The worker lives on, and the lock went with its parent. + CHECK(::kill(worker, 0) == 0); + CHECK(!dmi_store::ReadSpoolOwner(root, &owner_record)); + SpoolOwnerLock successor; + std::string error; + CHECK(SpoolOwnerLock::TryAdopt(root, &successor, &error) == + SpoolStatus::kOk); + ::kill(worker, SIGKILL); +} + +// (3c) A fork from another thread while a lock is being TAKEN: the child +// gets a copy of the lock file's descriptor between its open() and the +// flock -- the flock then locks the description both share. Registered +// with the fork handler only once the lock was held, that copy was never +// closed in the child, which kept the directory looking live after its +// owner died. Here the fork runs in that window, through the lock-open +// test seam. The worker says it has run before its owner is killed: the +// fork handler closes its copy only once the child is first scheduled, +// which on a loaded host can come tens of milliseconds after the owner +// is reaped, and until then the lock outlives its owner. +void TestAForkWhileALockIsTakenLeavesTheChildNothing() { + const std::string root = FreshRoot("fork-window") + "/spool"; + fs::create_directories(root); + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t owner = ::fork(); + if (owner == 0) { + ::close(ready[0]); + pid_t worker = -1; + dmi_store::SetLockOpenHookForTesting([&](const std::string&) { + if (worker >= 0) return; + worker = ::fork(); // no exec + if (worker == 0) { + // Past fork(), so past the fork handler. + const char ran = 'w'; + if (::write(ready[1], &ran, 1) != 1) ::_exit(3); + ::pause(); // outlives its parent until killed + ::_exit(0); + } + }); + SpoolOwnerLock lock; + std::string error; + const bool ok = + SpoolOwnerLock::TryAdopt(root, &lock, &error) == SpoolStatus::kOk; + dmi_store::SetLockOpenHookForTesting(nullptr); + char text[32]; + const int n = std::snprintf(text, sizeof(text), "%c%d\n", ok ? 'k' : 'x', + static_cast(worker)); + if (::write(ready[1], text, n) != n) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + // The worker's 'w' and the owner's "k\n", in any order. + std::string seen; + char byte = 0; + while (seen.find('\n') == std::string::npos || + seen.find('w') == std::string::npos) { + if (::read(ready[0], &byte, 1) != 1) break; + seen.push_back(byte); + } + ::close(ready[0]); + CHECK(seen.find('w') != std::string::npos); // the worker has run + const size_t outcome = seen.find_first_of("kx"); + CHECK(outcome != std::string::npos && seen[outcome] == 'k'); + const pid_t worker = static_cast(std::atoi( + outcome == std::string::npos ? "" : seen.c_str() + outcome + 1)); + CHECK(worker > 0); + dmi_store::SpoolOwner record; + CHECK(dmi_store::ReadSpoolOwner(root, &record)); + CHECK(record.pid == owner); + + ::kill(owner, SIGKILL); + int status = 0; + ::waitpid(owner, &status, 0); + CHECK(worker > 0 && ::kill(worker, 0) == 0); // the worker lives on + // And the lock went with its owner. + CHECK(!dmi_store::ReadSpoolOwner(root, nullptr)); + if (worker > 0) ::kill(worker, SIGKILL); +} + +// (4) Nesting, both ways. +void TestNestedDirectoriesAreRefused() { + const std::string base = FreshRoot("nested"); + std::string error; + SpoolOwnerLock outer; + CHECK(SpoolOwnerLock::Acquire(base + "/outer", false, &outer, &error) == + SpoolStatus::kOk); + SpoolOwnerLock inner; + CHECK(SpoolOwnerLock::Acquire(base + "/outer/inner", false, &inner, + &error) == SpoolStatus::kBadArgument); + CHECK(Contains(error, "nested")); + CHECK(Contains(error, base + "/outer")); + + SpoolOwnerLock deep; + CHECK(SpoolOwnerLock::Acquire(base + "/other/a/b", false, &deep, &error) == + SpoolStatus::kOk); + SpoolOwnerLock ancestor; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(base + "/other", false, &ancestor, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "contains")); + CHECK(Contains(error, base + "/other/a/b")); + Spool spool; + CHECK(Spool::Open({base + "/other", 1 << 20}, &spool, &error) == + SpoolStatus::kBadArgument); + // A refused take leaves nothing of its own behind: not the directory it + // created, nor a lock file it added to one that existed. + CHECK(!fs::exists(base + "/outer/inner")); + CHECK(!fs::exists(base + "/other/.owner.lock")); +} + +// (4b) An outer directory and one nested in it, taken at the same moment by +// two processes: each checks the other's lock only after publishing its +// own, so at most one of them wins. Checked first and locked second, both +// won most of the time (275 of 300 in the review's probe). +void TestAnOuterAndANestedTakeRacingNeverBothWin() { + const std::string base = FreshRoot("nest-race"); + int both = 0; + int outer_won = 0; + int inner_won = 0; + for (int trial = 0; trial < 200; ++trial) { + const std::string outer = base + "/t" + std::to_string(trial); + const std::string inner = outer + "/0123456789ab/r0-0a1b2c3d"; + if (trial % 2 == 0) fs::create_directories(outer); + int go[2], report[2], done[2]; + CHECK(::pipe(go) == 0 && ::pipe(report) == 0 && ::pipe(done) == 0); + pid_t children[2]; + for (int side = 0; side < 2; ++side) { + children[side] = ::fork(); + if (children[side] == 0) { + ::close(go[1]); + ::close(report[0]); + ::close(done[1]); + char byte = 0; + (void)!::read(go[0], &byte, 1); // EOF: the parent let both go + SpoolOwnerLock lock; + std::string error; + const bool won = + SpoolOwnerLock::Acquire(side == 0 ? outer : inner, false, &lock, + &error) == SpoolStatus::kOk; + byte = static_cast(side == 0 ? (won ? 'O' : 'o') + : (won ? 'I' : 'i')); + if (::write(report[1], &byte, 1) != 1) ::_exit(3); + (void)!::read(done[0], &byte, 1); // hold it until both reported + ::_exit(0); + } + } + ::close(go[0]); + ::close(report[1]); + ::close(done[0]); + ::close(go[1]); + char results[2] = {0, 0}; + CHECK(::read(report[0], &results[0], 1) == 1); + CHECK(::read(report[0], &results[1], 1) == 1); + const std::string seen(results, 2); + const bool o = seen.find('O') != std::string::npos; + const bool i = seen.find('I') != std::string::npos; + if (o && i) ++both; + if (o) ++outer_won; + if (i) ++inner_won; + ::close(done[1]); + ::close(report[0]); + for (const pid_t child : children) { + int status = 0; + ::waitpid(child, &status, 0); + } + } + if (both != 0) { + std::cerr << "outer and nested both acquired in " << both + << " of 200 trials\n"; + } + CHECK(both == 0); + CHECK(outer_won + inner_won > 0); +} + +// (4c) Nesting refuses only while the other directory's lock is HELD. A +// lock file nobody holds is a spool that was -- every take leaves its file +// behind -- and a dead spool directory inside another is never the outer +// one's to sweep, count or upload under its own keys: every walk of a +// spool passes over a subdirectory with its own lock file. So a spool_root +// that a sink-only run once owned still takes rank directories, and a +// spool_root that a crashed default-mode run left a rank directory in +// still takes the sink-only or explicit-record_sink modes (the rollback), +// which leave that directory to adoption. +void TestAStaleLockFileAboveDoesNotRefuseANestedDirectory() { + const std::string base = FreshRoot("stale-above"); + const std::string root = base + "/root"; + const std::string rank = root + "/0123456789ab/r0-0a1b2c3d"; + std::string error; + { + SpoolOwnerLock once; + CHECK(SpoolOwnerLock::Acquire(root, false, &once, &error) == + SpoolStatus::kOk); + } + CHECK(fs::exists(root + "/.owner.lock")); + SpoolOwnerLock inner; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(rank, false, &inner, &error) == + SpoolStatus::kOk); + CHECK(error.empty()); + SpoolOwnerLock outer; + CHECK(SpoolOwnerLock::Acquire(root, false, &outer, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "contains")); + CHECK(Contains(error, rank)); + CHECK(Contains(error, "pid " + std::to_string(::getpid()))); + // The inner one stages a pack and has a stage in flight, then dies. + { + SpoolConfig config{rank, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool writer; + CHECK(Spool::Open(config, &writer, &error) == SpoolStatus::kOk); + CHECK(StageOne(writer, 1, &error) == SpoolStatus::kOk); + } + const std::string in_flight = + rank + "/v1/.018f0000-0000-7000-8000-00000000beef.0badf00d.open"; + std::ofstream(in_flight) << "half a pack"; + inner.Release(); + // Its lock file stays, and nobody holds it: the outer take goes through, + // and nothing of the outer spool touches the dead one. + Spool flat; + error.clear(); + CHECK(Spool::Open({root, 150}, &flat, &error) == SpoolStatus::kOk); + CHECK(flat.Snapshot().bytes == 0); + std::vector staged; + CHECK(flat.Recover(&staged, &error) == SpoolStatus::kOk); + CHECK(staged.empty()); + CHECK(fs::exists(in_flight)); + CHECK(StageOne(flat, 2, &error) == SpoolStatus::kOk); // 100 of 150 + CHECK(flat.ListPending(&staged, &error) == SpoolStatus::kOk); + CHECK(staged.size() == 1 && staged[0].object_key.rfind("v1/", 0) == 0); + size_t dead_packs = 0; + for (const auto& entry : fs::recursive_directory_iterator(rank)) { + if (entry.path().extension() == ".ready") ++dead_packs; + } + CHECK(dead_packs == 1); + // And while the outer one holds the root, the rank directory cannot be + // taken: an adopter would first have to wait for it. + SpoolOwnerLock again; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(rank, false, &again, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "nested")); +} + +// (4e) An adopter takes a dead directory's lock with TryAdopt, which runs +// no nesting check, then Recovers it. A live spool nested inside that dead +// directory -- a root put there, which the nesting rule admits under an +// unheld lock -- had its in-flight .open files swept by that Recover, and +// its packs listed under the outer directory's keys. Every walk passes +// over a subdirectory with its own lock file, so the adoption drains only +// the dead directory's own packs and leaves the directory in place. +void TestAnAdopterLeavesASpoolNestedInADeadOneAlone() { + const std::string base = FreshRoot("nested-in-dead"); + const std::string dead = base + "/0123456789ab/r0-0000dead"; + const std::string nested = dead + "/inner"; + std::string error; + { + Spool gone; + CHECK(Spool::Open({dead, 1 << 20}, &gone, &error) == SpoolStatus::kOk); + CHECK(StageOne(gone, 1, &error) == SpoolStatus::kOk); + } // its owner died + SpoolOwnerLock live; + CHECK(SpoolOwnerLock::Acquire(nested, false, &live, &error) == + SpoolStatus::kOk); + { + SpoolConfig config{nested, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool writer; + CHECK(Spool::Open(config, &writer, &error) == SpoolStatus::kOk); + CHECK(StageOne(writer, 2, &error) == SpoolStatus::kOk); + } + const std::string in_flight = + nested + "/v1/.018f0000-0000-7000-8000-00000000beef.0badf00d.open"; + std::ofstream(in_flight) << "half a pack"; + + SpoolOwnerLock adopter; + CHECK(SpoolOwnerLock::TryAdopt(dead, &adopter, &error) == SpoolStatus::kOk); + SpoolConfig config{dead, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool adopted; + CHECK(Spool::Open(config, &adopted, &error) == SpoolStatus::kOk); + CHECK(adopted.Snapshot().bytes == 100); // its own pack only + std::vector staged; + CHECK(adopted.Recover(&staged, &error) == SpoolStatus::kOk); + CHECK(fs::exists(in_flight)); + CHECK(staged.size() == 1); + if (staged.size() == 1) { + CHECK(staged[0].object_key.rfind("v1/", 0) == 0); + CHECK(adopted.Remove(staged[0], &error) == SpoolStatus::kOk); + } + // Drained of its own, it still holds the live one: it stays. + CHECK(!adopter.ReleaseAndRemoveIfEmpty(&error)); + CHECK(fs::exists(in_flight)); + CHECK(live.held()); +} + +// (4d) A claim killed between creating its directory's staging copy +// (...creating, lock file inside) and renaming it into place +// leaves that copy behind. Nobody holds it, it holds no pack, and it must +// not refuse a take of the directories above it for good. A staging copy +// whose lock IS held is a claim in progress, and still refuses. +void TestAnUnheldClaimStagingDirectoryRefusesNothing() { + const std::string base = FreshRoot("staging"); + const std::string key = base + "/root/0123456789ab"; + const std::string leftover = key + "/.r0-0a1b2c3d.0badf00d.creating"; + CHECK(dmi_store::IsSpoolClaimStagingName(".r0-0a1b2c3d.0badf00d.creating")); + CHECK(!dmi_store::IsSpoolClaimStagingName("r0-0a1b2c3d")); + CHECK(!dmi_store::IsSpoolClaimStagingName(".creating")); + fs::create_directories(leftover); + std::ofstream(leftover + "/.owner.lock") << "host 1\n"; + std::string error; + { + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(base + "/root", false, &lock, &error) == + SpoolStatus::kOk); + CHECK(error.empty()); + } + { + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(key, false, &lock, &error) == + SpoolStatus::kOk); + CHECK(error.empty()); + } + // Held: a claim in the middle of creating its directory. + const std::string live = base + "/other/0123456789ab"; + SpoolOwnerLock claiming; + CHECK(SpoolOwnerLock::Acquire(live + "/.r1-0a1b2c3d.00c0ffee.creating", + false, &claiming, &error) == + SpoolStatus::kOk); + SpoolOwnerLock outer; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(base + "/other", false, &outer, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "contains")); +} + +// (5) The node-local check, through the test seam. +void TestSharedFilesystemsAreRefusedUnlessAllowed() { + const std::string root = FreshRoot("statfs") + "/spool"; + // Each one's flock does not keep out a process on another node: NFS and + // Lustre (the plan's two), BeeGFS (client-local unless + // tuneUseGlobalFileLocks), CIFS/SMB2, FUSE, which cannot tell sshfs, + // s3fs, gcsfuse or GlusterFS from a local filesystem, GPFS (IBM Storage + // Scale, whose flock is node-local), 9p, AFS (OpenAFS and kAFS) and + // OrangeFS -- network filesystems all, where a spool is never node-local. + const std::vector> shared = { + {0x6969, "NFS"}, {0x0BD00BD0, "Lustre"}, + {0x19830326, "BeeGFS"}, {0xFF534D42, "CIFS"}, + {0xFE534D42, "SMB2"}, {0x65735546, "FUSE"}, + {0x47504653, "GPFS"}, {0x01021997, "9p"}, + {0x5346414F, "AFS"}, {0x6B414653, "AFS"}, + {0x20030528, "OrangeFS"}}; + for (const auto& [magic, name] : shared) { + const char* named = dmi_store::SharedFilesystemName(magic); + CHECK(named != nullptr && std::string(named) == name); + } + CHECK(dmi_store::SharedFilesystemName(0xEF53) == nullptr); // ext4 + CHECK(dmi_store::SharedFilesystemName(0x58465342) == nullptr); // xfs + CHECK(dmi_store::SharedFilesystemName(0x794C7630) == nullptr); // overlayfs + CHECK(dmi_store::SharedFilesystemName(0x01021994) == nullptr); // tmpfs + CHECK(dmi_store::SharedFilesystemName(0x9123683E) == nullptr); // btrfs + CHECK(dmi_store::SharedFilesystemName(0x2FC12FC1) == nullptr); // zfs + + std::string error; + for (const auto& [magic, name] : shared) { + dmi_store::SetFilesystemTypeForTesting(magic); + SpoolOwnerLock lock; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, " is on " + name + " ")); + CHECK(Contains(error, "node-local")); + CHECK(!lock.held()); + Spool spool; + error.clear(); + CHECK(Spool::Open({root, 1 << 20}, &spool, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "node-local")); + // The explicit override. + SpoolConfig allowed{root, 1 << 20}; + allowed.allow_shared_filesystem = true; + CHECK(Spool::Open(allowed, &spool, &error) == SpoolStatus::kOk); + } + { + dmi_store::SetFilesystemTypeForTesting(0x6969); + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(root + "-allowed", true, &lock, &error) == + SpoolStatus::kOk); + } + dmi_store::SetFilesystemTypeForTesting(-1); + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(root + "-real", false, &lock, &error) == + SpoolStatus::kOk); +} + +// (6) Adoption's try-lock, and removing a drained directory. +void TestAdoptionLocksOnlyWhatExistsAndIsDead() { + const std::string base = FreshRoot("adopt"); + std::string error; + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::TryAdopt(base + "/missing", &lock, &error) != + SpoolStatus::kOk); + CHECK(!fs::exists(base + "/missing")); + + // A directory with no lock file at all: nobody owns it. + fs::create_directories(base + "/bare/v1"); + CHECK(SpoolOwnerLock::TryAdopt(base + "/bare", &lock, &error) == + SpoolStatus::kOk); + CHECK(lock.held()); + CHECK(lock.ReleaseAndRemoveIfEmpty(&error)); + CHECK(!lock.held()); + CHECK(!fs::exists(base + "/bare")); + + // A live owner is left alone. + SpoolOwnerLock live; + CHECK(SpoolOwnerLock::Acquire(base + "/live", false, &live, &error) == + SpoolStatus::kOk); + SpoolOwnerLock adopter; + CHECK(SpoolOwnerLock::TryAdopt(base + "/live", &adopter, &error) == + SpoolStatus::kOwned); + CHECK(!adopter.held()); + + // Anything but the lock file and empty directories keeps the directory. + SpoolOwnerLock kept; + CHECK(SpoolOwnerLock::Acquire(base + "/kept", false, &kept, &error) == + SpoolStatus::kOk); + fs::create_directories(base + "/kept/v1/tenant=t"); + std::ofstream(base + "/kept/v1/tenant=t/x.quarantined") << "bytes"; + CHECK(!kept.ReleaseAndRemoveIfEmpty(&error)); + CHECK(!kept.held()); + CHECK(fs::exists(base + "/kept/v1/tenant=t/x.quarantined")); + CHECK(fs::exists(base + "/kept/.owner.lock")); +} + +// (6b) A lock taken on a file a remover unlinked meanwhile guards nothing +// (ReleaseAndRemoveIfEmpty unlinks the lock file, then removes the +// directory), so it is let go and the file at the path is locked instead. +void TestALockOnAnUnlinkedFileIsTakenAgain() { + const std::string dir = FreshRoot("unlinked") + "/spool"; + fs::create_directories(dir); + std::ofstream(dir + "/.owner.lock") << ""; + int calls = 0; + dmi_store::SetLockOpenHookForTesting([&calls](const std::string& file) { + if (calls++ == 0) ::unlink(file.c_str()); // the remover's unlink + }); + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::TryAdopt(dir, &lock, &error) == SpoolStatus::kOk); + dmi_store::SetLockOpenHookForTesting(nullptr); + CHECK(calls == 2); + CHECK(lock.held()); + // What is held is the lock file at the path, which anyone else meets. + CHECK(dmi_store::ReadSpoolOwner(dir, nullptr)); + SpoolOwnerLock rival; + CHECK(SpoolOwnerLock::TryAdopt(dir, &rival, &error) == SpoolStatus::kOwned); +} + +// (6c) Liveness is not the lock file's alone. Something other than its +// holder -- an age-based cleaner such as systemd-tmpfiles, a person -- can +// unlink /.owner.lock while the owner lives: the owner's descriptor is +// then on the unlinked file, and whoever opens the path next meets a new +// one that nobody holds. Judged by that file alone the live directory read +// as dead, an adopter took it, and its Recover swept the owner's in-flight +// .open files. The owner also locks the directory itself, which cannot be +// unlinked while it holds anything. +void TestAReplacedLockFileLeavesTheDirectoryOwned() { + const std::string dir = FreshRoot("replaced") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t owner = ::fork(); + if (owner == 0) { + ::close(ready[0]); + SpoolOwnerLock lock; + std::string error; + const bool ok = SpoolOwnerLock::Acquire(dir, false, &lock, &error) == + SpoolStatus::kOk; + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + + CHECK(fs::remove(dir + "/.owner.lock")); + CHECK(dmi_store::ReadSpoolOwner(dir, nullptr)); // still owned + std::string error; + SpoolOwnerLock adopter; + CHECK(SpoolOwnerLock::TryAdopt(dir, &adopter, &error) == + SpoolStatus::kOwned); + CHECK(!adopter.held()); + Spool spool; + CHECK(Spool::Open({dir, 1 << 20}, &spool, &error) == SpoolStatus::kOwned); + + ::kill(owner, SIGKILL); + int status = 0; + ::waitpid(owner, &status, 0); + CHECK(!dmi_store::ReadSpoolOwner(dir, nullptr)); + CHECK(SpoolOwnerLock::TryAdopt(dir, &adopter, &error) == SpoolStatus::kOk); + + // In the owner's own process too: its held_by_caller Spools still open. + adopter.Release(); + SpoolOwnerLock mine; + CHECK(SpoolOwnerLock::Acquire(dir, false, &mine, &error) == + SpoolStatus::kOk); + CHECK(fs::remove(dir + "/.owner.lock")); + CHECK(dmi_store::SpoolOwnedByThisProcess(dir)); + SpoolConfig beside{dir, 1 << 20}; + beside.owner_lock = OwnerLock::kHeldByCaller; + Spool held; + CHECK(Spool::Open(beside, &held, &error) == SpoolStatus::kOk); +} + +void TestANewDirectoryAppearsWithItsLockHeld() { + // Created beside its lock file and renamed into place, so no scan of the + // parent can meet the directory before its owner holds it. + const std::string base = FreshRoot("atomic"); + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::Acquire(base + "/r0-0a1b2c3d", false, &lock, + &error) == SpoolStatus::kOk); + std::set names; + for (const auto& entry : fs::directory_iterator(base)) { + names.insert(entry.path().filename().string()); + } + CHECK(names == std::set{"r0-0a1b2c3d"}); + CHECK(lock.directory() == base + "/r0-0a1b2c3d"); + CHECK(fs::exists(base + "/r0-0a1b2c3d/.owner.lock")); + + // The window itself: a watcher -- an adopter's scan, in effect -- lists + // the parent over and over and probes each rank directory it has not yet + // seen held, while this process creates many and keeps holding every + // one, so their creation is the only moment one could read as unheld. + // Each take is slowed where it has opened a lock file and not locked it + // (the lock-open seam), so a directory there to be seen before its lock + // would be seen. One made first and locked after gives such a scan a + // dead-looking sibling, which an adopter would take, and remove from + // under its claimer. + const std::string parent = base + "/race"; + fs::create_directories(parent); + int report[2], stop[2]; + CHECK(::pipe(report) == 0 && ::pipe(stop) == 0); + const pid_t watcher = ::fork(); + if (watcher == 0) { + ::close(report[0]); + ::close(stop[1]); + ::fcntl(stop[0], F_SETFL, O_NONBLOCK); + std::set held, unheld; + char byte = 0; + while (::read(stop[0], &byte, 1) < 0 && errno == EAGAIN) { + std::error_code ec; + for (fs::directory_iterator it(parent, ec), end; !ec && it != end; + it.increment(ec)) { + const std::string name = it->path().filename().string(); + uint64_t rank = 0; + std::string incarnation; + if (held.count(name) != 0 || + !dmi_store::ParseSpoolRankDirectoryName(name, &rank, + &incarnation)) { + continue; + } + if (dmi_store::ReadSpoolOwner(it->path().string(), nullptr)) { + held.insert(name); + } else { + unheld.insert(name); + } + } + } + const int count = static_cast(unheld.size()); + if (::write(report[1], &count, sizeof(count)) != sizeof(count)) { + ::_exit(3); + } + ::_exit(0); + } + ::close(report[1]); + ::close(stop[0]); + dmi_store::SetLockOpenHookForTesting([](const std::string&) { + ::usleep(1000); + }); + std::vector claims(100); + int taken = 0; + for (size_t i = 0; i < claims.size(); ++i) { + char name[32]; + std::snprintf(name, sizeof(name), "/r0-%08zx", i); + if (SpoolOwnerLock::Acquire(parent + name, false, &claims[i], &error) == + SpoolStatus::kOk) { + ++taken; + } + } + dmi_store::SetLockOpenHookForTesting(nullptr); + ::close(stop[1]); // the watcher stops at EOF + int unheld = -1; + CHECK(::read(report[0], &unheld, sizeof(unheld)) == sizeof(unheld)); + ::close(report[0]); + int status = 0; + ::waitpid(watcher, &status, 0); + CHECK(taken == 100); + if (unheld != 0) { + std::cerr << "a scan met " << unheld << " unheld new directories\n"; + } + CHECK(unheld == 0); +} + +// (6e) An adopter recovers a dead spool a pack at a time. Recover() hashes +// every pack of what may be a large backlog in one call, and whatever +// waited for the adopter -- stop(), a flush() behind its cycle -- waited +// for all of it. BeginRecovery sweeps the dead owner's stale .open file and +// lists the ready packs, hashing none, the account left as it was; each +// ContinueRecovery validates the next one, quarantining a corrupt one at +// its turn, and the last rebuilds the account as Recover() does. Until +// then every other ready pack stays where it was: an adopter let go of +// midway leaves nothing lost. +void TestAnAdopterRecoversADeadSpoolAPackAtATime() { + const std::string dead = + FreshRoot("adopt-steps") + "/0123456789ab/r0-0000dead"; + std::string error; + { + Spool gone; + CHECK(Spool::Open({dead, 1 << 20}, &gone, &error) == SpoolStatus::kOk); + for (int n = 1; n <= 3; ++n) { + CHECK(StageOne(gone, n, &error) == SpoolStatus::kOk); + } + } // its owner died + const std::string stale = + dead + "/v1/.018f0000-0000-7000-8000-00000000dead.0badf00d.open"; + std::ofstream(stale) << "half a pack"; + const auto readies = [&dead] { + std::set found; + for (const auto& entry : fs::recursive_directory_iterator(dead)) { + const std::string name = entry.path().filename().string(); + if (name.size() > 6 && name.substr(name.size() - 6) == ".ready") { + found.insert(entry.path().string()); + } + } + return found; + }; + const std::set staged_before = readies(); + CHECK(staged_before.size() == 3); + // The second in listing order no longer matches its checksum. + const std::string corrupt = *std::next(staged_before.begin()); + std::ofstream(corrupt, std::ios::binary | std::ios::trunc) + << std::string(100, 'z'); + + SpoolOwnerLock adopter; + CHECK(SpoolOwnerLock::TryAdopt(dead, &adopter, &error) == SpoolStatus::kOk); + SpoolConfig config{dead, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool adopted; + CHECK(Spool::Open(config, &adopted, &error) == SpoolStatus::kOk); + const uint64_t bytes_before = adopted.Snapshot().bytes; + + dmi_store::SpoolRecovery recovery; + CHECK(adopted.BeginRecovery(&recovery, &error) == SpoolStatus::kOk); + CHECK(!fs::exists(stale)); + CHECK((std::set(recovery.listed.begin(), + recovery.listed.end()) == staged_before)); + CHECK(recovery.next == 0 && recovery.valid.empty()); + CHECK(readies() == staged_before); + CHECK(adopted.Snapshot().bytes == bytes_before); + + CHECK(!adopted.ContinueRecovery(&recovery)); // the first + CHECK(recovery.next == 1 && recovery.valid.size() == 1); + CHECK(recovery.valid[0].path == *staged_before.begin()); + CHECK(!adopted.ContinueRecovery(&recovery)); // the corrupt one + CHECK(recovery.next == 2 && recovery.valid.size() == 1); + CHECK(!fs::exists(corrupt)); + CHECK(fs::exists(corrupt.substr(0, corrupt.size() - 6) + ".quarantined")); + CHECK(readies().size() == 2); + CHECK(adopted.Snapshot().bytes == bytes_before); // not rebuilt yet + + CHECK(adopted.ContinueRecovery(&recovery)); // the last, and the account + CHECK(recovery.valid.size() == 2); + CHECK(recovery.valid[1].path == *staged_before.rbegin()); + CHECK(recovery.valid[1].object_key == + "v1/018f0000-0000-7000-8000-000000000003.dmi-pack"); + CHECK(adopted.Snapshot().bytes == 200); + CHECK(adopted.Snapshot().entries == 2); + CHECK(adopted.ContinueRecovery(&recovery)); // done stays done + CHECK(recovery.valid.size() == 2); + + // Nothing listed: done at the first step. + const std::string empty = + FreshRoot("adopt-steps-empty") + "/0123456789ab/r0-00000e00"; + { Spool gone; CHECK(Spool::Open({empty, 1 << 20}, &gone, &error) == + SpoolStatus::kOk); } + SpoolOwnerLock empty_adopter; + CHECK(SpoolOwnerLock::TryAdopt(empty, &empty_adopter, &error) == + SpoolStatus::kOk); + SpoolConfig empty_config{empty, 1 << 20}; + empty_config.owner_lock = OwnerLock::kHeldByCaller; + Spool empty_spool; + CHECK(Spool::Open(empty_config, &empty_spool, &error) == SpoolStatus::kOk); + dmi_store::SpoolRecovery nothing; + CHECK(empty_spool.BeginRecovery(¬hing, &error) == SpoolStatus::kOk); + CHECK(nothing.listed.empty()); + CHECK(empty_spool.ContinueRecovery(¬hing)); + CHECK(nothing.valid.empty()); +} + +// (8) The budget across incarnations. Every process start gets a fresh +// rank directory, so a spool that charged only its own directory let each +// crash-restart add a full max_bytes while uploads were blocked. With +// charge_dead_siblings, what the sibling rank directories hold counts +// against max_bytes too, while adoption can drain it: a dead one, or one +// this process's adoption holds -- not one another live process holds +// (that is its own budget), one this process holds for its own writing, +// or one an adopter left blocked. +void TestDeadSiblingsCountAgainstTheBudget() { + const std::string base = FreshRoot("budget"); + std::string error; + const auto stage_packs = [&](const std::string& dir, int first, int n) { + Spool spool; + CHECK(Spool::Open({dir, 1 << 20}, &spool, &error) == SpoolStatus::kOk); + for (int i = first; i < first + n; ++i) { + CHECK(StageOne(spool, i, &error) == SpoolStatus::kOk); + } + }; // the Spool, and with it its lock, goes: a dead incarnation + + // A dead incarnation left three 100-byte packs beside the new one. + const std::string key = base + "/0123456789ab"; + stage_packs(key + "/r0-0000000a", 1, 3); + SpoolConfig config{key + "/r0-0000000b", 450}; + config.charge_dead_siblings = true; + Spool spool; + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kOk); + CHECK(spool.Snapshot().sibling_bytes == 300); + CHECK(StageOne(spool, 4, &error) == SpoolStatus::kOk); // 300 + 100 + error.clear(); + CHECK(StageOne(spool, 5, &error) == SpoolStatus::kFull); // 300 + 200 + CHECK(Contains(error, "300 bytes")); + CHECK(Contains(error, "dead")); + // As adoption drains the dead one, the capacity comes back. + for (const auto& entry : + fs::recursive_directory_iterator(key + "/r0-0000000a")) { + if (entry.path().extension() == ".ready") fs::remove(entry.path()); + } + CHECK(StageOne(spool, 5, &error) == SpoolStatus::kOk); + CHECK(spool.Snapshot().sibling_bytes == 0); + + // Without the charge each incarnation had the whole budget to itself. + const std::string uncharged = base + "/ba9876543210"; + stage_packs(uncharged + "/r0-0000000a", 1, 3); + Spool alone; + CHECK(Spool::Open({uncharged + "/r0-0000000b", 450}, &alone, &error) == + SpoolStatus::kOk); + for (int i = 4; i < 8; ++i) { + CHECK(StageOne(alone, i, &error) == SpoolStatus::kOk); + } + + // A sibling another live process holds is its own budget. + const std::string shared = base + "/aaaaaaaaaaaa"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t child = ::fork(); + if (child == 0) { + ::close(ready[0]); + Spool live; + std::string child_error; + bool ok = Spool::Open({shared + "/r1-0000000d", 1 << 20}, &live, + &child_error) == SpoolStatus::kOk; + for (int i = 1; ok && i <= 3; ++i) { + ok = StageOne(live, i, &child_error) == SpoolStatus::kOk; + } + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + SpoolConfig beside{shared + "/r0-0000000e", 450}; + beside.charge_dead_siblings = true; + Spool next; + CHECK(Spool::Open(beside, &next, &error) == SpoolStatus::kOk); + CHECK(next.Snapshot().sibling_bytes == 0); + for (int i = 4; i < 8; ++i) { + CHECK(StageOne(next, i, &error) == SpoolStatus::kOk); + } + ::kill(child, SIGKILL); + int status = 0; + ::waitpid(child, &status, 0); + + // Only what adoption can drain is charged. One THIS process holds for + // its own writing -- an earlier engine's claim, kept owned while its + // unsealed sink might still stage -- is drained by no adoption of this + // process's (its service reads it as live), so charging it left every + // later sink in the process that much less budget until exit. + const std::string mixed = base + "/bbbbbbbbbbbb"; + const std::string sibling = mixed + "/r0-0000000f"; + SpoolOwnerLock held; + CHECK(SpoolOwnerLock::Acquire(sibling, false, &held, &error) == + SpoolStatus::kOk); + { + SpoolConfig kept{sibling, 1 << 20}; + kept.owner_lock = OwnerLock::kHeldByCaller; + Spool writer; + CHECK(Spool::Open(kept, &writer, &error) == SpoolStatus::kOk); + for (int i = 1; i <= 3; ++i) { + CHECK(StageOne(writer, i, &error) == SpoolStatus::kOk); + } + } + const auto charged = [&]() { + SpoolConfig own{mixed + "/r0-00000010", 450}; + own.charge_dead_siblings = true; + Spool mine; + CHECK(Spool::Open(own, &mine, &error) == SpoolStatus::kOk); + return mine.Snapshot().sibling_bytes; + }; + CHECK(charged() == 0); + { + SpoolConfig own{mixed + "/r0-00000010", 450}; + own.charge_dead_siblings = true; + Spool mine; + CHECK(Spool::Open(own, &mine, &error) == SpoolStatus::kOk); + for (int i = 4; i < 8; ++i) { + CHECK(StageOne(mine, i, &error) == SpoolStatus::kOk); + } + } + fs::remove_all(mixed + "/r0-00000010"); + + // One this process's adoption holds -- a dead one its service drains -- + // still counts: the room comes back as the adoption drains it. + held.Release(); + SpoolOwnerLock adopting; + CHECK(SpoolOwnerLock::TryAdopt(sibling, &adopting, &error) == + SpoolStatus::kOk); + dmi_store::SpoolOwner record; + CHECK(dmi_store::ReadSpoolOwner(sibling, &record)); + CHECK(record.adopting && record.blocked.empty()); + CHECK(record.pid == ::getpid()); + CHECK(charged() == 300); + + // One an adopter left for good -- blocked, its lock let go -- is drained + // by nobody here, and is not charged either; its lock file says why. + CHECK(adopting.MarkBlocked("it holds a pack this service can never upload")); + adopting.Release(); + CHECK(!dmi_store::ReadSpoolOwner(sibling, &record)); + CHECK(record.blocked == "it holds a pack this service can never upload"); + CHECK(!record.adopting); + CHECK(charged() == 0); + // A process that can adopt it after all takes it again, and the mark + // goes with its take. + CHECK(SpoolOwnerLock::TryAdopt(sibling, &adopting, &error) == + SpoolStatus::kOk); + CHECK(dmi_store::ReadSpoolOwner(sibling, &record)); + CHECK(record.adopting && record.blocked.empty()); + CHECK(charged() == 300); + adopting.Release(); + CHECK(charged() == 300); // dead: still to adopt + + // One this process cannot take and empty at all -- another user's -- is + // no adoption's here either. + fs::permissions(sibling + "/.owner.lock", fs::perms::owner_read); + fs::permissions(sibling, fs::perms::owner_read | fs::perms::owner_exec); + if (::access(sibling.c_str(), W_OK) != 0) { // not as root + CHECK(charged() == 0); + } + fs::permissions(sibling, fs::perms::owner_all); + fs::permissions(sibling + "/.owner.lock", + fs::perms::owner_read | fs::perms::owner_write); + CHECK(charged() == 300); +} + +// (7) The layout: //r-/. +dmi_store::SpoolDestination Destination() { + dmi_store::SpoolDestination destination; + destination.clickhouse_host = "ch"; + destination.clickhouse_port = 8123; + destination.database = "db"; + destination.table_prefix = "prefix"; + destination.s3_endpoint = "http://s3:9000"; + destination.s3_bucket = "bucket"; + destination.store_id = "s3"; + return destination; +} + +void TestTheDirectoryLayout() { + const std::string key = dmi_store::SpoolCatalogKey(Destination()); + CHECK(key == Sha256Hex("db/prefix/s3\nclickhouse ch:8123\n" + "s3 http://s3:9000/bucket").substr(0, 12)); + CHECK(dmi_store::IsSpoolCatalogKey(key)); + CHECK(!dmi_store::IsSpoolCatalogKey("0123456789aB")); + CHECK(!dmi_store::IsSpoolCatalogKey("0123456789a")); + // Every part of where the packs go is in the key: two deployments that + // share a spool_root and the default names, but not a server, never + // adopt each other's directories. + std::set keys{key}; + for (int field = 0; field < 7; ++field) { + dmi_store::SpoolDestination other = Destination(); + switch (field) { + case 0: other.clickhouse_host = "ch2"; break; + case 1: other.clickhouse_port = 8124; break; + case 2: other.database = "db2"; break; + case 3: other.table_prefix = "prefix2"; break; + case 4: other.s3_endpoint = "http://s3b:9000"; break; + case 5: other.s3_bucket = "bucket2"; break; + case 6: other.store_id = "s4"; break; + } + keys.insert(dmi_store::SpoolCatalogKey(other)); + } + CHECK(keys.size() == 8); + + CHECK(dmi_store::SpoolRankDirectoryName(3, "0a1b2c3d") == "r3-0a1b2c3d"); + uint64_t rank = 0; + std::string incarnation; + CHECK(dmi_store::ParseSpoolRankDirectoryName("r12-deadbeef", &rank, + &incarnation)); + CHECK(rank == 12 && incarnation == "deadbeef"); + for (const char* bad : {"r-1-deadbeef", "r1-DEADBEEF", "r1-deadbee", + "r1-deadbeef0", "rx-deadbeef", "1-deadbeef", + "r1deadbeef", ".r1-deadbeef", "r01-deadbeef"}) { + CHECK(!dmi_store::ParseSpoolRankDirectoryName(bad, &rank, &incarnation)); + } + std::set seen; + for (int i = 0; i < 64; ++i) { + const std::string fresh = dmi_store::NewSpoolIncarnation(); + CHECK(dmi_store::ParseSpoolRankDirectoryName("r0-" + fresh, &rank, + &incarnation)); + seen.insert(fresh); + } + CHECK(seen.size() == 64); + CHECK(dmi_store::SpoolRankDirectory("/b", Destination(), 2, "0a1b2c3d") == + "/b/" + key + "/r2-0a1b2c3d"); +} + +} // namespace + +int main() { + TestTwoTakesInOneProcessRefuseEachOther(); + TestHeldByCallerOpensBesideTheHolder(); + TestHeldByCallerBesideAnotherProcessIsRefused(); + TestHeldByCallerWhereTheKernelListsNoFlocks(); + TestHeldByCallerWithoutAHolderIsRefused(); + TestTheLockGoesWithItsSpool(); + TestASecondProcessIsRefusedUntilTheHolderDies(); + TestAForkedChildDoesNotKeepTheLockPastItsParent(); + TestAForkWhileALockIsTakenLeavesTheChildNothing(); + TestNestedDirectoriesAreRefused(); + TestAnOuterAndANestedTakeRacingNeverBothWin(); + TestAStaleLockFileAboveDoesNotRefuseANestedDirectory(); + TestAnUnheldClaimStagingDirectoryRefusesNothing(); + TestAnAdopterLeavesASpoolNestedInADeadOneAlone(); + TestSharedFilesystemsAreRefusedUnlessAllowed(); + TestAdoptionLocksOnlyWhatExistsAndIsDead(); + TestALockOnAnUnlinkedFileIsTakenAgain(); + TestAReplacedLockFileLeavesTheDirectoryOwned(); + TestANewDirectoryAppearsWithItsLockHeld(); + TestAnAdopterRecoversADeadSpoolAPackAtATime(); + TestTheDirectoryLayout(); + TestDeadSiblingsCountAgainstTheBudget(); + if (g_failures != 0) { + std::cerr << g_failures << " check(s) failed\n"; + return 1; + } + std::cout << "ok\n"; + return 0; +} diff --git a/tests/native/test_spool_reservations.cpp b/tests/native/test_spool_reservations.cpp index 43a7659c4..68abcfcbd 100644 --- a/tests/native/test_spool_reservations.cpp +++ b/tests/native/test_spool_reservations.cpp @@ -104,6 +104,16 @@ std::string FreshRoot(const char* tag) { return root; } +// The config for a SECOND Spool object on a root the first one opened. The +// first takes the directory's owner lock; in one process the second shares +// it (owner_lock=held_by_caller), as the engine's sink and storage service +// do -- two takes would refuse each other, since flock binds to an open file +// description rather than to the process. +dmi_store::SpoolConfig Beside(dmi_store::SpoolConfig config) { + config.owner_lock = dmi_store::OwnerLock::kHeldByCaller; + return config; +} + // (1) The serial uploader case: stage, remove through a second object, // stage again on the first object. Must succeed, and the accounting must // end at exactly one file. @@ -114,7 +124,7 @@ void TestSerialRemoveThroughAnotherSpoolIsReconciled() { std::string error; CHECK(dmi_store::Spool::Open(config, &writer, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &uploader, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &uploader, &error) == dmi_store::SpoolStatus::kOk); dmi_store::StagedPack first; @@ -246,7 +256,7 @@ void TestALostLinkRaceReleasesItsReservation() { std::string error; CHECK(dmi_store::Spool::Open(config, &loser, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &winner, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &winner, &error) == dmi_store::SpoolStatus::kOk); dmi_store::StagedPack winner_out; @@ -295,7 +305,7 @@ void TestARetryOfAnotherObjectsReadyFileIsAccounted() { std::string error; CHECK(dmi_store::Spool::Open(config, &writer, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &other, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &other, &error) == dmi_store::SpoolStatus::kOk); // `other` writes the ready file; `writer` has never seen it. @@ -345,7 +355,7 @@ void TestARetryCannotOversubscribeAnInflightReservation() { std::string error; CHECK(dmi_store::Spool::Open(config, &spool, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &other, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &other, &error) == dmi_store::SpoolStatus::kOk); dmi_store::SpoolStatus retry_status = dmi_store::SpoolStatus::kIo; @@ -405,7 +415,7 @@ void TestAnEexistLoserAccountsForTheWinnersFile() { std::string error; CHECK(dmi_store::Spool::Open(config, &loser, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &winner, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &winner, &error) == dmi_store::SpoolStatus::kOk); // The loser reserves, and while it is paused the winner links the ready @@ -484,7 +494,7 @@ void TestASerialRetryIsAdmittedEvenWhenAlreadyOverCap() { // retry now goes through a Spool that has NOT ledgered this path, so it // reconciles, sees 1000 > 900, and takes the capacity decision. dmi_store::Spool reopened; - dmi_store::SpoolConfig lowered{root, 900}; + dmi_store::SpoolConfig lowered = Beside({root, 900}); CHECK(dmi_store::Spool::Open(lowered, &reopened, &error) == dmi_store::SpoolStatus::kOk); @@ -530,7 +540,7 @@ void TestRecoverThenRemoveReleasesTheAccount() { { dmi_store::Spool writer; - CHECK(dmi_store::Spool::Open(config, &writer, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &writer, &error) == dmi_store::SpoolStatus::kOk); dmi_store::StagedPack out; CHECK(StageBytes(writer, 1, 1000, &out, &error) == @@ -596,7 +606,7 @@ void TestRemoveOfAnAlreadyDeletedFileStillUncounts() { std::string error; CHECK(dmi_store::Spool::Open(config, &writer, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &other, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &other, &error) == dmi_store::SpoolStatus::kOk); dmi_store::StagedPack staged; diff --git a/tests/test_native_capture_chain_live.py b/tests/test_native_capture_chain_live.py index 5004db901..7e2517356 100644 --- a/tests/test_native_capture_chain_live.py +++ b/tests/test_native_capture_chain_live.py @@ -210,24 +210,33 @@ def _run_chain(config, spool_root: Path, envelopes, *, sink_overrides=None, which is safe only while no sink writes there), then the sink; flush both, and return the snapshots and what the reader reads back. + Both open the spool as the engine opens them: under ONE owner lock this + process holds, each with owner_lock="held_by_caller" -- two takes in one + process would refuse each other. + The sink is the raw binding with `sink_overrides`, or, given a NativeSinkConfig, the one the engine builds from it.""" from dmi.storage.native_capture import ( NativeCaptureReader, NativeCaptureStorage, + _load_native_store_extension, ) + owner = _load_native_store_extension().SpoolOwnerLock(str(spool_root)) service = NativeCaptureStorage(config, spool_root=str(spool_root), - spool_max_bytes=1 << 40, sweep_spool=True) + spool_max_bytes=1 << 40, sweep_spool=True, + spool_owner_lock="held_by_caller") service.start() try: if sink_config is None: - sink, _lease = _open_sink(spool_root, **(sink_overrides or {})) + sink, _lease = _open_sink(spool_root, owner_lock="held_by_caller", + **(sink_overrides or {})) else: from dmi.storage.capture.native_sink import ( create_native_pack_sink, ) - sink = create_native_pack_sink(sink_config).native_sink + sink = create_native_pack_sink( + sink_config, owner_lock="held_by_caller").native_sink _lease = sink.attach() for envelope in envelopes: sink.submit_envelope(LAYOUT, envelope.rows, envelope.payload()) @@ -239,6 +248,7 @@ def _run_chain(config, spool_root: Path, envelopes, *, sink_overrides=None, service_snapshot = service.snapshot() finally: service.stop() + owner.release() reader = NativeCaptureReader(config) selection = reader.select(tenant_id="t") diff --git a/tests/test_native_capture_storage_gpu_e2e.py b/tests/test_native_capture_storage_gpu_e2e.py index b317ea184..7a74a9dc0 100644 --- a/tests/test_native_capture_storage_gpu_e2e.py +++ b/tests/test_native_capture_storage_gpu_e2e.py @@ -237,7 +237,8 @@ def test_close_alone_delivers_the_tail_to_the_catalog(fake_s3, tmp_path): """No flush_and_wait: close() must still seal the sink's open pack and drain it into the catalog. The 60 s linger means nothing but a flush can seal it, and 3 records against max_pack_records=2 leave one record in the - open pack when close() runs.""" + open pack when close() runs. The sink sealed, close() lets go of the + engine's spool directory and removes it.""" from dmi.api.v1 import HookPointV1, HookSpecV1, MonitoringEngine, TransportSpec from dmi.config import MonitoringConfig from dmi.storage.capture import CaptureRecordFormat @@ -283,6 +284,11 @@ def test_close_alone_delivers_the_tail_to_the_catalog(fake_s3, tmp_path): engine.close() # no flush_and_wait + # The ring's release sealed the sink (sealed_on_release), so close() + # let go of the engine's own spool directory, and removed it once + # drained: no rank directory is left for anyone to adopt. + assert list((tmp_path / "spool").glob("*/r*-*")) == [] + reader = NativeCaptureReader(storage) selection = reader.select(tenant_id="tenant-gpu") captures = {capture.descriptor["capture_id"]: capture diff --git a/tests/test_native_capture_storage_live.py b/tests/test_native_capture_storage_live.py index 7d89fa57b..9457919e8 100644 --- a/tests/test_native_capture_storage_live.py +++ b/tests/test_native_capture_storage_live.py @@ -23,6 +23,7 @@ import base64 import json import re +import shutil import socket import subprocess import threading @@ -113,12 +114,49 @@ def _storage_config(endpoint, prefix, **overrides): return NativeCaptureStorageConfig(**fields) +# Every spool directory a test points a service at is owned by the harness +# for the rest of the test: one SpoolOwnerLock held here, and each service +# opens the directory with owner_lock="held_by_caller". That is the engine's +# arrangement -- it holds the lock around its sink and its service -- and a +# test often builds a second service on a directory the first still has +# open; each taking the lock would refuse the others. +# +# The drivers are processes of their own, so they cannot open a directory +# this process owns: held_by_caller is refused unless the opening process +# holds the lock. They take it, as the plan has standalone callers do. The +# sink driver stages into a scratch directory it owns, and _stage moves its +# sealed packs into the spool, where a sink in this process would have +# staged them; the store driver uploads from a spool before any service of +# the test has opened it. +_HARNESS_LOCKS: dict = {} + + +@pytest.fixture(autouse=True) +def _harness_spool_locks(): + yield + for lock in _HARNESS_LOCKS.values(): + lock.release() + _HARNESS_LOCKS.clear() + + +def _held(spool_root) -> str: + """Hold spool_root's owner lock for the test; the mode to open it in.""" + from dmi.storage.native_capture import _load_native_store_extension + + key = str(spool_root) + if key not in _HARNESS_LOCKS: + _HARNESS_LOCKS[key] = _load_native_store_extension().SpoolOwnerLock( + key) + return "held_by_caller" + + def _service(config, spool_root: Path, *, sweep_spool=True): from dmi.storage.native_capture import NativeCaptureStorage return NativeCaptureStorage(config, spool_root=str(spool_root), spool_max_bytes=1 << 40, - sweep_spool=sweep_spool) + sweep_spool=sweep_spool, + spool_owner_lock=_held(spool_root)) def _record(index: int): @@ -141,12 +179,18 @@ def _record(index: int): def _stage(spool_root: Path, indexes, *, records_per_pack: int = 2): - """Stage records through the native sink core; return the tensors.""" + """Stage records through the native sink core; return the tensors. + + The driver takes a scratch directory of its own beside spool_root, and + once its sink is closed the sealed packs move into spool_root under the + same relative paths -- one rename each, so a service scanning the spool + meets a whole ready file or none.""" + scratch = spool_root.parent / f".stage-{uuid.uuid4().hex[:8]}" sink = _Driver(SINK_DRIVER) tensors = {} try: assert sink.call( - op="open", root=str(spool_root), max_bytes=1 << 40, + op="open", root=str(scratch), max_bytes=1 << 40, max_queue_records=256, max_queue_bytes=1 << 24, max_pack_bytes=8 << 20, max_pack_records=records_per_pack, max_linger_ns=1_000_000_000, overload="drop_newest", @@ -164,6 +208,11 @@ def _stage(spool_root: Path, indexes, *, records_per_pack: int = 2): assert snapshot["persisted_records"] == len(tensors), snapshot finally: sink.close() + for ready in sorted(scratch.rglob("*.dmi-pack.ready")): + target = spool_root / ready.relative_to(scratch) + target.parent.mkdir(parents=True, exist_ok=True) + ready.replace(target) + shutil.rmtree(scratch) return tensors @@ -307,6 +356,9 @@ def __init__(self, host: str, port: int): self.refused: list[float] = [] self._lock = threading.Lock() self._sockets: set[socket.socket] = set() + # close() has run: a request held back by a delay is not forwarded + # once it has. + self._closed = False threading.Thread(target=self._accept, daemon=True).start() @classmethod @@ -372,6 +424,7 @@ def _look_then_route(self, client): what delay_requests() says; anything else goes straight through.""" request = b"" continued = False + whole = False # the head and all Content-Length bytes of body try: client.settimeout(2.0) while True: @@ -379,6 +432,7 @@ def _look_then_route(self, client): if found: length = re.search(rb"(?i)content-length:\s*(\d+)", head) if length is None or len(body) >= int(length.group(1)): + whole = True break # libcurl holds a body over 1 KiB back until the server # says 100 Continue, or for a second; answer for it, so @@ -395,6 +449,12 @@ def _look_then_route(self, client): except OSError: client.close() return + if not whole: + # The client went away mid-request -- a cancel that landed + # after the head, say. Forwarded, the fake S3 stored the short + # body under the key, over what a later upload put there. + client.close() + return if continued: # The server must not answer 100 Continue a second time. head, _, body = request.partition(b"\r\n\r\n") @@ -418,6 +478,10 @@ def _look_then_route(self, client): and b"_publisher_lease" in request) if late: time.sleep(late_by) + if self._closed: + # Held back past close(): the test is done with this server. + client.close() + return try: upstream = socket.create_connection(self._target) upstream.sendall(request) @@ -487,7 +551,9 @@ def close(self): listener does not wake a thread blocked in accept(), which still takes one more queued connection: the stall settings go first, so that connection is refused rather than held open for its client's - whole timeout (a stop()'s lease release after a stall() did).""" + whole timeout (a stop()'s lease release after a stall() did). + A request a delay still holds back is dropped, not forwarded.""" + self._closed = True self._stalled = False self._stall_if = None self._slow_once = None @@ -497,6 +563,83 @@ def close(self): self._listener.close() +class _Recorder: + """A TCP server that keeps every byte it is sent, by connection.""" + + def __init__(self): + self._listener = socket.create_server(("127.0.0.1", 0)) + self.port = self._listener.getsockname()[1] + self.received: list[bytes] = [] + self._lock = threading.Lock() + threading.Thread(target=self._accept, daemon=True).start() + + def _accept(self): + while True: + try: + connection, _ = self._listener.accept() + except OSError: + return + threading.Thread(target=self._keep, args=(connection,), + daemon=True).start() + + def _keep(self, connection): + data = b"" + try: + while chunk := connection.recv(65536): + data += chunk + except OSError: + pass + connection.close() + with self._lock: + self.received.append(data) + + def close(self): + self._listener.close() + + +def test_the_switch_forwards_no_request_its_client_did_not_finish(tmp_path): + """The harness itself. A client that went away after a request's head + -- a flush deadline's cancel landing between libcurl's head and body, + which the switch had answered 100 Continue for -- was forwarded all + the same, its body short: the fake S3, whose signature check trusts + x-amz-content-sha256, then stored an empty object over what a later + upload put at the key. Nor does a request a delay still holds back + when close() runs reach the server after it.""" + recorder = _Recorder() + switch = _Switch("127.0.0.1", recorder.port) + switch.delay_requests(lambda request: 0.3) + head = (b"PUT /b/k HTTP/1.1\r\nHost: x\r\nContent-Length: 2048\r\n" + b"Expect: 100-continue\r\n\r\n") + try: + with socket.create_connection(("127.0.0.1", switch.port), + timeout=10) as client: + client.sendall(head) + assert client.recv(64).startswith(b"HTTP/1.1 100 Continue") + # The control: a whole request, held back and then forwarded. + with socket.create_connection(("127.0.0.1", switch.port), + timeout=10) as client: + client.sendall(head + bytes(2048)) + client.shutdown(socket.SHUT_WR) + time.sleep(1.0) + with recorder._lock: + forwarded = list(recorder.received) + assert len(forwarded) == 1, forwarded + assert forwarded[0].endswith(bytes(2048)), forwarded + + # Held back when close() runs, then dropped. + with socket.create_connection(("127.0.0.1", switch.port), + timeout=10) as client: + client.sendall(head + bytes(2048)) + time.sleep(0.1) + switch.close() + time.sleep(1.0) + with recorder._lock: + assert recorder.received == forwarded, recorder.received + finally: + switch.close() + recorder.close() + + def test_an_index_failure_keeps_the_pack_owed_until_it_lands(fake_s3, tmp_path): spool_root = tmp_path / "spool" switch = _Switch(CLICKHOUSE_HOST, CLICKHOUSE_HTTP_PORT) @@ -642,7 +785,9 @@ def test_a_failed_head_is_an_error_not_a_foreign_object(fake_s3, tmp_path): "etag": '"0"'} with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() - native.update(spool_root=str(tmp_path / "spool"), holder="head-test", + native.update(spool_root=str(tmp_path / "spool"), + spool_owner_lock=_held(tmp_path / "spool"), + holder="head-test", reconcile_prefix="fault/", s3_max_attempts=1) service = _load_native_store_extension().StorageService(native) service.start() # reconciles once @@ -692,7 +837,8 @@ def test_the_loop_backs_off_while_the_object_store_is_down(tmp_path): with _catalog() as (_client, catalog): native = _storage_config(dead, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="backoff-test", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="backoff-test", poll_interval_ns=20_000_000, max_backoff_ns=10_000_000_000, reconcile_on_start=False, s3_max_attempts=1, uploader_max_attempts=1) @@ -725,7 +871,8 @@ def test_a_pack_too_big_to_index_is_set_aside_and_the_rest_still_index( config = _storage_config(fake_s3, catalog.table_prefix) native = config._native_dict() native.update( - spool_root=str(spool_root), holder="poison-test", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="poison-test", poll_interval_ns=50_000_000, reconcile_on_start=False, # One descriptor fits, forty do not. indexer_max_estimated_bytes=4000) @@ -760,7 +907,8 @@ def test_a_batch_over_the_budget_splits_until_every_pack_indexes( config = _storage_config(fake_s3, catalog.table_prefix) native = config._native_dict() native.update( - spool_root=str(spool_root), holder="split-test", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="split-test", poll_interval_ns=50_000_000, reconcile_on_start=False, # A two-record pack renders ~720 bytes: two fit, ten do not. indexer_max_estimated_bytes=2000) @@ -795,7 +943,8 @@ def test_flush_returns_on_time_when_the_catalog_stops_answering( native = _storage_config( fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() - native.update(spool_root=str(spool_root), holder="stall-test", + native.update(spool_root=str(spool_root), + spool_owner_lock=_held(spool_root), holder="stall-test", poll_interval_ns=20_000_000, reconcile_on_start=False, clickhouse_request_timeout_s=5.0) service = _load_native_store_extension().StorageService(native) @@ -842,7 +991,8 @@ def test_a_flush_against_a_black_hole_catalog_overruns_by_one_request( fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="black-hole-test", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="black-hole-test", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, # Every cycle is due a reconcile. @@ -887,7 +1037,8 @@ def test_the_stop_after_a_flush_runs_no_further_cycle(fake_s3, tmp_path): fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="stop-after-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stop-after-flush", # Wakes while the flush below holds the cycle lock. poll_interval_ns=3_000_000_000, reconcile_on_start=False, # No renewal falls due while the test runs (a third of the TTL). @@ -951,7 +1102,8 @@ def test_a_flush_out_of_time_does_not_hash_the_spool(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="flush-out-of-time", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="flush-out-of-time", # The loop sleeps through the test, and nothing uploads. poll_interval_ns=60_000_000_000, reconcile_on_start=False, sweep_spool_on_start=False, @@ -1010,7 +1162,8 @@ def test_stop_cuts_a_listing_that_hashes_a_backlog(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="listing-stop", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="listing-stop", poll_interval_ns=20_000_000, reconcile_on_start=False, sweep_spool_on_start=False, uploader_max_in_flight_bytes=1 << 30) @@ -1046,7 +1199,8 @@ def test_a_flush_cuts_the_listing_its_deadline_passes_in(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="listing-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="listing-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, sweep_spool_on_start=False, @@ -1084,7 +1238,8 @@ def test_a_flush_returns_on_time_while_an_upload_stalls(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-upload-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-upload-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, s3_read_timeout_s=30) @@ -1137,12 +1292,18 @@ def test_a_flush_that_cuts_a_multipart_upload_waits_for_its_abort( outlasts the request timeout. Here the store holds the part and the abort both, so the flush pays the whole bound, and no more. The pack is a sparse file of zeros named for its checksum, over the client's - 64 MiB multipart threshold; it stays staged.""" + 64 MiB multipart threshold; it stays staged. The budget has to see the + part stalled before it ends: the listing and the upload each hash the + pack first, and a HEAD and the CreateMultipartUpload go before the + part, which on a loaded runner outlasted a 1 s budget -- the deadline + then cut the upload before its part was sent, and nothing was left to + abort.""" import hashlib from dmi.storage.native_capture import _load_native_store_extension size = 65 << 20 + budget = 4.0 digest = hashlib.sha256() zeros = bytes(1 << 20) for _ in range(size // len(zeros)): @@ -1157,7 +1318,8 @@ def test_a_flush_that_cuts_a_multipart_upload_waits_for_its_abort( with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="multipart-abort-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="multipart-abort-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, clickhouse_request_timeout_s=2.0, @@ -1167,7 +1329,7 @@ def test_a_flush_that_cuts_a_multipart_upload_waits_for_its_abort( try: s3.stall_requests(_multipart_part_or_abort) started = time.monotonic() - drained = service.flush(1.0) + drained = service.flush(budget) elapsed = time.monotonic() - started snapshot = service.snapshot() finally: @@ -1177,7 +1339,7 @@ def test_a_flush_that_cuts_a_multipart_upload_waits_for_its_abort( assert drained is False assert len(s3.stalled) == 2, s3.stalled # the part, then the abort # The deadline, up to a second for the cut, and the 5 s abort. - assert elapsed < 1.0 + 1.0 + 5.0 + 1.0, (elapsed, snapshot) + assert elapsed < budget + 1.0 + 5.0 + 1.0, (elapsed, snapshot) assert snapshot["cancelled_uploads"] == 1, snapshot assert snapshot["upload_failures"] == 0, snapshot assert _ready(spool_root) == [ready] @@ -1187,7 +1349,10 @@ def test_stop_returns_promptly_while_an_upload_stalls(fake_s3, tmp_path): """stop() joined a loop whose cycle was inside a PUT the store never answers, so it waited out the S3 read timeout on every attempt, with the lease held. It cancels the upload now: the pack stays in the spool, the - lease is released, and the next process uploads the pack.""" + lease is released, and the next process uploads the pack. The uploader + is allowed eight attempts: one whose own retries and backoff stop() did + not cut, on a client stop() did, would sleep out about 26 s of backoff + (four attempts' 1.75 s fit under the bound).""" from dmi.storage.native_capture import _load_native_store_extension spool_root = tmp_path / "spool" @@ -1195,9 +1360,10 @@ def test_stop_returns_promptly_while_an_upload_stalls(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-upload-stop", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-upload-stop", poll_interval_ns=20_000_000, reconcile_on_start=False, - s3_read_timeout_s=60) + s3_read_timeout_s=60, uploader_max_attempts=8) service = _load_native_store_extension().StorageService(native) service.start() stopper = None @@ -1259,7 +1425,8 @@ def test_a_flush_returns_on_time_while_the_index_reads_stall( with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-read-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-read-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, s3_read_timeout_s=3, @@ -1318,7 +1485,8 @@ def test_stop_returns_promptly_while_the_index_reads_stall(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-read-stop", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-read-stop", poll_interval_ns=20_000_000, reconcile_on_start=False, s3_read_timeout_s=3) service = _load_native_store_extension().StorageService(native) @@ -1397,7 +1565,8 @@ def _slow_guard(request: bytes) -> float: s3.url, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="stop-no-replay-guard", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stop-no-replay-guard", poll_interval_ns=20_000_000, reconcile_on_start=False, uploader_max_workers=1, s3_read_timeout_s=30) # Staged first, so the loop's first cycle lists all four. @@ -1463,7 +1632,8 @@ def test_an_object_store_read_outage_sets_no_pack_aside(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="read-outage", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="read-outage", poll_interval_ns=20_000_000, max_backoff_ns=200_000_000, reconcile_on_start=False, s3_read_timeout_s=1, s3_max_attempts=1, max_index_attempts=2) @@ -1520,7 +1690,8 @@ def test_a_flush_against_a_slow_catalog_indexes_one_batch_past_its_deadline( fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="slow-catalog-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="slow-catalog-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, indexer_max_packs=1, @@ -1584,7 +1755,8 @@ def test_a_close_whose_budget_ends_mid_upload_leaves_nothing_owed( with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="close-mid-upload", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="close-mid-upload", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, uploader_max_workers=1, indexer_max_packs=4) @@ -1637,7 +1809,8 @@ def test_the_one_batch_past_a_flushs_deadline_is_a_full_one(fake_s3, fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="full-batch-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="full-batch-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, indexer_max_packs=2, clickhouse_request_timeout_s=60.0) @@ -1676,7 +1849,8 @@ def test_only_the_loop_reconciles_never_a_flush(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="flush-no-reconcile", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="flush-no-reconcile", poll_interval_ns=2_000_000_000, reconcile_on_start=False, reconcile_interval_ns=1_000_000) service = _load_native_store_extension().StorageService(native) @@ -1706,7 +1880,8 @@ def test_a_service_started_again_after_stop_uploads_again(fake_s3, tmp_path): spool_root = tmp_path / "spool" with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() - native.update(spool_root=str(spool_root), holder="restarted", + native.update(spool_root=str(spool_root), + spool_owner_lock=_held(spool_root), holder="restarted", reconcile_on_start=False) service = _load_native_store_extension().StorageService(native) service.start() @@ -1738,7 +1913,8 @@ def test_stop_cuts_a_reconcile_whose_listing_stalls(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-reconcile-stop", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-reconcile-stop", poll_interval_ns=20_000_000, reconcile_on_start=False, reconcile_interval_ns=1_000_000, s3_read_timeout_s=30) service = _load_native_store_extension().StorageService(native) @@ -1795,7 +1971,8 @@ def _tick(): native = _storage_config( fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() - native.update(spool_root=str(spool_root), holder="gil-test", + native.update(spool_root=str(spool_root), + spool_owner_lock=_held(spool_root), holder="gil-test", poll_interval_ns=20_000_000, reconcile_on_start=False, clickhouse_request_timeout_s=2.0) service = _load_native_store_extension().StorageService(native) @@ -1838,7 +2015,8 @@ def test_the_lease_holds_through_an_object_store_outage(tmp_path): def _native(spool, holder): native = _storage_config(dead, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool), holder=holder, + spool_root=str(spool), spool_owner_lock=_held(spool), + holder=holder, poll_interval_ns=20_000_000, max_backoff_ns=10_000_000_000, lease_ttl_ns=3_000_000_000, publish_timeout_ns=1_000_000_000, clock_skew_ns=0, reconcile_on_start=False, diff --git a/tests/test_native_capture_storage_wiring.py b/tests/test_native_capture_storage_wiring.py index f37e686da..e119bc6a9 100644 --- a/tests/test_native_capture_storage_wiring.py +++ b/tests/test_native_capture_storage_wiring.py @@ -2,10 +2,12 @@ The service itself is C++ and runs against a real object store and catalog in test_native_capture_storage_live.py. This suite pins what the engine -promises around it, with the native modules faked: the service starts -before the sink opens the spool it sweeps, ``flush_and_wait`` waits for the -catalog as well as the spool, ``close`` drains and releases the lease, and a -failed attach does not leave the lease held. +promises around it, with the native modules faked: the engine takes its own +rank directory's owner lock before anything opens it, the service starts +before the sink opens the spool it sweeps, both open that directory +held_by_caller, ``flush_and_wait`` waits for the catalog as well as the +spool, ``close`` drains and releases the lease and only then the spool +lock, and a failed attach leaves neither the lease nor the lock held. """ from __future__ import annotations @@ -441,12 +443,38 @@ def rethrow_if_failed(self): pass -def _capture_engine(monkeypatch, tmp_path, *, fail_ring=False): +RANK_DIRECTORY = "{base}/0123456789ab/r{rank}-0a1b2c3d" + + +class _FakeSpoolLock: + """Stands in for _dmi_native_store.SpoolOwnerLock.""" + + def __init__(self, events, directory, allow_shared_filesystem=False): + self.events = events + self.directory = directory + self.allow_shared_filesystem = allow_shared_filesystem + self.held = True + events.append(("lock", "acquire", directory)) + + def release_and_remove_if_empty(self): + self.held = False + self.events.append(("lock", "release")) + return True + + def release(self): + self.release_and_remove_if_empty() + + +def _capture_engine(monkeypatch, tmp_path, *, fail_ring=False, + fail_start=False, spool_owner=None, storage=True): """An engine under storage_backend="persistent" with both native modules - faked. Returns (engine, events, services).""" + faked. Returns (engine, events, services); ``events`` also records the + spool locks and the sinks' keyword arguments (``sinks``).""" from dmi.storage.capture.native_sink import NativeSinkConfig events, services = [], [] + events_sinks: list = [] + locks: list = [] engine = MonitoringEngine(enable_ring_transport=False) engine._ring_transport = SimpleNamespace(null_offload=False, force_eager=False) @@ -455,7 +483,7 @@ def _capture_engine(monkeypatch, tmp_path, *, fail_ring=False): engine._storage_backend = "persistent" engine._capture_sink_config = NativeSinkConfig( spool_root=str(tmp_path / "spool"), spool_max_bytes=1 << 30) - engine._capture_storage_config = _storage_config() + engine._capture_storage_config = _storage_config() if storage else None class _Lease: def release(self): @@ -468,15 +496,33 @@ def _acquire_engine(self): class _NativePackSink(_RecordSink): def __init__(self, **kwargs): events.append(("sink", "open", kwargs["spool_root"])) + events_sinks.append(kwargs) def _service(config): service = _FakeService(events, config) + if fail_start: + def _refuse(): + events.append(("service", "start")) + raise RuntimeError("publisher lease held elsewhere") + service.start = _refuse services.append(service) return service + def _lock(directory, allow_shared_filesystem=False): + lock = _FakeSpoolLock(events, directory, allow_shared_filesystem) + locks.append(lock) + return lock + + def _rank_directory(base, destination, rank): + events.append(("layout", destination, rank)) + return RANK_DIRECTORY.format(base=base, rank=rank) + def _load_named_extension(name): if name == "_dmi_native_store": return SimpleNamespace(StorageService=_service, + SpoolOwnerLock=_lock, + spool_rank_directory=_rank_directory, + spool_owner=lambda directory: spool_owner, SEARCH_ITEM_COLUMNS=()) return SimpleNamespace(NativePackSink=_NativePackSink) @@ -527,6 +573,8 @@ def __init__(self, config, host): import dmi.transport monkeypatch.setattr(dmi.transport, "native", native, raising=False) + engine._test_sinks = events_sinks + engine._test_locks = locks return engine, events, services @@ -546,21 +594,85 @@ def encode(self, metadata, entry): def test_the_service_starts_before_the_sink_opens_the_spool(monkeypatch, tmp_path): + """The engine's own directory is locked first; the service starts (and + sweeps it) before the sink opens it; both open it under that lock.""" + monkeypatch.delenv("RANK", raising=False) engine, events, services = _capture_engine(monkeypatch, tmp_path) engine.create_record_runtime(_record_format()) - spool_root = str(tmp_path / "spool") - assert events[:3] == [ + directory = RANK_DIRECTORY.format(base=tmp_path / "spool", rank=0) + storage = _storage_config() + assert events[:5] == [ + ("layout", storage._spool_destination(), 0), + ("lock", "acquire", directory), ("service", "construct"), ("service", "start"), - ("sink", "open", spool_root), + ("sink", "open", directory), ] config = services[0].config - assert config["spool_root"] == spool_root + assert config["spool_root"] == directory assert config["spool_max_bytes"] == 1 << 30 assert config["sweep_spool_on_start"] is True + assert config["spool_owner_lock"] == "held_by_caller" + assert config["adopt_sibling_spools"] is True + assert config["spool_allow_shared_filesystem"] is False assert config["holder"] # a generated lease holder, never empty + (sink,) = engine._test_sinks + assert sink["owner_lock"] == "held_by_caller" + # Dead incarnations' packs beside it count against its budget. + assert sink["charge_dead_siblings"] is True + assert engine._test_locks[0].held + + +def test_the_spool_directory_is_named_for_the_rank(monkeypatch, tmp_path): + monkeypatch.setenv("RANK", "3") + engine, events, services = _capture_engine(monkeypatch, tmp_path) + + engine.create_record_runtime(_record_format()) + + assert services[0].config["spool_root"] == RANK_DIRECTORY.format( + base=tmp_path / "spool", rank=3) + + +@pytest.mark.parametrize("rank", ["", "-1", "x", "1.5"]) +def test_a_rank_that_is_not_a_rank_names_rank_zero(monkeypatch, tmp_path, rank): + monkeypatch.setenv("RANK", rank) + engine, _events, services = _capture_engine(monkeypatch, tmp_path) + + engine.create_record_runtime(_record_format()) + + assert services[0].config["spool_root"].endswith("/r0-0a1b2c3d") + + +def test_the_shared_filesystem_override_reaches_the_lock_and_both_spools( + monkeypatch, tmp_path): + from dmi.storage.capture.native_sink import NativeSinkConfig + + engine, _events, services = _capture_engine(monkeypatch, tmp_path) + engine._capture_sink_config = NativeSinkConfig( + spool_root=str(tmp_path / "spool"), + spool_allow_shared_filesystem=True) + + engine.create_record_runtime(_record_format()) + + assert engine._test_locks[0].allow_shared_filesystem is True + assert services[0].config["spool_allow_shared_filesystem"] is True + assert engine._test_sinks[0]["allow_shared_filesystem"] is True + + +def test_a_service_that_fails_to_start_releases_the_spool_lock( + monkeypatch, tmp_path): + engine, events, _services = _capture_engine(monkeypatch, tmp_path, + fail_start=True) + + with pytest.raises(RuntimeError, match="lease held elsewhere"): + engine.create_record_runtime(_record_format()) + + assert events[-1] == ("lock", "release") + assert not engine._test_locks[0].held + assert engine._capture_storage is None + assert not any(event[0] == "sink" for event in events) def test_an_explicit_sink_leaves_the_spool_unswept(monkeypatch, tmp_path): @@ -570,7 +682,65 @@ def test_an_explicit_sink_leaves_the_spool_unswept(monkeypatch, tmp_path): engine.create_record_runtime( _record_format(), record_sink=dmi.transport.native.RecordSink()) - assert services[0].config["sweep_spool_on_start"] is False + config = services[0].config + assert config["sweep_spool_on_start"] is False + # An explicit sink writes where it was built to: the configured root, + # which the engine neither lays out nor adopts siblings around. Nothing + # holds it here, so the service takes its lock. + assert config["spool_root"] == str(tmp_path / "spool") + assert config["adopt_sibling_spools"] is False + assert config["spool_owner_lock"] == "take" + assert not any(event[0] == "lock" for event in events) + + +def test_an_explicit_sink_holding_the_spool_shares_it_with_the_service( + monkeypatch, tmp_path): + """A NativePackSink the caller built takes the spool's owner lock; the + engine's service in the same process must open beside it, not take it + again (that would be refused, naming this very process).""" + import os + import socket + + engine, _events, services = _capture_engine( + monkeypatch, tmp_path, + spool_owner={"host": socket.gethostname(), "pid": os.getpid()}) + import dmi.transport + + engine.create_record_runtime( + _record_format(), record_sink=dmi.transport.native.RecordSink()) + + assert services[0].config["spool_owner_lock"] == "held_by_caller" + + +def test_an_explicit_sink_on_a_spool_another_process_owns_is_taken( + monkeypatch, tmp_path): + """Owned by another process: the service takes it, and so is refused + by the native spool naming that holder.""" + engine, _events, services = _capture_engine( + monkeypatch, tmp_path, spool_owner={"host": "elsewhere", "pid": 1}) + import dmi.transport + + engine.create_record_runtime( + _record_format(), record_sink=dmi.transport.native.RecordSink()) + + assert services[0].config["spool_owner_lock"] == "take" + + +def test_a_sink_without_a_service_owns_its_spool_itself(monkeypatch, tmp_path): + """No capture_storage_config: packs stay in the spool for something else + to drain, and the sink is the directory's one owner -- no layout, no + engine lock.""" + engine, events, services = _capture_engine(monkeypatch, tmp_path, + storage=False) + + engine.create_record_runtime(_record_format()) + + assert services == [] + (sink,) = engine._test_sinks + assert sink["spool_root"] == str(tmp_path / "spool") + assert sink["owner_lock"] == "take" + assert sink["charge_dead_siblings"] is False # no layout, no siblings + assert not any(event[0] in ("lock", "layout") for event in events) def test_a_failed_attach_stops_the_service_it_started(monkeypatch, tmp_path): @@ -580,8 +750,9 @@ def test_a_failed_attach_stops_the_service_it_started(monkeypatch, tmp_path): with pytest.raises(RuntimeError, match="ring init failed"): engine.create_record_runtime(_record_format()) - assert ("service", "stop") in events + assert events[-2:] == [("service", "stop"), ("lock", "release")] assert engine._capture_storage is None + assert not engine._test_locks[0].held def test_flush_waits_for_the_catalog_after_the_sink(monkeypatch, tmp_path): @@ -615,31 +786,127 @@ def test_close_flushes_the_sink_before_the_ring_stops(monkeypatch, tmp_path): engine.close() + # The spool lock last: after the sink and the service are both done. assert [event[:2] for event in events] == [ ("sink", "flush"), ("ring", "stop"), - ("service", "flush"), ("service", "stop")] + ("service", "flush"), ("service", "stop"), ("lock", "release")] # One budget for the whole drain: the service gets what the sink left. assert 59.0 <= events[0][2] <= 60.0 assert 0.0 <= events[2][2] <= 60.0 assert engine._capture_storage is None -def test_close_still_stops_when_the_sink_flush_fails(monkeypatch, tmp_path): - engine, events, _services = _capture_engine(monkeypatch, tmp_path) - engine.create_record_runtime(_record_format()) - +def _fail_the_sink_flush(engine, events): def _failing_flush(timeout_s): events.append(("sink", "flush", timeout_s)) raise TimeoutError("timed out waiting for durable record completion") engine._ring_transport.flush_records_and_wait = _failing_flush + + +def test_close_still_stops_when_the_sink_flush_fails(monkeypatch, tmp_path, + caplog): + """The service still stops. The spool lock does not go: a sink that did + not seal may still be staging -- it outlives close() through the user's + RecordRuntime, and its stagers and destructor write into the directory + -- so another process's adoption (or this one's next engine) must not + take the directory from under it. The kernel lets go at exit.""" + from dmi.storage import native_capture + + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + _fail_the_sink_flush(engine, events) + (lock,) = engine._test_locks events.clear() - engine.close() + with caplog.at_level("WARNING", logger="dmi.engine"): + engine.close() assert [event[:2] for event in events] == [ ("sink", "flush"), ("ring", "stop"), ("service", "flush"), ("service", "stop")] + assert lock.held + assert engine._spool_claim is None + # Kept alive for the process, so no garbage collection lets go of it. + assert any(claim._lock is lock + for claim in native_capture._HELD_SPOOL_CLAIMS) + assert lock.directory in caplog.text + assert "stays owned" in caplog.text + + +def test_close_keeps_the_lock_when_the_release_backstop_did_not_seal( + monkeypatch, tmp_path, caplog): + """The sink's flush went through, but the ring stopping drains what it + still queued into the sink, and the release backstop that stages it + timed out (sealed_on_release false): a stage is still on its way into + the directory, so the lock stays, as for a sink that never sealed.""" + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + engine._record_sink.sealed_on_release = False + (lock,) = engine._test_locks + events.clear() + + with caplog.at_level("WARNING", logger="dmi.engine"): + engine.close() + + assert [event[:2] for event in events] == [ + ("sink", "flush"), ("ring", "stop"), + ("service", "flush"), ("service", "stop")] + assert lock.held + assert "stays owned" in caplog.text + + +def test_close_releases_the_lock_once_the_release_backstop_sealed_the_sink( + monkeypatch, tmp_path): + """The sink's flush ran out of close()'s budget, but its release from + the stopping ring flushed it (sealed_on_release true), and a released + sink admits nothing more: nothing can stage into the directory, so the + lock goes, as for a sink whose flush went through.""" + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + _fail_the_sink_flush(engine, events) + engine._record_sink.sealed_on_release = True + (lock,) = engine._test_locks + events.clear() + + engine.close() + + assert [event[:2] for event in events] == [ + ("sink", "flush"), ("ring", "stop"), + ("service", "flush"), ("service", "stop"), ("lock", "release")] + assert not lock.held + + +def test_replacing_a_record_ring_asks_the_released_sink_whether_it_sealed( + monkeypatch, tmp_path): + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + engine._record_sink.sealed_on_release = False + (lock,) = engine._test_locks + events.clear() + + engine.enable_ring_transport(object()) + + assert [event[:2] for event in events] == [ + ("sink", "flush"), ("ring", "stop"), + ("service", "flush"), ("service", "stop"), ("ring", "create")] + assert lock.held + + +def test_replacing_a_record_ring_keeps_the_lock_when_the_sink_did_not_seal( + monkeypatch, tmp_path): + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + _fail_the_sink_flush(engine, events) + (lock,) = engine._test_locks + events.clear() + + engine.enable_ring_transport(object()) + + assert [event[:2] for event in events] == [ + ("sink", "flush"), ("ring", "stop"), + ("service", "flush"), ("service", "stop"), ("ring", "create")] + assert lock.held def test_close_releases_the_lease_even_when_the_drain_fails(monkeypatch, tmp_path): @@ -650,7 +917,9 @@ def test_close_releases_the_lease_even_when_the_drain_fails(monkeypatch, tmp_pat engine.close() - assert events[-1] == ("service", "stop") + # Whatever did not drain stays in the directory for the next process on + # the node to adopt; the lock goes either way. + assert events[-2:] == [("service", "stop"), ("lock", "release")] def test_replacing_a_record_ring_drains_and_stops_the_service( @@ -667,16 +936,18 @@ def test_replacing_a_record_ring_drains_and_stops_the_service( assert [event[:2] for event in events] == [ ("sink", "flush"), ("ring", "stop"), - ("service", "flush"), ("service", "stop"), ("ring", "create")] + ("service", "flush"), ("service", "stop"), ("lock", "release"), + ("ring", "create")] assert 59.0 <= events[0][2] <= 60.0 assert engine._capture_storage is None assert engine._record_mode is False - # A second record runtime starts its own service; nothing still holds - # the lease it takes. + # A second record runtime starts its own service in a directory of its + # own; nothing still holds the lease it takes, or the lock. engine.create_record_runtime(_record_format()) assert len(services) == 2 assert engine._capture_storage is not None + assert [lock.held for lock in engine._test_locks] == [False, True] # --- the publisher lease knobs ------------------------------------------------- diff --git a/tests/test_native_live_spool.py b/tests/test_native_live_spool.py index 658a0256d..bf05cc3ec 100644 --- a/tests/test_native_live_spool.py +++ b/tests/test_native_live_spool.py @@ -1,13 +1,26 @@ -"""An upload scan must not run crash cleanup against a live pack writer.""" -import select +"""A live pack writer owns its spool: no other process may sweep it. + +Crash cleanup (Recover) deletes every ``.open`` file its Spool is not +writing itself, so running it against a live writer in another process +deletes the writer's in-flight temp file and fails its stage with "cannot +link ready file". The spool's owner lock makes that impossible rather than +merely avoided: while the writer holds its directory, a second process's +open is refused, naming the writer. Once the writer dies -- even by +SIGKILL, mid-stage -- the lock goes with it, and the next process recovers +the directory: the stale temp file is swept and the sealed pack uploads. +""" +import os import shutil +import signal +import socket import subprocess import sys +import threading from pathlib import Path import pytest -from tests.test_native_s3_client import fake_s3 +from tests.test_native_s3_client import STATE, fake_s3 from tests.test_native_uploader import DriverSession, STORE_DRIVER, _upload_pending pytestmark = [ @@ -16,8 +29,11 @@ pytest.mark.skipif(not STORE_DRIVER.exists(), reason="native store driver not built"), ] +SPOOL_DRIVER = STORE_DRIVER.parent / "conformance_spool" + -def test_upload_scan_preserves_an_inflight_stage(fake_s3, tmp_path): +@pytest.fixture +def writer_binary(tmp_path) -> Path: compiler = shutil.which("g++") if compiler is None: pytest.skip("a C++17 compiler is required") @@ -31,22 +47,120 @@ def test_upload_scan_preserves_an_inflight_stage(fake_s3, tmp_path): "-o", str(executable)], check=True, capture_output=True, text=True, ) - spool = tmp_path / "spool" + return executable + + +def _start_writer(binary: Path, spool: Path, packs: int = 1): + """A writer paused with its last pack's temp file open.""" process = subprocess.Popen( - [str(executable), str(spool)], stdin=subprocess.PIPE, + [str(binary), str(spool), str(packs)], stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True, ) + # Blocking reads under a watchdog: select() on the pipe cannot see a line + # the text wrapper has already buffered behind an earlier one. + watchdog = threading.Timer(20, process.kill) + watchdog.start() + seen = [] + try: + for line in process.stdout: + if line.strip() == "OPEN": + return process + seen.append(line.strip()) # an earlier pack's stage status + finally: + watchdog.cancel() + raise AssertionError( + f"writer never opened its temp file: {seen} {process.stderr.read()}") + + +def _recover(spool: Path, **fields) -> dict: + import json + + proc = subprocess.run( + [str(SPOOL_DRIVER)], + input=json.dumps({"op": "recover", "root": str(spool), + "max_bytes": 1 << 30, **fields}) + "\n", + capture_output=True, text=True, timeout=30) + return json.loads(proc.stdout.strip()) + + +def test_an_upload_scan_is_refused_while_a_writer_holds_the_spool( + fake_s3, tmp_path, writer_binary): + spool = tmp_path / "spool" + process = _start_writer(writer_binary, spool) uploader = DriverSession(STORE_DRIVER) try: - assert select.select([process.stdout], [], [], 10)[0], "writer did not open its temp file" - assert process.stdout.readline().strip() == "OPEN" result = _upload_pending(uploader, fake_s3, spool) + # Refused before it touched anything, naming the holder. + assert not result["ok"], result + assert f"pid {process.pid}" in result["what"], result + assert socket.gethostname() in result["what"], result + assert len(list(spool.rglob("*.open"))) == 1 + output, error = process.communicate("continue\n", timeout=10) - assert result["ok"], result assert process.returncode == 0, output + error assert len(list(spool.rglob("*.dmi-pack.ready"))) == 1 + assert list(spool.rglob("*.open")) == [] + + # The writer is gone, and its lock with it. + result = _upload_pending(uploader, fake_s3, spool) + assert result["ok"], result + assert len(result["refs"]) == 1, result + assert list(spool.rglob("*.dmi-pack.ready")) == [] finally: uploader.close() if process.poll() is None: process.kill() process.communicate() + + +def test_held_by_caller_cannot_sweep_a_live_writers_spool(tmp_path, writer_binary): + """owner_lock=held_by_caller opens without a lock of its own, for a + second Spool in the process that holds the directory. From any other + process it is refused, naming the writer: it used to pass on "something + holds the lock", and its Recover then deleted the writer's .open file, + failing the writer's stage with "cannot link ready file".""" + spool = tmp_path / "spool" + process = _start_writer(writer_binary, spool) + try: + result = _recover(spool, owner_lock="held_by_caller") + assert not result["ok"], result + assert "held_by_caller" in result["what"], result + assert f"pid {process.pid}" in result["what"], result + assert len(list(spool.rglob("*.open"))) == 1 + + output, error = process.communicate("continue\n", timeout=10) + assert process.returncode == 0, output + error + assert len(list(spool.rglob("*.dmi-pack.ready"))) == 1 + finally: + if process.poll() is None: + process.kill() + process.communicate() + + +def test_a_sigkilled_writers_spool_is_recovered_by_the_next_process( + fake_s3, tmp_path, writer_binary): + spool = tmp_path / "spool" + # The first pack is sealed; the second is mid-stage when the writer dies. + process = _start_writer(writer_binary, spool, packs=2) + assert len(list(spool.rglob("*.dmi-pack.ready"))) == 1 + (stale,) = spool.rglob("*.open") + os.kill(process.pid, signal.SIGKILL) + process.communicate(timeout=10) + assert process.returncode == -signal.SIGKILL + + recovered = _recover(spool) + assert recovered["ok"], recovered + assert len(recovered["staged"]) == 1, recovered + assert not stale.exists() + + uploader = DriverSession(STORE_DRIVER) + try: + result = _upload_pending(uploader, fake_s3, spool) + finally: + uploader.close() + assert result["ok"], result + assert [ref["pack_id"] for ref in result["refs"]] == [ + recovered["staged"][0]["pack_id"]] + assert list(spool.rglob("*.dmi-pack.ready")) == [] + with STATE.lock: + assert result["refs"][0]["object_key"] in STATE.objects diff --git a/tests/test_native_s3_client.py b/tests/test_native_s3_client.py index 52178e84c..5d50ec0a3 100644 --- a/tests/test_native_s3_client.py +++ b/tests/test_native_s3_client.py @@ -141,9 +141,22 @@ def _reject_unsigned(self, body: bytes) -> bool: return True return False - def _read_body(self) -> bytes: + def _read_body(self): + """The request's body, or None when it is not the one signed for: + shorter than its Content-Length (the client went away mid-body), or + not what x-amz-content-sha256 hashes. S3 stores neither + (IncompleteBody, XAmzContentSHA256Mismatch), and a body short of + its length passed the signature check here, which trusts that + header -- so a PUT cut off mid-body stored an empty object.""" length = int(self.headers.get("Content-Length", 0)) - return self.rfile.read(length) if length else b"" + body = self.rfile.read(length) if length else b"" + if len(body) < length: + return None + claimed = self.headers.get("x-amz-content-sha256", "") + if (len(claimed) == 64 + and hashlib.sha256(body).hexdigest() != claimed.lower()): + return None + return body def _split(self): from urllib.parse import unquote @@ -216,6 +229,15 @@ def _route(self): return key = segments[1] body = self._read_body() + if body is None: + # Nobody may be listening; the connection goes, and nothing is + # recorded or stored. + self.close_connection = True + try: + self._send(400, {}, b"incomplete or mismatched body") + except OSError: + pass + return self._record(body) if self._reject_unsigned(body): return @@ -987,3 +1009,78 @@ def test_an_uncancelled_request_is_untouched_by_the_cancel_hook(fake_s3): assert put["ok"], put assert put.get("cancelled") is False, put assert STATE.objects["packs/on-time"]["body"] == b"data" + + +def _signed_put_head(endpoint: str, key: str, body_hash: str, + length: int) -> bytes: + """A SigV4-signed PUT's head as the native client sends one, for a body + of `length` bytes that x-amz-content-sha256 says hashes to `body_hash`.""" + import datetime as datetime_module + + host = urlsplit(endpoint).netloc + credentials = botocore.credentials.Credentials(ACCESS, SECRET, None) + request = botocore.awsrequest.AWSRequest( + method="PUT", url=f"http://{host}/{BUCKET}/{key}", data=b"", + headers={"x-amz-content-sha256": body_hash}) + frozen = datetime_module.datetime.now(datetime_module.timezone.utc).replace( + microsecond=0, tzinfo=None) + with mock.patch.object( + botocore.auth, "get_current_datetime", return_value=frozen + ): + botocore.auth.SigV4Auth(credentials, "s3", REGION).add_auth(request) + head = f"PUT /{BUCKET}/{key} HTTP/1.1\r\nHost: {host}\r\n" + for name in ("x-amz-content-sha256", "X-Amz-Date", "Authorization"): + head += f"{name}: {request.headers[name]}\r\n" + head += f"Content-Length: {length}\r\n\r\n" + return head.encode() + + +def _send_raw(endpoint: str, request: bytes) -> bytes: + """Send `request`, say no more (shutdown for writing), and return the + answer, if any.""" + import socket + + parts = urlsplit(endpoint) + with socket.create_connection((parts.hostname, parts.port), + timeout=10) as connection: + connection.sendall(request) + connection.shutdown(socket.SHUT_WR) + answer = b"" + try: + while chunk := connection.recv(65536): + answer += chunk + except OSError: + pass + return answer + + +def test_the_fake_stores_no_put_that_is_not_the_body_it_signed_for(fake_s3): + """The harness itself. Its signature check takes the payload hash from + x-amz-content-sha256, as S3 does, so a PUT whose client went away after + the head -- a cancel landing between libcurl's head and its body -- + passed it with a short body and was stored, empty, over what a later + upload put at the key. S3 stores neither a body short of its + Content-Length nor one its x-amz-content-sha256 does not hash.""" + body = bytes(range(256)) * 8 + digest = hashlib.sha256(body).hexdigest() + + # The control: the same head with its whole body is stored. + answer = _send_raw(fake_s3, _signed_put_head( + fake_s3, "packs/whole", digest, len(body)) + body) + assert answer.startswith(b"HTTP/1.1 200"), answer + assert STATE.objects["packs/whole"]["body"] == body + + # The head, and half the body, then nothing more. + _send_raw(fake_s3, _signed_put_head( + fake_s3, "packs/short", digest, len(body)) + body[:1024]) + # The head alone. + _send_raw(fake_s3, _signed_put_head( + fake_s3, "packs/headless", digest, len(body))) + # A whole body, but not the one the head hashes. + answer = _send_raw(fake_s3, _signed_put_head( + fake_s3, "packs/other", digest, len(body)) + bytes(len(body))) + assert answer.startswith(b"HTTP/1.1 400"), answer + with STATE.lock: + assert set(STATE.objects) == {"packs/whole"}, sorted(STATE.objects) + assert [call["path"] for call in STATE.calls] == [ + f"/{BUCKET}/packs/whole"], STATE.calls diff --git a/tests/test_native_sink_release.py b/tests/test_native_sink_release.py index 84ac52fee..a6a9bc8de 100644 --- a/tests/test_native_sink_release.py +++ b/tests/test_native_sink_release.py @@ -237,6 +237,51 @@ def test_a_zero_release_timeout_leaves_the_open_pack_to_the_pipeline( del lease +def test_sealed_on_release_says_whether_the_sink_can_still_stage( + native_sink_module, tmp_path): + """The engine lets go of its spool directory's owner lock only once + nothing can stage into the directory any more. sealed_on_release says + so: true once a release's flush has persisted everything admitted -- + a released sink admits nothing -- and false while the sink is + attached, after a release whose flush timed out (a stage is still on + its way, and lands later), and with the backstop off.""" + sink = _binding_sink(native_sink_module, tmp_path / "sealed") + assert sink.sealed_on_release is False + lease = sink.attach() + sink.submit_envelope(LAYOUT, *_envelope(0)) + _wait_until_admitted(sink, 1) + assert sink.sealed_on_release is False # attached: it may take more + del lease + assert sink.sealed_on_release is True + assert len(_ready(tmp_path / "sealed")) == 1 + lease = sink.attach() + assert sink.sealed_on_release is False # attached again + del lease + assert sink.sealed_on_release is True # nothing left to flush + + wedged = _binding_sink(native_sink_module, tmp_path / "wedged", + release_flush_timeout_s=0.5) + release_stages = wedged._hold_stages_for_testing(10.0) + lease = wedged.attach() + wedged.submit_envelope(LAYOUT, *_envelope(1)) + _wait_until_admitted(wedged, 1) + del lease + assert wedged.sealed_on_release is False + assert _ready(tmp_path / "wedged") == [] + release_stages() # the stage the release gave up on still lands + deadline = time.monotonic() + 10.0 + while not _ready(tmp_path / "wedged"): + assert time.monotonic() < deadline, wedged.snapshot() + time.sleep(0.01) + assert wedged.sealed_on_release is False + + off = _binding_sink(native_sink_module, tmp_path / "off", + release_flush_timeout_s=0.0) + lease = off.attach() + del lease + assert off.sealed_on_release is False + + @pytest.mark.parametrize("timeout_s", [-1.0, float("nan"), float("inf")]) def test_the_release_timeout_must_be_finite_and_not_negative( native_sink_module, tmp_path, timeout_s): diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py new file mode 100644 index 000000000..fceccff69 --- /dev/null +++ b/tests/test_native_spool_adoption_live.py @@ -0,0 +1,1136 @@ +"""B6 live: a SIGKILLed capture process's spool is adopted by its successor. + +A capture process -- here a child running what the engine composes: one +SpoolOwnerLock on its own rank directory of the section 2.3 layout, the +storage service and the REAL native pack sink both opened held_by_caller -- +captures, gets part of it indexed, stages the rest, and is SIGKILLed with +packs still in its spool and the publisher lease still live. A new +incarnation on the same node and catalog, in a directory of its own, then +starts: it waits out the dead lease, sweeps its own directory, takes the +dead directory's owner lock (the kernel dropped it with the process), +sweeps its stale .open file, uploads and indexes every ready pack, and +removes the directory. Every capture of the dead process reads back from +the catalog with its bytes. A live sibling -- another process's spool, +lock held -- is left alone. + +When the object store is down at start, the dead directory stays as it +was (its packs are durable there), flush() does not report drained, and +the loop adopts it once the store is back. A sibling whose owner is still +alive at start and dies later is adopted by a later pass. stop() cuts an +adoption as it cuts the service's own work -- its uploads, and its listing +of a dead backlog -- and leaves the dead directory with every pack in it; +a flush waits for the adoption step in flight, not for a whole slice, nor +for a dead backlog's whole listing. + +Needs ClickHouse on 127.0.0.1:8123/9000 and the native sink and store +modules: make -C native build/_dmi_native_sink build/_dmi_native_store +PYTHON=/bin/python +""" + +from __future__ import annotations + +import json +import os +import signal +import subprocess +import sys +import threading +import time +from pathlib import Path + +import pytest + +# Module-level so the fake-S3 fixture registers in this module. +from tests.test_native_s3_client import ( # noqa: E402 + ACCESS, BUCKET, REGION, SECRET, STATE, fake_s3, +) + +REPO = Path(__file__).resolve().parents[1] +BUILD = REPO / "native" / "build" +SINK_BUILT = bool(sorted(BUILD.glob("_dmi_native_sink*.so"))) +STORE_BUILT = bool(sorted(BUILD.glob("_dmi_native_store*.so"))) + +pytestmark = [ + pytest.mark.manual, + pytest.mark.clickhouse, + pytest.mark.skipif( + not (SINK_BUILT and STORE_BUILT), + reason="the native sink and store modules are not built; run " + "`make -C native build/_dmi_native_sink build/_dmi_native_store " + "PYTHON=/bin/python`", + ), +] + +CLICKHOUSE_HOST = os.environ.get("DMI_CLICKHOUSE_HOST", "127.0.0.1") +CLICKHOUSE_HTTP_PORT = int(os.environ.get("DMI_CLICKHOUSE_HTTP_PORT", "8123")) +DATABASE = os.environ.get("DMI_CLICKHOUSE_DATABASE", "default") +LAYOUT = "capture_pack_reference_v1" +# Short, so the successor's wait for the dead process's lease stays short. +LEASE = dict(lease_ttl_s=3.0, publish_timeout_s=1.0) +INDEXED_BY_THE_DEAD = range(0, 4) # flushed to the catalog before the kill +STAGED_BY_THE_DEAD = range(4, 10) # only in its spool when it dies +STAGED_BY_THE_LIVE = range(20, 24) # a live sibling's, never adopted +RECORDS_PER_PACK = 2 + + +def _storage_config(endpoint: str, prefix: str, **overrides): + from dmi.storage.native_capture import NativeCaptureStorageConfig + + fields = dict( + s3_endpoint=endpoint, s3_bucket=BUCKET, s3_region=REGION, + s3_access_key=ACCESS, s3_secret_key=SECRET, + s3_allow_insecure_http=True, clickhouse_host=CLICKHOUSE_HOST, + clickhouse_port=CLICKHOUSE_HTTP_PORT, database=DATABASE, + table_prefix=prefix, poll_interval_s=0.05, **LEASE) + fields.update(overrides) + return NativeCaptureStorageConfig(**fields) + + +def _store(): + from dmi.storage.native_capture import _load_native_store_extension + + return _load_native_store_extension() + + +def _claim(base: Path, config): + """What the engine does first: a fresh rank directory, locked.""" + store = _store() + directory = store.spool_rank_directory( + str(base), config._spool_destination(), 0) + return store.SpoolOwnerLock(directory) + + +def _service(config, directory: str, **options): + from dmi.storage.native_capture import NativeCaptureStorage + + return NativeCaptureStorage( + config, spool_root=directory, spool_max_bytes=1 << 30, + sweep_spool=True, spool_owner_lock="held_by_caller", + adopt_sibling_spools=True, **options) + + +def _envelope(indexes, width: int = 6): + from tests.test_native_capture_chain_live import _Envelope + + import torch + + envelope = _Envelope() + for index in indexes: + envelope.add(index, torch.arange(width, dtype=torch.float16) + index) + return envelope + + +def _dead_capture_process(base: str, endpoint: str, prefix: str) -> None: + """The child: capture, index some, stage the rest, then wait to die.""" + import torch # noqa: F401 -- the sink extension links against it + + sys.path.insert(0, str(BUILD)) + import _dmi_native_sink + + # The loop must not upload the second batch before the kill: with this + # poll interval only the explicit flush below runs a cycle. The lease + # thread renews meanwhile, so the process dies holding the lease. + config = _storage_config(endpoint, prefix, poll_interval_s=3600.0) + lock = _claim(Path(base), config) + service = _service(config, lock.directory) + service.start() + sink = _dmi_native_sink.NativePackSink( + spool_root=lock.directory, layout=LAYOUT, + max_pack_records=RECORDS_PER_PACK, max_linger_ns=600 * 10**9, + owner_lock="held_by_caller") + _lease = sink.attach() + first = _envelope(INDEXED_BY_THE_DEAD) + sink.submit_envelope(LAYOUT, first.rows, first.payload()) + assert sink.flush_and_wait(60.0) + service.flush(60.0) + second = _envelope(STAGED_BY_THE_DEAD) + sink.submit_envelope(LAYOUT, second.rows, second.payload()) + assert sink.flush_and_wait(60.0) + sink.rethrow_if_failed() + print(json.dumps({"directory": lock.directory}), flush=True) + time.sleep(3600) + + +def _spawn_dead_process(base: Path, endpoint: str, prefix: str): + env = dict(os.environ, PYTHONPATH=str(REPO / "src") + os.pathsep + str(REPO), + CUDA_VISIBLE_DEVICES="") + child = subprocess.Popen( + [sys.executable, "-m", "tests.test_native_spool_adoption_live", + str(base), endpoint, prefix], + cwd=str(REPO), env=env, stdout=subprocess.PIPE, stderr=subprocess.PIPE, + text=True) + line = child.stdout.readline() + if not line: + child.wait(timeout=30) + raise AssertionError(f"the capture process failed: {child.stderr.read()}") + return child, Path(json.loads(line)["directory"]) + + +def _sigkill(child) -> None: + os.kill(child.pid, signal.SIGKILL) + child.communicate(timeout=30) + assert child.returncode == -signal.SIGKILL + + +def _stale_open_file(directory: Path) -> Path: + """A stage the dead process had in flight: its temp file, left behind.""" + (ready,) = sorted(directory.rglob("*.dmi-pack.ready"))[:1] + stale = ready.parent / ".018f0000-0000-7000-8000-00000000dead.0badf00d.open" + stale.write_bytes(b"half a pack") + return stale + + +def _claim_staging(parent: Path, name: str, *, age_s: float) -> Path: + """The staging copy a claim killed before its rename leaves behind.""" + staging = parent / f".{name}.0badf00d.creating" + staging.mkdir() + (staging / ".owner.lock").write_text("host 1\n") + then = time.time() - age_s + os.utime(staging, (then, then)) + return staging + + +def _wait_for(predicate, timeout_s: float = 30.0) -> None: + deadline = time.monotonic() + timeout_s + while not predicate(): + if time.monotonic() >= deadline: + raise AssertionError("timed out waiting for the service") + time.sleep(0.05) + + +def _adopted(service, spools: int): + """Whether `spools` dead siblings have been adopted and nothing more is + to adopt, as of the last cycle's end.""" + def done(): + snapshot = service.snapshot() + return (snapshot["adopted_spools"] == spools + and not snapshot["adoption_owed"]) + return done + + +def _read_all(config) -> dict: + from dmi.storage.native_capture import NativeCaptureReader + + reader = NativeCaptureReader(config) + selection = reader.select(tenant_id="t") + return {capture.descriptor["capture_id"]: capture.payload + for capture in reader.read(selection, byte_limit=1 << 24)} + + +def _expected() -> dict: + expected = {} + for indexes in (INDEXED_BY_THE_DEAD, STAGED_BY_THE_DEAD): + envelope = _envelope(indexes) + for capture_id, tensor in envelope.expected.items(): + expected[capture_id] = tensor.contiguous().view(-1).numpy().tobytes() + return expected + + +def _catalog(): + from tests.test_native_capture_chain_live import _catalog as chain_catalog + + return chain_catalog() + + +def test_a_sigkilled_process_spool_is_adopted_by_its_successor( + fake_s3, tmp_path): + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix) + child, dead = _spawn_dead_process(base, fake_s3, prefix) + try: + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert len(staged) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + finally: + _sigkill(child) + stale = _stale_open_file(dead) + assert _store().spool_owner(str(dead)) is None # died with it + # A claim killed before it renamed its directory into place leaves + # the staging copy (lock file inside). An old one is cleared; a + # fresh one may be a claim in progress, and is left alone. + old_claim = _claim_staging(dead.parent, "r0-0badf00d", age_s=3600) + fresh_claim = _claim_staging(dead.parent, "r0-00c0ffee", age_s=0) + + # A live sibling bound for the same catalog: ready packs its sink + # staged, and a stage it has in flight. Its lock is held by this + # process, so an adopter that took it for dead would get past the + # held_by_caller check and sweep it; only its liveness saves it. + live = _claim(base, config) + live_directory = Path(live.directory) + _stage_into(live.directory, STAGED_BY_THE_LIVE) + live_packs = sorted(live_directory.rglob("*.dmi-pack.ready")) + assert len(live_packs) == len(STAGED_BY_THE_LIVE) // RECORDS_PER_PACK + live_open = live_packs[0].parent / ( + ".018f0000-0000-7000-8000-00000000beef.0badf00d.open") + live_open.write_bytes(b"half a pack") + + lock = _claim(base, config) + service = _service(config, lock.directory) + started = time.monotonic() + service.start() # waits out the dead lease; the loop adopts + try: + _wait_for(_adopted(service, 1)) + snapshot = service.snapshot() + assert time.monotonic() - started < 30 + assert snapshot["adopted_spools"] == 1, snapshot + assert snapshot["adopted_packs"] == len(staged), snapshot + assert snapshot["adoption_owed"] is False, snapshot + assert not stale.exists() + assert not dead.exists() + assert not old_claim.exists() + assert fresh_claim.exists() + assert snapshot["live_siblings"] == 1, snapshot + service.flush(60.0) + # Only the dead process's captures reach the catalog. + assert _read_all(config) == _expected() + finally: + service.stop() + # The live sibling is as it was: nothing swept, nothing uploaded. + assert sorted(live_directory.rglob("*.dmi-pack.ready")) == live_packs + assert live_open.read_bytes() == b"half a pack" + live_ids = {path.name.split(".")[0] for path in live_packs} + with STATE.lock: + uploaded = list(STATE.objects) + assert not [key for key in uploaded + if any(pack_id in key for pack_id in live_ids)] + assert live.held + assert lock.release_and_remove_if_empty() + live.release() + + +def test_a_dead_spool_waits_in_place_while_the_object_store_is_down( + fake_s3, tmp_path): + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + # Both processes reach the store through the switch: the endpoint is + # part of the catalog key, so a successor spelling it differently + # would not see the dead directory as its sibling. + switch = _Switch.to_url(fake_s3) + child, dead = _spawn_dead_process(base, switch.url, prefix) + _sigkill(child) + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert staged + + # Let the dead process's lease lapse, so start() has none to wait + # out and its time is its own. + time.sleep(LEASE["lease_ttl_s"] + 1.0) + switch.cut() + config = _storage_config(switch.url, prefix) + lock = _claim(base, config) + service = _service(config, lock.directory) + started = time.monotonic() + service.start() + try: + # start() leaves the dead backlog to the loop: it does not sit + # through every dead pack's retry chain against a store that + # refuses them (about 8 s a round of four packs). + assert time.monotonic() - started < 3.0 + # Nor does flush(), which is about this process's own records: + # it returns within its deadline, drained or not. + flushed = time.monotonic() + try: + service.flush(0.5) + except TimeoutError: + pass # a loop cycle still in an upload round holds the cycle + assert time.monotonic() - flushed < 2.0 + _wait_for(lambda: service.snapshot()["upload_failures"] > 0) + snapshot = service.snapshot() + assert snapshot["adoption_owed"] is True, snapshot + assert snapshot["adopted_spools"] == 0, snapshot + # Nothing left the dead spool: its packs are durable there. + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + + switch.restore() + _wait_for(_adopted(service, 1), 90.0) + snapshot = service.snapshot() + assert snapshot["adoption_owed"] is False, snapshot + assert not dead.exists() + service.flush(60.0) + assert _read_all(config) == _expected() + finally: + service.stop() + lock.release_and_remove_if_empty() + + +def _stage_into(directory: str, indexes, width: int = 6) -> None: + """Stage records into a directory this process holds, as its own sink + would: the REAL native pack sink, held_by_caller.""" + import torch # noqa: F401 -- the sink extension links against it + + sys.path.insert(0, str(BUILD)) + try: + import _dmi_native_sink + finally: + sys.path.remove(str(BUILD)) + sink = _dmi_native_sink.NativePackSink( + spool_root=directory, layout=LAYOUT, + max_pack_records=RECORDS_PER_PACK, max_linger_ns=600 * 10**9, + owner_lock="held_by_caller") + lease = sink.attach() + envelope = _envelope(indexes, width) + sink.submit_envelope(LAYOUT, envelope.rows, envelope.payload()) + assert sink.flush_and_wait(60.0) + sink.rethrow_if_failed() + del lease, sink + + +def test_no_dead_spool_is_uploaded_while_an_adopted_pack_is_owed( + fake_s3, tmp_path): + """Owner decision 7, inside adoption. Two dead siblings; the catalog + goes away just after the first one's packs reach the object store, so + they cannot be indexed and are owed in memory. The second sibling's + packs must then stay in its spool, where a crash cannot lose them -- + not be uploaded into a list only this process remembers. Once the + catalog is back, both are indexed.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + # The servers are part of the catalog key, so the siblings are + # claimed through the same switch URLs as the service reaches. + clickhouse = _Switch(CLICKHOUSE_HOST, CLICKHOUSE_HTTP_PORT) + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, + clickhouse_host="127.0.0.1", + clickhouse_port=clickhouse.port) + halves = (range(0, 4), range(4, 10)) + siblings = [] + for indexes in halves: + sibling = _claim(base, config) + siblings.append(Path(sibling.directory)) + _stage_into(sibling.directory, indexes) + sibling.release() # its owner is gone + first, second = sorted(siblings) # adopted in this order + second_packs = sorted(second.rglob("*.dmi-pack.ready")) + assert second_packs + + cut = threading.Event() + + def _cut_the_catalog_at_the_first_upload(request: bytes) -> float: + if request.startswith(b"PUT ") and not cut.is_set(): + cut.set() + clickhouse.cut() + return 0.0 + + store.delay_requests(_cut_the_catalog_at_the_first_upload) + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + _wait_for(lambda: service.snapshot()["pending_index"] > 0) + snapshot = service.snapshot() + assert cut.is_set() + assert not sorted(first.rglob("*.dmi-pack.ready")), snapshot + # Nothing of the second left its spool while the first's packs + # were owed. + assert sorted(second.rglob("*.dmi-pack.ready")) == second_packs + assert snapshot["adoption_owed"] is True, snapshot + with pytest.raises(TimeoutError): + service.flush(2.0) + assert sorted(second.rglob("*.dmi-pack.ready")) == second_packs + + clickhouse.restore() + service.flush(60.0) + _wait_for(_adopted(service, 2), 60.0) + snapshot = service.snapshot() + assert snapshot["adoption_owed"] is False, snapshot + assert not first.exists() and not second.exists() + service.flush(60.0) + expected = {} + for indexes in halves: + for capture_id, tensor in _envelope(indexes).expected.items(): + expected[capture_id] = ( + tensor.contiguous().view(-1).numpy().tobytes()) + assert _read_all(config) == expected + finally: + service.stop() + lock.release_and_remove_if_empty() + + +def test_a_dead_spool_this_service_can_never_upload_is_left_and_reported( + fake_s3, tmp_path): + """A pack this service can never upload -- here larger than its + uploader_max_in_flight_bytes, which a crashed run with a larger bound + left behind -- blocks its dead directory for good: retrying it only + re-hashes it and keeps the loop backing off. It is reported once and + left in place for a process that can (or a person), the siblings after + it are adopted, and nothing is owed or retried.""" + base = tmp_path / "spool" + with _catalog() as prefix: + # A 4096-byte bound: the wide captures' packs exceed it, the narrow + # ones' fit. + config = _storage_config(fake_s3, prefix, + uploader_max_in_flight_bytes=4096) + directories = [] + for indexes, width in ((range(0, 4), 4096), (range(4, 10), 6)): + sibling = _claim(base, config) + directories.append(Path(sibling.directory)) + _stage_into(sibling.directory, indexes, width) + sibling.release() # its owner is gone + blocked, adoptable = directories + blocked_packs = sorted(blocked.rglob("*.dmi-pack.ready")) + assert all(path.stat().st_size > 4096 for path in blocked_packs) + + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + _wait_for(_adopted(service, 1)) + snapshot = service.snapshot() + assert snapshot["blocked_siblings"] == [str(blocked)], snapshot + assert "in-flight byte limit" in snapshot["last_error"], snapshot + assert sorted(blocked.rglob("*.dmi-pack.ready")) == blocked_packs + # Let go of, and marked in its lock file -- for a person, and for + # the sinks on the node, which charge a dead directory against + # their budgets only while an adoption can drain it. + assert _store().spool_owner(str(blocked)) is None + record = (blocked / ".owner.lock").read_text() + assert ("\nblocked: it holds a pack this service can never " + "upload") in record, record + assert "in-flight byte limit" in record, record + assert not adoptable.exists() + # Not retried: no more failed uploads, and no backoff -- the + # loop keeps its poll interval. + failures, cycles = (snapshot["upload_failures"], + snapshot["cycles"]) + time.sleep(1.0) + snapshot = service.snapshot() + assert snapshot["upload_failures"] == failures, snapshot + assert snapshot["cycles"] >= cycles + 5, snapshot + assert snapshot["adoption_owed"] is False, snapshot + started = time.monotonic() + service.flush(10.0) + assert time.monotonic() - started < 2.0 + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in _envelope( + range(4, 10)).expected.items()} + assert _read_all(config) == expected + finally: + service.stop() + lock.release_and_remove_if_empty() + + +def test_packs_left_outside_the_layout_are_adopted_once_moved_as_told( + fake_s3, tmp_path, caplog): + """A sink-only run (or an engine from before the layout) leaves its + packs in spool_root itself, where nothing adopts them. The claim's + warning names a directory of the layout to move them into; moved there + with their paths kept, they are adopted like a dead process's.""" + import logging + import shutil + + import torch # noqa: F401 -- the sink extension links against it + + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix) + sys.path.insert(0, str(BUILD)) + try: + import _dmi_native_sink + finally: + sys.path.remove(str(BUILD)) + sink = _dmi_native_sink.NativePackSink( + spool_root=str(base), layout=LAYOUT, + max_pack_records=RECORDS_PER_PACK, max_linger_ns=600 * 10**9) + lease = sink.attach() + envelope = _envelope(STAGED_BY_THE_DEAD) + sink.submit_envelope(LAYOUT, envelope.rows, envelope.payload()) + assert sink.flush_and_wait(60.0) + del lease, sink + flat = sorted((base / "v1").rglob("*.dmi-pack.ready")) + assert len(flat) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + + with caplog.at_level(logging.WARNING, + logger="dmi.storage.native_capture"): + claim = claim_spool_directory( + NativeSinkConfig(spool_root=str(base)), config) + (warning,) = [record.getMessage() for record in caplog.records + if "outside" in record.getMessage()] + key = Path(claim.directory).parent.name + target = base / key / "r0-00000000" + assert f"{len(flat)} ready pack(s)" in warning + assert str(target) in warning + + target.mkdir() + shutil.move(str(base / "v1"), str(target / "v1")) + service = _service(config, claim.directory) + service.start() + try: + _wait_for(_adopted(service, 1)) + assert not target.exists() + service.flush(60.0) + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in envelope.expected.items()} + assert _read_all(config) == expected + finally: + service.stop() + claim.release() + + +def test_a_sibling_whose_owner_dies_after_start_is_adopted_by_a_recheck( + fake_s3, tmp_path): + """A sibling still owned when the service starts -- a predecessor still + inside its close(), another rank that dies later -- is left alone then. + The service looks again every adoption_recheck_interval_ns while it + has such a sibling, and adopts it once its owner is gone, rather than + leaving its packs for the next restart on the node.""" + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix) + sibling = _claim(base, config) + sibling_directory = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + staged = sorted(sibling_directory.rglob("*.dmi-pack.ready")) + assert len(staged) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + + lock = _claim(base, config) + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="recheck-test", poll_interval_ns=50_000_000, + sweep_spool_on_start=True, spool_owner_lock="held_by_caller", + adopt_sibling_spools=True, + adoption_recheck_interval_ns=200_000_000, + **config._lease_native()) + service = _store().StorageService(native) + service.start() + try: + _wait_for(lambda: not service.snapshot()["adoption_owed"]) + snapshot = service.snapshot() + assert snapshot["live_siblings"] == 1, snapshot + assert snapshot["adopted_spools"] == 0, snapshot + assert snapshot["adoption_owed"] is False, snapshot + assert sorted(sibling_directory.rglob( + "*.dmi-pack.ready")) == staged + + sibling.release() # its owner is gone + deadline = time.monotonic() + 30 + while ((service.snapshot()["adopted_spools"] == 0 + or service.snapshot()["live_siblings"] != 0) + and time.monotonic() < deadline): + time.sleep(0.05) + snapshot = service.snapshot() + assert snapshot["adopted_spools"] == 1, snapshot + assert snapshot["adopted_packs"] == len(staged), snapshot + assert snapshot["live_siblings"] == 0, snapshot + assert not sibling_directory.exists() + assert service.flush(60.0) + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in _envelope( + STAGED_BY_THE_DEAD).expected.items()} + assert _read_all(config) == expected + finally: + service.stop() + lock.release_and_remove_if_empty() + + +def test_flush_adopts_nothing_while_the_loop_is_idle(fake_s3, tmp_path): + """flush() covers this process's records, and its cycles never adopt: + a dead backlog is the loop's. With the object store down, a flush that + adopted sat through a dead pack's retry chain, past its deadline by a + round (minutes against a stalled store). Here the loop is idle -- its + first round has failed and its poll interval is an hour -- so the flush + holds the cycle itself and would be seen adopting: uploading (and + failing) dead packs and overrunning its deadline.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, poll_interval_s=3600.0) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + sibling.release() # its owner is gone + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert staged + store.cut() + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + # The loop's first cycle, which start() kicks, begins adopting + # and fails its first round; then it waits out its interval. + _wait_for(lambda: service.snapshot()["upload_failures"] > 0 + and service.snapshot()["cycles"] >= 1, 60.0) + before = service.snapshot() + assert before["adoption_owed"] is True, before + flushed = time.monotonic() + service.flush(0.5) # nothing of its own: drained + assert time.monotonic() - flushed < 2.0 + after = service.snapshot() + assert after["upload_failures"] == before["upload_failures"], after + assert after["adopted_packs"] == 0, after + assert after["adoption_owed"] is True, after + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + finally: + service.stop() + lock.release_and_remove_if_empty() + + +def test_stop_cuts_an_adoption_stalled_on_its_uploads(fake_s3, tmp_path): + """An adoption uploads through the service's own upload client and + Cancellation, so stop() cuts it as it cuts the service's own uploads. + Here every PUT is held open and never answered: before, the adoption's + uploader went round a pack's retries on a client stop() did not cancel, + and stop() sat through read timeouts. Now stop() returns at once, every + pack of the dead spool is still in it (none was uploaded, none lost), + and the directory is left, unlocked, for the next incarnation -- which + adopts it. The uploader is allowed eight attempts: an adoption uploader + whose own retries and backoff stop() did not cut, on a client stop() + did, would sleep out about 26 s of backoff (four attempts' 1.75 s fit + under the bound).""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, s3_read_timeout_s=30, + reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + sibling.release() # its owner is gone + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert len(staged) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + store.stall_requests(lambda request: request.startswith(b"PUT ")) + lock = _claim(base, config) + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="adoption-stall-stop", poll_interval_ns=50_000_000, + sweep_spool_on_start=True, reconcile_on_start=False, + spool_owner_lock="held_by_caller", adopt_sibling_spools=True, + uploader_max_attempts=8, **config._lease_native()) + service = _store().StorageService(native) + service.start() + try: + _wait_for(lambda: store.stalled, 30.0) # an adopted PUT in flight + stopping = time.monotonic() + service.stop() + assert time.monotonic() - stopping < 5.0 + snapshot = service.snapshot() + assert snapshot["cancelled_uploads"] >= 1, snapshot + assert snapshot["upload_failures"] == 0, snapshot + assert snapshot["adopted_packs"] == 0, snapshot + assert "adopting dead spool" not in snapshot["last_error"], snapshot + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + assert _store().spool_owner(str(dead)) is None + finally: + service.stop() + lock.release_and_remove_if_empty() + + store.restore() + successor_lock = _claim(base, config) + successor = _service(config, successor_lock.directory) + successor.start() + try: + _wait_for(_adopted(successor, 1), 60.0) + assert not dead.exists() + successor.flush(60.0) + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in _envelope( + STAGED_BY_THE_DEAD).expected.items()} + assert _read_all(config) == expected + finally: + successor.stop() + store.close() + successor_lock.release_and_remove_if_empty() + + +def _sparse_backlog(directory: Path, count: int, size: int) -> list[Path]: + """`count` ready packs of `size` zero bytes, sparse and named for their + checksum: validating them hashes every byte, which takes seconds, while + they take no disk.""" + import hashlib + import uuid + + digest = hashlib.sha256() + zeros = bytes(1 << 20) + for _ in range(size // len(zeros)): + digest.update(zeros) + packs = directory / "v1" + packs.mkdir(parents=True, exist_ok=True) + backlog = [] + for _ in range(count): + ready = packs / f"{uuid.uuid4()}.1.1.{digest.hexdigest()}.dmi-pack.ready" + with open(ready, "wb") as sparse: + sparse.truncate(size) + backlog.append(ready) + return sorted(backlog) + + +def test_stop_cuts_an_adoption_listing_a_dead_backlog(fake_s3, tmp_path): + """An adoption lists a dead spool once, validating every pack, which + over a backlog takes seconds: 32 sparse packs of 256 MiB, here. That + listing stops between packs once stop() cancels the uploads, as the + service's own listing does, so stop() does not wait for the rest of + it; the directory keeps every pack, and nothing of it was uploaded.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + sibling.release() # its owner is gone + backlog = _sparse_backlog(dead, 32, 256 << 20) + # Were the listing not cut, no pack may reach the fake store's + # memory: every PUT is held back, and stop() cuts those. + store.stall_requests(lambda request: request.startswith(b"PUT ")) + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + # The loop's first cycle, which start() kicks, begins the + # listing; the whole of it takes several seconds. + _wait_for(lambda: _store().spool_owner(str(dead)) is not None, + 10.0) + time.sleep(0.3) + stopping = time.monotonic() + service.stop() + elapsed = time.monotonic() - stopping + assert elapsed < 2.5, elapsed + snapshot = service.snapshot() + assert snapshot["adopted_packs"] == 0, snapshot + assert store.stalled == [], store.stalled + assert sorted(dead.rglob("*.dmi-pack.ready")) == backlog + assert _store().spool_owner(str(dead)) is None + finally: + service.stop() + store.close() + lock.release_and_remove_if_empty() + + +def _no_tenant_in_the_path(request: bytes) -> bool: + """A PUT of a pack outside the capture layout: the sparse backlog's.""" + line = request.split(b"\r\n", 1)[0] + return line.startswith(b"PUT ") and b"tenant" not in line + + +def test_a_flush_does_not_wait_for_an_adoption_to_list_a_dead_backlog( + fake_s3, tmp_path): + """An adoption validates a dead spool's packs before it uploads any, + hashing every byte: over a backlog -- an object-store outage's, up to + the dead spool's own budget -- for as long as the backlog is big. A + flush of this process's own records waited for all of it, since the + listing was one step of the adoption, and timed out with its own pack + still staged: engine.close() then reported capture storage undrained. + The listing now validates one pack per step, and a flush waits for the + one in flight. Here the backlog is 512 sparse packs of 64 MiB, which + take their listing 20 s and more; the flush drains the pack staged + meanwhile well inside its budget, with the listing still going.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + sibling.release() # its owner is gone + backlog = _sparse_backlog(dead, 512, 64 << 20) + # Were the listing to end, its packs would go no further than the + # switch: the service's own go through. + store.stall_requests(_no_tenant_in_the_path) + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + # The loop's first cycle, which start() kicks, begins listing. + _wait_for(lambda: _store().spool_owner(str(dead)) is not None, + 10.0) + time.sleep(0.3) + _stage_into(lock.directory, range(0, 2)) # as close() would + flushing = time.monotonic() + service.flush(8.0) # raises TimeoutError when not drained + elapsed = time.monotonic() - flushing + snapshot = service.snapshot() + assert snapshot["uploaded_packs"] == 1, snapshot + assert snapshot["indexed_packs"] == 1, snapshot + assert not sorted(Path(lock.directory).rglob("*.dmi-pack.ready")) + # Drained with the listing still going: the flush did not wait + # it out. + assert snapshot["adopted_packs"] == 0, snapshot + assert snapshot["adoption_owed"] is True, snapshot + owner = _store().spool_owner(str(dead)) + assert owner is not None and owner["pid"] == os.getpid(), ( + owner, elapsed) + assert store.stalled == [], store.stalled + stopping = time.monotonic() + service.stop() + assert time.monotonic() - stopping < 5.0 + assert sorted(dead.rglob("*.dmi-pack.ready")) == backlog + assert _store().spool_owner(str(dead)) is None + finally: + service.stop() + store.close() + lock.release_and_remove_if_empty() + + +def test_a_flush_waits_for_the_adoption_step_in_flight_not_its_slice( + fake_s3, tmp_path): + """flush() never adopts, and it does not wait out the loop's adoption + either: a cycle adopting lets go of the cycle at its next step while a + flush is running. Here each adopted PUT takes a second and the slice is + a minute, so a cycle would hold the cycle for the whole dead backlog, + five rounds; a flush(4.0) of a process with nothing of its own used to + time out behind it. It now returns drained after the round in flight, + and the adoption goes on afterwards.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + backlog = range(100, 140) # 20 packs: five rounds of four + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, backlog) + sibling.release() # its owner is gone + assert len(sorted(dead.rglob("*.dmi-pack.ready"))) == 20 + store.delay_requests( + lambda request: 1.0 if request.startswith(b"PUT ") else 0.0) + lock = _claim(base, config) + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="flush-yield-test", poll_interval_ns=50_000_000, + sweep_spool_on_start=True, reconcile_on_start=False, + spool_owner_lock="held_by_caller", adopt_sibling_spools=True, + adoption_slice_ns=60_000_000_000, **config._lease_native()) + service = _store().StorageService(native) + service.start() + try: + _wait_for(lambda: service.snapshot()["adopted_packs"] >= 4, 30.0) + flushing = time.monotonic() + assert service.flush(4.0) + elapsed = time.monotonic() - flushing + assert elapsed < 2.5, elapsed + snapshot = service.snapshot() + assert snapshot["adopted_packs"] < 20, snapshot + assert snapshot["adoption_owed"] is True, snapshot + + store.restore() + _wait_for(_adopted(service, 1), 60.0) + assert not dead.exists() + assert service.flush(60.0) + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in _envelope(backlog).expected.items()} + assert _read_all(config) == expected + finally: + service.stop() + store.close() + lock.release_and_remove_if_empty() + + +def test_flushes_that_keep_coming_slow_an_adoption_down_but_never_stop_it( + fake_s3, tmp_path): + """A cycle adopting lets go of the cycle at its next step while a flush + is waiting for it -- but only after one step at least, or flushes that + keep coming, as a capture loop's do, would stop adoption altogether: + every adopting cycle would find one waiting and take no step. Here two + threads flush back to back while a sibling that was alive at start + dies; it is still adopted, a step a cycle. + + Every cycle -- a flush's and the loop's -- first lists the service's + own spool, which here hashes a 32 MiB pack that fails validation and + cannot be set aside (a directory holds its quarantine name). So by the + time the loop's cycle comes to adopt, a flush is waiting for it. + Python threads flushing on their own leave gaps between their flushes, + and a cycle that took no step while a flush waited still adopted + through those gaps.""" + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert len(staged) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + lock = _claim(base, config) + (slow,) = _sparse_backlog(Path(lock.directory) / "slow", 1, 32 << 20) + with open(slow, "r+b") as corrupt: + corrupt.write(b"x") # no longer the zeros its name hashes + slow.with_name(slow.name[:-len(".ready")] + ".quarantined").mkdir() + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="flush-storm-test", poll_interval_ns=50_000_000, + sweep_spool_on_start=True, reconcile_on_start=False, + spool_owner_lock="held_by_caller", adopt_sibling_spools=True, + adoption_recheck_interval_ns=100_000_000, + **config._lease_native()) + service = _store().StorageService(native) + done = threading.Event() + flushed = [0, 0] + failures = [] + + def _flush_back_to_back(index: int) -> None: + try: + while not done.is_set(): + service.flush(10.0) + flushed[index] += 1 + # Not straight back for the lock the cycle just let go + # of: the loop, woken for it, takes its turn. + time.sleep(0.005) + except Exception as exc: # noqa: BLE001 -- reported below + failures.append(exc) + + flushers = [threading.Thread(target=_flush_back_to_back, args=(i,), + daemon=True) for i in range(2)] + service.start() + try: + # The first look finds the sibling alive, and leaves it. + _wait_for(lambda: not service.snapshot()["adoption_owed"]) + assert service.snapshot()["live_siblings"] == 1 + for flusher in flushers: + flusher.start() + _wait_for(lambda: min(flushed) >= 3) + sibling.release() # its owner is gone, the flushes still coming + before = list(flushed) + _wait_for(_adopted(service, 1), 60.0) + snapshot = service.snapshot() + # The flushes kept coming all the while. + assert all(now > then for now, then in zip(flushed, before)), ( + flushed, before) + assert not failures, failures + assert snapshot["adopted_packs"] == len(staged), snapshot + assert not dead.exists() + assert slow.exists() # still listed, and still refused + finally: + done.set() + for flusher in flushers: + if flusher.ident is not None: # started + flusher.join(timeout=60.0) + service.stop() + lock.release_and_remove_if_empty() + + +def test_a_cycle_adopts_nothing_while_its_own_packs_are_not_all_up( + fake_s3, tmp_path): + """A cycle adopts only once every pack this process staged has gone + up: its own records come first. Here the service's own spool holds a + pack the store refuses every time (its key is under the fake store's + fault/always-500/), so each cycle's own upload fails; the dead sibling + beside it waits, every pack in place, however many cycles pass. Once + that pack is gone, the next cycle adopts the sibling.""" + import hashlib + import uuid + + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + sibling.release() # its owner is gone + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert staged + lock = _claim(base, config) + refused = Path(lock.directory) / "fault" / "always-500" + refused.mkdir(parents=True) + body = b"never goes up" + own = refused / (f"{uuid.uuid4()}.1.1." + f"{hashlib.sha256(body).hexdigest()}.dmi-pack.ready") + own.write_bytes(body) + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="own-first-test", poll_interval_ns=50_000_000, + max_backoff_ns=200_000_000, sweep_spool_on_start=True, + reconcile_on_start=False, spool_owner_lock="held_by_caller", + adopt_sibling_spools=True, uploader_max_attempts=1, + s3_max_attempts=1, **config._lease_native()) + service = _store().StorageService(native) + service.start() + try: + _wait_for(lambda: service.snapshot()["upload_failures"] >= 5) + snapshot = service.snapshot() + assert snapshot["adopted_packs"] == 0, snapshot + assert snapshot["adoption_owed"] is True, snapshot + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + assert _store().spool_owner(str(dead)) is None # not even taken + + own.unlink() + _wait_for(_adopted(service, 1)) + snapshot = service.snapshot() + assert snapshot["adopted_packs"] == len(staged), snapshot + assert not dead.exists() + finally: + service.stop() + lock.release_and_remove_if_empty() + + +def test_a_latched_service_lets_go_of_the_sibling_it_was_adopting( + fake_s3, tmp_path): + """A service whose catalog another publisher keeps for 2 x TTL latches, + and its loop stops for good -- but the dead sibling it was half-way + through adopting stayed locked by this process, with nobody working on + it, until the engine's close(), possibly hours of capture later. Every + other process on the node read it as live meanwhile. The latched loop + lets go of it, at once.""" + from dmi.storage.native_capture import NativeCaptureStorage + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + # The servers are part of the catalog key: the sibling is claimed + # through the same switch URLs as the service reaches. + clickhouse = _Switch(CLICKHOUSE_HOST, CLICKHOUSE_HTTP_PORT) + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, + clickhouse_host="127.0.0.1", + clickhouse_port=clickhouse.port, + reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + sibling.release() # its owner is gone + staged = sorted(dead.rglob("*.dmi-pack.ready")) + store.cut() # so the adoption stays half done + lock = _claim(base, config) + service = _service(config, lock.directory) + rival = NativeCaptureStorage( + _storage_config(fake_s3, prefix, holder="rival-publisher", + start_lease_wait_s=10.0, + reconcile_on_start=False), + spool_root=str(tmp_path / "rival"), spool_max_bytes=1 << 30, + sweep_spool=True) + service.start() + try: + _wait_for(lambda: service.snapshot()["upload_failures"] > 0, 60.0) + owner = _store().spool_owner(str(dead)) + assert owner is not None and owner["pid"] == os.getpid(), owner + + clickhouse.cut() # the service can renew no more + rival.start() # waits out the lease the cut service cannot renew + clickhouse.restore() + _wait_for(lambda: service.snapshot()["failed"], 30.0) + _wait_for(lambda: _store().spool_owner(str(dead)) is None, 5.0) + snapshot = service.snapshot() + assert snapshot["adoption_owed"] is False, snapshot + assert snapshot["adopted_packs"] == 0, snapshot + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + finally: + service.stop() + rival.stop() + clickhouse.close() + store.close() + lock.release_and_remove_if_empty() + + +if __name__ == "__main__": + _dead_capture_process(*sys.argv[1:4]) diff --git a/tests/test_native_spool_owner_lock.py b/tests/test_native_spool_owner_lock.py new file mode 100644 index 000000000..8c813ba29 --- /dev/null +++ b/tests/test_native_spool_owner_lock.py @@ -0,0 +1,271 @@ +"""B6: one owner per spool directory, across processes. + +A spool's crash cleanup (Recover) deletes every ``.open`` file its own object +is not writing, so it is only safe while no other process writes there. The +C++ spool therefore takes an owner lock -- flock on ``/.owner.lock`` -- +when it opens a directory with ``owner_lock="take"`` (the default), and a +second process that tries is refused, told who holds it. What that lock +must also refuse, and what it must leave alone: + +* a directory nested under, or containing, a held one: Scan walks + recursively, so the outer spool's cleanup would reach into the inner one + while its owner writes it -- and every walk passes over a subdirectory + with a lock file of its own, so a dead spool directory inside another is + left for its successor to adopt; +* ``owner_lock="held_by_caller"`` unless THIS process holds the lock: that + mode opens without a lock of its own, for a second Spool in the process + that holds one. Beside another process's lock it is refused, naming the + holder -- else any process could open a live writer's directory that way + and sweep its in-flight ``.open`` files; +* ``/_refs/``, where the upload handoff's ref files will live: no scan + may sweep, quarantine or list anything under it. + +The drivers are separate processes, so these are real cross-process locks. +The in-process cases -- two Spool objects in one process, the node-local +check, the directory layout -- are pinned by tests/native/ +test_spool_owner_lock.cpp (test_native_spool_owner_lock_unit.py). + +The Python spool (dmi.storage.capture.spool) takes no lock; the C++ spool is +deliberately stricter, and that is not ported to the reference. Skipping +``_refs/`` is C++-only as well: the reference's constructor counts, and its +recover() sweeps and quarantines, ``.open`` and ``.ready`` files there, and +that divergence is deliberate, not ported either, like passing over +nested spool directories. + +Build: make -C native build/conformance_spool build/conformance_sink +""" + +from __future__ import annotations + +import json +import socket +import subprocess +from pathlib import Path + +import pytest + +REPO_ROOT = Path(__file__).resolve().parents[1] +BUILD = REPO_ROOT / "native" / "build" +SPOOL_DRIVER = BUILD / "conformance_spool" +SINK_DRIVER = BUILD / "conformance_sink" + +pytestmark = [ + pytest.mark.cpu, + pytest.mark.skipif( + not (SPOOL_DRIVER.exists() and SINK_DRIVER.exists()), + reason="native spool/sink drivers are not built; run " + "`make -C native build/conformance_spool build/conformance_sink`", + ), +] + +READY_NAME = ( + "018f0000-0000-7000-8000-000000000001.1700000000000000000.1." + + "0" * 64 + ".dmi-pack.ready" +) + + +class _Holder: + """A conformance_sink process whose open PackSink holds a spool.""" + + def __init__(self, root: Path, **fields): + self.proc = subprocess.Popen( + [str(SINK_DRIVER)], stdin=subprocess.PIPE, stdout=subprocess.PIPE, + text=True, bufsize=1) + self.opened = self.call( + op="open", root=str(root), max_bytes=1 << 30, + max_queue_records=16, max_queue_bytes=1 << 20, + max_pack_bytes=1 << 20, max_pack_records=16, + max_linger_ns=1_000_000_000, overload="drop_newest", + admission_timeout=-1, **fields) + + @property + def pid(self) -> int: + return self.proc.pid + + def call(self, **fields) -> dict: + self.proc.stdin.write(json.dumps(fields) + "\n") + self.proc.stdin.flush() + return json.loads(self.proc.stdout.readline()) + + def close(self): + if self.opened.get("ok"): + self.call(op="close", timeout=10) + try: + self.proc.stdin.close() + except BrokenPipeError: + pass + self.proc.wait(timeout=30) + + +def _spool(**fields) -> dict: + """One conformance_spool op in a fresh process (it opens per op).""" + fields.setdefault("max_bytes", 1 << 30) + proc = subprocess.run( + [str(SPOOL_DRIVER)], input=json.dumps(fields) + "\n", + capture_output=True, text=True, timeout=60) + lines = [line for line in proc.stdout.splitlines() if line.strip()] + assert lines, proc.stderr + return json.loads(lines[0]) + + +def test_a_second_process_is_refused_and_told_who_holds_the_spool(tmp_path): + root = tmp_path / "spool" + holder = _Holder(root) + try: + assert holder.opened["ok"], holder.opened + response = _spool(op="recover", root=str(root)) + assert not response["ok"], response + assert response["status"] == "open", response + what = response["what"] + assert f"pid {holder.pid}" in what, what + assert socket.gethostname() in what, what + assert str(root) in what, what + finally: + holder.close() + # The lock goes with its holder: the next process opens the directory. + assert _spool(op="recover", root=str(root))["ok"] + + +def test_the_lock_file_records_the_holder(tmp_path): + root = tmp_path / "spool" + holder = _Holder(root) + try: + assert holder.opened["ok"], holder.opened + record = (root / ".owner.lock").read_text() + assert record.split() == [socket.gethostname(), str(holder.pid)] + finally: + holder.close() + + +def test_a_directory_nested_under_an_owned_one_is_refused(tmp_path): + outer = tmp_path / "spool" + holder = _Holder(outer) + try: + assert holder.opened["ok"], holder.opened + response = _spool(op="recover", root=str(outer / "inner")) + assert not response["ok"], response + assert "nested" in response["what"], response + assert str(outer) in response["what"], response + finally: + holder.close() + + +def test_a_directory_containing_an_owned_one_is_refused(tmp_path): + outer = tmp_path / "spool" + holder = _Holder(outer / "a" / "inner") + try: + assert holder.opened["ok"], holder.opened + response = _spool(op="recover", root=str(outer)) + assert not response["ok"], response + assert "contains" in response["what"], response + assert str(outer / "a" / "inner") in response["what"], response + finally: + holder.close() + + +def test_a_dead_spool_directory_inside_another_is_left_alone(tmp_path): + """Nesting refuses only while the other directory's lock is HELD. Above: + every take leaves its lock file behind, so a spool_root some spool once + owned would otherwise refuse every directory under it for good. Below: + a dead directory's packs are for its successor to adopt, not for the + outer spool to sweep and upload under its own keys -- so every walk of + a spool passes over a subdirectory with a lock file of its own, and the + outer take goes through. Refused instead, as it was, a spool_root that + a crashed default-mode run left a rank directory in refused the + sink-only and explicit-record_sink modes, the rollback included.""" + outer = tmp_path / "spool" + assert _spool(op="recover", root=str(outer))["ok"] + assert (outer / ".owner.lock").exists() + inner = outer / "inner" + assert _spool(op="recover", root=str(inner))["ok"] + assert (inner / ".owner.lock").exists() + (inner / "v1").mkdir() + stale = inner / "v1" / ".018f0000-0000-7000-8000-000000000001.abcd1234.open" + stale.write_bytes(b"in progress") + bogus = inner / "v1" / READY_NAME # the wrong checksum: quarantined if met + bogus.write_bytes(b"not a pack") + + response = _spool(op="recover", root=str(outer)) + + assert response["ok"], response + assert response["staged"] == [] + assert response["snapshot"] == {"entries": 0, "bytes": 0} + assert stale.exists() and bogus.exists() + assert sorted(p.name for p in (inner / "v1").iterdir()) == sorted( + [stale.name, bogus.name]) + + +def test_held_by_caller_is_refused_beside_another_processs_holder(tmp_path): + """A process that does not hold the lock cannot open held_by_caller: + the holder is another process, so the open is refused, naming it, and + its Recover never runs.""" + root = tmp_path / "spool" + holder = _Holder(root) + try: + assert holder.opened["ok"], holder.opened + response = _spool(op="recover", root=str(root), + owner_lock="held_by_caller") + assert not response["ok"], response + assert response["status"] == "open", response + what = response["what"] + assert "held_by_caller" in what, what + assert f"pid {holder.pid}" in what, what + assert socket.gethostname() in what, what + finally: + holder.close() + + +def test_held_by_caller_is_refused_when_nothing_holds_the_lock(tmp_path): + root = tmp_path / "spool" + response = _spool(op="recover", root=str(root), + owner_lock="held_by_caller") + assert not response["ok"], response + assert "held_by_caller" in response["what"], response + # Once some spool has taken and let go of the lock, the file exists and + # is unlocked: still refused. + assert _spool(op="recover", root=str(root))["ok"] + response = _spool(op="recover", root=str(root), + owner_lock="held_by_caller") + assert not response["ok"], response + + +def test_an_unknown_owner_lock_mode_is_refused(tmp_path): + response = _spool(op="recover", root=str(tmp_path / "spool"), + owner_lock="share") + assert not response["ok"], response + assert "owner_lock" in response["what"], response + + +def test_recovery_leaves_the_refs_directory_alone(tmp_path): + """/_refs/ will hold the upload handoff's ref files (E2a). A + sweep that deleted, quarantined or listed them would break it. C++ + only: the Python reference spool still walks _refs/, deliberately + unported.""" + root = tmp_path / "spool" + refs = root / "_refs" + refs.mkdir(parents=True) + stale = refs / ".018f0000-0000-7000-8000-000000000001.abcd1234.open" + stale.write_bytes(b"in progress") + bogus = refs / READY_NAME # the wrong checksum: would be quarantined + bogus.write_bytes(b"not a pack") + ref = refs / "018f0000-0000-7000-8000-000000000001.uploaded" + ref.write_text("{}") + + response = _spool(op="recover", root=str(root)) + + assert response["ok"], response + assert response["staged"] == [] + assert response["snapshot"] == {"entries": 0, "bytes": 0} + assert stale.exists() and bogus.exists() and ref.exists() + assert sorted(p.name for p in refs.iterdir()) == sorted( + [stale.name, bogus.name, ref.name]) + + +def test_the_open_accounting_skips_the_refs_directory(tmp_path): + root = tmp_path / "spool" + refs = root / "_refs" + refs.mkdir(parents=True) + (refs / READY_NAME).write_bytes(b"x" * 100) + (refs / ".018f0000-0000-7000-8000-000000000001.abcd1234.open" + ).write_bytes(b"y" * 50) + assert _spool(op="snapshot", root=str(root))["snapshot"]["bytes"] == 0 diff --git a/tests/test_native_spool_owner_lock_unit.py b/tests/test_native_spool_owner_lock_unit.py new file mode 100644 index 000000000..cb89b4ad7 --- /dev/null +++ b/tests/test_native_spool_owner_lock_unit.py @@ -0,0 +1,76 @@ +"""The spool owner lock in process: tests/native/test_spool_owner_lock.cpp. + +Compiles the C++ test against the real spool.cpp and runs it. It pins what +the cross-process driver tests (test_native_spool_owner_lock.py) cannot +reach: two Spool objects in ONE process that both take a directory refuse +each other -- the regression the engine avoids by holding one lock and +opening its sink's and service's Spools with held_by_caller -- the node-local +check through its statfs test seam, adoption's try-lock, and the directory +layout. One case forks, so the refusal across processes is covered here too. + +Needs a C++17 compiler and libcrypto (the same as conformance_spool). +""" + +from __future__ import annotations + +import os +import shutil +import subprocess +from pathlib import Path + +import pytest + + +@pytest.mark.cpu +def test_spool_owner_lock_unit(tmp_path): + compiler = shutil.which("g++") or shutil.which("c++") + if compiler is None: + pytest.skip("a C++17 compiler is required") + + root = Path(__file__).resolve().parents[1] + csrc = root / "native" / "csrc" + source = root / "tests" / "native" / "test_spool_owner_lock.cpp" + executable = tmp_path / "test_spool_owner_lock" + extra: list[str] = [] + # Homebrew OpenSSL is not on the default search path on macOS; Linux + # hosts (CI) find libcrypto without help. + for candidate in ("/opt/homebrew/opt/openssl@3", "/usr/local/opt/openssl@3"): + if Path(candidate, "include", "openssl", "sha.h").exists(): + extra += [f"-I{candidate}/include", f"-L{candidate}/lib"] + break + compile_result = subprocess.run( + [ + compiler, + "-std=c++17", + "-O0", + "-pthread", + f"-I{csrc}", + *extra, + str(source), + str(csrc / "store" / "spool.cpp"), + "-o", + str(executable), + "-lcrypto", + ], + capture_output=True, + text=True, + check=False, + ) + if compile_result.returncode != 0 and "openssl" in ( + compile_result.stdout + compile_result.stderr + ): + pytest.skip("libcrypto headers are unavailable: " + + compile_result.stderr[-400:]) + assert compile_result.returncode == 0, ( + compile_result.stdout + compile_result.stderr + ) + + run_result = subprocess.run( + [str(executable)], + capture_output=True, + text=True, + check=False, + env={**os.environ, "SPOOL_TEST_ROOT": str(tmp_path)}, + timeout=120, + ) + assert run_result.returncode == 0, run_result.stdout + run_result.stderr diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py new file mode 100644 index 000000000..937244512 --- /dev/null +++ b/tests/test_native_spool_ownership.py @@ -0,0 +1,719 @@ +"""B6 through the Python surface: one owner per spool directory. + +``_dmi_native_store`` binds the spool's owner lock (``SpoolOwnerLock``), the +section 2.3 layout (``spool_rank_directory``) and who owns a directory +(``spool_owner``); the storage service and the native pack sink each open +their spool with ``owner_lock="take"`` or ``"held_by_caller"``. + +The regression this pins is the in-process one: a sink and a storage +service on ONE directory in ONE process. Were each to take the lock, the +second would be refused by the first -- flock binds to an open file +description, not to the process -- so the engine holds one SpoolOwnerLock +and opens both held_by_caller. The service is constructed, not started: +start() needs a catalog, which the live suites bring +(test_native_spool_adoption_live.py, test_native_capture_chain_live.py). +""" + +from __future__ import annotations + +import hashlib +import json +import os +import signal +import socket +import subprocess +import sys +from pathlib import Path + +import pytest + +REPO = Path(__file__).resolve().parents[1] +BUILD = REPO / "native" / "build" +STORE_BUILT = bool(sorted(BUILD.glob("_dmi_native_store*.so"))) +SINK_BUILT = bool(sorted(BUILD.glob("_dmi_native_sink*.so"))) + +pytestmark = [ + pytest.mark.cpu, + pytest.mark.skipif( + not STORE_BUILT, + reason="the native store module is not built; run `make -C native " + "build/_dmi_native_store PYTHON=/bin/python`"), +] + +LAYOUT = "capture_pack_reference_v1" +NFS_SUPER_MAGIC = 0x6969 + + +def _store(): + from dmi.storage.native_capture import _load_native_store_extension + + return _load_native_store_extension() + + +def _config(**overrides): + from dmi.storage.native_capture import NativeCaptureStorageConfig + + fields = dict( + s3_endpoint="http://127.0.0.1:9", s3_bucket="bucket", + s3_access_key="AKIA-test", s3_secret_key="secret-test", + s3_allow_insecure_http=True, clickhouse_port=9) + fields.update(overrides) + return NativeCaptureStorageConfig(**fields) + + +def _rank_directory(base: Path, config, rank: int = 0) -> str: + return _store().spool_rank_directory( + str(base), config._spool_destination(), rank) + + +# Where a spool's packs go: the catalog's server and names, the store's. +DESTINATION = dict( + clickhouse_host="ch", clickhouse_port=8123, database="db", + table_prefix="prefix", s3_endpoint="http://s3:9000", s3_bucket="bucket", + store_id="s3") + + +def _service(config, spool_root, **options): + from dmi.storage.native_capture import NativeCaptureStorage + + return NativeCaptureStorage(config, spool_root=str(spool_root), + spool_max_bytes=1 << 30, sweep_spool=True, + **options) + + +def _sink(spool_root, **options): + sys.path.insert(0, str(BUILD)) + try: + import _dmi_native_sink + finally: + sys.path.remove(str(BUILD)) + return _dmi_native_sink.NativePackSink( + spool_root=str(spool_root), layout=LAYOUT, max_pack_records=4, + **options) + + +def _stage_one(sink) -> None: + """One float16 capture through the sink's ring-facing submit.""" + import torch + from dmi.storage.capture import CaptureMetadata + + tensor = torch.arange(6, dtype=torch.float16) + metadata = CaptureMetadata( + capture_id="own-0", tenant_id="t", experiment_id="e", run_id="r", + session_id="s", request_id="q", sequence_id="n", model_id="m", + model_revision="mr", adapter_revision=None, + capture_policy_version="v", hook_name="resid_post", layer_number=0, + producer_rank=0, step_number=0, token_start=0, token_end=1, + batch_position=0, dtype="float16", shape=(6,), + captured_at_ns=1_700_000_000_000_000_000, + ).to_mapping() + lease = sink.attach() + sink.submit_envelope(LAYOUT, [{ + "metadata_json": json.dumps(metadata), "offset": 0, "length": 12, + "dtype": 5, "shape": [6]}], tensor.view(torch.uint8)) + assert sink.flush_and_wait(30.0) + sink.rethrow_if_failed() + del lease + + +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_the_real_sink_and_service_share_one_spool_in_one_process(tmp_path): + pytest.importorskip("torch") + config = _config() + directory = _rank_directory(tmp_path / "spool", config) + with _store().SpoolOwnerLock(directory) as lock: + service = _service(config, directory, + spool_owner_lock="held_by_caller", + adopt_sibling_spools=True) + sink = _sink(directory, owner_lock="held_by_caller") + _stage_one(sink) + assert len(list(Path(directory).rglob("*.dmi-pack.ready"))) == 1 + snapshot = service.snapshot() + assert snapshot["adopted_spools"] == 0 + assert snapshot["adoption_owed"] is False + # Neither Spool took a lock of its own: the one holder is this + # process, through `lock`. + assert _store().spool_owner(directory)["pid"] == os.getpid() + del sink, service + assert lock.held + assert _store().spool_owner(directory) is None + + +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_two_takes_in_one_process_refuse_each_other(tmp_path): + """What a per-Spool lock would do to the engine's sink and service.""" + pytest.importorskip("torch") + directory = tmp_path / "spool" + service = _service(_config(), directory) # take + with pytest.raises(RuntimeError, match=f"owned by pid {os.getpid()}"): + _sink(directory) # take + del service + _sink(directory) # the service's lock went with it + + +def test_a_second_process_is_refused_naming_the_holder(tmp_path): + directory = tmp_path / "spool" + with _store().SpoolOwnerLock(str(directory)): + probe = ( + "import sys; sys.path.insert(0, sys.argv[1]);" + "import _dmi_native_store as m\n" + "try:\n" + " m.SpoolOwnerLock(sys.argv[2])\n" + "except m.SpoolOwnedError as e:\n" + " print(e)\n") + result = subprocess.run( + [sys.executable, "-c", probe, str(BUILD), str(directory)], + capture_output=True, text=True, timeout=60) + assert result.returncode == 0, result.stderr + assert f"owned by pid {os.getpid()} on host {socket.gethostname()}" in ( + result.stdout), result.stdout + + +def test_a_forked_worker_does_not_keep_a_dead_owners_lock(tmp_path): + """A child forked without exec -- a fork-started DataLoader or + multiprocessing worker -- shares the lock's open file description. It + used to keep the lock after its parent was SIGKILLed, so the dead + parent's directory read as owned, naming the dead pid, and was never + adopted. + + The worker says it has run before its owner is killed: the fork + handler closes its copy of the lock only once the child is first + scheduled, which on a loaded host can come tens of milliseconds after + the owner is reaped, and until then the lock outlives its owner.""" + directory = tmp_path / "spool" + script = ( + "import os, sys, time; sys.path.insert(0, sys.argv[1]);" + "import _dmi_native_store as m\n" + "lock = m.SpoolOwnerLock(sys.argv[2])\n" + "worker = os.fork()\n" + "if worker == 0:\n" + " print('worker ran', flush=True)\n" # past fork(), and its handler + " time.sleep(3600)\n" + " os._exit(0)\n" + "print(worker, flush=True)\n" + "time.sleep(3600)\n") + owner = subprocess.Popen( + [sys.executable, "-c", script, str(BUILD), str(directory)], + stdout=subprocess.PIPE, text=True) + worker = None + try: + # The worker's line and the owner's, in either order. + lines = {owner.stdout.readline().strip(), + owner.stdout.readline().strip()} + assert "worker ran" in lines, lines + (worker,) = [int(line) for line in lines if line != "worker ran"] + assert _store().spool_owner(str(directory))["pid"] == owner.pid + os.kill(owner.pid, signal.SIGKILL) + owner.wait(timeout=30) + os.kill(worker, 0) # the worker outlives its parent + assert _store().spool_owner(str(directory)) is None + with _store().SpoolOwnerLock(str(directory)): + pass + finally: + if worker is not None: + try: + os.kill(worker, signal.SIGKILL) + except ProcessLookupError: + pass + if owner.poll() is None: + owner.kill() + owner.wait(timeout=30) + + +def test_a_removed_lock_file_leaves_a_live_directory_owned(tmp_path): + """Something other than the owner -- an age-based cleaner such as + systemd-tmpfiles, a person -- can remove .owner.lock from under a live + owner. Judged by that file alone, another process then saw nobody + holding the directory, took it, and could sweep the owner's stages in + flight. The owner locks the directory itself as well.""" + directory = tmp_path / "spool" + with _store().SpoolOwnerLock(str(directory)) as lock: + (directory / ".owner.lock").unlink() + probe = ( + "import sys; sys.path.insert(0, sys.argv[1]);" + "import _dmi_native_store as m\n" + "print(m.spool_owner(sys.argv[2]) is not None)\n" + "try:\n" + " m.SpoolOwnerLock(sys.argv[2])\n" + " print('taken')\n" + "except m.SpoolOwnedError:\n" + " print('owned')\n") + result = subprocess.run( + [sys.executable, "-c", probe, str(BUILD), str(directory)], + capture_output=True, text=True, timeout=60) + assert result.returncode == 0, result.stderr + assert result.stdout.split() == ["True", "owned"], result.stdout + assert lock.held + assert _store().spool_owner(str(directory)) is None + with _store().SpoolOwnerLock(str(directory)): + pass + + +def test_spool_owner_names_the_holder_while_it_holds(tmp_path): + store = _store() + directory = str(tmp_path / "spool") + assert store.spool_owner(directory) is None + lock = store.SpoolOwnerLock(directory) + assert lock.held + assert store.spool_owner(directory) == { + "host": socket.gethostname(), "pid": os.getpid()} + lock.release() + assert not lock.held + assert store.spool_owner(directory) is None + + +def test_the_owned_error_is_a_runtime_error(tmp_path): + store = _store() + assert issubclass(store.SpoolOwnedError, RuntimeError) + with store.SpoolOwnerLock(str(tmp_path / "spool")): + with pytest.raises(store.SpoolOwnedError): + store.SpoolOwnerLock(str(tmp_path / "spool")) + + +def test_a_nested_spool_directory_is_refused(tmp_path): + store = _store() + with store.SpoolOwnerLock(str(tmp_path / "outer")): + with pytest.raises(ValueError, match="nested"): + store.SpoolOwnerLock(str(tmp_path / "outer" / "inner")) + with pytest.raises(RuntimeError, match="nested"): + _service(_config(), tmp_path / "outer" / "inner") + + +def test_a_spool_root_a_sink_once_owned_still_takes_rank_directories(tmp_path): + """A sink-only run (or the explicit-record_sink rollback) takes + spool_root itself and leaves its lock file there. The next default run + claims a rank directory under it, which that unheld lock file must not + refuse as nested.""" + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + root = tmp_path / "root" + _store().SpoolOwnerLock(str(root)).release() + assert (root / ".owner.lock").exists() + claim = claim_spool_directory(NativeSinkConfig(spool_root=str(root)), + _config()) + try: + assert claim.held + assert Path(claim.directory).parent.parent == root + finally: + claim.release() + + +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_a_crashed_default_run_leaves_spool_root_to_the_other_modes(tmp_path): + """The reverse of the stale lock above: a default-mode run that dies + (or whose close did not drain) leaves its rank directory, lock file + and all, under spool_root. The sink-only mode (the sink takes + spool_root) and the explicit-record_sink rollback (the service takes + it) were then refused as containing an "owned" directory nobody held, + until a default-mode start adopted it -- the rollback blocked just when + it is most likely needed. Both now take spool_root, pass over the dead + directory, and leave it to adoption; a live one still refuses them.""" + pytest.importorskip("torch") + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + root = tmp_path / "root" + sink_config = NativeSinkConfig(spool_root=str(root)) + crashed = claim_spool_directory(sink_config, _config()) + dead = Path(crashed.directory) + (dead / "v1").mkdir() + left = dead / "v1" / "left.dmi-pack.ready" + left.write_bytes(b"pack") + in_flight = dead / "v1" / ".left.0badf00d.open" + in_flight.write_bytes(b"half") + crashed._lock.release() # killed: the kernel let go, the files stay + assert _store().spool_owner(str(dead)) is None + + service = _service(_config(), root, spool_owner_lock="take") # rollback + del service + sink = _sink(root) # sink-only + _stage_one(sink) + del sink + flat = sorted((root / "v1").rglob("*.dmi-pack.ready")) + assert len(flat) == 1 + assert left.read_bytes() == b"pack" and in_flight.exists() + + live = claim_spool_directory(sink_config, _config()) + try: + with pytest.raises(RuntimeError, match="contains the spool directory"): + _service(_config(), root, spool_owner_lock="take") + with pytest.raises(RuntimeError, match=f"pid {os.getpid()}"): + _sink(root) + finally: + live.release() + + +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_a_dead_siblings_bytes_count_against_the_sinks_budget(tmp_path): + """Each process start claims a fresh rank directory. A sink that + charged only its own let every crash-restart add a whole + spool_max_bytes while uploads were blocked; charge_dead_siblings (what + the engine passes) counts what the dead incarnations beside it still + hold, as the one directory every restart reused did before.""" + pytest.importorskip("torch") + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + sink_config = NativeSinkConfig(spool_root=str(tmp_path / "root")) + dead = claim_spool_directory(sink_config, _config()) + dead_directory = Path(dead.directory) + (dead_directory / "v1").mkdir() + (dead_directory / "v1" / "left.dmi-pack.ready").write_bytes( + b"\0" * (1 << 20)) + dead._lock.release() # killed: the kernel let go, the packs stay + budget = (1 << 20) + 256 # room for the dead bytes, not for a pack + + claim = claim_spool_directory(sink_config, _config()) + try: + sink = _sink(claim.directory, owner_lock="held_by_caller", + spool_max_bytes=budget, max_pack_bytes=budget, + charge_dead_siblings=True) + with pytest.raises(RuntimeError, match="spool byte limit exceeded"): + _stage_one(sink) + del sink + uncharged = _sink(claim.directory, owner_lock="held_by_caller", + spool_max_bytes=budget, max_pack_bytes=budget) + _stage_one(uncharged) # the whole budget, as if nothing were left + del uncharged + finally: + claim.release() + + +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_a_directory_this_process_keeps_is_not_charged_to_its_next_sink( + tmp_path): + """An engine whose sink did not seal keeps its claim held until the + process exits (the sink may still stage), and the same process's next + create_record_runtime claims a fresh directory beside it. No adoption + in this process can drain the kept one -- its service reads it as live + -- so charging it cut every later sink's budget until exit, and its + refusal called the bytes "still to be adopted".""" + pytest.importorskip("torch") + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + sink_config = NativeSinkConfig(spool_root=str(tmp_path / "root")) + kept = claim_spool_directory(sink_config, _config()) + (Path(kept.directory) / "v1").mkdir() + (Path(kept.directory) / "v1" / "left.dmi-pack.ready").write_bytes( + b"\0" * (1 << 20)) + budget = (1 << 20) + 256 # room for the kept bytes, not for a pack + claim = claim_spool_directory(sink_config, _config()) + try: + assert _store().spool_owner(kept.directory)["pid"] == os.getpid() + sink = _sink(claim.directory, owner_lock="held_by_caller", + spool_max_bytes=budget, max_pack_bytes=budget, + charge_dead_siblings=True) + _stage_one(sink) # the whole budget + del sink + finally: + claim.release() + kept.release() + + +def test_a_claim_warns_of_packs_left_outside_the_layout(tmp_path, caplog): + """Before the per-process layout the engine spooled into + /v1/... and its next start swept and uploaded whatever a + crashed run left there; a sink-only run still writes there. Nothing + adopts packs outside //r-/ now, so a claim + says so -- how many, where, and how to hand them to adoption -- rather + than leave them silently.""" + import logging + + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + root = tmp_path / "root" + sink = NativeSinkConfig(spool_root=str(root)) + # Packs inside the layout are adoption's, and no reason to warn. + first = claim_spool_directory(sink, _config()) + (Path(first.directory) / "v1").mkdir() + (Path(first.directory) / "v1" / "in.dmi-pack.ready").write_bytes(b"x") + first._lock.release() + with caplog.at_level(logging.WARNING, logger="dmi.storage.native_capture"): + claim = claim_spool_directory(sink, _config()) + claim.release() + assert not [r for r in caplog.records if "outside" in r.getMessage()] + + flat = root / "v1" / "tenant=t" / "date=2026-09-01" + flat.mkdir(parents=True) + for name in ("a", "b"): + (flat / f"{name}.dmi-pack.ready").write_bytes(b"pack") + (flat / ".c.0badf00d.open").write_bytes(b"half") # not a pack + caplog.clear() + with caplog.at_level(logging.WARNING, logger="dmi.storage.native_capture"): + claim = claim_spool_directory(sink, _config()) + key = Path(claim.directory).parent.name + claim.release() + (warning,) = [r.getMessage() for r in caplog.records + if "outside" in r.getMessage()] + assert "2 ready pack(s)" in warning + assert str(flat) in warning + assert f"{root}/{key}/r0-00000000" in warning + + +def test_a_dropped_spool_claim_keeps_its_directory_owned(tmp_path): + """Only release() lets go of a claim. An engine dropped without close() + drops its claim, while the ring and the sink it activated may still be + capturing into the directory: garbage collection must not unlock it for + another process to adopt. The kernel lets go when the process exits.""" + import gc + + from dmi.storage import native_capture + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + claim = claim_spool_directory( + NativeSinkConfig(spool_root=str(tmp_path / "root")), _config()) + directory = claim.directory + del claim + gc.collect() + try: + assert _store().spool_owner(directory)["pid"] == os.getpid() + finally: + for kept in list(native_capture._HELD_SPOOL_CLAIMS): + if kept.directory == directory: + kept.release() + assert _store().spool_owner(directory) is None + + +def test_releasing_a_claim_removes_its_directory_once_drained(tmp_path): + """SpoolClaim.release() is how the engine lets go of its directory: a + drained one is removed with its lock file, one still holding a pack + stays for the next process on the node to adopt.""" + from dmi.storage import native_capture + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + sink = NativeSinkConfig(spool_root=str(tmp_path / "root")) + claim = claim_spool_directory(sink, _config()) + drained = Path(claim.directory) + (drained / "v1" / "tenant=t").mkdir(parents=True) + assert claim.release() is True + assert not drained.exists() + assert not claim.held + assert claim not in native_capture._HELD_SPOOL_CLAIMS + + claim = claim_spool_directory(sink, _config()) + kept = Path(claim.directory) + (kept / "v1").mkdir() + (kept / "v1" / "left.dmi-pack.ready").write_bytes(b"pack") + assert claim.release() is False + assert (kept / "v1" / "left.dmi-pack.ready").exists() + assert _store().spool_owner(str(kept)) is None + assert claim not in native_capture._HELD_SPOOL_CLAIMS + + +@pytest.mark.parametrize("f_type, name", [ + (0x6969, "NFS"), (0x0BD00BD0, "Lustre"), (0x19830326, "BeeGFS"), + (0xFF534D42, "CIFS"), (0xFE534D42, "SMB2"), (0x65735546, "FUSE"), + (0x47504653, "GPFS"), (0x01021997, "9p"), (0x5346414F, "AFS"), + (0x6B414653, "AFS"), (0x20030528, "OrangeFS")]) +def test_each_shared_filesystem_is_refused_by_name(tmp_path, f_type, name): + store = _store() + store._set_spool_filesystem_type_for_testing(f_type) + try: + with pytest.raises(ValueError, match=f"is on {name} .*node-local"): + store.SpoolOwnerLock(str(tmp_path / "shared")) + with store.SpoolOwnerLock(str(tmp_path / "shared"), + allow_shared_filesystem=True): + pass + finally: + store._set_spool_filesystem_type_for_testing(None) + + +def test_a_shared_filesystem_is_refused_unless_allowed(tmp_path): + store = _store() + store._set_spool_filesystem_type_for_testing(NFS_SUPER_MAGIC) + try: + with pytest.raises(ValueError, match="NFS.*node-local"): + store.SpoolOwnerLock(str(tmp_path / "nfs")) + with pytest.raises(RuntimeError, match="node-local"): + _service(_config(), tmp_path / "nfs") + with store.SpoolOwnerLock(str(tmp_path / "nfs"), + allow_shared_filesystem=True): + pass + _service(_config(), tmp_path / "nfs-allowed", + spool_allow_shared_filesystem=True) + finally: + store._set_spool_filesystem_type_for_testing(None) + with store.SpoolOwnerLock(str(tmp_path / "local")): + pass + + +def test_the_rank_directory_layout(tmp_path): + store = _store() + key = hashlib.sha256( + b"db/prefix/s3\nclickhouse ch:8123\ns3 http://s3:9000/bucket" + ).hexdigest()[:12] + assert store.spool_catalog_key(DESTINATION) == key + assert store.spool_rank_directory( + "/base", DESTINATION, 3, "0a1b2c3d") == f"/base/{key}/r3-0a1b2c3d" + fresh = {store.spool_rank_directory("/base", DESTINATION, 0) + for _ in range(16)} + assert len(fresh) == 16 # a fresh incarnation each time + for path in fresh: + assert path.startswith(f"/base/{key}/r0-") + with pytest.raises(ValueError, match="incarnation"): + store.spool_rank_directory("/base", DESTINATION, 0, "XYZ") + incomplete = dict(DESTINATION) + del incomplete["s3_bucket"] + with pytest.raises(KeyError, match="s3_bucket"): + store.spool_catalog_key(incomplete) + assert _config()._spool_destination() == { + name: getattr(_config(), name) for name in DESTINATION} + + +def test_the_catalog_key_names_the_servers_not_only_the_names(tmp_path): + """The key was sha256(database/table_prefix/store_id), so two + deployments on one node with the default names (default, dmi, s3) but + different ClickHouse servers and buckets shared a key: whichever started + first adopted the other's dead directories, uploading its packs to its + own bucket and indexing them into its own catalog. Every server and + name the packs go to is in the key now.""" + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + store = _store() + key = store.spool_catalog_key(DESTINATION) + for field, value in (("clickhouse_host", "ch2"), ("clickhouse_port", 8124), + ("s3_endpoint", "http://s3b:9000"), + ("s3_bucket", "bucket2")): + assert store.spool_catalog_key({**DESTINATION, field: value}) != key + + staging = _config(clickhouse_host="ch-staging", s3_bucket="staging") + production = _config(clickhouse_host="ch-production", + s3_bucket="production") + sink = NativeSinkConfig(spool_root=str(tmp_path / "spool")) + staged = claim_spool_directory(sink, staging) + produced = claim_spool_directory(sink, production) + try: + assert (Path(staged.directory).parent + != Path(produced.directory).parent) + # A staging service cannot adopt from under production's key. + with pytest.raises(ValueError, match="adopt_sibling_spools"): + _service(staging, produced.directory, + spool_owner_lock="held_by_caller", + adopt_sibling_spools=True) + finally: + staged.release() + produced.release() + + +def test_adoption_needs_a_rank_directory_under_this_catalogs_key(tmp_path): + config = _config() + with pytest.raises(ValueError, match="adopt_sibling_spools"): + _service(config, tmp_path / "spool", adopt_sibling_spools=True) + # A rank directory, but under another catalog's key: adopting its + # siblings would index that catalog's packs into this one. + other = _store().spool_rank_directory( + str(tmp_path / "base"), + {**config._spool_destination(), "table_prefix": "another_prefix"}, 0) + with pytest.raises(ValueError, match="adopt_sibling_spools"): + _service(config, other, adopt_sibling_spools=True) + mine = _rank_directory(tmp_path / "base", config) + assert _service(config, mine, adopt_sibling_spools=True).snapshot()[ + "adopted_packs"] == 0 + + +def test_an_unknown_owner_lock_mode_is_refused(tmp_path): + with pytest.raises(ValueError, match="spool_owner_lock"): + _service(_config(), tmp_path / "spool", spool_owner_lock="share") + + +def test_held_by_caller_with_nothing_held_is_refused(tmp_path): + """The directory exists -- so this is the check that nothing holds its + lock, not a failure to resolve a path -- first with no lock file, then + with one nobody holds.""" + directory = tmp_path / "spool" + directory.mkdir() + with pytest.raises(RuntimeError, match="held_by_caller, but nothing holds"): + _service(_config(), directory, spool_owner_lock="held_by_caller") + _store().SpoolOwnerLock(str(directory)).release() + assert (directory / ".owner.lock").exists() + with pytest.raises(RuntimeError, match="held_by_caller, but nothing holds"): + _service(_config(), directory, spool_owner_lock="held_by_caller") + + +def test_held_by_caller_still_refuses_a_shared_filesystem(tmp_path): + """The node-local check applies to a Spool opened held_by_caller too: + the caller's lock was taken on a local disk, but what the Spool opens + is judged again.""" + store = _store() + directory = tmp_path / "spool" + with store.SpoolOwnerLock(str(directory)): + store._set_spool_filesystem_type_for_testing(NFS_SUPER_MAGIC) + try: + with pytest.raises(RuntimeError, match="is on NFS .*node-local"): + _service(_config(), directory, + spool_owner_lock="held_by_caller") + _service(_config(), directory, spool_owner_lock="held_by_caller", + spool_allow_shared_filesystem=True) + finally: + store._set_spool_filesystem_type_for_testing(None) + + +class _OtherProcessHolder: + """Another process holding a directory's SpoolOwnerLock until closed.""" + + def __init__(self, directory): + script = ( + "import sys; sys.path.insert(0, sys.argv[1]);" + "import _dmi_native_store as m\n" + "lock = m.SpoolOwnerLock(sys.argv[2])\n" + "print('held', flush=True)\n" + "sys.stdin.read()\n") + self.proc = subprocess.Popen( + [sys.executable, "-c", script, str(BUILD), str(directory)], + stdin=subprocess.PIPE, stdout=subprocess.PIPE, text=True) + assert self.proc.stdout.readline().strip() == "held" + + @property + def pid(self) -> int: + return self.proc.pid + + def close(self) -> None: + self.proc.stdin.close() + self.proc.wait(timeout=30) + + +def test_held_by_caller_beside_another_processs_lock_is_refused(tmp_path): + """held_by_caller is for the process that holds the lock. The service + and the sink opened that way beside ANOTHER process's lock are refused, + naming it, rather than sweeping and uploading from its directory.""" + directory = tmp_path / "spool" + holder = _OtherProcessHolder(directory) + try: + with pytest.raises(RuntimeError, + match=f"held_by_caller.*pid {holder.pid}"): + _service(_config(), directory, spool_owner_lock="held_by_caller") + if SINK_BUILT: + pytest.importorskip("torch") + with pytest.raises(RuntimeError, + match=f"held_by_caller.*pid {holder.pid}"): + _sink(directory, owner_lock="held_by_caller") + finally: + holder.close() + + +def test_a_drained_directory_is_removed_with_its_lock(tmp_path): + store = _store() + directory = tmp_path / "spool" + lock = store.SpoolOwnerLock(str(directory)) + (directory / "v1" / "tenant=t").mkdir(parents=True) + assert lock.release_and_remove_if_empty() + assert not directory.exists() + lock = store.SpoolOwnerLock(str(directory)) + (directory / "kept.quarantined").write_bytes(b"x") + assert not lock.release_and_remove_if_empty() + assert (directory / "kept.quarantined").exists() + assert not lock.held diff --git a/tests/test_native_uploader.py b/tests/test_native_uploader.py index c714f6a9f..acb82ac9f 100644 --- a/tests/test_native_uploader.py +++ b/tests/test_native_uploader.py @@ -381,12 +381,15 @@ def test_a_cancel_stops_the_listing_between_packs(fake_s3, tmp_path): """UploadPending lists the spool before it uploads anything, and a listing re-hashes every staged pack: over a backlog, seconds a GiB of it, all before a single worker looked at the cancel. The listing now - stops between packs once cancelled, and nothing is tried. 16 sparse - packs of 256 MiB, zeros named for their checksum, take it about 3 s.""" + stops between packs once cancelled, and nothing is tried. 128 sparse + packs of 64 MiB, zeros named for their checksum, take it about 6 s + idle; one of them takes a small part of the bound below even on a + loaded runner, where 256 MiB packs took it over, the whole listing + slowing with them.""" import hashlib import uuid - size = 256 << 20 + size = 64 << 20 digest = hashlib.sha256() zeros = bytes(1 << 20) for _ in range(size // len(zeros)): @@ -394,7 +397,7 @@ def test_a_cancel_stops_the_listing_between_packs(fake_s3, tmp_path): root = tmp_path / "spool" root.mkdir() backlog = [] - for _ in range(16): + for _ in range(128): ready = root / f"{uuid.uuid4()}.1.1.{digest.hexdigest()}.dmi-pack.ready" with open(ready, "wb") as sparse: sparse.truncate(size) @@ -411,9 +414,9 @@ def test_a_cancel_stops_the_listing_between_packs(fake_s3, tmp_path): assert result["refs"] == [] and result["failures"] == [], result assert result["snapshot"]["cancelled_packs"] == 0, result # Between packs, not after the listing: the cut ends within one pack's - # hash of the cancel, about 0.5 s here, where the whole listing takes - # over 3 s. The bound leaves room for a loaded runner, whose hashing - # slows with it (about 1.1 s at 4x oversubscription). + # hash of the cancel, about 0.1 s here, where the whole listing takes + # about 6 s. The bound leaves room for a loaded runner, whose hashing + # slows with it (a pack about 0.6 s at 2.5x oversubscription). assert elapsed < 2.0, elapsed assert sorted(root.rglob("*.dmi-pack.ready")) == sorted(backlog) assert STATE.calls == [] @@ -523,6 +526,8 @@ def test_pack_over_the_byte_gate_fails_fast(fake_s3, tmp_path): assert result["snapshot"]["failed_packs"] == 1 assert result["failures"][0]["attempts"] == 0 assert "in-flight" in result["failures"][0]["error"] + # No retry by this uploader can admit it. + assert result["failures"][0]["retryable"] is False finally: sink.close() store.close() @@ -619,8 +624,10 @@ def test_a_pack_conflict_reports_why_it_will_not_overwrite(fake_s3, tmp_path): assert staged["object_key"] in failure["error"], failure assert "do not overwrite" in failure["error"], failure # Counted like every other attempted failure: the HEAD that found - # the conflict was an attempt, and it is not retried. + # the conflict was an attempt, and it is not retried -- now or by a + # later batch. assert failure["attempts"] == 1, failure + assert failure["retryable"] is False, failure # The refusal is real: nothing was written over the existing object, # and the staged pack is still there to inspect. @@ -876,6 +883,8 @@ def test_a_64_mib_pack_uploads_multipart_over_https_with_a_private_ca( assert untrusted["refs"][0]["pack_id"] == "", untrusted assert failure["pack_id"] == staged["pack_id"], failure assert "certificate" in failure["error"].lower(), failure + # A transport failure: a later try (with the CA, here) can succeed. + assert failure["retryable"] is True, failure assert STATE.calls == [] and STATE.objects == {} assert Path(staged["path"]).exists()