From 7b6c9c585afcc5a3fb9e23cfadd4f3fb4e2b793b Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Mon, 28 Sep 2026 19:43:39 -0400 Subject: [PATCH 01/43] Give each spool directory one owner process: an flock on /.owner.lock Recover() deletes every .open file its own Spool object is not writing. Nothing kept a second process from running it against a live writer: the old test_native_live_spool.py only showed that an upload SCAN (ListPending) spares the writer's temp file, and one process per spool was left to the caller (storage_service.h said so). A second process pointed at the same directory -- two jobs on one node, a restart racing its predecessor -- swept the writer's in-flight temp file and failed its stage with "cannot link ready file". The spool now has an owner lock, flock(LOCK_EX|LOCK_NB) on /.owner.lock, recording " " in the file: * SpoolConfig.owner_lock = take (the default) takes it in Open(), before the accounting walk and so before any Recover(), and holds it for the Spool's life. A second process is refused with kOwned: "spool directory X is owned by pid N on host H (it holds X/.owner.lock)". The kernel drops the lock with its holder, SIGKILL included. * held_by_caller takes none, for a process that runs a sink and a service on one directory: two takes in ONE process refuse each other (flock binds to an open file description), so such a process holds one SpoolOwnerLock and opens both Spools with held_by_caller. Open() refuses it when nothing holds the lock. * A directory nested under, or containing, an owned one (a .owner.lock, held or not) is refused: Scan walks recursively. * NFS and Lustre are refused by statfs f_type (allow_shared_filesystem overrides; SetFilesystemTypeForTesting stands in for statfs), since flock there does not keep out a process on another node. * Scan, Open's accounting and the capacity reconcile skip /_refs/, where the upload handoff's ref files will live (plan section 2.4). SpoolOwnerLock is the lock on its own: Acquire creates a new directory beside its already-held lock file and renames it into place, so no scan of the parent meets it unowned; TryAdopt locks an existing directory whose owner is gone; ReleaseAndRemoveIfEmpty removes a drained one. spool.h also gains the section 2.3 layout, //r-/, for the service's adoption and the engine, next. The drivers accept "owner_lock". Every Spool takes by default, so two in one process refuse each other: test_spool_reservations.cpp's second Spool object on a root now opens held_by_caller, and the storage service's and the pack sink's Spools -- which test_native_capture_chain_live.py and the storage live suite run in one process -- refuse each other until the next commit gives both an owner-lock mode. The Python DurablePackSpool (spool.py:208-214) stays unlocked and its recover() still deletes every .open file: the C++ spool is deliberately stricter, and that is not ported to the reference. Tests, each new one red before the change for the reason it names: * test_native_spool_owner_lock.py (drivers, across processes): a second process was admitted ({"ok": true}) where it is now refused naming the holder's pid and host; no lock file was written; nested and containing roots were admitted; held_by_caller with nothing held, and an unknown mode, were admitted; recovery deleted /_refs/*.open; Open counted 150 bytes under _refs where it now counts 0. * test_native_spool_owner_lock_unit.py compiles tests/native/test_spool_owner_lock.cpp (red: the API did not exist): two takes in one process refuse each other; held_by_caller beside a SpoolOwnerLock stages and lists; a forked holder is named and its lock freed by SIGKILL; nesting both ways; NFS/Lustre refused through the test seam, admitted with the override; adoption's try-lock; the layout. * test_native_live_spool.py is rewritten: the upload scan beside a paused live writer is now refused, naming the writer (red: it ran); and a writer SIGKILLed mid-stage is recovered by the next process, its stale .open swept and its sealed pack uploaded (green before too: it guards that the lock does not outlive its holder). * tests/test_native_spool.py (all 20) and test_native_spool_reservations.py pass; so does every other cpu test under tests/test_native_*.py. --- docs/capture-storage-design.md | 3 + native/csrc/sink/conformance_sink.cpp | 10 +- native/csrc/sink/pack_sink.cpp | 2 + native/csrc/sink/pack_sink.h | 5 + native/csrc/store/conformance_spool.cpp | 13 + native/csrc/store/conformance_store.cpp | 17 +- native/csrc/store/spool.cpp | 595 ++++++++++++++++++++- native/csrc/store/spool.h | 158 +++++- tests/native/live_spool_stage.cpp | 62 ++- tests/native/test_spool_owner_lock.cpp | 390 ++++++++++++++ tests/native/test_spool_reservations.cpp | 26 +- tests/test_native_live_spool.py | 108 +++- tests/test_native_spool_owner_lock.py | 231 ++++++++ tests/test_native_spool_owner_lock_unit.py | 76 +++ 14 files changed, 1646 insertions(+), 50 deletions(-) create mode 100644 tests/native/test_spool_owner_lock.cpp create mode 100644 tests/test_native_spool_owner_lock.py create mode 100644 tests/test_native_spool_owner_lock_unit.py diff --git a/docs/capture-storage-design.md b/docs/capture-storage-design.md index 8fdc5c4ed..f63fc389b 100644 --- a/docs/capture-storage-design.md +++ b/docs/capture-storage-design.md @@ -1827,6 +1827,9 @@ bytes, and retained failure details all have explicit caps. writer will be native. See *Phase 6 -- Decision: the production writer is native*. - One process owns a spool directory; cross-process locking is not implemented. + (The native C++ spool does lock it: an owner lock on `/.owner.lock`, + B6, documented in `native/csrc/store/spool.h`. The Python reference + deliberately stays unlocked.) - Durable mode stages synchronously and uploads through a separate explicit uploader, so remote backpressure is isolated from local commit. - The pipeline remains opt-in and is not connected to Ring². diff --git a/native/csrc/sink/conformance_sink.cpp b/native/csrc/sink/conformance_sink.cpp index 47196a0ff..042f6b9e3 100644 --- a/native/csrc/sink/conformance_sink.cpp +++ b/native/csrc/sink/conformance_sink.cpp @@ -4,7 +4,7 @@ // {"op":"open","root":"...","max_bytes":N,"max_queue_records":N, // "max_queue_bytes":N,"max_pack_bytes":N,"max_pack_records":N, // "max_linger_ns":N,"overload":"block"|"drop_newest", -// "admission_timeout":-1} +// "admission_timeout":-1,"owner_lock":"take"|"held_by_caller" (optional)} // -> {"ok":true} // {"op":"submit","metadata":{...canonical field names...},"payload_b64":"..."} // -> {"ok":true,"admission":"accepted"|...} @@ -209,6 +209,14 @@ int main() { jc::FindString(line, "overload") == "block" ? dmi_sink::Overload::kBlock : dmi_sink::Overload::kDropNewest; + const std::string owner_lock = jc::FindString(line, "owner_lock"); + if (!owner_lock.empty() && + !dmi_store::ParseOwnerLock(owner_lock, &config.spool_owner_lock)) { + std::string out = "{\"ok\":false,\"what\":"; + jc::EscapeJson("unknown owner_lock: " + owner_lock, &out); + std::cout << out << "}\n"; + continue; + } const int64_t workers = Integer(line, "num_workers"); config.num_workers = static_cast(workers > 0 ? workers : 1); // Before the sink exists: a limit that cannot be represented must not diff --git a/native/csrc/sink/pack_sink.cpp b/native/csrc/sink/pack_sink.cpp index adee3592e..6398bdb6c 100644 --- a/native/csrc/sink/pack_sink.cpp +++ b/native/csrc/sink/pack_sink.cpp @@ -96,6 +96,8 @@ std::string PackSink::Start(std::string* spool_error) { dmi_store::SpoolConfig spool_config; spool_config.root = config_.spool_root; spool_config.max_bytes = config_.spool_max_bytes; + spool_config.owner_lock = config_.spool_owner_lock; + spool_config.allow_shared_filesystem = config_.spool_allow_shared_filesystem; std::string error; const dmi_store::SpoolStatus st = dmi_store::Spool::Open(spool_config, &spool_, &error); diff --git a/native/csrc/sink/pack_sink.h b/native/csrc/sink/pack_sink.h index daed49088..ab57613e6 100644 --- a/native/csrc/sink/pack_sink.h +++ b/native/csrc/sink/pack_sink.h @@ -68,6 +68,11 @@ struct SinkConfig { double admission_timeout_s = -1.0; std::string spool_root; uint64_t spool_max_bytes = 1ull << 40; + // The spool directory's owner lock (spool.h). kTake owns the directory + // for the sink's life; a process that also runs a storage service on it + // holds one SpoolOwnerLock and opens both with kHeldByCaller. + dmi_store::OwnerLock spool_owner_lock = dmi_store::OwnerLock::kTake; + bool spool_allow_shared_filesystem = false; // Pack assembler workers. Records route by scope hash // (tenant, session, producer_rank), so one scope always lands on one // worker: per-scope ordering and single-scope packs are preserved at any diff --git a/native/csrc/store/conformance_spool.cpp b/native/csrc/store/conformance_spool.cpp index 9112ce919..b8883b847 100644 --- a/native/csrc/store/conformance_spool.cpp +++ b/native/csrc/store/conformance_spool.cpp @@ -11,6 +11,9 @@ // -> {"ok":true} // {"op":"snapshot","root":"...","max_bytes":N} // -> {"ok":true,"snapshot":{"entries":N,"bytes":N,"peak_bytes":N,"max_bytes":N}} +// Every op re-opens the spool, and takes "owner_lock":"take" (the default) +// or "held_by_caller"; an open another owner refuses answers +// {"ok":false,"status":"open","what":"...owned by pid N on host H..."}. // Errors: {"ok":false,"status":"...","what":"..."}. #include "spool.h" @@ -103,8 +106,18 @@ int main() { continue; } if (config.max_bytes == 0) config.max_bytes = 1ull << 40; + // Optional: "take" (the default) or "held_by_caller". + const std::string owner_lock = jc::FindString(line, "owner_lock"); dmi_store::Spool spool; std::string error; + if (!owner_lock.empty() && + !dmi_store::ParseOwnerLock(owner_lock, &config.owner_lock)) { + std::string out = "{\"ok\":false,\"status\":\"open\",\"what\":"; + jc::EscapeJson("unknown owner_lock: " + owner_lock, &out); + out += "}\n"; + std::cout << out; + continue; + } if (dmi_store::Spool::Open(config, &spool, &error) != dmi_store::SpoolStatus::kOk) { std::string out = "{\"ok\":false,\"status\":\"open\",\"what\":"; diff --git a/native/csrc/store/conformance_store.cpp b/native/csrc/store/conformance_store.cpp index 3456bf043..952856486 100644 --- a/native/csrc/store/conformance_store.cpp +++ b/native/csrc/store/conformance_store.cpp @@ -15,6 +15,9 @@ // {"op":"list",...,"prefix":"...","delimiter":"...","max_keys":N, // "continuation":"..."} -> {"ok":true,"truncated":bool,"next_token":"...", // "objects":[{"key":"...","size":N,"etag":"..."}...],"attempts":N} +// {"op":"upload_one"|"upload_pending",...,"root":"...", +// "owner_lock":"take"|"held_by_caller" (optional, take by default)} +// open the spool at root, owner lock included, for the one op. // Errors: {"ok":false,"what":"..."}. #include "s3_client.h" @@ -294,7 +297,12 @@ int main() { if (spool_config.max_bytes == 0) spool_config.max_bytes = 1ull << 40; dmi_store::Spool spool; std::string spool_error; - if (dmi_store::Spool::Open(spool_config, &spool, &spool_error) != + const std::string owner_lock = jc::FindString(line, "owner_lock"); + if (!owner_lock.empty() && + !dmi_store::ParseOwnerLock(owner_lock, &spool_config.owner_lock)) { + out += "false,\"what\":"; + jc::EscapeJson("spool open: unknown owner_lock: " + owner_lock, &out); + } else if (dmi_store::Spool::Open(spool_config, &spool, &spool_error) != dmi_store::SpoolStatus::kOk) { out += "false,\"what\":"; jc::EscapeJson("spool open: " + spool_error, &out); @@ -337,7 +345,12 @@ int main() { if (spool_config.max_bytes == 0) spool_config.max_bytes = 1ull << 40; dmi_store::Spool spool; std::string spool_error; - if (dmi_store::Spool::Open(spool_config, &spool, &spool_error) != + const std::string owner_lock = jc::FindString(line, "owner_lock"); + if (!owner_lock.empty() && + !dmi_store::ParseOwnerLock(owner_lock, &spool_config.owner_lock)) { + out += "false,\"what\":"; + jc::EscapeJson("spool open: unknown owner_lock: " + owner_lock, &out); + } else if (dmi_store::Spool::Open(spool_config, &spool, &spool_error) != dmi_store::SpoolStatus::kOk) { out += "false,\"what\":"; jc::EscapeJson("spool open: " + spool_error, &out); diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index d0df3aa80..ecf2757ab 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -3,12 +3,20 @@ #include #include +#include +#include +#include +#include #include #include #include #include +#include #include +#include +#include +#include #include namespace dmi_store { @@ -18,6 +26,35 @@ namespace { constexpr const char* kReadySuffix = ".dmi-pack.ready"; constexpr const char* kOpenSuffix = ".open"; +constexpr const char* kOwnerLockFile = ".owner.lock"; +// /_refs/: the upload handoff's ref files (plan section 2.4). Not a +// legal object-key component (those start with an alphanumeric), so no pack +// is ever staged under it. +constexpr const char* kRefsDirectory = "_refs"; + +// statfs(2) f_type values (linux/magic.h has NFS; Lustre's is its own). +constexpr uint32_t kNfsSuperMagic = 0x6969; +constexpr uint32_t kLustreSuperMagic = 0x0BD00BD0; + +std::atomic g_filesystem_type_for_testing{-1}; + +// Whether a recursive walk of a spool root is at /_refs, which no scan +// enters. +bool AtRefsDirectory(const fs::recursive_directory_iterator& it) { + std::error_code ec; + return it.depth() == 0 && it->path().filename() == kRefsDirectory && + it->is_directory(ec); +} + +std::string Hostname() { + char host[256] = {0}; + if (::gethostname(host, sizeof(host) - 1) != 0 || host[0] == '\0') { + return "unknown-host"; + } + return host; +} + +std::string Errno(int error) { return std::strerror(error); } bool HasSuffix(const std::string& name, const char* suffix) { const size_t n = std::strlen(suffix); @@ -236,14 +273,542 @@ std::string ReadyName(const std::string& pack_id, uint64_t created, std::to_string(records) + "." + checksum + ".dmi-pack.ready"; } +void FsyncParent(const std::string& path) { + FsyncDir(fs::path(path).parent_path().string(), nullptr); +} + +// " ": whoever holds the lock records itself, so a refused +// process can say who holds the directory. +void WriteOwnerRecord(int fd) { + const std::string record = + Hostname() + " " + std::to_string(::getpid()) + "\n"; + if (::ftruncate(fd, 0) == 0) { + (void)!::pwrite(fd, record.data(), record.size(), 0); + } +} + +void ReadOwnerRecord(int fd, SpoolOwner* owner) { + char buffer[512]; + const ssize_t n = ::pread(fd, buffer, sizeof(buffer) - 1, 0); + owner->host.clear(); + owner->pid = 0; + if (n <= 0) return; + std::string record(buffer, static_cast(n)); + while (!record.empty() && std::isspace(static_cast( + record.back()))) { + record.pop_back(); + } + const size_t space = record.rfind(' '); + if (space == std::string::npos) { + owner->host = record; + return; + } + owner->host = record.substr(0, space); + const std::string pid = record.substr(space + 1); + if (!pid.empty() && pid.size() <= 18 && + std::all_of(pid.begin(), pid.end(), + [](char c) { return c >= '0' && c <= '9'; })) { + owner->pid = std::strtoll(pid.c_str(), nullptr, 10); + } +} + +std::string OwnedMessage(const std::string& dir, const SpoolOwner& owner) { + const std::string file = dir + "/" + kOwnerLockFile; + if (owner.pid <= 0) { + return "spool directory " + dir + " is owned by another process, which " + "holds " + file + " and has not recorded itself yet; a spool " + "directory has one owner process"; + } + return "spool directory " + dir + " is owned by pid " + + std::to_string(owner.pid) + " on host " + owner.host + + " (it holds " + file + "); a spool directory has one owner process"; +} + +// Whether `fd` is still the file at `path`: a remover unlinks the lock file +// before it removes a drained directory, and a lock taken on the unlinked +// file guards nothing. +bool IsFileAt(int fd, const std::string& path) { + struct stat by_fd{}, by_path{}; + return ::fstat(fd, &by_fd) == 0 && ::stat(path.c_str(), &by_path) == 0 && + by_fd.st_dev == by_path.st_dev && by_fd.st_ino == by_path.st_ino; +} + +// A path that may not exist yet, absolute, with its existing prefix's +// symlinks resolved. +std::string CanonicalPath(const std::string& path, std::string* error) { + std::error_code ec; + const fs::path absolute = fs::absolute(path, ec); + if (ec) { + if (error) *error = "cannot resolve " + path + ": " + ec.message(); + return ""; + } + const fs::path canonical = fs::weakly_canonical(absolute, ec); + if (ec) { + if (error) *error = "cannot resolve " + path + ": " + ec.message(); + return ""; + } + std::string out = canonical.string(); + while (out.size() > 1 && out.back() == '/') out.pop_back(); + return out; +} + +// A spool directory must not be nested under, or contain, another owned +// directory -- one with a lock file, held or not: Scan walks recursively, +// so the outer spool's Recover would sweep the inner one's .open files and +// upload its packs under the outer spool's keys. +SpoolStatus CheckNotNested(const std::string& dir, bool exists, + std::string* error) { + fs::path ancestor(dir); + while (ancestor.has_parent_path() && ancestor.parent_path() != ancestor) { + ancestor = ancestor.parent_path(); + std::error_code ec; + if (fs::exists(ancestor / kOwnerLockFile, ec)) { + if (error) { + *error = "spool directory " + dir + " is nested under the owned " + "spool directory " + ancestor.string() + " (it has " + + kOwnerLockFile + "), whose recovery would sweep this one; " + "use a directory outside it"; + } + return SpoolStatus::kBadArgument; + } + } + if (!exists) return SpoolStatus::kOk; + std::error_code ec; + for (auto it = fs::recursive_directory_iterator( + dir, fs::directory_options::skip_permission_denied, ec); + !ec && it != fs::recursive_directory_iterator(); it.increment(ec)) { + if (it->path().filename() != kOwnerLockFile) continue; + const fs::path owned = it->path().parent_path(); + if (owned == fs::path(dir)) continue; + if (error) { + *error = "spool directory " + dir + " contains the owned spool " + "directory " + owned.string() + " (it has " + kOwnerLockFile + + "), which this one's recovery would sweep; use a directory " + "that does not contain it"; + } + return SpoolStatus::kBadArgument; + } + return SpoolStatus::kOk; +} + +} // namespace + +const char* OwnerLockName(OwnerLock mode) { + return mode == OwnerLock::kHeldByCaller ? "held_by_caller" : "take"; +} + +bool ParseOwnerLock(const std::string& text, OwnerLock* mode) { + if (text == "take") { + *mode = OwnerLock::kTake; + return true; + } + if (text == "held_by_caller") { + *mode = OwnerLock::kHeldByCaller; + return true; + } + return false; +} + +const char* SharedFilesystemName(int64_t f_type) { + switch (static_cast(f_type)) { + case kNfsSuperMagic: return "NFS"; + case kLustreSuperMagic: return "Lustre"; + default: return nullptr; + } +} + +void SetFilesystemTypeForTesting(int64_t f_type) { + g_filesystem_type_for_testing.store(f_type); +} + +SpoolStatus CheckNodeLocal(const std::string& dir, + bool allow_shared_filesystem, std::string* error) { + int64_t f_type = g_filesystem_type_for_testing.load(); + if (f_type < 0) { + struct statfs info{}; + if (::statfs(dir.c_str(), &info) != 0) { + if (error) *error = "cannot statfs " + dir + ": " + Errno(errno); + return SpoolStatus::kIo; + } + f_type = static_cast(static_cast(info.f_type)); + } + const char* shared = SharedFilesystemName(f_type); + if (shared == nullptr || allow_shared_filesystem) return SpoolStatus::kOk; + if (error) { + char magic[32]; + std::snprintf(magic, sizeof(magic), "0x%llx", + static_cast(f_type)); + *error = "spool directory " + dir + " is on " + shared + + " (statfs f_type " + magic + "): a spool must be node-local, " + "since its owner lock (flock) does not keep out a process on " + "another node there. Use a local disk, or set " + "allow_shared_filesystem if no process on another node can " + "reach this directory"; + } + return SpoolStatus::kBadArgument; +} + +bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner) { + const std::string file = dir + "/" + kOwnerLockFile; + const int fd = ::open(file.c_str(), O_RDONLY | O_CLOEXEC); + if (fd < 0) return false; + // A probe: if the lock can be taken nobody holds it, and it is let go at + // once. (A take racing the probe retries, see LockInPlace.) + if (::flock(fd, LOCK_EX | LOCK_NB) == 0) { + ::flock(fd, LOCK_UN); + ::close(fd); + return false; + } + const bool held = errno == EWOULDBLOCK; + if (held && owner != nullptr) ReadOwnerRecord(fd, owner); + ::close(fd); + return held; +} + +namespace { + +// Locks the lock file of an existing directory, creating the file if it has +// none. Retries a lock lost to a remover's unlink, and a refusal as brief as +// another process's ReadSpoolOwner probe. +SpoolStatus LockInPlace(const std::string& dir, int* fd_out, + std::string* error) { + const std::string file = dir + "/" + kOwnerLockFile; + for (int attempt = 0; attempt < 8; ++attempt) { + const int fd = ::open(file.c_str(), O_RDWR | O_CREAT | O_CLOEXEC, 0644); + if (fd < 0) { + if (error) *error = "cannot open " + file + ": " + Errno(errno); + return SpoolStatus::kIo; + } + if (::flock(fd, LOCK_EX | LOCK_NB) != 0) { + const int failure = errno; + SpoolOwner owner; + ReadOwnerRecord(fd, &owner); + ::close(fd); + if (failure != EWOULDBLOCK) { + if (error) *error = "cannot lock " + file + ": " + Errno(failure); + return SpoolStatus::kIo; + } + if (attempt < 2) { + std::this_thread::sleep_for(std::chrono::milliseconds(2)); + continue; + } + if (error) *error = OwnedMessage(dir, owner); + return SpoolStatus::kOwned; + } + if (!IsFileAt(fd, file)) { + ::close(fd); + if (!fs::is_directory(dir)) { + if (error) *error = "spool directory " + dir + " was removed while " + "it was being locked"; + return SpoolStatus::kIo; + } + continue; + } + WriteOwnerRecord(fd); + *fd_out = fd; + return SpoolStatus::kOk; + } + if (error) *error = "cannot lock " + file + ": it keeps being replaced"; + return SpoolStatus::kIo; +} + +// Creates `dir` with its lock file already held: built under a hidden name +// beside it, then renamed into place, so a scan of the parent never meets +// the directory unowned (an adopter would otherwise take a brand-new +// sibling for a dead one). Falls back to LockInPlace if `dir` appears +// meanwhile. +SpoolStatus CreateLocked(const std::string& dir, int* fd_out, + std::string* error) { + const fs::path target(dir); + const std::string parent = target.parent_path().string(); + const std::string name = target.filename().string(); + std::random_device random; + for (int attempt = 0; attempt < 8; ++attempt) { + char suffix[16]; + std::snprintf(suffix, sizeof(suffix), "%08x", + static_cast(random())); + const std::string staging = + parent + "/." + name + "." + suffix + ".creating"; + if (::mkdir(staging.c_str(), 0755) != 0) { + if (errno == EEXIST) continue; + if (error) *error = "cannot create " + staging + ": " + Errno(errno); + return SpoolStatus::kIo; + } + const std::string file = staging + "/" + kOwnerLockFile; + const int fd = + ::open(file.c_str(), O_RDWR | O_CREAT | O_EXCL | O_CLOEXEC, 0644); + if (fd < 0 || ::flock(fd, LOCK_EX | LOCK_NB) != 0) { + const int failure = errno; + if (fd >= 0) ::close(fd); + ::unlink(file.c_str()); + ::rmdir(staging.c_str()); + if (error) *error = "cannot lock " + file + ": " + Errno(failure); + return SpoolStatus::kIo; + } + WriteOwnerRecord(fd); + ::fsync(fd); + FsyncDir(staging, nullptr); +#ifdef RENAME_NOREPLACE + const int renamed = ::renameat2(AT_FDCWD, staging.c_str(), AT_FDCWD, + dir.c_str(), RENAME_NOREPLACE); +#else + const int renamed = ::rename(staging.c_str(), dir.c_str()); +#endif + if (renamed != 0) { + const int failure = errno; + ::close(fd); + ::unlink(file.c_str()); + ::rmdir(staging.c_str()); + if (failure == EEXIST || failure == ENOTEMPTY) { + return LockInPlace(dir, fd_out, error); + } + if (error) { + *error = "cannot create spool directory " + dir + ": " + + Errno(failure); + } + return SpoolStatus::kIo; + } + FsyncDir(parent, nullptr); + *fd_out = fd; + return SpoolStatus::kOk; + } + if (error) *error = "cannot create spool directory " + dir; + return SpoolStatus::kIo; +} + +} // namespace + +SpoolOwnerLock::~SpoolOwnerLock() { Release(); } + +SpoolOwnerLock::SpoolOwnerLock(SpoolOwnerLock&& other) noexcept + : fd_(other.fd_), dir_(std::move(other.dir_)) { + other.fd_ = -1; + other.dir_.clear(); +} + +SpoolOwnerLock& SpoolOwnerLock::operator=(SpoolOwnerLock&& other) noexcept { + if (this != &other) { + Release(); + fd_ = other.fd_; + dir_ = std::move(other.dir_); + other.fd_ = -1; + other.dir_.clear(); + } + return *this; +} + +void SpoolOwnerLock::Release() { + if (fd_ >= 0) ::close(fd_); // closing the last descriptor unlocks + fd_ = -1; + dir_.clear(); +} + +SpoolStatus SpoolOwnerLock::Acquire(const std::string& dir, + bool allow_shared_filesystem, + SpoolOwnerLock* out, std::string* error) { + out->Release(); + if (dir.empty()) { + if (error) *error = "spool directory must not be empty"; + return SpoolStatus::kBadArgument; + } + const std::string canonical = CanonicalPath(dir, error); + if (canonical.empty()) return SpoolStatus::kIo; + std::error_code ec; + const bool exists = fs::is_directory(canonical, ec); + if (!exists && fs::exists(canonical, ec)) { + if (error) *error = "spool directory " + canonical + " is not a directory"; + return SpoolStatus::kBadArgument; + } + const std::string parent = fs::path(canonical).parent_path().string(); + if (!exists) { + fs::create_directories(parent, ec); + if (ec) { + if (error) *error = "cannot create " + parent + ": " + ec.message(); + return SpoolStatus::kIo; + } + } + SpoolStatus status = + CheckNodeLocal(exists ? canonical : parent, allow_shared_filesystem, + error); + if (status != SpoolStatus::kOk) return status; + status = CheckNotNested(canonical, exists, error); + if (status != SpoolStatus::kOk) return status; + int fd = -1; + status = exists ? LockInPlace(canonical, &fd, error) + : CreateLocked(canonical, &fd, error); + if (status != SpoolStatus::kOk) return status; + out->fd_ = fd; + out->dir_ = canonical; + return SpoolStatus::kOk; +} + +SpoolStatus SpoolOwnerLock::TryAdopt(const std::string& dir, + SpoolOwnerLock* out, + std::string* error) { + out->Release(); + char resolved[4096]; + if (::realpath(dir.c_str(), resolved) == nullptr || + !fs::is_directory(resolved)) { + if (error) *error = "no spool directory to adopt at " + dir; + return SpoolStatus::kBadArgument; + } + int fd = -1; + const SpoolStatus status = LockInPlace(resolved, &fd, error); + if (status != SpoolStatus::kOk) return status; + out->fd_ = fd; + out->dir_ = resolved; + return SpoolStatus::kOk; +} + +bool SpoolOwnerLock::ReleaseAndRemoveIfEmpty(std::string* error) { + if (!held()) return false; + const std::string dir = dir_; + const fs::path lock_file = fs::path(dir) / kOwnerLockFile; + std::vector subdirectories; + std::error_code ec; + for (auto it = fs::recursive_directory_iterator(dir, ec); + !ec && it != fs::recursive_directory_iterator(); it.increment(ec)) { + std::error_code type_ec; + if (it->is_directory(type_ec) && !it->is_symlink(type_ec)) { + subdirectories.push_back(it->path()); + } else if (it->path() != lock_file) { + Release(); // something is left: the directory stays as it is + return false; + } + } + if (ec) { + if (error) *error = "cannot list " + dir + ": " + ec.message(); + Release(); + return false; + } + // Deepest first, so each is empty when its turn comes. + std::sort(subdirectories.begin(), subdirectories.end(), + [](const fs::path& a, const fs::path& b) { + return a.string().size() > b.string().size(); + }); + for (const fs::path& subdirectory : subdirectories) { + ::rmdir(subdirectory.c_str()); + } + // Unlinked while held: a process that opens the file from here on creates + // a new one (and the rmdir below then fails, leaving it the directory); one + // that opened the old file first finds, once it locks it, that the file + // is no longer at the path (IsFileAt), and lets it go. + ::unlink(lock_file.c_str()); + const bool removed = ::rmdir(dir.c_str()) == 0; + if (!removed && error) { + *error = "cannot remove " + dir + ": " + Errno(errno); + } + if (removed) FsyncParent(dir); + Release(); + return removed; +} + +std::string SpoolCatalogKey(const std::string& database, + const std::string& table_prefix, + const std::string& store_id) { + const std::string text = database + "/" + table_prefix + "/" + store_id; + unsigned char digest[SHA256_DIGEST_LENGTH]; + SHA256(reinterpret_cast(text.data()), text.size(), + digest); + static const char* kHex = "0123456789abcdef"; + std::string out; + for (int i = 0; i < 6; ++i) { + out.push_back(kHex[digest[i] >> 4]); + out.push_back(kHex[digest[i] & 0xF]); + } + return out; +} + +namespace { +bool IsLowerHex(const std::string& text, size_t size) { + return text.size() == size && + std::all_of(text.begin(), text.end(), [](char c) { + return (c >= '0' && c <= '9') || (c >= 'a' && c <= 'f'); + }); +} } // namespace +bool IsSpoolCatalogKey(const std::string& name) { return IsLowerHex(name, 12); } + +std::string SpoolRankDirectoryName(uint64_t producer_rank, + const std::string& incarnation) { + return "r" + std::to_string(producer_rank) + "-" + incarnation; +} + +bool ParseSpoolRankDirectoryName(const std::string& name, + uint64_t* producer_rank, + std::string* incarnation) { + if (name.size() < 4 || name[0] != 'r') return false; + const size_t dash = name.find('-'); + if (dash == std::string::npos || dash < 2) return false; + const std::string digits = name.substr(1, dash - 1); + if (digits.size() > 19 || (digits.size() > 1 && digits[0] == '0') || + !std::all_of(digits.begin(), digits.end(), + [](char c) { return c >= '0' && c <= '9'; })) { + return false; + } + const std::string tail = name.substr(dash + 1); + if (!IsLowerHex(tail, 8)) return false; + *producer_rank = std::strtoull(digits.c_str(), nullptr, 10); + *incarnation = tail; + return true; +} + +std::string NewSpoolIncarnation() { + std::random_device random; + char out[16]; + std::snprintf(out, sizeof(out), "%08x", static_cast(random())); + return out; +} + +std::string SpoolRankDirectory(const std::string& base, + const std::string& database, + const std::string& table_prefix, + const std::string& store_id, + uint64_t producer_rank, + const std::string& incarnation) { + std::string root = base; + while (root.size() > 1 && root.back() == '/') root.pop_back(); + return root + "/" + SpoolCatalogKey(database, table_prefix, store_id) + + "/" + SpoolRankDirectoryName(producer_rank, incarnation); +} + SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { if (config.max_bytes == 0) { if (error) *error = "max_bytes must be positive"; return SpoolStatus::kBadArgument; } + if (config.root.empty()) { + if (error) *error = "spool root must not be empty"; + return SpoolStatus::kBadArgument; + } + // A re-opened object gives up the directory it owned first. + out->owner_lock_.Release(); std::error_code ec; + if (config.owner_lock == OwnerLock::kTake) { + // Before anything reads the directory: the accounting walk below, and + // above all Recover(), belong to its one owner. Creates the root. + const SpoolStatus locked = SpoolOwnerLock::Acquire( + config.root, config.allow_shared_filesystem, &out->owner_lock_, + error); + if (locked != SpoolStatus::kOk) return locked; + } else { + // The caller took the lock, so the directory and its lock file exist. + char held[4096]; + if (::realpath(config.root.c_str(), held) == nullptr || + !ReadSpoolOwner(held, nullptr)) { + if (error) { + *error = "spool owner_lock=held_by_caller, but nothing holds " + + config.root + "/" + kOwnerLockFile + + ": take a SpoolOwnerLock on the directory first, or open " + "it with owner_lock=take"; + } + return SpoolStatus::kBadArgument; + } + const SpoolStatus local = + CheckNodeLocal(held, config.allow_shared_filesystem, error); + if (local != SpoolStatus::kOk) return local; + } fs::create_directories(config.root, ec); if (ec) { if (error) *error = "cannot create spool root: " + ec.message(); @@ -270,9 +835,13 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { // files plus stale .open files both count until Recover() runs, and only // the ready PATHS are remembered (_accounted_ready = ready_bytes), so a // later retry of one of them is recognised as already counted. - for (const auto& entry : - fs::recursive_directory_iterator(out->root_, ec)) { - if (ec) break; + for (auto it = fs::recursive_directory_iterator(out->root_, ec); + it != fs::recursive_directory_iterator(); ++it) { + if (AtRefsDirectory(it)) { + it.disable_recursion_pending(); + continue; + } + const fs::directory_entry& entry = *it; if (!entry.is_regular_file()) continue; const std::string name = entry.path().filename().string(); const bool is_ready = HasSuffix(name, kReadySuffix); @@ -328,9 +897,13 @@ void Spool::ReconcileCommittedLocked() { uint64_t ready_count = 0; std::unordered_map seen_ready; std::error_code walk_ec; - for (const auto& entry : - fs::recursive_directory_iterator(root_, walk_ec)) { - if (walk_ec) break; + for (auto it = fs::recursive_directory_iterator(root_, walk_ec); + it != fs::recursive_directory_iterator(); ++it) { + if (AtRefsDirectory(it)) { + it.disable_recursion_pending(); + continue; + } + const fs::directory_entry& entry = *it; if (!entry.is_regular_file()) continue; const std::string name = entry.path().filename().string(); if (HasSuffix(name, kReadySuffix)) { @@ -635,9 +1208,13 @@ SpoolStatus Spool::Scan(std::vector* out, bool discard_open_files, std::vector readies; uint64_t bytes = 0; std::unordered_map seen_ready; - for (const auto& entry : - fs::recursive_directory_iterator(root_, ec)) { - if (ec) break; + for (auto it = fs::recursive_directory_iterator(root_, ec); + it != fs::recursive_directory_iterator(); ++it) { + if (AtRefsDirectory(it)) { + it.disable_recursion_pending(); + continue; + } + const fs::directory_entry& entry = *it; if (!entry.is_regular_file()) continue; const std::string path = entry.path().string(); const std::string name = entry.path().filename().string(); diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index ad3558653..317d62fe2 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -11,6 +11,33 @@ // the directory chain to the root. Idempotent: re-staging validates the // existing ready file (name, size, sha256) and returns it; a different pack // under the same pack_id is a conflict. +// +// One owner per directory (B6). Recover() deletes every .open file this +// object is not writing, so it is only safe while no other PROCESS writes +// there. A spool directory therefore has an owner lock -- flock(LOCK_EX) on +// /.owner.lock, whose content records the holder's host and pid -- +// and a second process that tries to take it is refused, told who holds +// it. The lock goes with its holder, even one killed with SIGKILL. +// - owner_lock=kTake (the default) takes it in Open(), before anything +// reads the directory, and holds it for the Spool object's life. +// - owner_lock=kHeldByCaller takes none: the caller holds a +// SpoolOwnerLock on the directory already. Two Spools in ONE process +// that both take refuse each other (flock binds to an open file +// description, not to the process), so a process running a sink and a +// storage service on one directory holds one SpoolOwnerLock and opens +// both Spools with kHeldByCaller. Open() refuses it when nothing holds +// the lock; it cannot tell who does. +// A directory nested under, or containing, an owned directory (one with a +// .owner.lock file, held or not) is refused: Scan walks recursively, so the +// outer spool's Recover would reach into the inner one. /_refs/ is +// never scanned: the upload handoff's ref files live there (plan section +// 2.4). The spool must be node-local: NFS and Lustre are refused by statfs +// f_type unless allow_shared_filesystem is set, since neither guarantees a +// flock that excludes a process on another node. +// +// The Python DurablePackSpool (spool.py) takes no lock, and its recover() +// deletes every .open file under its root; the C++ spool is deliberately +// stricter, and that is not ported to the reference. #ifndef DMI_STORE_SPOOL_H_ #define DMI_STORE_SPOOL_H_ @@ -25,9 +52,22 @@ namespace dmi_store { +enum class OwnerLock { + kTake = 0, // Open() takes /.owner.lock for the Spool's life + kHeldByCaller, // the caller holds a SpoolOwnerLock on +}; + +// "take" / "held_by_caller". +const char* OwnerLockName(OwnerLock mode); +bool ParseOwnerLock(const std::string& text, OwnerLock* mode); + struct SpoolConfig { std::string root; uint64_t max_bytes = 0; + OwnerLock owner_lock = OwnerLock::kTake; + // Admit a root on NFS or Lustre. Only safe when every process that could + // open the directory runs on this node. + bool allow_shared_filesystem = false; }; struct StagedPack { @@ -54,6 +94,7 @@ enum class SpoolStatus { kIntegrity, // ready file fails validation kIo, // filesystem error kBadArgument, + kOwned, // another owner holds the directory's owner lock }; inline const char* SpoolStatusName(SpoolStatus s) { @@ -64,14 +105,118 @@ inline const char* SpoolStatusName(SpoolStatus s) { case SpoolStatus::kIntegrity: return "ready pack failed validation"; case SpoolStatus::kIo: return "spool filesystem error"; case SpoolStatus::kBadArgument: return "invalid argument"; + case SpoolStatus::kOwned: return "spool directory is owned by another process"; } return "unknown"; } +// The statfs f_type names of the shared filesystems a spool refuses: "NFS", +// "Lustre", or nullptr for any other. The list needs maintenance as +// deployments meet new ones. +const char* SharedFilesystemName(int64_t f_type); + +// Refuses `dir` (which must exist) on a shared filesystem unless +// `allow_shared_filesystem`. +SpoolStatus CheckNodeLocal(const std::string& dir, + bool allow_shared_filesystem, std::string* error); + +// Test seam: every node-local check in this binary reads `f_type` instead of +// calling statfs(2). A negative value restores statfs. +void SetFilesystemTypeForTesting(int64_t f_type); + +// The holder recorded in a directory's owner lock file. +struct SpoolOwner { + std::string host; + int64_t pid = 0; +}; + +// Whether /.owner.lock is held right now (by any process, this one +// included), and if so who recorded themselves in it. A holder that has +// locked but not yet written its record reads as an empty host and pid 0. +bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner); + +// The owner lock of one spool directory: flock(LOCK_EX) on /.owner.lock, +// released with the object (or Release()), and by the kernel when the +// process dies. The descriptor is close-on-exec; a child forked WITHOUT exec +// shares it, and keeps the lock held for as long as it lives. +class SpoolOwnerLock { + public: + static constexpr const char* kFileName = ".owner.lock"; + + SpoolOwnerLock() = default; + ~SpoolOwnerLock(); + SpoolOwnerLock(SpoolOwnerLock&& other) noexcept; + SpoolOwnerLock& operator=(SpoolOwnerLock&& other) noexcept; + SpoolOwnerLock(const SpoolOwnerLock&) = delete; + SpoolOwnerLock& operator=(const SpoolOwnerLock&) = delete; + + // Takes the lock on `dir`, creating it when it does not exist -- beside + // its lock file, already held, and renamed into place, so no scan of the + // parent ever meets the directory before its owner holds it -- and records + // this host and pid in it. kOwned, naming the holder, when another holder + // has it; kBadArgument for a shared filesystem or a directory nested + // under, or containing, an owned one. + static SpoolStatus Acquire(const std::string& dir, + bool allow_shared_filesystem, + SpoolOwnerLock* out, std::string* error); + + // Adoption's try-lock: takes the lock of an EXISTING directory whose owner + // is gone, creating its lock file if it has none. kOwned while its owner + // lives; never creates the directory. + static SpoolStatus TryAdopt(const std::string& dir, SpoolOwnerLock* out, + std::string* error); + + bool held() const { return fd_ >= 0; } + // The canonical path of the directory, while held. + const std::string& directory() const { return dir_; } + + void Release(); + + // Releases the lock, first removing the directory if it holds nothing but + // its lock file and empty subdirectories. Anything else -- a pack, a + // quarantined file, a ref -- keeps the directory, which the next owner or + // adopter meets as it was left. Returns whether the directory was removed. + bool ReleaseAndRemoveIfEmpty(std::string* error); + + private: + int fd_ = -1; + std::string dir_; +}; + +// The spool layout of the plan's section 2.3. A capture process spools into +// //r-/ +// where catalog_key is the first 12 hex digits of +// sha256("//"), and incarnation is 8 hex +// digits fresh for every process start, so no two processes -- two jobs on +// one node, or a restart of the same rank -- ever share a directory. The +// directories under one catalog key are siblings: packs bound for one +// catalog and store, which a successor on the node adopts once their owner +// has died (CaptureStorageService, adopt_sibling_spools). +std::string SpoolCatalogKey(const std::string& database, + const std::string& table_prefix, + const std::string& store_id); +bool IsSpoolCatalogKey(const std::string& name); +std::string SpoolRankDirectoryName(uint64_t producer_rank, + const std::string& incarnation); +// "r-<8 lowercase hex>", the rank in canonical decimal. +bool ParseSpoolRankDirectoryName(const std::string& name, + uint64_t* producer_rank, + std::string* incarnation); +std::string NewSpoolIncarnation(); +std::string SpoolRankDirectory(const std::string& base, + const std::string& database, + const std::string& table_prefix, + const std::string& store_id, + uint64_t producer_rank, + const std::string& incarnation); + class Spool { public: - // Opens (creating) the root. Recovery of pre-existing files is explicit - // via Recover(), matching the Python constructor + recover() split. + // Opens (creating) the root, after the node-local check and, with kTake, + // after taking its owner lock (kOwned when another holder has it); + // kHeldByCaller is refused when nothing holds it. Recovery of + // pre-existing files is explicit via Recover(), matching the Python + // constructor + recover() split. static SpoolStatus Open(SpoolConfig config, Spool* out, std::string* error); Spool() = default; @@ -86,7 +231,9 @@ class Spool { // Startup cleanup: delete abandoned "*.open" files, validate ready packs, // quarantine failures, and rebuild accounting. Other writers on this root - // must be stopped; use ListPending() while they are running. + // must be stopped; use ListPending() while they are running. The owner + // lock keeps other processes out; writers in this process sharing it + // (kHeldByCaller) are the caller's to order. SpoolStatus Recover(std::vector* out, std::string* error); // Validate and list ready packs without deleting in-progress writes. @@ -97,6 +244,9 @@ class Spool { SpoolSnapshot Snapshot() const; + // The canonical root, once opened. + const std::string& root() const { return root_; } + // Test seam: called by Stage() after its capacity reservation is taken and // before the temp file is written, outside the lock. Lets a test hold one // stager at exactly the point where its reservation exists but nothing is @@ -128,6 +278,8 @@ class Spool { std::string root_; uint64_t max_bytes_ = 0; + // Held for the object's life under OwnerLock::kTake; empty otherwise. + SpoolOwnerLock owner_lock_; mutable std::mutex mutex_; // Two accounts, kept apart on purpose: // committed_bytes_/committed_entries_ -- ready files this object knows diff --git a/tests/native/live_spool_stage.cpp b/tests/native/live_spool_stage.cpp index 53047cc94..ab69b629b 100644 --- a/tests/native/live_spool_stage.cpp +++ b/tests/native/live_spool_stage.cpp @@ -1,15 +1,29 @@ -// Hold a real Stage after opening its temp file, while an uploader scans. +// Hold a real Stage after opening its temp file, while another process tries +// the same spool. +// +// live_spool_stage [packs] +// +// Opens with the default owner lock (take) and stages `packs` packs +// (default 1). When the LAST one has created its temp file it prints OPEN +// and waits for a line on stdin, so the test can act -- or SIGKILL it -- +// with a live writer and an in-flight .open file on disk. #include "store/spool.h" #include #include #include +#include #include #include #include #include #include +namespace { +int g_pause_at = 1; +int g_opened = 0; +} // namespace + extern "C" int __real_open(const char*, int, ...); extern "C" int __wrap_open(const char* path, int flags, ...) { mode_t mode = 0; @@ -20,7 +34,8 @@ extern "C" int __wrap_open(const char* path, int flags, ...) { va_end(args); } const int fd = __real_open(path, flags, mode); - if (fd >= 0 && (flags & O_CREAT) && std::strstr(path, ".open")) { + if (fd >= 0 && (flags & O_CREAT) && std::strstr(path, ".open") && + ++g_opened == g_pause_at) { std::cout << "OPEN" << std::endl; std::string release; std::getline(std::cin, release); @@ -29,23 +44,34 @@ extern "C" int __wrap_open(const char* path, int flags, ...) { } int main(int argc, char** argv) { - if (argc != 2) return 2; + if (argc != 2 && argc != 3) return 2; + const int packs = argc == 3 ? std::atoi(argv[2]) : 1; + if (packs < 1) return 2; + g_pause_at = packs; dmi_store::Spool spool; std::string error; - if (dmi_store::Spool::Open({argv[1], 5000}, &spool, &error) != - dmi_store::SpoolStatus::kOk) return 3; - const std::vector data(1000, 42); - unsigned char digest[SHA256_DIGEST_LENGTH]; - SHA256(data.data(), data.size(), digest); - char checksum[65]; - for (int i = 0; i < SHA256_DIGEST_LENGTH; ++i) { - std::snprintf(checksum + 2 * i, 3, "%02x", digest[i]); + if (dmi_store::Spool::Open({argv[1], 50000}, &spool, &error) != + dmi_store::SpoolStatus::kOk) { + std::cout << "open failed: " << error << std::endl; + return 3; + } + for (int n = 1; n <= packs; ++n) { + const std::vector data(1000, static_cast(41 + n)); + unsigned char digest[SHA256_DIGEST_LENGTH]; + SHA256(data.data(), data.size(), digest); + char checksum[65]; + for (int i = 0; i < SHA256_DIGEST_LENGTH; ++i) { + std::snprintf(checksum + 2 * i, 3, "%02x", digest[i]); + } + char id[40]; + std::snprintf(id, sizeof(id), "018f0000-0000-7000-8000-%012d", n); + dmi_store::StagedPack staged; + const auto status = spool.Stage( + id, 1700000000000000000ull + n, 1, checksum, + std::string(id) + ".dmi-pack", data.data(), data.size(), &staged, + &error); + std::cout << dmi_store::SpoolStatusName(status) << ": " << error << '\n'; + if (status != dmi_store::SpoolStatus::kOk) return 1; } - const std::string id = "018f0000-0000-7000-8000-000000000001"; - dmi_store::StagedPack staged; - const auto status = spool.Stage( - id, 1700000000000000000ull, 1, checksum, id + ".dmi-pack", - data.data(), data.size(), &staged, &error); - std::cout << dmi_store::SpoolStatusName(status) << ": " << error << '\n'; - return status == dmi_store::SpoolStatus::kOk ? 0 : 1; + return 0; } diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp new file mode 100644 index 000000000..320b77ebe --- /dev/null +++ b/tests/native/test_spool_owner_lock.cpp @@ -0,0 +1,390 @@ +// B6: the spool owner lock, in process and across a fork. +// +// 1. Two Spool objects that both TAKE one directory refuse each other, +// even in one process: flock binds to an open file description, not to +// the process. That is why the engine holds one SpoolOwnerLock and both +// of its Spools (sink and service) open with held_by_caller. +// 2. held_by_caller opens beside a holder, and is refused when nothing +// holds the lock. +// 3. A second process is refused, told the holder's pid and host; the +// lock goes with its holder, even one killed with SIGKILL. +// 4. Nesting: a directory under, or containing, an owned one is refused. +// 5. The node-local check refuses NFS and Lustre by statfs f_type, unless +// explicitly allowed (a test seam stands in for statfs). +// 6. Adoption's try-lock never creates a directory, and a released +// directory that holds nothing but its lock file can be removed. +// 7. The directory layout of the plan's section 2.3. +// +// Built and run by tests/test_native_spool_owner_lock_unit.py. + +#include +#include +#include +#include + +#include +#include +#include +#include +#include +#include +#include +#include + +#include "store/spool.h" + +namespace fs = std::filesystem; +using dmi_store::OwnerLock; +using dmi_store::Spool; +using dmi_store::SpoolConfig; +using dmi_store::SpoolOwnerLock; +using dmi_store::SpoolStatus; + +namespace { + +int g_failures = 0; + +#define CHECK(cond) \ + do { \ + if (!(cond)) { \ + std::cerr << __FILE__ << ":" << __LINE__ << ": CHECK failed: " #cond \ + << "\n"; \ + ++g_failures; \ + } \ + } while (0) + +bool Contains(const std::string& text, const std::string& part) { + return text.find(part) != std::string::npos; +} + +std::string FreshRoot(const char* tag) { + const char* base = std::getenv("SPOOL_TEST_ROOT"); + const std::string root = + std::string(base != nullptr ? base : "/tmp") + "/owner-" + tag; + fs::remove_all(root); + fs::create_directories(root); + return fs::canonical(root).string(); +} + +std::string Hostname() { + char host[256] = {0}; + ::gethostname(host, sizeof(host) - 1); + return host; +} + +std::string Sha256Hex(const std::string& data) { + unsigned char digest[SHA256_DIGEST_LENGTH]; + SHA256(reinterpret_cast(data.data()), data.size(), + digest); + static const char* kHex = "0123456789abcdef"; + std::string out(64, '0'); + for (int i = 0; i < 32; ++i) { + out[2 * i] = kHex[digest[i] >> 4]; + out[2 * i + 1] = kHex[digest[i] & 0xF]; + } + return out; +} + +SpoolStatus StageOne(Spool& spool, int n, std::string* error) { + char id[40]; + std::snprintf(id, sizeof(id), "018f0000-0000-7000-8000-%012d", n); + const std::string data(100, static_cast('a' + n)); + dmi_store::StagedPack staged; + return spool.Stage(id, 1700000000000000000ull + n, 1, Sha256Hex(data), + std::string("v1/") + id + ".dmi-pack", + reinterpret_cast(data.data()), + data.size(), &staged, error); +} + +// (1) The regression a per-Spool lock would cause in the engine. +void TestTwoTakesInOneProcessRefuseEachOther() { + const std::string root = FreshRoot("two-takes") + "/spool"; + Spool service, sink; + std::string error; + CHECK(Spool::Open({root, 1 << 20}, &service, &error) == SpoolStatus::kOk); + error.clear(); + CHECK(Spool::Open({root, 1 << 20}, &sink, &error) == SpoolStatus::kOwned); + CHECK(Contains(error, "pid " + std::to_string(::getpid()))); + CHECK(Contains(error, Hostname())); +} + +// (2) held_by_caller beside a holder -- the engine's shape: one lock, two +// Spools, both staging and listing. +void TestHeldByCallerOpensBesideTheHolder() { + const std::string root = FreshRoot("held") + "/spool"; + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk); + CHECK(lock.held()); + SpoolConfig config{root, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool service, sink; + CHECK(Spool::Open(config, &service, &error) == SpoolStatus::kOk); + CHECK(Spool::Open(config, &sink, &error) == SpoolStatus::kOk); + CHECK(StageOne(sink, 1, &error) == SpoolStatus::kOk); + std::vector pending; + CHECK(service.ListPending(&pending, &error) == SpoolStatus::kOk); + CHECK(pending.size() == 1); + // A take is still refused while the engine's lock is held. + Spool rival; + CHECK(Spool::Open({root, 1 << 20}, &rival, &error) == SpoolStatus::kOwned); +} + +void TestHeldByCallerWithoutAHolderIsRefused() { + const std::string root = FreshRoot("unheld") + "/spool"; + SpoolConfig config{root, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool spool; + std::string error; + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kBadArgument); + CHECK(Contains(error, "held_by_caller")); + // A lock file nobody holds is refused too. + { + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk); + } + error.clear(); + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kBadArgument); + CHECK(Contains(error, "held_by_caller")); +} + +void TestTheLockGoesWithItsSpool() { + const std::string root = FreshRoot("scope") + "/spool"; + std::string error; + { + Spool first; + CHECK(Spool::Open({root, 1 << 20}, &first, &error) == SpoolStatus::kOk); + } + Spool second; + CHECK(Spool::Open({root, 1 << 20}, &second, &error) == SpoolStatus::kOk); +} + +// (3) Across a fork: the child takes the lock, the parent is refused and +// told who holds it; the lock goes with the child, even on SIGKILL. +void TestASecondProcessIsRefusedUntilTheHolderDies() { + const std::string root = FreshRoot("fork") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t child = ::fork(); + if (child == 0) { + ::close(ready[0]); + SpoolOwnerLock lock; + std::string error; + const bool ok = SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk; + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOwned); + CHECK(Contains(error, "pid " + std::to_string(child))); + CHECK(Contains(error, Hostname())); + CHECK(Contains(error, root)); + dmi_store::SpoolOwner owner; + CHECK(dmi_store::ReadSpoolOwner(root, &owner)); + CHECK(owner.pid == child); + CHECK(owner.host == Hostname()); + Spool spool; + CHECK(Spool::Open({root, 1 << 20}, &spool, &error) == SpoolStatus::kOwned); + + ::kill(child, SIGKILL); + int status = 0; + ::waitpid(child, &status, 0); + CHECK(!dmi_store::ReadSpoolOwner(root, &owner)); + error.clear(); + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk); + CHECK(dmi_store::ReadSpoolOwner(root, &owner)); + CHECK(owner.pid == ::getpid()); +} + +// (4) Nesting, both ways. +void TestNestedDirectoriesAreRefused() { + const std::string base = FreshRoot("nested"); + std::string error; + SpoolOwnerLock outer; + CHECK(SpoolOwnerLock::Acquire(base + "/outer", false, &outer, &error) == + SpoolStatus::kOk); + SpoolOwnerLock inner; + CHECK(SpoolOwnerLock::Acquire(base + "/outer/inner", false, &inner, + &error) == SpoolStatus::kBadArgument); + CHECK(Contains(error, "nested")); + CHECK(Contains(error, base + "/outer")); + + SpoolOwnerLock deep; + CHECK(SpoolOwnerLock::Acquire(base + "/other/a/b", false, &deep, &error) == + SpoolStatus::kOk); + SpoolOwnerLock ancestor; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(base + "/other", false, &ancestor, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "contains")); + CHECK(Contains(error, base + "/other/a/b")); + Spool spool; + CHECK(Spool::Open({base + "/other", 1 << 20}, &spool, &error) == + SpoolStatus::kBadArgument); +} + +// (5) The node-local check, through the test seam. +void TestSharedFilesystemsAreRefusedUnlessAllowed() { + const std::string root = FreshRoot("statfs") + "/spool"; + CHECK(std::string(dmi_store::SharedFilesystemName(0x6969)) == "NFS"); + CHECK(std::string(dmi_store::SharedFilesystemName(0x0BD00BD0)) == + "Lustre"); + CHECK(dmi_store::SharedFilesystemName(0xEF53) == nullptr); // ext4 + CHECK(dmi_store::SharedFilesystemName(0x58465342) == nullptr); // xfs + + std::string error; + for (const int64_t magic : {int64_t{0x6969}, int64_t{0x0BD00BD0}}) { + dmi_store::SetFilesystemTypeForTesting(magic); + SpoolOwnerLock lock; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, magic == 0x6969 ? "NFS" : "Lustre")); + CHECK(Contains(error, "node-local")); + CHECK(!lock.held()); + Spool spool; + error.clear(); + CHECK(Spool::Open({root, 1 << 20}, &spool, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "node-local")); + // The explicit override. + SpoolConfig allowed{root, 1 << 20}; + allowed.allow_shared_filesystem = true; + CHECK(Spool::Open(allowed, &spool, &error) == SpoolStatus::kOk); + } + { + dmi_store::SetFilesystemTypeForTesting(0x6969); + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(root + "-allowed", true, &lock, &error) == + SpoolStatus::kOk); + } + dmi_store::SetFilesystemTypeForTesting(-1); + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(root + "-real", false, &lock, &error) == + SpoolStatus::kOk); +} + +// (6) Adoption's try-lock, and removing a drained directory. +void TestAdoptionLocksOnlyWhatExistsAndIsDead() { + const std::string base = FreshRoot("adopt"); + std::string error; + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::TryAdopt(base + "/missing", &lock, &error) != + SpoolStatus::kOk); + CHECK(!fs::exists(base + "/missing")); + + // A directory with no lock file at all: nobody owns it. + fs::create_directories(base + "/bare/v1"); + CHECK(SpoolOwnerLock::TryAdopt(base + "/bare", &lock, &error) == + SpoolStatus::kOk); + CHECK(lock.held()); + CHECK(lock.ReleaseAndRemoveIfEmpty(&error)); + CHECK(!lock.held()); + CHECK(!fs::exists(base + "/bare")); + + // A live owner is left alone. + SpoolOwnerLock live; + CHECK(SpoolOwnerLock::Acquire(base + "/live", false, &live, &error) == + SpoolStatus::kOk); + SpoolOwnerLock adopter; + CHECK(SpoolOwnerLock::TryAdopt(base + "/live", &adopter, &error) == + SpoolStatus::kOwned); + CHECK(!adopter.held()); + + // Anything but the lock file and empty directories keeps the directory. + SpoolOwnerLock kept; + CHECK(SpoolOwnerLock::Acquire(base + "/kept", false, &kept, &error) == + SpoolStatus::kOk); + fs::create_directories(base + "/kept/v1/tenant=t"); + std::ofstream(base + "/kept/v1/tenant=t/x.quarantined") << "bytes"; + CHECK(!kept.ReleaseAndRemoveIfEmpty(&error)); + CHECK(!kept.held()); + CHECK(fs::exists(base + "/kept/v1/tenant=t/x.quarantined")); + CHECK(fs::exists(base + "/kept/.owner.lock")); +} + +void TestANewDirectoryAppearsWithItsLockHeld() { + // Created beside its lock file and renamed into place, so no scan of the + // parent can meet the directory before its owner holds it. + const std::string base = FreshRoot("atomic"); + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::Acquire(base + "/r0-0a1b2c3d", false, &lock, + &error) == SpoolStatus::kOk); + std::set names; + for (const auto& entry : fs::directory_iterator(base)) { + names.insert(entry.path().filename().string()); + } + CHECK(names == std::set{"r0-0a1b2c3d"}); + CHECK(lock.directory() == base + "/r0-0a1b2c3d"); + CHECK(fs::exists(base + "/r0-0a1b2c3d/.owner.lock")); +} + +// (7) The layout: //r-/. +void TestTheDirectoryLayout() { + const std::string key = dmi_store::SpoolCatalogKey("db", "prefix", "s3"); + CHECK(key == Sha256Hex("db/prefix/s3").substr(0, 12)); + CHECK(dmi_store::IsSpoolCatalogKey(key)); + CHECK(!dmi_store::IsSpoolCatalogKey("0123456789aB")); + CHECK(!dmi_store::IsSpoolCatalogKey("0123456789a")); + CHECK(dmi_store::SpoolCatalogKey("db", "prefix", "s3") != + dmi_store::SpoolCatalogKey("db", "prefix", "s4")); + + CHECK(dmi_store::SpoolRankDirectoryName(3, "0a1b2c3d") == "r3-0a1b2c3d"); + uint64_t rank = 0; + std::string incarnation; + CHECK(dmi_store::ParseSpoolRankDirectoryName("r12-deadbeef", &rank, + &incarnation)); + CHECK(rank == 12 && incarnation == "deadbeef"); + for (const char* bad : {"r-1-deadbeef", "r1-DEADBEEF", "r1-deadbee", + "r1-deadbeef0", "rx-deadbeef", "1-deadbeef", + "r1deadbeef", ".r1-deadbeef", "r01-deadbeef"}) { + CHECK(!dmi_store::ParseSpoolRankDirectoryName(bad, &rank, &incarnation)); + } + std::set seen; + for (int i = 0; i < 64; ++i) { + const std::string fresh = dmi_store::NewSpoolIncarnation(); + CHECK(dmi_store::ParseSpoolRankDirectoryName("r0-" + fresh, &rank, + &incarnation)); + seen.insert(fresh); + } + CHECK(seen.size() == 64); + CHECK(dmi_store::SpoolRankDirectory("/b", "db", "prefix", "s3", 2, + "0a1b2c3d") == + "/b/" + key + "/r2-0a1b2c3d"); +} + +} // namespace + +int main() { + TestTwoTakesInOneProcessRefuseEachOther(); + TestHeldByCallerOpensBesideTheHolder(); + TestHeldByCallerWithoutAHolderIsRefused(); + TestTheLockGoesWithItsSpool(); + TestASecondProcessIsRefusedUntilTheHolderDies(); + TestNestedDirectoriesAreRefused(); + TestSharedFilesystemsAreRefusedUnlessAllowed(); + TestAdoptionLocksOnlyWhatExistsAndIsDead(); + TestANewDirectoryAppearsWithItsLockHeld(); + TestTheDirectoryLayout(); + if (g_failures != 0) { + std::cerr << g_failures << " check(s) failed\n"; + return 1; + } + std::cout << "ok\n"; + return 0; +} diff --git a/tests/native/test_spool_reservations.cpp b/tests/native/test_spool_reservations.cpp index 43a7659c4..68abcfcbd 100644 --- a/tests/native/test_spool_reservations.cpp +++ b/tests/native/test_spool_reservations.cpp @@ -104,6 +104,16 @@ std::string FreshRoot(const char* tag) { return root; } +// The config for a SECOND Spool object on a root the first one opened. The +// first takes the directory's owner lock; in one process the second shares +// it (owner_lock=held_by_caller), as the engine's sink and storage service +// do -- two takes would refuse each other, since flock binds to an open file +// description rather than to the process. +dmi_store::SpoolConfig Beside(dmi_store::SpoolConfig config) { + config.owner_lock = dmi_store::OwnerLock::kHeldByCaller; + return config; +} + // (1) The serial uploader case: stage, remove through a second object, // stage again on the first object. Must succeed, and the accounting must // end at exactly one file. @@ -114,7 +124,7 @@ void TestSerialRemoveThroughAnotherSpoolIsReconciled() { std::string error; CHECK(dmi_store::Spool::Open(config, &writer, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &uploader, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &uploader, &error) == dmi_store::SpoolStatus::kOk); dmi_store::StagedPack first; @@ -246,7 +256,7 @@ void TestALostLinkRaceReleasesItsReservation() { std::string error; CHECK(dmi_store::Spool::Open(config, &loser, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &winner, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &winner, &error) == dmi_store::SpoolStatus::kOk); dmi_store::StagedPack winner_out; @@ -295,7 +305,7 @@ void TestARetryOfAnotherObjectsReadyFileIsAccounted() { std::string error; CHECK(dmi_store::Spool::Open(config, &writer, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &other, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &other, &error) == dmi_store::SpoolStatus::kOk); // `other` writes the ready file; `writer` has never seen it. @@ -345,7 +355,7 @@ void TestARetryCannotOversubscribeAnInflightReservation() { std::string error; CHECK(dmi_store::Spool::Open(config, &spool, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &other, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &other, &error) == dmi_store::SpoolStatus::kOk); dmi_store::SpoolStatus retry_status = dmi_store::SpoolStatus::kIo; @@ -405,7 +415,7 @@ void TestAnEexistLoserAccountsForTheWinnersFile() { std::string error; CHECK(dmi_store::Spool::Open(config, &loser, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &winner, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &winner, &error) == dmi_store::SpoolStatus::kOk); // The loser reserves, and while it is paused the winner links the ready @@ -484,7 +494,7 @@ void TestASerialRetryIsAdmittedEvenWhenAlreadyOverCap() { // retry now goes through a Spool that has NOT ledgered this path, so it // reconciles, sees 1000 > 900, and takes the capacity decision. dmi_store::Spool reopened; - dmi_store::SpoolConfig lowered{root, 900}; + dmi_store::SpoolConfig lowered = Beside({root, 900}); CHECK(dmi_store::Spool::Open(lowered, &reopened, &error) == dmi_store::SpoolStatus::kOk); @@ -530,7 +540,7 @@ void TestRecoverThenRemoveReleasesTheAccount() { { dmi_store::Spool writer; - CHECK(dmi_store::Spool::Open(config, &writer, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &writer, &error) == dmi_store::SpoolStatus::kOk); dmi_store::StagedPack out; CHECK(StageBytes(writer, 1, 1000, &out, &error) == @@ -596,7 +606,7 @@ void TestRemoveOfAnAlreadyDeletedFileStillUncounts() { std::string error; CHECK(dmi_store::Spool::Open(config, &writer, &error) == dmi_store::SpoolStatus::kOk); - CHECK(dmi_store::Spool::Open(config, &other, &error) == + CHECK(dmi_store::Spool::Open(Beside(config), &other, &error) == dmi_store::SpoolStatus::kOk); dmi_store::StagedPack staged; diff --git a/tests/test_native_live_spool.py b/tests/test_native_live_spool.py index 658a0256d..e0a5669ec 100644 --- a/tests/test_native_live_spool.py +++ b/tests/test_native_live_spool.py @@ -1,13 +1,26 @@ -"""An upload scan must not run crash cleanup against a live pack writer.""" -import select +"""A live pack writer owns its spool: no other process may sweep it. + +Crash cleanup (Recover) deletes every ``.open`` file its Spool is not +writing itself, so running it against a live writer in another process +deletes the writer's in-flight temp file and fails its stage with "cannot +link ready file". The spool's owner lock makes that impossible rather than +merely avoided: while the writer holds its directory, a second process's +open is refused, naming the writer. Once the writer dies -- even by +SIGKILL, mid-stage -- the lock goes with it, and the next process recovers +the directory: the stale temp file is swept and the sealed pack uploads. +""" +import os import shutil +import signal +import socket import subprocess import sys +import threading from pathlib import Path import pytest -from tests.test_native_s3_client import fake_s3 +from tests.test_native_s3_client import STATE, fake_s3 from tests.test_native_uploader import DriverSession, STORE_DRIVER, _upload_pending pytestmark = [ @@ -16,8 +29,11 @@ pytest.mark.skipif(not STORE_DRIVER.exists(), reason="native store driver not built"), ] +SPOOL_DRIVER = STORE_DRIVER.parent / "conformance_spool" + -def test_upload_scan_preserves_an_inflight_stage(fake_s3, tmp_path): +@pytest.fixture +def writer_binary(tmp_path) -> Path: compiler = shutil.which("g++") if compiler is None: pytest.skip("a C++17 compiler is required") @@ -31,22 +47,96 @@ def test_upload_scan_preserves_an_inflight_stage(fake_s3, tmp_path): "-o", str(executable)], check=True, capture_output=True, text=True, ) - spool = tmp_path / "spool" + return executable + + +def _start_writer(binary: Path, spool: Path, packs: int = 1): + """A writer paused with its last pack's temp file open.""" process = subprocess.Popen( - [str(executable), str(spool)], stdin=subprocess.PIPE, + [str(binary), str(spool), str(packs)], stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True, ) + # Blocking reads under a watchdog: select() on the pipe cannot see a line + # the text wrapper has already buffered behind an earlier one. + watchdog = threading.Timer(20, process.kill) + watchdog.start() + seen = [] + try: + for line in process.stdout: + if line.strip() == "OPEN": + return process + seen.append(line.strip()) # an earlier pack's stage status + finally: + watchdog.cancel() + raise AssertionError( + f"writer never opened its temp file: {seen} {process.stderr.read()}") + + +def _recover(spool: Path) -> dict: + import json + + proc = subprocess.run( + [str(SPOOL_DRIVER)], + input=json.dumps({"op": "recover", "root": str(spool), + "max_bytes": 1 << 30}) + "\n", + capture_output=True, text=True, timeout=30) + return json.loads(proc.stdout.strip()) + + +def test_an_upload_scan_is_refused_while_a_writer_holds_the_spool( + fake_s3, tmp_path, writer_binary): + spool = tmp_path / "spool" + process = _start_writer(writer_binary, spool) uploader = DriverSession(STORE_DRIVER) try: - assert select.select([process.stdout], [], [], 10)[0], "writer did not open its temp file" - assert process.stdout.readline().strip() == "OPEN" result = _upload_pending(uploader, fake_s3, spool) + # Refused before it touched anything, naming the holder. + assert not result["ok"], result + assert f"pid {process.pid}" in result["what"], result + assert socket.gethostname() in result["what"], result + assert len(list(spool.rglob("*.open"))) == 1 + output, error = process.communicate("continue\n", timeout=10) - assert result["ok"], result assert process.returncode == 0, output + error assert len(list(spool.rglob("*.dmi-pack.ready"))) == 1 + assert list(spool.rglob("*.open")) == [] + + # The writer is gone, and its lock with it. + result = _upload_pending(uploader, fake_s3, spool) + assert result["ok"], result + assert len(result["refs"]) == 1, result + assert list(spool.rglob("*.dmi-pack.ready")) == [] finally: uploader.close() if process.poll() is None: process.kill() process.communicate() + + +def test_a_sigkilled_writers_spool_is_recovered_by_the_next_process( + fake_s3, tmp_path, writer_binary): + spool = tmp_path / "spool" + # The first pack is sealed; the second is mid-stage when the writer dies. + process = _start_writer(writer_binary, spool, packs=2) + assert len(list(spool.rglob("*.dmi-pack.ready"))) == 1 + (stale,) = spool.rglob("*.open") + os.kill(process.pid, signal.SIGKILL) + process.communicate(timeout=10) + assert process.returncode == -signal.SIGKILL + + recovered = _recover(spool) + assert recovered["ok"], recovered + assert len(recovered["staged"]) == 1, recovered + assert not stale.exists() + + uploader = DriverSession(STORE_DRIVER) + try: + result = _upload_pending(uploader, fake_s3, spool) + finally: + uploader.close() + assert result["ok"], result + assert [ref["pack_id"] for ref in result["refs"]] == [ + recovered["staged"][0]["pack_id"]] + assert list(spool.rglob("*.dmi-pack.ready")) == [] + with STATE.lock: + assert result["refs"][0]["object_key"] in STATE.objects diff --git a/tests/test_native_spool_owner_lock.py b/tests/test_native_spool_owner_lock.py new file mode 100644 index 000000000..5732e2854 --- /dev/null +++ b/tests/test_native_spool_owner_lock.py @@ -0,0 +1,231 @@ +"""B6: one owner per spool directory, across processes. + +A spool's crash cleanup (Recover) deletes every ``.open`` file its own object +is not writing, so it is only safe while no other process writes there. The +C++ spool therefore takes an owner lock -- flock on ``/.owner.lock`` -- +when it opens a directory with ``owner_lock="take"`` (the default), and a +second process that tries is refused, told who holds it. What that lock +must also refuse, and what it must leave alone: + +* a directory nested under, or containing, an owned directory: Scan walks + recursively, so the outer spool's cleanup would reach into the inner one; +* ``owner_lock="held_by_caller"`` with nothing holding the lock: that mode + opens without a lock of its own, on the caller's word that one is held; +* ``/_refs/``, where the upload handoff's ref files will live: no scan + may sweep, quarantine or list anything under it. + +The drivers are separate processes, so these are real cross-process locks. +The in-process cases -- two Spool objects in one process, the node-local +check, the directory layout -- are pinned by tests/native/ +test_spool_owner_lock.cpp (test_native_spool_owner_lock_unit.py). + +The Python spool (dmi.storage.capture.spool) takes no lock; the C++ spool is +deliberately stricter, and that is not ported to the reference. + +Build: make -C native build/conformance_spool build/conformance_sink +""" + +from __future__ import annotations + +import json +import socket +import subprocess +from pathlib import Path + +import pytest + +REPO_ROOT = Path(__file__).resolve().parents[1] +BUILD = REPO_ROOT / "native" / "build" +SPOOL_DRIVER = BUILD / "conformance_spool" +SINK_DRIVER = BUILD / "conformance_sink" + +pytestmark = [ + pytest.mark.cpu, + pytest.mark.skipif( + not (SPOOL_DRIVER.exists() and SINK_DRIVER.exists()), + reason="native spool/sink drivers are not built; run " + "`make -C native build/conformance_spool build/conformance_sink`", + ), +] + +READY_NAME = ( + "018f0000-0000-7000-8000-000000000001.1700000000000000000.1." + + "0" * 64 + ".dmi-pack.ready" +) + + +class _Holder: + """A conformance_sink process whose open PackSink holds a spool.""" + + def __init__(self, root: Path, **fields): + self.proc = subprocess.Popen( + [str(SINK_DRIVER)], stdin=subprocess.PIPE, stdout=subprocess.PIPE, + text=True, bufsize=1) + self.opened = self.call( + op="open", root=str(root), max_bytes=1 << 30, + max_queue_records=16, max_queue_bytes=1 << 20, + max_pack_bytes=1 << 20, max_pack_records=16, + max_linger_ns=1_000_000_000, overload="drop_newest", + admission_timeout=-1, **fields) + + @property + def pid(self) -> int: + return self.proc.pid + + def call(self, **fields) -> dict: + self.proc.stdin.write(json.dumps(fields) + "\n") + self.proc.stdin.flush() + return json.loads(self.proc.stdout.readline()) + + def close(self): + if self.opened.get("ok"): + self.call(op="close", timeout=10) + try: + self.proc.stdin.close() + except BrokenPipeError: + pass + self.proc.wait(timeout=30) + + +def _spool(**fields) -> dict: + """One conformance_spool op in a fresh process (it opens per op).""" + fields.setdefault("max_bytes", 1 << 30) + proc = subprocess.run( + [str(SPOOL_DRIVER)], input=json.dumps(fields) + "\n", + capture_output=True, text=True, timeout=60) + lines = [line for line in proc.stdout.splitlines() if line.strip()] + assert lines, proc.stderr + return json.loads(lines[0]) + + +def test_a_second_process_is_refused_and_told_who_holds_the_spool(tmp_path): + root = tmp_path / "spool" + holder = _Holder(root) + try: + assert holder.opened["ok"], holder.opened + response = _spool(op="recover", root=str(root)) + assert not response["ok"], response + assert response["status"] == "open", response + what = response["what"] + assert f"pid {holder.pid}" in what, what + assert socket.gethostname() in what, what + assert str(root) in what, what + finally: + holder.close() + # The lock goes with its holder: the next process opens the directory. + assert _spool(op="recover", root=str(root))["ok"] + + +def test_the_lock_file_records_the_holder(tmp_path): + root = tmp_path / "spool" + holder = _Holder(root) + try: + assert holder.opened["ok"], holder.opened + record = (root / ".owner.lock").read_text() + assert record.split() == [socket.gethostname(), str(holder.pid)] + finally: + holder.close() + + +def test_a_directory_nested_under_an_owned_one_is_refused(tmp_path): + outer = tmp_path / "spool" + holder = _Holder(outer) + try: + assert holder.opened["ok"], holder.opened + response = _spool(op="recover", root=str(outer / "inner")) + assert not response["ok"], response + assert "nested" in response["what"], response + assert str(outer) in response["what"], response + finally: + holder.close() + + +def test_a_directory_containing_an_owned_one_is_refused(tmp_path): + outer = tmp_path / "spool" + holder = _Holder(outer / "a" / "inner") + try: + assert holder.opened["ok"], holder.opened + response = _spool(op="recover", root=str(outer)) + assert not response["ok"], response + assert "contains" in response["what"], response + assert str(outer / "a" / "inner") in response["what"], response + finally: + holder.close() + + +def test_a_stale_lock_file_still_marks_an_owned_directory(tmp_path): + """Nesting is judged by the lock FILE, not by a live holder: an outer + directory some spool once owned is still a spool directory, and the + next process to open it would sweep the inner one.""" + outer = tmp_path / "spool" + assert _spool(op="recover", root=str(outer))["ok"] + assert (outer / ".owner.lock").exists() + response = _spool(op="recover", root=str(outer / "inner")) + assert not response["ok"], response + assert "nested" in response["what"], response + + +def test_held_by_caller_opens_beside_the_holder(tmp_path): + root = tmp_path / "spool" + holder = _Holder(root) + try: + assert holder.opened["ok"], holder.opened + response = _spool(op="recover", root=str(root), + owner_lock="held_by_caller") + assert response["ok"], response + finally: + holder.close() + + +def test_held_by_caller_is_refused_when_nothing_holds_the_lock(tmp_path): + root = tmp_path / "spool" + response = _spool(op="recover", root=str(root), + owner_lock="held_by_caller") + assert not response["ok"], response + assert "held_by_caller" in response["what"], response + # Once some spool has taken and let go of the lock, the file exists and + # is unlocked: still refused. + assert _spool(op="recover", root=str(root))["ok"] + response = _spool(op="recover", root=str(root), + owner_lock="held_by_caller") + assert not response["ok"], response + + +def test_an_unknown_owner_lock_mode_is_refused(tmp_path): + response = _spool(op="recover", root=str(tmp_path / "spool"), + owner_lock="share") + assert not response["ok"], response + assert "owner_lock" in response["what"], response + + +def test_recovery_leaves_the_refs_directory_alone(tmp_path): + """/_refs/ will hold the upload handoff's ref files (E2a). A + sweep that deleted, quarantined or listed them would break it.""" + root = tmp_path / "spool" + refs = root / "_refs" + refs.mkdir(parents=True) + stale = refs / ".018f0000-0000-7000-8000-000000000001.abcd1234.open" + stale.write_bytes(b"in progress") + bogus = refs / READY_NAME # the wrong checksum: would be quarantined + bogus.write_bytes(b"not a pack") + ref = refs / "018f0000-0000-7000-8000-000000000001.uploaded" + ref.write_text("{}") + + response = _spool(op="recover", root=str(root)) + + assert response["ok"], response + assert response["staged"] == [] + assert response["snapshot"] == {"entries": 0, "bytes": 0} + assert stale.exists() and bogus.exists() and ref.exists() + assert sorted(p.name for p in refs.iterdir()) == sorted( + [stale.name, bogus.name, ref.name]) + + +def test_the_open_accounting_skips_the_refs_directory(tmp_path): + root = tmp_path / "spool" + refs = root / "_refs" + refs.mkdir(parents=True) + (refs / READY_NAME).write_bytes(b"x" * 100) + (refs / ".018f0000-0000-7000-8000-000000000001.abcd1234.open" + ).write_bytes(b"y" * 50) + assert _spool(op="snapshot", root=str(root))["snapshot"]["bytes"] == 0 diff --git a/tests/test_native_spool_owner_lock_unit.py b/tests/test_native_spool_owner_lock_unit.py new file mode 100644 index 000000000..cb89b4ad7 --- /dev/null +++ b/tests/test_native_spool_owner_lock_unit.py @@ -0,0 +1,76 @@ +"""The spool owner lock in process: tests/native/test_spool_owner_lock.cpp. + +Compiles the C++ test against the real spool.cpp and runs it. It pins what +the cross-process driver tests (test_native_spool_owner_lock.py) cannot +reach: two Spool objects in ONE process that both take a directory refuse +each other -- the regression the engine avoids by holding one lock and +opening its sink's and service's Spools with held_by_caller -- the node-local +check through its statfs test seam, adoption's try-lock, and the directory +layout. One case forks, so the refusal across processes is covered here too. + +Needs a C++17 compiler and libcrypto (the same as conformance_spool). +""" + +from __future__ import annotations + +import os +import shutil +import subprocess +from pathlib import Path + +import pytest + + +@pytest.mark.cpu +def test_spool_owner_lock_unit(tmp_path): + compiler = shutil.which("g++") or shutil.which("c++") + if compiler is None: + pytest.skip("a C++17 compiler is required") + + root = Path(__file__).resolve().parents[1] + csrc = root / "native" / "csrc" + source = root / "tests" / "native" / "test_spool_owner_lock.cpp" + executable = tmp_path / "test_spool_owner_lock" + extra: list[str] = [] + # Homebrew OpenSSL is not on the default search path on macOS; Linux + # hosts (CI) find libcrypto without help. + for candidate in ("/opt/homebrew/opt/openssl@3", "/usr/local/opt/openssl@3"): + if Path(candidate, "include", "openssl", "sha.h").exists(): + extra += [f"-I{candidate}/include", f"-L{candidate}/lib"] + break + compile_result = subprocess.run( + [ + compiler, + "-std=c++17", + "-O0", + "-pthread", + f"-I{csrc}", + *extra, + str(source), + str(csrc / "store" / "spool.cpp"), + "-o", + str(executable), + "-lcrypto", + ], + capture_output=True, + text=True, + check=False, + ) + if compile_result.returncode != 0 and "openssl" in ( + compile_result.stdout + compile_result.stderr + ): + pytest.skip("libcrypto headers are unavailable: " + + compile_result.stderr[-400:]) + assert compile_result.returncode == 0, ( + compile_result.stdout + compile_result.stderr + ) + + run_result = subprocess.run( + [str(executable)], + capture_output=True, + text=True, + check=False, + env={**os.environ, "SPOOL_TEST_ROOT": str(tmp_path)}, + timeout=120, + ) + assert run_result.returncode == 0, run_result.stdout + run_result.stderr From dd12fb782f38b593bd19d54abcb7ad0901b181e2 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Mon, 28 Sep 2026 20:08:34 -0400 Subject: [PATCH 02/43] Open a sink and a service on one spool under one lock, and adopt dead siblings With the spool's owner lock (previous commit) every Spool takes its directory by default, and the storage service's and the pack sink's Spools became two takes: in ONE process they refuse each other, since flock binds to an open file description. Run at the previous commit, test_native_capture_chain_live.py -- the real NativePackSink and the real CaptureStorageService on one root in one process, as the engine runs them -- fails every case with exactly that: NativePackSink: sink start failed: cannot open spool: spool directory .../spool is owned by pid on host ... Both now take the directory's owner-lock mode: StorageServiceConfig .spool_owner_lock and SinkConfig.spool_owner_lock, "take" or "held_by_caller" (with the shared-filesystem override beside it), bound as the service dict's spool_owner_lock and NativePackSink(owner_lock=...). _dmi_native_store binds the lock itself -- SpoolOwnerLock (a context manager; refused with SpoolOwnedError, a RuntimeError naming the holder, or ValueError for a shared filesystem or a nested directory), spool_owner (who holds a directory), spool_rank_directory / spool_catalog_key (the section 2.3 layout) and a statfs test seam -- so a process holds one lock and opens both Spools held_by_caller. The chain suite now does exactly that, which is the engine's arrangement (next commit). adopt_sibling_spools: the service's directory is one rank directory //r-/ of the layout -- refused at construction unless it is, and unless the key is THIS catalog's, since adopted packs are indexed here -- and start(), after the lease and the sweep of its own directory and before the reconcile, adopts every sibling whose owner is dead (its lock can be taken; the kernel dropped it with the process): Recover sweeps its .open files, the uploader sends its .ready packs under the keys their paths give, they are indexed like the service's own (or owed in pending_index_), and the directory is removed once nothing but its lock file is left. A live sibling is skipped. Adoption follows the cycle's upload rules -- no upload without the lease or while a pack is owed -- so while the service cannot index, a dead spool's packs stay where they are durable, and a sibling left undrained is retried by the loop and keeps flush() from reporting drained. The snapshot says adopted_spools, adopted_packs and adoption_owed. E2b extends this with ref replay. The storage live suite stages and uploads through driver processes while a service has the same directory open, and builds second services on directories the first still holds. Its harness now holds each directory's lock for the test and opens every service and driver held_by_caller -- the engine's arrangement across processes. No assertion changed. Tests: * test_native_spool_ownership.py (cpu, new): the real sink and the real service share one rank directory under one SpoolOwnerLock and stage a pack; two takes refuse each other; a second process is refused naming this pid and host; spool_owner; nesting; NFS refused through the seam and admitted with the override; the layout; adoption refused for a directory outside the layout or under another catalog's key; a drained directory removed with its lock, a non-empty one kept. Red at the previous commit (the bindings did not exist), but for the case showing two takes refusing each other, which the previous commit introduced. * test_native_spool_adoption_live.py (ClickHouse + fake S3, new): a child running the engine's composition gets part of its capture indexed, stages the rest and is SIGKILLed holding the lease; a new incarnation waits the lease out, adopts the dead directory (its stale .open swept), and every capture of the dead process reads back byte-equal, while a live sibling is left alone. With the object store cut the dead directory stays as it was, flush() times out, and the loop adopts it once the store is back. Red with adopt_siblings() stubbed out: adopted_spools 0 != 1, and adoption_owed False. * test_native_capture_chain_live.py, test_native_capture_storage_live.py (48) and the new live suite pass; the lease byte-identity suite is untouched. --- native/csrc/catalog/bindings_store.cpp | 122 ++++++++++ native/csrc/catalog/storage_service.cpp | 230 +++++++++++++++--- native/csrc/catalog/storage_service.h | 60 ++++- native/csrc/sink/bindings_sink.cpp | 19 +- src/dmi/storage/capture/native_sink.py | 34 ++- src/dmi/storage/native_capture.py | 44 +++- tests/test_native_capture_chain_live.py | 16 +- tests/test_native_capture_storage_live.py | 72 ++++-- tests/test_native_spool_adoption_live.py | 279 ++++++++++++++++++++++ tests/test_native_spool_ownership.py | 268 +++++++++++++++++++++ 10 files changed, 1082 insertions(+), 62 deletions(-) create mode 100644 tests/test_native_spool_adoption_live.py create mode 100644 tests/test_native_spool_ownership.py diff --git a/native/csrc/catalog/bindings_store.cpp b/native/csrc/catalog/bindings_store.cpp index d07593772..50dcaf05f 100644 --- a/native/csrc/catalog/bindings_store.cpp +++ b/native/csrc/catalog/bindings_store.cpp @@ -2,6 +2,10 @@ // // StorageService spool -> object store -> catalog, on a background thread // CaptureReader search / select / hydrate against the catalog + store +// SpoolOwnerLock the owner lock of one spool directory (store/spool.h), +// which the engine holds around its sink and service +// spool_rank_directory, spool_catalog_key, spool_owner +// the section 2.3 spool layout, and who owns a directory // // Pure C++ plus libcurl and libcrypto. It uses pybind11's headers but links // nothing from torch, and registers no ring types, so it loads beside @@ -10,12 +14,14 @@ #include #include +#include #include #include #include "catalog/hydration.h" #include "catalog/reader.h" #include "catalog/storage_service.h" +#include "store/spool.h" namespace py = pybind11; namespace dc = dmi_catalog; @@ -75,10 +81,25 @@ dc::ClickHouseConnection clickhouse_connection(const py::dict& d) { return c; } +dmi_store::OwnerLock owner_lock(const std::string& text) { + dmi_store::OwnerLock mode = dmi_store::OwnerLock::kTake; + if (!dmi_store::ParseOwnerLock(text, &mode)) { + throw py::value_error("spool_owner_lock must be 'take' or " + "'held_by_caller', got '" + text + "'"); + } + return mode; +} + dc::StorageServiceConfig service_config(const py::dict& d) { dc::StorageServiceConfig c; c.spool_root = get(d, "spool_root", ""); c.spool_max_bytes = get(d, "spool_max_bytes", c.spool_max_bytes); + c.spool_owner_lock = + owner_lock(get(d, "spool_owner_lock", "take")); + c.spool_allow_shared_filesystem = get( + d, "spool_allow_shared_filesystem", c.spool_allow_shared_filesystem); + c.adopt_sibling_spools = + get(d, "adopt_sibling_spools", c.adopt_sibling_spools); c.s3 = s3_config(d); c.uploader.store_id = get(d, "store_id", c.uploader.store_id); c.uploader.max_workers = get(d, "uploader_max_workers", c.uploader.max_workers); @@ -128,6 +149,9 @@ py::dict snapshot_dict(const dc::StorageServiceSnapshot& s) { out["swept_on_start"] = s.swept_on_start; out["pending_index"] = s.pending_index; out["rejected_packs"] = s.rejected_packs; + out["adopted_spools"] = s.adopted_spools; + out["adopted_packs"] = s.adopted_packs; + out["adoption_owed"] = s.adoption_owed; out["failed"] = s.failed; out["lease_state"] = s.lease_state; // Seconds on the monotonic clock, comparable with time.monotonic() (both @@ -274,6 +298,23 @@ class CaptureReader { dc::NativeCaptureReader reader_; }; +// A refused SpoolStatus as the Python exception that says what it is: a +// directory another process owns, a refused configuration, or an I/O error. +PyObject* g_spool_owned_error = nullptr; // SpoolOwnedError, set at import + +[[noreturn]] void raise_spool_status(dmi_store::SpoolStatus status, + const std::string& error) { + if (status == dmi_store::SpoolStatus::kOwned) { + PyErr_SetString(g_spool_owned_error, error.c_str()); + throw py::error_already_set(); + } + if (status == dmi_store::SpoolStatus::kBadArgument) { + throw py::value_error(error); + } + PyErr_SetString(PyExc_OSError, error.c_str()); + throw py::error_already_set(); +} + } // namespace PYBIND11_MODULE(_dmi_native_store, m) { @@ -295,6 +336,87 @@ PYBIND11_MODULE(_dmi_native_store, m) { [](const dc::CaptureStorageService& s) { return snapshot_dict(s.snapshot()); }) .def("rethrow_if_failed", &dc::CaptureStorageService::rethrow_if_failed); + // Raised when another owner holds a spool directory; a RuntimeError, so + // callers catching the service's refusals keep catching it. + static py::exception spool_owned_error( + m, "SpoolOwnedError", PyExc_RuntimeError); + g_spool_owned_error = spool_owned_error.ptr(); + + py::class_(m, "SpoolOwnerLock") + .def(py::init([](const std::string& directory, + bool allow_shared_filesystem) { + auto lock = std::make_unique(); + std::string error; + const dmi_store::SpoolStatus status = + dmi_store::SpoolOwnerLock::Acquire( + directory, allow_shared_filesystem, lock.get(), &error); + if (status != dmi_store::SpoolStatus::kOk) { + raise_spool_status(status, error); + } + return lock; + }), + py::arg("directory"), py::arg("allow_shared_filesystem") = false) + .def_property_readonly("directory", &dmi_store::SpoolOwnerLock::directory) + .def_property_readonly("held", &dmi_store::SpoolOwnerLock::held) + .def("release", &dmi_store::SpoolOwnerLock::Release) + .def("release_and_remove_if_empty", + [](dmi_store::SpoolOwnerLock& self) { + std::string error; + return self.ReleaseAndRemoveIfEmpty(&error); + }) + .def("__enter__", + [](dmi_store::SpoolOwnerLock& self) -> dmi_store::SpoolOwnerLock& { + return self; + }, py::return_value_policy::reference) + .def("__exit__", [](dmi_store::SpoolOwnerLock& self, const py::args&) { + self.Release(); + return false; + }); + + m.def("spool_owner", + [](const std::string& directory) -> py::object { + dmi_store::SpoolOwner owner; + if (!dmi_store::ReadSpoolOwner(directory, &owner)) return py::none(); + py::dict out; + out["host"] = owner.host; + out["pid"] = owner.pid; + return out; + }, + py::arg("directory"), + "Who holds a spool directory's owner lock, or None when nothing " + "does."); + m.def("spool_catalog_key", &dmi_store::SpoolCatalogKey, py::arg("database"), + py::arg("table_prefix"), py::arg("store_id")); + m.def("spool_rank_directory", + [](const std::string& base, const std::string& database, + const std::string& table_prefix, const std::string& store_id, + uint64_t producer_rank, std::optional incarnation) { + const std::string fresh = + incarnation ? *incarnation : dmi_store::NewSpoolIncarnation(); + uint64_t parsed_rank = 0; + std::string parsed; + if (!dmi_store::ParseSpoolRankDirectoryName( + dmi_store::SpoolRankDirectoryName(producer_rank, fresh), + &parsed_rank, &parsed)) { + throw py::value_error("incarnation must be 8 lowercase hex " + "digits"); + } + return dmi_store::SpoolRankDirectory(base, database, table_prefix, + store_id, producer_rank, fresh); + }, + py::arg("base"), py::arg("database"), py::arg("table_prefix"), + py::arg("store_id"), py::arg("producer_rank"), + py::arg("incarnation") = py::none(), + "//r-, the section " + "2.3 spool layout; a fresh incarnation when none is given."); + m.def("_set_spool_filesystem_type_for_testing", + [](std::optional f_type) { + dmi_store::SetFilesystemTypeForTesting(f_type ? *f_type : -1); + }, + py::arg("f_type"), + "Test seam: node-local checks in this module read f_type instead " + "of statfs(2); None restores statfs."); + py::class_(m, "CaptureReader") .def(py::init(), py::arg("config")) .def("search", &CaptureReader::search, py::arg("filters")) diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index df57f79ca..af73f1265 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -4,6 +4,7 @@ #include #include #include +#include #include #include #include @@ -155,10 +156,34 @@ CaptureStorageService::CaptureStorageService(StorageServiceConfig config) " ms, or raise lease_ttl_ns"); } std::string error; - if (dmi_store::Spool::Open({config_.spool_root, config_.spool_max_bytes}, - &spool_, &error) != dmi_store::SpoolStatus::kOk) { + dmi_store::SpoolConfig spool_config{config_.spool_root, + config_.spool_max_bytes}; + spool_config.owner_lock = config_.spool_owner_lock; + spool_config.allow_shared_filesystem = config_.spool_allow_shared_filesystem; + if (dmi_store::Spool::Open(spool_config, &spool_, &error) != + dmi_store::SpoolStatus::kOk) { throw std::runtime_error("storage service: cannot open spool: " + error); } + if (config_.adopt_sibling_spools) { + // Siblings are adopted INTO this catalog, so this directory must sit + // under this catalog's key: a directory under another catalog's key + // would index that catalog's packs here. + const std::filesystem::path own(spool_.root()); + const std::string key = dmi_store::SpoolCatalogKey( + config_.writer.database, config_.writer.table_prefix, + config_.uploader.store_id); + uint64_t rank = 0; + std::string incarnation; + if (!dmi_store::ParseSpoolRankDirectoryName(own.filename().string(), + &rank, &incarnation) || + own.parent_path().filename().string() != key) { + throw std::invalid_argument( + "storage service: adopt_sibling_spools needs spool_root to be a " + "rank directory /" + key + "/r- (this " + "catalog's key for database, table_prefix and store_id), got " + + spool_.root()); + } + } uploader_ = std::make_unique(&spool_, &s3_, config_.uploader); } @@ -224,12 +249,10 @@ void CaptureStorageService::start() { } void CaptureStorageService::sweep_and_reconcile_at_start() { - // After the lease, never before, so a second process pointed at this spool - // usually learns that the catalog is held before it can delete a live - // sink's .open files. Only usually: a holder that is quarantined has let - // its row lapse, and a second process can take the lease in that gap. The - // spool itself is not locked; one process per spool is the caller's job - // until the spool gets an owner lock. + // After the lease: a start refused the catalog never touches the spool. + // The spool's owner lock (taken at construction, by this service or its + // caller) is what keeps another process's writer out of the directory + // this deletes .open files in. if (config_.sweep_spool_on_start) { std::vector recovered; std::string error; @@ -241,6 +264,27 @@ void CaptureStorageService::sweep_and_reconcile_at_start() { state_.swept_on_start = recovered.size(); } + // Dead siblings next, after this directory's own sweep and before the + // reconcile, which then finds their packs committed. Nothing here fails + // start(): a sibling left undrained is owed, and the loop retries it. + if (config_.adopt_sibling_spools) { + try { + adopt_siblings(); + } catch (const CatalogError& exc) { + adoption_owed_ = true; + record_error(std::string("adopting dead spools at start ") + + (is_lease_refusal(exc) ? "lost the publisher lease: " + : "failed: ") + + exc.what()); + } catch (const std::exception& exc) { + adoption_owed_ = true; + record_error(std::string("adopting dead spools at start failed: ") + + exc.what()); + } + std::lock_guard lock(state_mutex_); + state_.adoption_owed = adoption_owed_; + } + // A failed pass is not fatal -- the bucket is still there next time. Nor // is a lease lost while it runs, to a quarantine or to another holder: // that is the running service's case, and the loop handles it as it does @@ -388,24 +432,6 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { LeaseScope lease(this); catalog = ensure_publisher_lease(); } - // Indexes refs, keeping whatever does not index owed: it is already gone - // from the spool, so pending_index_ is the only record of it in-process. - const auto index_or_owe = [this, catalog](std::vector refs) { - if (!catalog) { - pending_index_.insert(pending_index_.end(), refs.begin(), refs.end()); - return; - } - std::vector unindexed; - try { - if (!refs.empty()) index_bounded(std::move(refs), &unindexed); - } catch (...) { - pending_index_.insert(pending_index_.end(), unindexed.begin(), - unindexed.end()); - throw; - } - pending_index_.insert(pending_index_.end(), unindexed.begin(), - unindexed.end()); - }; try { // 1. Retry what earlier cycles uploaded but could not index. While any // of it is still owed, the catalog is down or refusing: upload @@ -414,7 +440,7 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { if (catalog && !pending_index_.empty()) { std::vector owed; owed.swap(pending_index_); - index_or_owe(std::move(owed)); + index_or_owe(std::move(owed), catalog); } // 2. Upload everything the sink has staged -- but only with the lease @@ -454,7 +480,15 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { } // 3. Index them. - index_or_owe(std::move(to_index)); + index_or_owe(std::move(to_index), catalog); + + // 3a. A dead sibling an earlier adoption pass left undrained, under the + // same rule as the uploads above: only with the lease, nothing owed + // and nothing of our own failing. + if (catalog && adoption_owed_ && pending_index_.empty() && + upload_failures == 0) { + adopt_siblings(); + } // 4. Reconcile on its interval, or when the pass at start() lost the // lease before it finished. The lease thread keeps the lease alive. @@ -477,8 +511,8 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { // would be re-hashed on every cycle of an outage. // Without the lease nothing can be confirmed in the catalog, so the // cycle is not drained, and it counts towards the backoff. - outcome.failed = - !catalog || upload_failures != 0 || !pending_index_.empty(); + outcome.failed = !catalog || upload_failures != 0 || + !pending_index_.empty() || adoption_owed_; bool nothing_pending = !batch.refs.empty(); if (batch.refs.empty() && !outcome.failed) { std::vector pending; @@ -507,11 +541,147 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { { std::lock_guard lock(state_mutex_); state_.pending_index = pending_index_.size(); + state_.adoption_owed = adoption_owed_; } failure_streak_ = outcome.failed ? std::min(failure_streak_ + 1, 32) : 0; return outcome; } +void CaptureStorageService::index_or_owe(std::vector refs, + bool catalog) { + if (!catalog) { + pending_index_.insert(pending_index_.end(), refs.begin(), refs.end()); + return; + } + std::vector unindexed; + try { + if (!refs.empty()) index_bounded(std::move(refs), &unindexed); + } catch (...) { + pending_index_.insert(pending_index_.end(), unindexed.begin(), + unindexed.end()); + throw; + } + pending_index_.insert(pending_index_.end(), unindexed.begin(), + unindexed.end()); +} + +void CaptureStorageService::adopt_siblings() { + namespace fs = std::filesystem; + // Owed until the pass completes: a lost lease propagates from the middle. + adoption_owed_ = true; + const fs::path own(spool_.root()); + std::vector siblings; + std::error_code ec; + for (fs::directory_iterator it(own.parent_path(), ec), end; + !ec && it != end; it.increment(ec)) { + uint64_t rank = 0; + std::string incarnation; + std::error_code type_ec; + if (it->path() == own || it->is_symlink(type_ec) || + !it->is_directory(type_ec) || + !dmi_store::ParseSpoolRankDirectoryName( + it->path().filename().string(), &rank, &incarnation)) { + continue; + } + siblings.push_back(it->path().string()); + } + if (ec) { + record_error("adoption: cannot list " + own.parent_path().string() + + ": " + ec.message()); + return; + } + std::sort(siblings.begin(), siblings.end()); + bool owed = false; + for (const std::string& sibling : siblings) { + if (!adopt_sibling(sibling)) owed = true; + } + adoption_owed_ = owed; +} + +bool CaptureStorageService::adopt_sibling(const std::string& directory) { + dmi_store::SpoolOwnerLock lock; + std::string error; + const dmi_store::SpoolStatus locked = + dmi_store::SpoolOwnerLock::TryAdopt(directory, &lock, &error); + if (locked == dmi_store::SpoolStatus::kOwned) return true; // it lives + if (locked != dmi_store::SpoolStatus::kOk) { + // Another adopter drained and removed it meanwhile: nothing is owed. + if (!std::filesystem::exists(directory)) return true; + record_error("adopting dead spool " + directory + ": " + error); + return false; + } + // The rules of the cycle's uploads: none without the lease, and none while + // an uploaded pack is still owed to the catalog. + bool catalog = false; + { + LeaseScope lease(this); + catalog = writer_.held_lease() != nullptr; + } + if (!catalog || !pending_index_.empty()) return false; + + dmi_store::SpoolConfig config{directory, config_.spool_max_bytes}; + config.owner_lock = dmi_store::OwnerLock::kHeldByCaller; // `lock` + config.allow_shared_filesystem = config_.spool_allow_shared_filesystem; + dmi_store::Spool spool; + std::vector ready; + if (dmi_store::Spool::Open(config, &spool, &error) != + dmi_store::SpoolStatus::kOk || + spool.Recover(&ready, &error) != dmi_store::SpoolStatus::kOk) { + record_error("adopting dead spool " + directory + ": " + error); + return false; + } + // Each pack's identity and object key come from the pack and its path in + // the dead directory, exactly as its owner would have uploaded it. + dmi_store::SpoolUploader uploader(&spool, &s3_, config_.uploader); + const dmi_store::UploadBatchResult batch = uploader.UploadPending(-1); + std::vector to_index; + uint64_t uploaded_bytes = 0; + size_t failures = 0; + for (size_t i = 0; i < batch.refs.size(); ++i) { + const dmi_store::PackRef& ref = batch.refs[i]; + if (ref.pack_id.empty()) { + ++failures; + if (i < batch.failures.size()) { + record_error("adopting dead spool " + directory + ": upload failed " + "for " + batch.failures[i].object_key + ": " + + batch.failures[i].error); + } + continue; + } + to_index.push_back({ref.pack_id, ref.store_id, ref.object_key, + ref.object_bytes, ref.checksum, ref.record_count}); + uploaded_bytes += ref.object_bytes; + } + { + std::lock_guard state(state_mutex_); + state_.uploaded_packs += to_index.size(); + state_.uploaded_bytes += uploaded_bytes; + state_.upload_failures += failures; + state_.adopted_packs += to_index.size(); + } + // Uploaded, so gone from the dead spool: indexed now, or owed in + // pending_index_ like any pack of this service's own. + index_or_owe(std::move(to_index), true); + if (failures != 0) return false; // they stay in the dead spool + std::vector left; + if (spool.ListPending(&left, &error) != dmi_store::SpoolStatus::kOk || + !left.empty()) { + return false; + } + { + std::lock_guard state(state_mutex_); + ++state_.adopted_spools; + } + if (!lock.ReleaseAndRemoveIfEmpty(&error)) { + // Nothing to upload is left, only files that are not packs (a + // quarantined one, say): the directory stays for someone to look at, + // and is not owed. + record_error("adopted dead spool " + directory + " was drained but " + "still holds files that are not packs; left in place"); + } + return true; +} + void CaptureStorageService::index_bounded(std::vector refs, std::vector* unindexed) { // Batches the indexer can take: at most max_packs, and halved again when the diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index 74d697ee3..7a29dff3f 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -20,6 +20,14 @@ // already committed, and indexes the rest. It runs at start(), and // periodically when reconcile_interval_ns is non-zero. // +// One owner per spool directory. The service's spool is opened under the +// directory's owner lock (store/spool.h), so a second process on it is +// refused at construction, naming the holder. With adopt_sibling_spools +// the service's directory is one rank directory of the plan's section 2.3 +// layout, and at start() it adopts the siblings whose owner has died: a +// crashed process's spool is recovered by the next process on the node for +// the same catalog, whatever run it belongs to. +// // Deployment shape: the service holds the catalog's single publisher lease, so // run ONE service per (database, table_prefix). A second one waits up to // start_lease_wait_ns for the lease at start(), then fails naming the holder. @@ -84,6 +92,27 @@ struct StorageServiceConfig { // sink writes. std::string spool_root; uint64_t spool_max_bytes = 1ull << 40; + // The spool directory's owner lock (store/spool.h). kTake owns it for the + // service's life, and refuses a directory another process owns. A process + // that also runs the sink on it -- the engine -- holds one SpoolOwnerLock + // and opens both with kHeldByCaller: two takes in one process refuse each + // other. + dmi_store::OwnerLock spool_owner_lock = dmi_store::OwnerLock::kTake; + bool spool_allow_shared_filesystem = false; + // Adopt the spools of dead processes bound for this catalog. spool_root + // must then be a rank directory of the plan's section 2.3 layout, + // //r-/ + // under THIS catalog's key (SpoolCatalogKey of writer.database, + // writer.table_prefix and uploader.store_id), or construction throws. + // start(), after the lease and the sweep of its own directory, tries the + // owner lock of every sibling rank directory; each one whose owner is + // gone has its .open files swept, its .ready packs uploaded and indexed, + // and is removed once nothing but its lock file is left. A sibling that + // could not be drained (no lease, an upload that failed, a pack still + // owed) is retried by the loop's cycles, and flush() does not report + // drained until it has been. Live siblings -- another rank or job on this + // node -- are left alone. + bool adopt_sibling_spools = false; dmi_store::S3Config s3; dmi_store::UploaderConfig uploader; // uploader.store_id names the store @@ -143,8 +172,9 @@ struct StorageServiceConfig { // past its row: until it resumes and next checks, or -- after a system // suspend, which the steady clock the deadline runs on does not count -- // until a renewal or publish is refused. Two on different (database, - // table_prefix) pairs each hold a lease and upload freely. The spool - // itself is not locked; one process per spool is the caller's job. + // table_prefix) pairs each hold a lease and upload freely. The spool's + // owner lock (spool_owner_lock) is what keeps a second process off the + // directory itself: it is refused at construction, before any of this. bool sweep_spool_on_start = true; bool reconcile_on_start = true; }; @@ -167,6 +197,11 @@ struct StorageServiceSnapshot { uint64_t swept_on_start = 0; // ready packs Recover() found at start uint64_t pending_index = 0; // uploaded packs awaiting a retried index uint64_t rejected_packs = 0; // set aside: cannot be indexed (see flush) + // adopt_sibling_spools: dead siblings drained, the ready packs of theirs + // that were uploaded, and whether one is still owed a retry. + uint64_t adopted_spools = 0; + uint64_t adopted_packs = 0; + bool adoption_owed = false; // A foreign lease outlived 2 x TTL: the service stopped for good. bool failed = false; // "none" before start, "held", "quarantined" (an unknown outcome set the @@ -198,8 +233,9 @@ class CaptureStorageService { // Ensure the catalog schema, take the publisher lease (waiting up to // start_lease_wait_ns for another holder's to expire, or for a claim that // timed out to go through -- past it, once, to wait out the quarantine - // such a claim left), sweep the spool, reconcile once, then start the - // background cycle. The lease renews from the moment it is taken. + // such a claim left), sweep the spool, adopt dead siblings + // (adopt_sibling_spools), reconcile once, then start the background + // cycle. The lease renews from the moment it is taken. // Throws if the lease is still held by another publisher when the wait // ends, or its claim still times out. A lease lost while the reconcile // runs does not fail start(): the loop takes a fresh one, as it would @@ -247,9 +283,19 @@ class CaptureStorageService { class LeaseScope; void loop(); - // start()'s spool sweep and reconcile, with the lease held and the lease - // thread renewing it. Requires cycle_mutex_. + // start()'s spool sweep, adoption and reconcile, with the lease held and + // the lease thread renewing it. Requires cycle_mutex_. void sweep_and_reconcile_at_start(); + // One adoption pass over the sibling rank directories; sets + // adoption_owed_ to whether one was left undrained. Requires + // cycle_mutex_. Only a lost lease propagates. + void adopt_siblings(); + // Adopts one sibling; false when it is owed another try. + bool adopt_sibling(const std::string& directory); + // Indexes refs that are gone from their spool, keeping whatever does not + // index in pending_index_ -- the only record of it in-process. With no + // catalog, keeps them all. Requires cycle_mutex_. + void index_or_owe(std::vector refs, bool catalog); // Stops the lease thread and waits for it. void stop_lease_thread(); CycleOutcome run_cycle(); // requires cycle_mutex_ @@ -327,6 +373,8 @@ class CaptureStorageService { // The reconcile at start() lost the lease before it finished; the loop // runs one once it holds a lease again. Guarded by cycle_mutex_. bool reconcile_owed_ = false; + // An adoption pass left a dead sibling undrained. Guarded by cycle_mutex_. + bool adoption_owed_ = false; int failure_streak_ = 0; // consecutive failed cycles, for the backoff // Uploaded, so gone from the spool, but not yet in the catalog. std::vector pending_index_; diff --git a/native/csrc/sink/bindings_sink.cpp b/native/csrc/sink/bindings_sink.cpp index 696256494..915fc499e 100644 --- a/native/csrc/sink/bindings_sink.cpp +++ b/native/csrc/sink/bindings_sink.cpp @@ -157,8 +157,17 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { uint64_t max_pack_bytes, uint64_t max_pack_records, uint64_t max_linger_ns, uint64_t spool_max_bytes, const std::string& overload, - std::optional admission_timeout_s) { + std::optional admission_timeout_s, + const std::string& owner_lock, + bool allow_shared_filesystem) { dmi_sink::SinkConfig config; + if (!dmi_store::ParseOwnerLock(owner_lock, + &config.spool_owner_lock)) { + throw py::value_error( + "owner_lock must be 'take' or 'held_by_caller', got '" + + owner_lock + "'"); + } + config.spool_allow_shared_filesystem = allow_shared_filesystem; config.overload = ParseOverload(overload); config.admission_timeout_s = ParseAdmissionTimeout(admission_timeout_s); @@ -185,7 +194,13 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { // SinkConfig's own defaults: the Python NativeSinkConfig, which // the ring-fed sink is built from, picks block with 2 s. py::arg("overload") = "drop_newest", - py::arg("admission_timeout_s") = py::none()) + py::arg("admission_timeout_s") = py::none(), + // The spool directory's owner lock (store/spool.h): "take" owns + // it for the sink's life; "held_by_caller" when the caller holds + // a SpoolOwnerLock on it, as the engine does around its sink and + // storage service. + py::arg("owner_lock") = "take", + py::arg("allow_shared_filesystem") = false) .def("attach", [](std::shared_ptr self) { // Simulates engine ownership for tests (the real engine takes diff --git a/src/dmi/storage/capture/native_sink.py b/src/dmi/storage/capture/native_sink.py index bc34b942c..3482cc4a3 100644 --- a/src/dmi/storage/capture/native_sink.py +++ b/src/dmi/storage/capture/native_sink.py @@ -86,12 +86,28 @@ def _load_native_sink_extension() -> Any: class NativePackSinkHandle: """Owns a native pack sink; mirrors CapturePackReferenceSink's surface - (``record_format`` + ``native_sink``) so call sites switch by factory.""" - - def __init__(self, config: NativeSinkConfig) -> None: + (``record_format`` + ``native_sink``) so call sites switch by factory. + + ``spool_root`` overrides ``config.spool_root`` (the engine passes its + own rank directory under it), and ``owner_lock`` is the spool + directory's owner-lock mode: ``"take"`` owns the directory for the + sink's life, ``"held_by_caller"`` when the caller holds a + ``SpoolOwnerLock`` on it -- as the engine does around its sink and its + storage service, which would otherwise refuse each other. + """ + + def __init__( + self, + config: NativeSinkConfig, + *, + spool_root: str | None = None, + owner_lock: str = "take", + ) -> None: module = _load_native_sink_extension() self._native_sink = module.NativePackSink( - spool_root=config.spool_root, + spool_root=config.spool_root if spool_root is None else spool_root, + owner_lock=owner_lock, + allow_shared_filesystem=config.spool_allow_shared_filesystem, layout=LAYOUT_NAME, num_workers=config.num_workers, max_queue_records=config.max_queue_records, @@ -120,12 +136,18 @@ def native_sink(self) -> Any: return self._native_sink -def create_native_pack_sink(config: NativeSinkConfig) -> NativePackSinkHandle: +def create_native_pack_sink( + config: NativeSinkConfig, + *, + spool_root: str | None = None, + owner_lock: str = "take", +) -> NativePackSinkHandle: """Select the native capture writer for one record runtime.""" if not isinstance(config, NativeSinkConfig): raise TypeError("config must be a NativeSinkConfig") - return NativePackSinkHandle(config) + return NativePackSinkHandle(config, spool_root=spool_root, + owner_lock=owner_lock) __all__ = [ diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index ebd9f76c4..7c119ffc8 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -124,6 +124,11 @@ class NativeSinkConfig: A record larger than ``max_queue_bytes`` or ``max_pack_bytes`` can never be admitted; ``validate_capture_bounds`` refuses such a bound at attach, before any forward runs. + + ``spool_root`` must be node-local, and a spool directory has one owner + process, held by an flock on its ``.owner.lock``. A root on NFS or + Lustre is refused unless ``spool_allow_shared_filesystem``: flock there + does not keep out a process on another node. """ spool_root: str @@ -136,10 +141,13 @@ class NativeSinkConfig: max_linger_ns: int = 1_000_000_000 overload: str = "block" admission_timeout_s: Optional[float] = 2.0 + spool_allow_shared_filesystem: bool = False def __post_init__(self) -> None: if not self.spool_root: raise ValueError("spool_root is required") + if type(self.spool_allow_shared_filesystem) is not bool: + raise TypeError("spool_allow_shared_filesystem must be bool") for name in ( "spool_max_bytes", "num_workers", @@ -554,8 +562,25 @@ def validate_capture_bounds( "uploader_max_in_flight_bytes") +# The spool directory's owner-lock modes (native/csrc/store/spool.h). +SPOOL_OWNER_LOCKS = ("take", "held_by_caller") + + class NativeCaptureStorage: - """The in-process storage service: spool -> object store -> catalog.""" + """The in-process storage service: spool -> object store -> catalog. + + The spool directory has one owner process, held by an flock on + ``/.owner.lock``. ``spool_owner_lock="take"`` makes this + service its owner for the service's life, and refuses a directory + another process owns, naming it. A process that also runs the sink on + the directory -- the engine -- holds one ``SpoolOwnerLock`` and passes + ``"held_by_caller"`` here and to the sink: two takes in one process + refuse each other. + + ``adopt_sibling_spools`` needs ``spool_root`` to be a rank directory of + the spool layout (``spool_rank_directory``); ``start`` then drains the + sibling directories whose owners have died into this catalog. + """ def __init__( self, @@ -564,9 +589,22 @@ def __init__( spool_root: str, spool_max_bytes: int, sweep_spool: bool, + spool_owner_lock: str = "take", + adopt_sibling_spools: bool = False, + spool_allow_shared_filesystem: bool = False, ) -> None: if not isinstance(config, NativeCaptureStorageConfig): raise TypeError("config must be a NativeCaptureStorageConfig") + if spool_owner_lock not in SPOOL_OWNER_LOCKS: + raise ValueError( + f"spool_owner_lock must be one of {SPOOL_OWNER_LOCKS}, got " + f"{spool_owner_lock!r}") + for name, value in ( + ("adopt_sibling_spools", adopt_sibling_spools), + ("spool_allow_shared_filesystem", + spool_allow_shared_filesystem)): + if type(value) is not bool: + raise TypeError(f"{name} must be bool") module = _load_native_store_extension() native = config._native_dict() native.update( @@ -579,6 +617,9 @@ def __init__( reconcile_prefix=config.reconcile_prefix, reconcile_interval_ns=int(config.reconcile_interval_s * 1e9), sweep_spool_on_start=sweep_spool, + spool_owner_lock=spool_owner_lock, + adopt_sibling_spools=adopt_sibling_spools, + spool_allow_shared_filesystem=spool_allow_shared_filesystem, **config._lease_native(), ) self._config = config @@ -812,6 +853,7 @@ def read( __all__ = [ "PACK_FRAMING_RESERVE_BYTES", "SINK_OVERLOAD_POLICIES", + "SPOOL_OWNER_LOCKS", "NativeSinkConfig", "NativeCapture", "NativeCapturePage", diff --git a/tests/test_native_capture_chain_live.py b/tests/test_native_capture_chain_live.py index 5004db901..7e2517356 100644 --- a/tests/test_native_capture_chain_live.py +++ b/tests/test_native_capture_chain_live.py @@ -210,24 +210,33 @@ def _run_chain(config, spool_root: Path, envelopes, *, sink_overrides=None, which is safe only while no sink writes there), then the sink; flush both, and return the snapshots and what the reader reads back. + Both open the spool as the engine opens them: under ONE owner lock this + process holds, each with owner_lock="held_by_caller" -- two takes in one + process would refuse each other. + The sink is the raw binding with `sink_overrides`, or, given a NativeSinkConfig, the one the engine builds from it.""" from dmi.storage.native_capture import ( NativeCaptureReader, NativeCaptureStorage, + _load_native_store_extension, ) + owner = _load_native_store_extension().SpoolOwnerLock(str(spool_root)) service = NativeCaptureStorage(config, spool_root=str(spool_root), - spool_max_bytes=1 << 40, sweep_spool=True) + spool_max_bytes=1 << 40, sweep_spool=True, + spool_owner_lock="held_by_caller") service.start() try: if sink_config is None: - sink, _lease = _open_sink(spool_root, **(sink_overrides or {})) + sink, _lease = _open_sink(spool_root, owner_lock="held_by_caller", + **(sink_overrides or {})) else: from dmi.storage.capture.native_sink import ( create_native_pack_sink, ) - sink = create_native_pack_sink(sink_config).native_sink + sink = create_native_pack_sink( + sink_config, owner_lock="held_by_caller").native_sink _lease = sink.attach() for envelope in envelopes: sink.submit_envelope(LAYOUT, envelope.rows, envelope.payload()) @@ -239,6 +248,7 @@ def _run_chain(config, spool_root: Path, envelopes, *, sink_overrides=None, service_snapshot = service.snapshot() finally: service.stop() + owner.release() reader = NativeCaptureReader(config) selection = reader.select(tenant_id="t") diff --git a/tests/test_native_capture_storage_live.py b/tests/test_native_capture_storage_live.py index 01218b5ba..068241a3d 100644 --- a/tests/test_native_capture_storage_live.py +++ b/tests/test_native_capture_storage_live.py @@ -113,12 +113,43 @@ def _storage_config(endpoint, prefix, **overrides): return NativeCaptureStorageConfig(**fields) +# Every spool directory a test points a service or a driver at is owned by +# the harness for the rest of the test: one SpoolOwnerLock held here, and +# each service, sink driver and store driver opens the directory with +# owner_lock="held_by_caller". That is the engine's arrangement -- it holds +# the lock around its sink and its service -- stretched over the processes a +# test uses: the drivers stage and upload from processes of their own while +# a service is up, and a test often builds a second service on a directory +# the first still has open. Each taking the lock would refuse the others. +_HARNESS_LOCKS: dict = {} + + +@pytest.fixture(autouse=True) +def _harness_spool_locks(): + yield + for lock in _HARNESS_LOCKS.values(): + lock.release() + _HARNESS_LOCKS.clear() + + +def _held(spool_root) -> str: + """Hold spool_root's owner lock for the test; the mode to open it in.""" + from dmi.storage.native_capture import _load_native_store_extension + + key = str(spool_root) + if key not in _HARNESS_LOCKS: + _HARNESS_LOCKS[key] = _load_native_store_extension().SpoolOwnerLock( + key) + return "held_by_caller" + + def _service(config, spool_root: Path, *, sweep_spool=True): from dmi.storage.native_capture import NativeCaptureStorage return NativeCaptureStorage(config, spool_root=str(spool_root), spool_max_bytes=1 << 40, - sweep_spool=sweep_spool) + sweep_spool=sweep_spool, + spool_owner_lock=_held(spool_root)) def _record(index: int): @@ -150,7 +181,7 @@ def _stage(spool_root: Path, indexes, *, records_per_pack: int = 2): max_queue_records=256, max_queue_bytes=1 << 24, max_pack_bytes=8 << 20, max_pack_records=records_per_pack, max_linger_ns=1_000_000_000, overload="drop_newest", - admission_timeout=-1)["ok"] + admission_timeout=-1, owner_lock=_held(spool_root))["ok"] for index in indexes: metadata, tensor = _record(index) response = sink.call( @@ -587,7 +618,8 @@ def test_a_pack_uploaded_but_never_indexed_is_reconciled_at_start( region=REGION, access=ACCESS, secret=SECRET, token=None, insecure=True, connect_timeout=5, read_timeout=15, max_attempts=4, store_id="s3", root=str(spool_root), spool_max_bytes=1 << 40, - limit=-1, max_workers=4, max_in_flight_bytes=1 << 30) + limit=-1, max_workers=4, max_in_flight_bytes=1 << 30, + owner_lock=_held(spool_root)) assert uploaded["ok"], uploaded # And a foreign object where packs live, which must not be indexed. foreign = store.call( @@ -632,7 +664,9 @@ def test_a_failed_head_is_an_error_not_a_foreign_object(fake_s3, tmp_path): "etag": '"0"'} with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() - native.update(spool_root=str(tmp_path / "spool"), holder="head-test", + native.update(spool_root=str(tmp_path / "spool"), + spool_owner_lock=_held(tmp_path / "spool"), + holder="head-test", reconcile_prefix="fault/", s3_max_attempts=1) service = _load_native_store_extension().StorageService(native) service.start() # reconciles once @@ -682,7 +716,8 @@ def test_the_loop_backs_off_while_the_object_store_is_down(tmp_path): with _catalog() as (_client, catalog): native = _storage_config(dead, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="backoff-test", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="backoff-test", poll_interval_ns=20_000_000, max_backoff_ns=10_000_000_000, reconcile_on_start=False, s3_max_attempts=1, uploader_max_attempts=1) @@ -715,7 +750,8 @@ def test_a_pack_too_big_to_index_is_set_aside_and_the_rest_still_index( config = _storage_config(fake_s3, catalog.table_prefix) native = config._native_dict() native.update( - spool_root=str(spool_root), holder="poison-test", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="poison-test", poll_interval_ns=50_000_000, reconcile_on_start=False, # One descriptor fits, forty do not. indexer_max_estimated_bytes=4000) @@ -750,7 +786,8 @@ def test_a_batch_over_the_budget_splits_until_every_pack_indexes( config = _storage_config(fake_s3, catalog.table_prefix) native = config._native_dict() native.update( - spool_root=str(spool_root), holder="split-test", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="split-test", poll_interval_ns=50_000_000, reconcile_on_start=False, # A two-record pack renders ~720 bytes: two fit, ten do not. indexer_max_estimated_bytes=2000) @@ -785,7 +822,8 @@ def test_flush_returns_on_time_when_the_catalog_stops_answering( native = _storage_config( fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() - native.update(spool_root=str(spool_root), holder="stall-test", + native.update(spool_root=str(spool_root), + spool_owner_lock=_held(spool_root), holder="stall-test", poll_interval_ns=20_000_000, reconcile_on_start=False, clickhouse_request_timeout_s=5.0) service = _load_native_store_extension().StorageService(native) @@ -835,7 +873,8 @@ def _tick(): native = _storage_config( fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() - native.update(spool_root=str(spool_root), holder="gil-test", + native.update(spool_root=str(spool_root), + spool_owner_lock=_held(spool_root), holder="gil-test", poll_interval_ns=20_000_000, reconcile_on_start=False, clickhouse_request_timeout_s=2.0) service = _load_native_store_extension().StorageService(native) @@ -878,7 +917,8 @@ def test_the_lease_holds_through_an_object_store_outage(tmp_path): def _native(spool, holder): native = _storage_config(dead, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool), holder=holder, + spool_root=str(spool), spool_owner_lock=_held(spool), + holder=holder, poll_interval_ns=20_000_000, max_backoff_ns=10_000_000_000, lease_ttl_ns=3_000_000_000, publish_timeout_ns=1_000_000_000, clock_skew_ns=0, reconcile_on_start=False, @@ -1437,7 +1477,8 @@ def test_a_pass_whose_lease_changed_while_it_read_rereads_the_replay_guard( region=REGION, access=ACCESS, secret=SECRET, token=None, insecure=True, connect_timeout=5, read_timeout=15, max_attempts=4, store_id="s3", root=str(spool_root), spool_max_bytes=1 << 40, - limit=-1, max_workers=4, max_in_flight_bytes=1 << 30) + limit=-1, max_workers=4, max_in_flight_bytes=1 << 30, + owner_lock=_held(spool_root)) assert uploaded["ok"], uploaded finally: store.close() @@ -1507,7 +1548,8 @@ def test_a_conflicted_publish_reports_the_conflict_unless_the_lease_was_lost( region=REGION, access=ACCESS, secret=SECRET, token=None, insecure=True, connect_timeout=5, read_timeout=15, max_attempts=4, store_id="s3", root=str(spool_root), spool_max_bytes=1 << 40, - limit=-1, max_workers=4, max_in_flight_bytes=1 << 30) + limit=-1, max_workers=4, max_in_flight_bytes=1 << 30, + owner_lock=_held(spool_root)) assert uploaded["ok"], uploaded finally: store.close() @@ -1688,7 +1730,8 @@ def _upload_behind_the_service(spool_root: Path, endpoint: str) -> list: region=REGION, access=ACCESS, secret=SECRET, token=None, insecure=True, connect_timeout=5, read_timeout=15, max_attempts=4, store_id="s3", root=str(spool_root), spool_max_bytes=1 << 40, - limit=-1, max_workers=4, max_in_flight_bytes=1 << 30) + limit=-1, max_workers=4, max_in_flight_bytes=1 << 30, + owner_lock=_held(spool_root)) assert uploaded["ok"], uploaded finally: store.close() @@ -2368,7 +2411,8 @@ def test_a_64_mib_pack_over_https_with_a_private_ca_hydrates_exactly( max_queue_records=records, max_queue_bytes=2 * records * record_bytes, max_pack_bytes=2 * records * record_bytes, max_pack_records=records, max_linger_ns=60_000_000_000, - overload="drop_newest", admission_timeout=-1)["ok"] + overload="drop_newest", admission_timeout=-1, + owner_lock=_held(spool_root))["ok"] for index in range(records): metadata = CaptureMetadata( capture_id=f"tls-{index:04d}", tenant_id="t", diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py new file mode 100644 index 000000000..867f293cb --- /dev/null +++ b/tests/test_native_spool_adoption_live.py @@ -0,0 +1,279 @@ +"""B6 live: a SIGKILLed capture process's spool is adopted by its successor. + +A capture process -- here a child running what the engine composes: one +SpoolOwnerLock on its own rank directory of the section 2.3 layout, the +storage service and the REAL native pack sink both opened held_by_caller -- +captures, gets part of it indexed, stages the rest, and is SIGKILLed with +packs still in its spool and the publisher lease still live. A new +incarnation on the same node and catalog, in a directory of its own, then +starts: it waits out the dead lease, sweeps its own directory, takes the +dead directory's owner lock (the kernel dropped it with the process), +sweeps its stale .open file, uploads and indexes every ready pack, and +removes the directory. Every capture of the dead process reads back from +the catalog with its bytes. A live sibling -- another process's spool, +lock held -- is left alone. + +When the object store is down at start, the dead directory stays as it +was (its packs are durable there), flush() does not report drained, and +the loop adopts it once the store is back. + +Needs ClickHouse on 127.0.0.1:8123/9000 and the native sink and store +modules: make -C native build/_dmi_native_sink build/_dmi_native_store +PYTHON=/bin/python +""" + +from __future__ import annotations + +import json +import os +import signal +import subprocess +import sys +import time +from pathlib import Path + +import pytest + +# Module-level so the fake-S3 fixture registers in this module. +from tests.test_native_s3_client import ( # noqa: E402 + ACCESS, BUCKET, REGION, SECRET, fake_s3, +) + +REPO = Path(__file__).resolve().parents[1] +BUILD = REPO / "native" / "build" +SINK_BUILT = bool(sorted(BUILD.glob("_dmi_native_sink*.so"))) +STORE_BUILT = bool(sorted(BUILD.glob("_dmi_native_store*.so"))) + +pytestmark = [ + pytest.mark.manual, + pytest.mark.clickhouse, + pytest.mark.skipif( + not (SINK_BUILT and STORE_BUILT), + reason="the native sink and store modules are not built; run " + "`make -C native build/_dmi_native_sink build/_dmi_native_store " + "PYTHON=/bin/python`", + ), +] + +CLICKHOUSE_HOST = os.environ.get("DMI_CLICKHOUSE_HOST", "127.0.0.1") +CLICKHOUSE_HTTP_PORT = int(os.environ.get("DMI_CLICKHOUSE_HTTP_PORT", "8123")) +DATABASE = os.environ.get("DMI_CLICKHOUSE_DATABASE", "default") +LAYOUT = "capture_pack_reference_v1" +# Short, so the successor's wait for the dead process's lease stays short. +LEASE = dict(lease_ttl_s=3.0, publish_timeout_s=1.0) +INDEXED_BY_THE_DEAD = range(0, 4) # flushed to the catalog before the kill +STAGED_BY_THE_DEAD = range(4, 10) # only in its spool when it dies +RECORDS_PER_PACK = 2 + + +def _storage_config(endpoint: str, prefix: str, **overrides): + from dmi.storage.native_capture import NativeCaptureStorageConfig + + fields = dict( + s3_endpoint=endpoint, s3_bucket=BUCKET, s3_region=REGION, + s3_access_key=ACCESS, s3_secret_key=SECRET, + s3_allow_insecure_http=True, clickhouse_host=CLICKHOUSE_HOST, + clickhouse_port=CLICKHOUSE_HTTP_PORT, database=DATABASE, + table_prefix=prefix, poll_interval_s=0.05, **LEASE) + fields.update(overrides) + return NativeCaptureStorageConfig(**fields) + + +def _store(): + from dmi.storage.native_capture import _load_native_store_extension + + return _load_native_store_extension() + + +def _claim(base: Path, config): + """What the engine does first: a fresh rank directory, locked.""" + store = _store() + directory = store.spool_rank_directory( + str(base), config.database, config.table_prefix, config.store_id, 0) + return store.SpoolOwnerLock(directory) + + +def _service(config, directory: str, **options): + from dmi.storage.native_capture import NativeCaptureStorage + + return NativeCaptureStorage( + config, spool_root=directory, spool_max_bytes=1 << 30, + sweep_spool=True, spool_owner_lock="held_by_caller", + adopt_sibling_spools=True, **options) + + +def _envelope(indexes): + from tests.test_native_capture_chain_live import _Envelope + + import torch + + envelope = _Envelope() + for index in indexes: + envelope.add(index, torch.arange(6, dtype=torch.float16) + index) + return envelope + + +def _dead_capture_process(base: str, endpoint: str, prefix: str) -> None: + """The child: capture, index some, stage the rest, then wait to die.""" + import torch # noqa: F401 -- the sink extension links against it + + sys.path.insert(0, str(BUILD)) + import _dmi_native_sink + + # The loop must not upload the second batch before the kill: with this + # poll interval only the explicit flush below runs a cycle. The lease + # thread renews meanwhile, so the process dies holding the lease. + config = _storage_config(endpoint, prefix, poll_interval_s=3600.0) + lock = _claim(Path(base), config) + service = _service(config, lock.directory) + service.start() + sink = _dmi_native_sink.NativePackSink( + spool_root=lock.directory, layout=LAYOUT, + max_pack_records=RECORDS_PER_PACK, max_linger_ns=600 * 10**9, + owner_lock="held_by_caller") + _lease = sink.attach() + first = _envelope(INDEXED_BY_THE_DEAD) + sink.submit_envelope(LAYOUT, first.rows, first.payload()) + assert sink.flush_and_wait(60.0) + service.flush(60.0) + second = _envelope(STAGED_BY_THE_DEAD) + sink.submit_envelope(LAYOUT, second.rows, second.payload()) + assert sink.flush_and_wait(60.0) + sink.rethrow_if_failed() + print(json.dumps({"directory": lock.directory}), flush=True) + time.sleep(3600) + + +def _spawn_dead_process(base: Path, endpoint: str, prefix: str): + env = dict(os.environ, PYTHONPATH=str(REPO / "src") + os.pathsep + str(REPO), + CUDA_VISIBLE_DEVICES="") + child = subprocess.Popen( + [sys.executable, "-m", "tests.test_native_spool_adoption_live", + str(base), endpoint, prefix], + cwd=str(REPO), env=env, stdout=subprocess.PIPE, stderr=subprocess.PIPE, + text=True) + line = child.stdout.readline() + if not line: + child.wait(timeout=30) + raise AssertionError(f"the capture process failed: {child.stderr.read()}") + return child, Path(json.loads(line)["directory"]) + + +def _sigkill(child) -> None: + os.kill(child.pid, signal.SIGKILL) + child.communicate(timeout=30) + assert child.returncode == -signal.SIGKILL + + +def _stale_open_file(directory: Path) -> Path: + """A stage the dead process had in flight: its temp file, left behind.""" + (ready,) = sorted(directory.rglob("*.dmi-pack.ready"))[:1] + stale = ready.parent / ".018f0000-0000-7000-8000-00000000dead.0badf00d.open" + stale.write_bytes(b"half a pack") + return stale + + +def _read_all(config) -> dict: + from dmi.storage.native_capture import NativeCaptureReader + + reader = NativeCaptureReader(config) + selection = reader.select(tenant_id="t") + return {capture.descriptor["capture_id"]: capture.payload + for capture in reader.read(selection, byte_limit=1 << 24)} + + +def _expected() -> dict: + expected = {} + for indexes in (INDEXED_BY_THE_DEAD, STAGED_BY_THE_DEAD): + envelope = _envelope(indexes) + for capture_id, tensor in envelope.expected.items(): + expected[capture_id] = tensor.contiguous().view(-1).numpy().tobytes() + return expected + + +def _catalog(): + from tests.test_native_capture_chain_live import _catalog as chain_catalog + + return chain_catalog() + + +def test_a_sigkilled_process_spool_is_adopted_by_its_successor( + fake_s3, tmp_path): + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix) + child, dead = _spawn_dead_process(base, fake_s3, prefix) + try: + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert len(staged) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + finally: + _sigkill(child) + stale = _stale_open_file(dead) + assert _store().spool_owner(str(dead)) is None # died with it + + # Another process's live spool, bound for the same catalog. + live = _claim(base, config) + (Path(live.directory) / "marker").write_text("live") + + lock = _claim(base, config) + service = _service(config, lock.directory) + started = time.monotonic() + service.start() # waits out the dead lease, then adopts + try: + snapshot = service.snapshot() + assert time.monotonic() - started < 30 + assert snapshot["adopted_spools"] == 1, snapshot + assert snapshot["adopted_packs"] == len(staged), snapshot + assert snapshot["adoption_owed"] is False, snapshot + assert not stale.exists() + assert not dead.exists() + service.flush(60.0) + assert _read_all(config) == _expected() + finally: + service.stop() + assert (Path(live.directory) / "marker").read_text() == "live" + assert live.held + assert lock.release_and_remove_if_empty() + live.release() + + +def test_a_dead_spool_waits_in_place_while_the_object_store_is_down( + fake_s3, tmp_path): + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + child, dead = _spawn_dead_process(base, fake_s3, prefix) + _sigkill(child) + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert staged + + switch = _Switch.to_url(fake_s3) + switch.cut() + config = _storage_config(switch.url, prefix) + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + snapshot = service.snapshot() + assert snapshot["adoption_owed"] is True, snapshot + assert snapshot["adopted_spools"] == 0, snapshot + # Nothing left the dead spool: its packs are durable there. + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + with pytest.raises(TimeoutError): + service.flush(2.0) + + switch.restore() + service.flush(60.0) + snapshot = service.snapshot() + assert snapshot["adoption_owed"] is False, snapshot + assert snapshot["adopted_spools"] == 1, snapshot + assert not dead.exists() + assert _read_all(config) == _expected() + finally: + service.stop() + lock.release_and_remove_if_empty() + + +if __name__ == "__main__": + _dead_capture_process(*sys.argv[1:4]) diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py new file mode 100644 index 000000000..fa3962f61 --- /dev/null +++ b/tests/test_native_spool_ownership.py @@ -0,0 +1,268 @@ +"""B6 through the Python surface: one owner per spool directory. + +``_dmi_native_store`` binds the spool's owner lock (``SpoolOwnerLock``), the +section 2.3 layout (``spool_rank_directory``) and who owns a directory +(``spool_owner``); the storage service and the native pack sink each open +their spool with ``owner_lock="take"`` or ``"held_by_caller"``. + +The regression this pins is the in-process one: a sink and a storage +service on ONE directory in ONE process. Were each to take the lock, the +second would be refused by the first -- flock binds to an open file +description, not to the process -- so the engine holds one SpoolOwnerLock +and opens both held_by_caller. The service is constructed, not started: +start() needs a catalog, which the live suites bring +(test_native_spool_adoption_live.py, test_native_capture_chain_live.py). +""" + +from __future__ import annotations + +import hashlib +import json +import os +import socket +import subprocess +import sys +from pathlib import Path + +import pytest + +REPO = Path(__file__).resolve().parents[1] +BUILD = REPO / "native" / "build" +STORE_BUILT = bool(sorted(BUILD.glob("_dmi_native_store*.so"))) +SINK_BUILT = bool(sorted(BUILD.glob("_dmi_native_sink*.so"))) + +pytestmark = [ + pytest.mark.cpu, + pytest.mark.skipif( + not STORE_BUILT, + reason="the native store module is not built; run `make -C native " + "build/_dmi_native_store PYTHON=/bin/python`"), +] + +LAYOUT = "capture_pack_reference_v1" +NFS_SUPER_MAGIC = 0x6969 + + +def _store(): + from dmi.storage.native_capture import _load_native_store_extension + + return _load_native_store_extension() + + +def _config(**overrides): + from dmi.storage.native_capture import NativeCaptureStorageConfig + + fields = dict( + s3_endpoint="http://127.0.0.1:9", s3_bucket="bucket", + s3_access_key="AKIA-test", s3_secret_key="secret-test", + s3_allow_insecure_http=True, clickhouse_port=9) + fields.update(overrides) + return NativeCaptureStorageConfig(**fields) + + +def _rank_directory(base: Path, config, rank: int = 0) -> str: + return _store().spool_rank_directory( + str(base), config.database, config.table_prefix, config.store_id, + rank) + + +def _service(config, spool_root, **options): + from dmi.storage.native_capture import NativeCaptureStorage + + return NativeCaptureStorage(config, spool_root=str(spool_root), + spool_max_bytes=1 << 30, sweep_spool=True, + **options) + + +def _sink(spool_root, **options): + sys.path.insert(0, str(BUILD)) + try: + import _dmi_native_sink + finally: + sys.path.remove(str(BUILD)) + return _dmi_native_sink.NativePackSink( + spool_root=str(spool_root), layout=LAYOUT, max_pack_records=4, + **options) + + +def _stage_one(sink) -> None: + """One float16 capture through the sink's ring-facing submit.""" + import torch + from dmi.storage.capture import CaptureMetadata + + tensor = torch.arange(6, dtype=torch.float16) + metadata = CaptureMetadata( + capture_id="own-0", tenant_id="t", experiment_id="e", run_id="r", + session_id="s", request_id="q", sequence_id="n", model_id="m", + model_revision="mr", adapter_revision=None, + capture_policy_version="v", hook_name="resid_post", layer_number=0, + producer_rank=0, step_number=0, token_start=0, token_end=1, + batch_position=0, dtype="float16", shape=(6,), + captured_at_ns=1_700_000_000_000_000_000, + ).to_mapping() + lease = sink.attach() + sink.submit_envelope(LAYOUT, [{ + "metadata_json": json.dumps(metadata), "offset": 0, "length": 12, + "dtype": 5, "shape": [6]}], tensor.view(torch.uint8)) + assert sink.flush_and_wait(30.0) + sink.rethrow_if_failed() + del lease + + +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_the_real_sink_and_service_share_one_spool_in_one_process(tmp_path): + pytest.importorskip("torch") + config = _config() + directory = _rank_directory(tmp_path / "spool", config) + with _store().SpoolOwnerLock(directory) as lock: + service = _service(config, directory, + spool_owner_lock="held_by_caller", + adopt_sibling_spools=True) + sink = _sink(directory, owner_lock="held_by_caller") + _stage_one(sink) + assert len(list(Path(directory).rglob("*.dmi-pack.ready"))) == 1 + snapshot = service.snapshot() + assert snapshot["adopted_spools"] == 0 + assert snapshot["adoption_owed"] is False + # Neither Spool took a lock of its own: the one holder is this + # process, through `lock`. + assert _store().spool_owner(directory)["pid"] == os.getpid() + del sink, service + assert lock.held + assert _store().spool_owner(directory) is None + + +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_two_takes_in_one_process_refuse_each_other(tmp_path): + """What a per-Spool lock would do to the engine's sink and service.""" + pytest.importorskip("torch") + directory = tmp_path / "spool" + service = _service(_config(), directory) # take + with pytest.raises(RuntimeError, match=f"owned by pid {os.getpid()}"): + _sink(directory) # take + del service + _sink(directory) # the service's lock went with it + + +def test_a_second_process_is_refused_naming_the_holder(tmp_path): + directory = tmp_path / "spool" + with _store().SpoolOwnerLock(str(directory)): + probe = ( + "import sys; sys.path.insert(0, sys.argv[1]);" + "import _dmi_native_store as m\n" + "try:\n" + " m.SpoolOwnerLock(sys.argv[2])\n" + "except m.SpoolOwnedError as e:\n" + " print(e)\n") + result = subprocess.run( + [sys.executable, "-c", probe, str(BUILD), str(directory)], + capture_output=True, text=True, timeout=60) + assert result.returncode == 0, result.stderr + assert f"owned by pid {os.getpid()} on host {socket.gethostname()}" in ( + result.stdout), result.stdout + + +def test_spool_owner_names_the_holder_while_it_holds(tmp_path): + store = _store() + directory = str(tmp_path / "spool") + assert store.spool_owner(directory) is None + lock = store.SpoolOwnerLock(directory) + assert lock.held + assert store.spool_owner(directory) == { + "host": socket.gethostname(), "pid": os.getpid()} + lock.release() + assert not lock.held + assert store.spool_owner(directory) is None + + +def test_the_owned_error_is_a_runtime_error(tmp_path): + store = _store() + assert issubclass(store.SpoolOwnedError, RuntimeError) + with store.SpoolOwnerLock(str(tmp_path / "spool")): + with pytest.raises(store.SpoolOwnedError): + store.SpoolOwnerLock(str(tmp_path / "spool")) + + +def test_a_nested_spool_directory_is_refused(tmp_path): + store = _store() + with store.SpoolOwnerLock(str(tmp_path / "outer")): + with pytest.raises(ValueError, match="nested"): + store.SpoolOwnerLock(str(tmp_path / "outer" / "inner")) + with pytest.raises(RuntimeError, match="nested"): + _service(_config(), tmp_path / "outer" / "inner") + + +def test_a_shared_filesystem_is_refused_unless_allowed(tmp_path): + store = _store() + store._set_spool_filesystem_type_for_testing(NFS_SUPER_MAGIC) + try: + with pytest.raises(ValueError, match="NFS.*node-local"): + store.SpoolOwnerLock(str(tmp_path / "nfs")) + with pytest.raises(RuntimeError, match="node-local"): + _service(_config(), tmp_path / "nfs") + with store.SpoolOwnerLock(str(tmp_path / "nfs"), + allow_shared_filesystem=True): + pass + _service(_config(), tmp_path / "nfs-allowed", + spool_allow_shared_filesystem=True) + finally: + store._set_spool_filesystem_type_for_testing(None) + with store.SpoolOwnerLock(str(tmp_path / "local")): + pass + + +def test_the_rank_directory_layout(tmp_path): + store = _store() + key = hashlib.sha256(b"db/prefix/s3").hexdigest()[:12] + assert store.spool_catalog_key("db", "prefix", "s3") == key + assert store.spool_rank_directory( + "/base", "db", "prefix", "s3", 3, "0a1b2c3d") == ( + f"/base/{key}/r3-0a1b2c3d") + fresh = {store.spool_rank_directory("/base", "db", "prefix", "s3", 0) + for _ in range(16)} + assert len(fresh) == 16 # a fresh incarnation each time + for path in fresh: + assert path.startswith(f"/base/{key}/r0-") + with pytest.raises(ValueError, match="incarnation"): + store.spool_rank_directory("/base", "db", "prefix", "s3", 0, "XYZ") + + +def test_adoption_needs_a_rank_directory_under_this_catalogs_key(tmp_path): + config = _config() + with pytest.raises(ValueError, match="adopt_sibling_spools"): + _service(config, tmp_path / "spool", adopt_sibling_spools=True) + # A rank directory, but under another catalog's key: adopting its + # siblings would index that catalog's packs into this one. + other = _store().spool_rank_directory( + str(tmp_path / "base"), config.database, "another_prefix", + config.store_id, 0) + with pytest.raises(ValueError, match="adopt_sibling_spools"): + _service(config, other, adopt_sibling_spools=True) + mine = _rank_directory(tmp_path / "base", config) + assert _service(config, mine, adopt_sibling_spools=True).snapshot()[ + "adopted_packs"] == 0 + + +def test_an_unknown_owner_lock_mode_is_refused(tmp_path): + with pytest.raises(ValueError, match="spool_owner_lock"): + _service(_config(), tmp_path / "spool", spool_owner_lock="share") + + +def test_held_by_caller_with_nothing_held_is_refused(tmp_path): + with pytest.raises(RuntimeError, match="held_by_caller"): + _service(_config(), tmp_path / "spool", + spool_owner_lock="held_by_caller") + + +def test_a_drained_directory_is_removed_with_its_lock(tmp_path): + store = _store() + directory = tmp_path / "spool" + lock = store.SpoolOwnerLock(str(directory)) + (directory / "v1" / "tenant=t").mkdir(parents=True) + assert lock.release_and_remove_if_empty() + assert not directory.exists() + lock = store.SpoolOwnerLock(str(directory)) + (directory / "kept.quarantined").write_bytes(b"x") + assert not lock.release_and_remove_if_empty() + assert (directory / "kept.quarantined").exists() + assert not lock.held From f5b34b09ad64cefd32dc08d48c02c424dea029c6 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Mon, 28 Sep 2026 20:08:58 -0400 Subject: [PATCH 03/43] Hold the spool owner lock in the engine around its sink and its service Under storage_backend="persistent" with capture_storage_config, the engine pointed its storage service and its native pack sink at capture_sink_config.spool_root itself. Now it claims a directory of its own there, the plan's section 2.3 layout: //r-/ catalog_key being the first 12 hex digits of sha256 of database/table_prefix/store_id, rank torchrun's RANK (0 when unset or not a rank: it only labels the directory), and the incarnation fresh for every create_record_runtime, so two jobs on a node, or a restart racing its predecessor, never share a directory. claim_spool_directory creates it with its owner lock already held, BEFORE the service is built; the service starts (sweeping the directory, then adopting the dead siblings under the same catalog key) and only then does the sink open it, both held_by_caller. The claim is let go last -- after the sink is sealed, the ring stopped and the service drained and stopped -- on close(), on enable_ring_transport replacing the record ring, and on a create_record_runtime that fails after the service started (or whose service fails to start). A drained directory is removed then; one still holding packs stays for the next process on the node to adopt. Without capture_storage_config the sink takes spool_root itself, as the one owner of packs something else will drain. With an explicit record_sink the service drains spool_root as that sink writes it -- unswept, as before, and adopting nothing -- opening it held_by_caller when this process already holds its lock (a NativePackSink the caller built took it; a second take would refuse this very process), and taking it otherwise, which another process's lock refuses by name. NativeSinkConfig.spool_allow_shared_filesystem carries the NFS/Lustre override to the lock and both Spools. flush_and_wait's TimeoutError now says when a dead process's spool is what is still owed. The v1 contract describes the directory, its owner, adoption and the node-local rule. Tests (test_native_capture_storage_wiring.py, the native modules faked), each red against the previous engine for the reason it names: * the order is now layout, lock, service construct, service start, sink open -- on the rank directory, both held_by_caller, adoption on (was: service first, on spool_root, no lock); * RANK names the directory, and a RANK that is not a rank names rank 0; * the shared-filesystem override reaches the lock and both Spools (the field did not exist); * a service that fails to start leaves no lock and opens no sink; * close(), a failed drain, a failed attach and replacing the record ring each release the lock after the service stops; * an explicit sink: spool_root, unswept, no adoption, take -- or held_by_caller when this process holds the lock; * with no storage config the sink takes spool_root itself; * a flush that times out on an owed adoption says so (the message named only the unindexed packs). The full cpu tier passes. --- docs/integration-api-v1.md | 18 ++ src/dmi/engine.py | 94 ++++++-- src/dmi/storage/native_capture.py | 101 ++++++++- tests/test_native_capture_storage_wiring.py | 225 ++++++++++++++++++-- 4 files changed, 400 insertions(+), 38 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index b8d63bb75..daceeecf5 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -185,6 +185,24 @@ refuses the persistent backend with `ConfigurationError` rather than generating with nothing stored. The catalog takes one publisher per `(database, table_prefix)`, so a second engine on the same catalog is refused at `create_record_runtime`. +The engine spools into a directory of its own under +`capture_sink_config.spool_root`, +`//r-/` (the key is the first 12 +hex digits of sha256 of `database/table_prefix/store_id`, the rank torchrun's +`RANK`, 0 when unset, and the incarnation fresh on every +`create_record_runtime`), and owns it: an flock on its `.owner.lock`, taken +before the service starts and let go after the sink and the service are done, +when a drained directory is removed. A second process on a directory is +refused, naming the holder's pid and host. At start the service adopts the +directories under the same catalog key whose owners have died: their stale +`.open` files are swept, their ready packs uploaded and indexed, and the +directory removed, so a crashed process's packs reach the catalog through the +next one on the node, whatever run it belongs to. The spool root must be +node-local: NFS and Lustre are refused unless +`NativeSinkConfig.spool_allow_shared_filesystem`. Without +`capture_storage_config` the sink owns `spool_root` itself. With an explicit +`record_sink`, the service drains `spool_root` as that sink writes it, +unswept and adopting nothing. To reach a secured catalog, set `clickhouse_scheme="https"` (and the server's TLS HTTP port, usually 8443) on `NativeCaptureStorageConfig`. The client always verifies the server's certificate and name, against libcurl's built-in CA diff --git a/src/dmi/engine.py b/src/dmi/engine.py index 211ff3ebf..9dd85d9d7 100644 --- a/src/dmi/engine.py +++ b/src/dmi/engine.py @@ -162,6 +162,9 @@ def __init__( "NativeCaptureStorageConfig") # The running storage service, while a record runtime is attached. self._capture_storage: Optional[Any] = None + # This process's own spool directory and its owner lock, held from + # before the service starts until the sink and the service are done. + self._spool_claim: Optional[Any] = None host_configured = host_engine is not None or db_config is not None if self._storage_backend == "in-memory" and not host_configured: raise ValueError( @@ -422,7 +425,7 @@ def create_record_runtime( # spool: its start sweeps a crashed sink's stale .open files, which # is safe only while nothing writes there. An explicit record_sink # may already hold the spool open, so it is left unswept. - storage = self._start_capture_storage(sweep_spool=record_sink is None) + storage = self._start_capture_storage(record_sink) try: runtime = self._attach_record_runtime( record_format, record_schema, record_sink, @@ -431,7 +434,10 @@ def create_record_runtime( except BaseException: if storage is not None: self._capture_storage = None - storage.stop() + try: + storage.stop() + finally: + self._release_spool_claim() raise return runtime @@ -486,23 +492,75 @@ def _refuse_an_unbounded_sink_admission( "NativePackSink (storage_backend='persistent', or record_sink=) " "to use a budget") - def _start_capture_storage(self, *, sweep_spool: bool) -> Optional[Any]: + def _start_capture_storage(self, record_sink: Optional[Any]) -> Optional[Any]: + """Start the storage service, and claim the spool it drains. + + With the default sink, this process spools into a directory of its + own under ``capture_sink_config.spool_root`` -- + ``//r-/``, fresh on every + start -- and owns it: its owner lock is taken here, before the + service opens it, and held until the sink and the service are done + (``_release_spool_claim``). The service and the sink both open it + ``held_by_caller``; two takes in one process refuse each other. The + service adopts the directories of dead processes beside it. + + An explicit ``record_sink`` writes where it was built to: the + service drains ``spool_root`` itself, unswept and with no siblings + to adopt, beside the sink's lock if this process holds one. + """ config = self._capture_storage_config if config is None or self._storage_backend != "persistent": return None - from .storage.native_capture import NativeCaptureStorage + from .storage.native_capture import ( + NativeCaptureStorage, + claim_spool_directory, + spool_owner_lock_beside, + ) sink_config = self._capture_sink_config - storage = NativeCaptureStorage( - config, - spool_root=sink_config.spool_root, - spool_max_bytes=sink_config.spool_max_bytes, - sweep_spool=sweep_spool, - ) - storage.start() + shared = sink_config.spool_allow_shared_filesystem + if record_sink is not None: + storage = NativeCaptureStorage( + config, + spool_root=sink_config.spool_root, + spool_max_bytes=sink_config.spool_max_bytes, + sweep_spool=False, + spool_owner_lock=spool_owner_lock_beside( + sink_config.spool_root), + spool_allow_shared_filesystem=shared, + ) + storage.start() + self._capture_storage = storage + return storage + claim = claim_spool_directory(sink_config, config) + try: + storage = NativeCaptureStorage( + config, + spool_root=claim.directory, + spool_max_bytes=sink_config.spool_max_bytes, + sweep_spool=True, + spool_owner_lock="held_by_caller", + adopt_sibling_spools=True, + spool_allow_shared_filesystem=shared, + ) + storage.start() + except BaseException: + claim.release() + raise + self._spool_claim = claim self._capture_storage = storage return storage + def _release_spool_claim(self) -> None: + """Let go of this process's spool directory, once nothing writes it. + + A drained directory is removed; one still holding packs stays, for + the next process on the node for this catalog to adopt. + """ + claim, self._spool_claim = self._spool_claim, None + if claim is not None: + claim.release() + def _attach_record_runtime( self, record_format: "RecordFormat[MetadataT]", @@ -526,8 +584,13 @@ def _attach_record_runtime( ): from .storage.capture.native_sink import create_native_pack_sink + # Into the directory the engine claimed for the service, under + # its lock; without a service the sink owns the spool root. + claim = self._spool_claim record_sink = create_native_pack_sink( - self._capture_sink_config + self._capture_sink_config, + spool_root=None if claim is None else claim.directory, + owner_lock="take" if claim is None else "held_by_caller", ).native_sink _native_engine = _native_module() @@ -850,7 +913,12 @@ def _retire_capture_storage(self, storage: Any, deadline: float) -> None: except Exception as exc: _LOG.warning("capture storage did not drain: %s", exc) finally: - storage.stop() + try: + storage.stop() + finally: + # Last: the ring is stopped and the sink sealed by now, and + # the service has stopped touching the directory. + self._release_spool_claim() def close(self) -> None: """Tear down backend resources.""" diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index 7c119ffc8..d42147bde 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -21,6 +21,12 @@ ``table_prefix``), so run one capture process per catalog. A second engine on the same catalog waits ``start_lease_wait_s`` for the lease, then fails at ``create_record_runtime`` with the lease held, naming the holder. + +Each spool directory has one owner process (an flock on its +``.owner.lock``). The engine claims a fresh directory of its own under +``spool_root`` (:func:`claim_spool_directory`) and its service adopts the +spools of dead processes beside it, so a crashed run's packs reach the +catalog through the next process on the node for the same catalog. """ from __future__ import annotations @@ -125,10 +131,15 @@ class NativeSinkConfig: never be admitted; ``validate_capture_bounds`` refuses such a bound at attach, before any forward runs. - ``spool_root`` must be node-local, and a spool directory has one owner - process, held by an flock on its ``.owner.lock``. A root on NFS or - Lustre is refused unless ``spool_allow_shared_filesystem``: flock there - does not keep out a process on another node. + ``spool_root`` must be node-local, and each spool directory has one + owner process, held by an flock on its ``.owner.lock``. With + ``capture_storage_config`` the engine spools into a directory of its own + under it, ``//r-/`` (see + :func:`claim_spool_directory`), and its storage service adopts the + directories of dead processes beside it. Without one, the sink owns + ``spool_root`` itself. A root on NFS or Lustre is refused unless + ``spool_allow_shared_filesystem``: flock there does not keep out a + process on another node. """ spool_root: str @@ -566,6 +577,76 @@ def validate_capture_bounds( SPOOL_OWNER_LOCKS = ("take", "held_by_caller") +def _spool_producer_rank() -> int: + """The rank a spool directory is named for: torchrun's global ``RANK``, + 0 for a single process or anything that is not a rank. It only labels + the directory; the incarnation is what keeps two processes apart.""" + text = os.environ.get("RANK", "") + return int(text) if text.isdigit() else 0 + + +class SpoolClaim: + """This process's own spool directory, owned through its lock. + + From :func:`claim_spool_directory`. Hold it for as long as anything in + the process writes or reads the directory -- the engine holds it from + before its storage service starts until the sink and the service are + done -- then :meth:`release` it. + """ + + def __init__(self, lock: Any) -> None: + self._lock = lock + self.directory: str = lock.directory + + @property + def held(self) -> bool: + return bool(self._lock.held) + + def release(self) -> bool: + """Let go of the directory, removing it if nothing but its lock + file is left. Whatever did not drain stays, and the next process on + the node for this catalog adopts it. Returns whether it was + removed.""" + return bool(self._lock.release_and_remove_if_empty()) + + +def claim_spool_directory( + sink_config: NativeSinkConfig, + storage_config: NativeCaptureStorageConfig, +) -> SpoolClaim: + """Create and lock this process's spool directory. + + ``//r-/``: the catalog key is + the first 12 hex digits of sha256 of ``database/table_prefix/store_id``, + so every directory under it holds packs for this catalog and store; the + incarnation is fresh for every call, so no two processes -- two jobs on + one node, or a restart -- share a directory; the rank is torchrun's + ``RANK`` (0 when unset). The directory is created with its owner lock + already held. Raises ``SpoolOwnedError`` (a ``RuntimeError``) if another + process holds it, and ``ValueError`` for a shared filesystem or a + directory nested in another spool. + """ + module = _load_native_store_extension() + directory = module.spool_rank_directory( + sink_config.spool_root, storage_config.database, + storage_config.table_prefix, storage_config.store_id, + _spool_producer_rank()) + return SpoolClaim(module.SpoolOwnerLock( + directory, + allow_shared_filesystem=sink_config.spool_allow_shared_filesystem)) + + +def spool_owner_lock_beside(spool_root: str) -> str: + """The owner-lock mode for a second Spool on a directory: ``held_by_caller`` + when this process already holds its lock (a sink the caller built took + it), else ``take``, which another process's lock refuses by name.""" + owner = _load_native_store_extension().spool_owner(spool_root) + if (owner is not None and owner["pid"] == os.getpid() + and owner["host"] == socket.gethostname()): + return "held_by_caller" + return "take" + + class NativeCaptureStorage: """The in-process storage service: spool -> object store -> catalog. @@ -645,10 +726,15 @@ def flush(self, timeout_s: float) -> None: """ if not self._service.flush(float(timeout_s)): snapshot = self._service.snapshot() + # A dead process's spool this service has still to adopt keeps + # it undrained too (adopt_sibling_spools). + adopting = (", and a dead process's spool still to adopt" + if snapshot.get("adoption_owed") else "") raise TimeoutError( "timed out waiting for staged packs to reach the catalog " - f"({snapshot['pending_index']} uploaded but unindexed); last " - f"error: {snapshot['last_error'] or 'none'}") + f"({snapshot['pending_index']} uploaded but unindexed" + f"{adopting}); last error: " + f"{snapshot['last_error'] or 'none'}") def stop(self) -> None: """Stop the background thread and release the lease. No flush.""" @@ -861,5 +947,8 @@ def read( "NativeCaptureSelection", "NativeCaptureStorage", "NativeCaptureStorageConfig", + "SpoolClaim", + "claim_spool_directory", + "spool_owner_lock_beside", "validate_capture_bounds", ] diff --git a/tests/test_native_capture_storage_wiring.py b/tests/test_native_capture_storage_wiring.py index 21cead936..de34cb4b8 100644 --- a/tests/test_native_capture_storage_wiring.py +++ b/tests/test_native_capture_storage_wiring.py @@ -2,10 +2,12 @@ The service itself is C++ and runs against a real object store and catalog in test_native_capture_storage_live.py. This suite pins what the engine -promises around it, with the native modules faked: the service starts -before the sink opens the spool it sweeps, ``flush_and_wait`` waits for the -catalog as well as the spool, ``close`` drains and releases the lease, and a -failed attach does not leave the lease held. +promises around it, with the native modules faked: the engine takes its own +rank directory's owner lock before anything opens it, the service starts +before the sink opens the spool it sweeps, both open that directory +held_by_caller, ``flush_and_wait`` waits for the catalog as well as the +spool, ``close`` drains and releases the lease and only then the spool +lock, and a failed attach leaves neither the lease nor the lock held. """ from __future__ import annotations @@ -417,12 +419,38 @@ def rethrow_if_failed(self): pass -def _capture_engine(monkeypatch, tmp_path, *, fail_ring=False): +RANK_DIRECTORY = "{base}/0123456789ab/r{rank}-0a1b2c3d" + + +class _FakeSpoolLock: + """Stands in for _dmi_native_store.SpoolOwnerLock.""" + + def __init__(self, events, directory, allow_shared_filesystem=False): + self.events = events + self.directory = directory + self.allow_shared_filesystem = allow_shared_filesystem + self.held = True + events.append(("lock", "acquire", directory)) + + def release_and_remove_if_empty(self): + self.held = False + self.events.append(("lock", "release")) + return True + + def release(self): + self.release_and_remove_if_empty() + + +def _capture_engine(monkeypatch, tmp_path, *, fail_ring=False, + fail_start=False, spool_owner=None, storage=True): """An engine under storage_backend="persistent" with both native modules - faked. Returns (engine, events, services).""" + faked. Returns (engine, events, services); ``events`` also records the + spool locks and the sinks' keyword arguments (``sinks``).""" from dmi.storage.capture.native_sink import NativeSinkConfig events, services = [], [] + events_sinks: list = [] + locks: list = [] engine = MonitoringEngine(enable_ring_transport=False) engine._ring_transport = SimpleNamespace(null_offload=False, force_eager=False) @@ -431,7 +459,7 @@ def _capture_engine(monkeypatch, tmp_path, *, fail_ring=False): engine._storage_backend = "persistent" engine._capture_sink_config = NativeSinkConfig( spool_root=str(tmp_path / "spool"), spool_max_bytes=1 << 30) - engine._capture_storage_config = _storage_config() + engine._capture_storage_config = _storage_config() if storage else None class _Lease: def release(self): @@ -444,15 +472,33 @@ def _acquire_engine(self): class _NativePackSink(_RecordSink): def __init__(self, **kwargs): events.append(("sink", "open", kwargs["spool_root"])) + events_sinks.append(kwargs) def _service(config): service = _FakeService(events, config) + if fail_start: + def _refuse(): + events.append(("service", "start")) + raise RuntimeError("publisher lease held elsewhere") + service.start = _refuse services.append(service) return service + def _lock(directory, allow_shared_filesystem=False): + lock = _FakeSpoolLock(events, directory, allow_shared_filesystem) + locks.append(lock) + return lock + + def _rank_directory(base, database, table_prefix, store_id, rank): + events.append(("layout", database, table_prefix, store_id, rank)) + return RANK_DIRECTORY.format(base=base, rank=rank) + def _load_named_extension(name): if name == "_dmi_native_store": return SimpleNamespace(StorageService=_service, + SpoolOwnerLock=_lock, + spool_rank_directory=_rank_directory, + spool_owner=lambda directory: spool_owner, SEARCH_ITEM_COLUMNS=()) return SimpleNamespace(NativePackSink=_NativePackSink) @@ -503,6 +549,8 @@ def __init__(self, config, host): import dmi.transport monkeypatch.setattr(dmi.transport, "native", native, raising=False) + engine._test_sinks = events_sinks + engine._test_locks = locks return engine, events, services @@ -522,21 +570,84 @@ def encode(self, metadata, entry): def test_the_service_starts_before_the_sink_opens_the_spool(monkeypatch, tmp_path): + """The engine's own directory is locked first; the service starts (and + sweeps it) before the sink opens it; both open it under that lock.""" + monkeypatch.delenv("RANK", raising=False) engine, events, services = _capture_engine(monkeypatch, tmp_path) engine.create_record_runtime(_record_format()) - spool_root = str(tmp_path / "spool") - assert events[:3] == [ + directory = RANK_DIRECTORY.format(base=tmp_path / "spool", rank=0) + storage = _storage_config() + assert events[:5] == [ + ("layout", storage.database, storage.table_prefix, storage.store_id, + 0), + ("lock", "acquire", directory), ("service", "construct"), ("service", "start"), - ("sink", "open", spool_root), + ("sink", "open", directory), ] config = services[0].config - assert config["spool_root"] == spool_root + assert config["spool_root"] == directory assert config["spool_max_bytes"] == 1 << 30 assert config["sweep_spool_on_start"] is True + assert config["spool_owner_lock"] == "held_by_caller" + assert config["adopt_sibling_spools"] is True + assert config["spool_allow_shared_filesystem"] is False assert config["holder"] # a generated lease holder, never empty + (sink,) = engine._test_sinks + assert sink["owner_lock"] == "held_by_caller" + assert engine._test_locks[0].held + + +def test_the_spool_directory_is_named_for_the_rank(monkeypatch, tmp_path): + monkeypatch.setenv("RANK", "3") + engine, events, services = _capture_engine(monkeypatch, tmp_path) + + engine.create_record_runtime(_record_format()) + + assert services[0].config["spool_root"] == RANK_DIRECTORY.format( + base=tmp_path / "spool", rank=3) + + +@pytest.mark.parametrize("rank", ["", "-1", "x", "1.5"]) +def test_a_rank_that_is_not_a_rank_names_rank_zero(monkeypatch, tmp_path, rank): + monkeypatch.setenv("RANK", rank) + engine, _events, services = _capture_engine(monkeypatch, tmp_path) + + engine.create_record_runtime(_record_format()) + + assert services[0].config["spool_root"].endswith("/r0-0a1b2c3d") + + +def test_the_shared_filesystem_override_reaches_the_lock_and_both_spools( + monkeypatch, tmp_path): + from dmi.storage.capture.native_sink import NativeSinkConfig + + engine, _events, services = _capture_engine(monkeypatch, tmp_path) + engine._capture_sink_config = NativeSinkConfig( + spool_root=str(tmp_path / "spool"), + spool_allow_shared_filesystem=True) + + engine.create_record_runtime(_record_format()) + + assert engine._test_locks[0].allow_shared_filesystem is True + assert services[0].config["spool_allow_shared_filesystem"] is True + assert engine._test_sinks[0]["allow_shared_filesystem"] is True + + +def test_a_service_that_fails_to_start_releases_the_spool_lock( + monkeypatch, tmp_path): + engine, events, _services = _capture_engine(monkeypatch, tmp_path, + fail_start=True) + + with pytest.raises(RuntimeError, match="lease held elsewhere"): + engine.create_record_runtime(_record_format()) + + assert events[-1] == ("lock", "release") + assert not engine._test_locks[0].held + assert engine._capture_storage is None + assert not any(event[0] == "sink" for event in events) def test_an_explicit_sink_leaves_the_spool_unswept(monkeypatch, tmp_path): @@ -546,7 +657,64 @@ def test_an_explicit_sink_leaves_the_spool_unswept(monkeypatch, tmp_path): engine.create_record_runtime( _record_format(), record_sink=dmi.transport.native.RecordSink()) - assert services[0].config["sweep_spool_on_start"] is False + config = services[0].config + assert config["sweep_spool_on_start"] is False + # An explicit sink writes where it was built to: the configured root, + # which the engine neither lays out nor adopts siblings around. Nothing + # holds it here, so the service takes its lock. + assert config["spool_root"] == str(tmp_path / "spool") + assert config["adopt_sibling_spools"] is False + assert config["spool_owner_lock"] == "take" + assert not any(event[0] == "lock" for event in events) + + +def test_an_explicit_sink_holding_the_spool_shares_it_with_the_service( + monkeypatch, tmp_path): + """A NativePackSink the caller built takes the spool's owner lock; the + engine's service in the same process must open beside it, not take it + again (that would be refused, naming this very process).""" + import os + import socket + + engine, _events, services = _capture_engine( + monkeypatch, tmp_path, + spool_owner={"host": socket.gethostname(), "pid": os.getpid()}) + import dmi.transport + + engine.create_record_runtime( + _record_format(), record_sink=dmi.transport.native.RecordSink()) + + assert services[0].config["spool_owner_lock"] == "held_by_caller" + + +def test_an_explicit_sink_on_a_spool_another_process_owns_is_taken( + monkeypatch, tmp_path): + """Owned by another process: the service takes it, and so is refused + by the native spool naming that holder.""" + engine, _events, services = _capture_engine( + monkeypatch, tmp_path, spool_owner={"host": "elsewhere", "pid": 1}) + import dmi.transport + + engine.create_record_runtime( + _record_format(), record_sink=dmi.transport.native.RecordSink()) + + assert services[0].config["spool_owner_lock"] == "take" + + +def test_a_sink_without_a_service_owns_its_spool_itself(monkeypatch, tmp_path): + """No capture_storage_config: packs stay in the spool for something else + to drain, and the sink is the directory's one owner -- no layout, no + engine lock.""" + engine, events, services = _capture_engine(monkeypatch, tmp_path, + storage=False) + + engine.create_record_runtime(_record_format()) + + assert services == [] + (sink,) = engine._test_sinks + assert sink["spool_root"] == str(tmp_path / "spool") + assert sink["owner_lock"] == "take" + assert not any(event[0] in ("lock", "layout") for event in events) def test_a_failed_attach_stops_the_service_it_started(monkeypatch, tmp_path): @@ -556,8 +724,9 @@ def test_a_failed_attach_stops_the_service_it_started(monkeypatch, tmp_path): with pytest.raises(RuntimeError, match="ring init failed"): engine.create_record_runtime(_record_format()) - assert ("service", "stop") in events + assert events[-2:] == [("service", "stop"), ("lock", "release")] assert engine._capture_storage is None + assert not engine._test_locks[0].held def test_flush_waits_for_the_catalog_after_the_sink(monkeypatch, tmp_path): @@ -581,6 +750,19 @@ def test_flush_reports_packs_that_did_not_reach_the_catalog(monkeypatch, tmp_pat engine.flush_and_wait(1.0) +def test_flush_says_when_a_dead_spool_is_still_to_adopt(monkeypatch, tmp_path): + engine, _events, services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + services[0].flush_results = [False] + services[0].snapshot = lambda: { + "pending_index": 0, "adoption_owed": True, + "last_error": "adopting dead spool /x: upload failed"} + + with pytest.raises(TimeoutError, + match="dead process's spool still to adopt.*upload"): + engine.flush_and_wait(1.0) + + def test_close_flushes_the_sink_before_the_ring_stops(monkeypatch, tmp_path): """The sink's open pack is in memory until a flush seals it, and stopping the ring releases the sink without one. So close() flushes the sink @@ -591,9 +773,10 @@ def test_close_flushes_the_sink_before_the_ring_stops(monkeypatch, tmp_path): engine.close() + # The spool lock last: after the sink and the service are both done. assert [event[:2] for event in events] == [ ("sink", "flush"), ("ring", "stop"), - ("service", "flush"), ("service", "stop")] + ("service", "flush"), ("service", "stop"), ("lock", "release")] # One budget for the whole drain: the service gets what the sink left. assert 59.0 <= events[0][2] <= 60.0 assert 0.0 <= events[2][2] <= 60.0 @@ -615,7 +798,7 @@ def _failing_flush(timeout_s): assert [event[:2] for event in events] == [ ("sink", "flush"), ("ring", "stop"), - ("service", "flush"), ("service", "stop")] + ("service", "flush"), ("service", "stop"), ("lock", "release")] def test_close_releases_the_lease_even_when_the_drain_fails(monkeypatch, tmp_path): @@ -626,7 +809,9 @@ def test_close_releases_the_lease_even_when_the_drain_fails(monkeypatch, tmp_pat engine.close() - assert events[-1] == ("service", "stop") + # Whatever did not drain stays in the directory for the next process on + # the node to adopt; the lock goes either way. + assert events[-2:] == [("service", "stop"), ("lock", "release")] def test_replacing_a_record_ring_drains_and_stops_the_service( @@ -643,16 +828,18 @@ def test_replacing_a_record_ring_drains_and_stops_the_service( assert [event[:2] for event in events] == [ ("sink", "flush"), ("ring", "stop"), - ("service", "flush"), ("service", "stop"), ("ring", "create")] + ("service", "flush"), ("service", "stop"), ("lock", "release"), + ("ring", "create")] assert 59.0 <= events[0][2] <= 60.0 assert engine._capture_storage is None assert engine._record_mode is False - # A second record runtime starts its own service; nothing still holds - # the lease it takes. + # A second record runtime starts its own service in a directory of its + # own; nothing still holds the lease it takes, or the lock. engine.create_record_runtime(_record_format()) assert len(services) == 2 assert engine._capture_storage is not None + assert [lock.held for lock in engine._test_locks] == [False, True] # --- the publisher lease knobs ------------------------------------------------- From 244d33e42210b9c84becc086d0314dcdec2638a3 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 03:02:55 -0400 Subject: [PATCH 04/43] Refuse held_by_caller unless this process holds the spool's owner lock Spool::Open with owner_lock=held_by_caller only probed that SOMEONE held /.owner.lock. That passes in exactly the dangerous case: another live process owns the directory. So any process could open a live writer's spool held_by_caller and Recover it, deleting the writer's in-flight .open file -- the "cannot link ready file" race B6 exists to make impossible. Reproduced with tests/native/live_spool_stage.cpp paused mid-stage (take): conformance_spool recover with take was refused naming the writer, the same op with held_by_caller returned ok, the .open file was gone and the writer exited 1 with "cannot link ready file". held_by_caller now also requires that one of THIS process's descriptors holds the lock (SpoolOwnedByThisProcess): /proc/self/fdinfo lists the flocks each open file description holds, so the kernel answers, not the owner record, which is written after the lock is taken and names a pid that means nothing across pid namespaces. The record (host, pid) is the fallback only where /proc cannot be read. Beside another process's lock the open is refused with kOwned (SpoolOwnedError from the store module), naming the holder; with nothing holding the lock it stays kBadArgument. The tests had encoded the race as intended behaviour, and change with it: * test_native_spool_owner_lock.py: held_by_caller beside a conformance_sink process's lock is now refused, naming its pid and host (it asserted the recover succeeded); * test_native_live_spool.py: the finding's repro -- held_by_caller beside the paused writer is refused, its .open file survives, and the writer's stage completes (was: ok, file gone, writer rc=1); * test_native_spool_ownership.py: the service and the sink opened held_by_caller beside another process's SpoolOwnerLock are refused naming it (both opened); * test_spool_owner_lock.cpp: the same across a fork, in process (kOk). Each was red against the previous spool. The storage live harness ran its conformance_sink and conformance_store drivers held_by_caller beside the test process's lock, which is what this refuses. The drivers now take, as the plan has standalone callers do: the sink driver stages into a scratch directory it owns and _stage renames the sealed packs into the spool (one rename per ready file, so a scanning service meets whole files), and the store driver uploads from a spool no service of the test has opened yet, which is already true at all four call sites. Services keep holding the harness's lock held_by_caller, in process. --- native/csrc/store/spool.cpp | 66 ++++++++++++++++++++++- native/csrc/store/spool.h | 16 ++++-- tests/native/test_spool_owner_lock.cpp | 46 +++++++++++++++- tests/test_native_capture_storage_live.py | 55 +++++++++++-------- tests/test_native_live_spool.py | 28 +++++++++- tests/test_native_spool_owner_lock.py | 19 +++++-- tests/test_native_spool_ownership.py | 43 +++++++++++++++ 7 files changed, 239 insertions(+), 34 deletions(-) diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index ecf2757ab..6bd82c3d2 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -8,11 +8,13 @@ #include #include #include +#include #include #include #include #include +#include #include #include #include @@ -465,6 +467,54 @@ bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner) { return held; } +bool SpoolOwnedByThisProcess(const std::string& dir) { + const std::string file = dir + "/" + kOwnerLockFile; + struct stat target{}; + if (::stat(file.c_str(), &target) != 0) return false; + // /proc/self/fdinfo/ lists the flocks each open file description + // holds ("lock: 1: FLOCK ADVISORY WRITE ..."), so the kernel says + // whether one of this process's descriptors on the file holds the lock -- + // a SpoolOwnerLock's, or one a Spool took with kTake. The record in the + // file is only a fallback: it is written after the lock is taken, and a + // pid says nothing across pid namespaces. + DIR* fds = ::opendir("/proc/self/fd"); + if (fds == nullptr) { + SpoolOwner owner; + return ReadSpoolOwner(dir, &owner) && owner.pid == ::getpid() && + owner.host == Hostname(); + } + const int listing = ::dirfd(fds); + bool held = false; + while (!held) { + const dirent* entry = ::readdir(fds); + if (entry == nullptr) break; + char* end = nullptr; + const long fd = std::strtol(entry->d_name, &end, 10); + if (end == entry->d_name || *end != '\0' || fd == listing) continue; + struct stat by_fd{}; + if (::fstat(static_cast(fd), &by_fd) != 0 || + by_fd.st_dev != target.st_dev || by_fd.st_ino != target.st_ino) { + continue; + } + const std::string info = + std::string("/proc/self/fdinfo/") + entry->d_name; + std::FILE* in = std::fopen(info.c_str(), "re"); + if (in == nullptr) continue; + char line[512]; + while (std::fgets(line, sizeof(line), in) != nullptr) { + if (std::strncmp(line, "lock:", 5) == 0 && + std::strstr(line, " FLOCK ") != nullptr && + std::strstr(line, " WRITE ") != nullptr) { + held = true; + break; + } + } + std::fclose(in); + } + ::closedir(fds); + return held; +} + namespace { // Locks the lock file of an existing directory, creating the file if it has @@ -795,8 +845,9 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { } else { // The caller took the lock, so the directory and its lock file exist. char held[4096]; + SpoolOwner owner; if (::realpath(config.root.c_str(), held) == nullptr || - !ReadSpoolOwner(held, nullptr)) { + !ReadSpoolOwner(held, &owner)) { if (error) { *error = "spool owner_lock=held_by_caller, but nothing holds " + config.root + "/" + kOwnerLockFile + @@ -805,6 +856,19 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { } return SpoolStatus::kBadArgument; } + // Held, but by THIS process? "Someone holds it" passes exactly when + // another live process owns the directory, and this Spool's Recover + // would then delete that owner's in-flight .open files. + if (!SpoolOwnedByThisProcess(held)) { + if (error) { + *error = "spool owner_lock=held_by_caller, but this process does " + "not hold the owner lock of " + std::string(held) + ": " + + OwnedMessage(held, owner) + ". held_by_caller is for a " + "second Spool in the process that holds the directory's " + "SpoolOwnerLock"; + } + return SpoolStatus::kOwned; + } const SpoolStatus local = CheckNodeLocal(held, config.allow_shared_filesystem, error); if (local != SpoolStatus::kOk) return local; diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 317d62fe2..7ffa6dd35 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -20,13 +20,16 @@ // it. The lock goes with its holder, even one killed with SIGKILL. // - owner_lock=kTake (the default) takes it in Open(), before anything // reads the directory, and holds it for the Spool object's life. -// - owner_lock=kHeldByCaller takes none: the caller holds a +// - owner_lock=kHeldByCaller takes none: the calling process holds a // SpoolOwnerLock on the directory already. Two Spools in ONE process // that both take refuse each other (flock binds to an open file // description, not to the process), so a process running a sink and a // storage service on one directory holds one SpoolOwnerLock and opens -// both Spools with kHeldByCaller. Open() refuses it when nothing holds -// the lock; it cannot tell who does. +// both Spools with kHeldByCaller. Open() refuses it unless one of THIS +// process's descriptors holds the lock (SpoolOwnedByThisProcess): kOwned, +// naming the holder, beside another process's lock, and kBadArgument +// when nothing holds it. Standalone callers -- the drivers, adoption -- +// take. // A directory nested under, or containing, an owned directory (one with a // .owner.lock file, held or not) is refused: Scan walks recursively, so the // outer spool's Recover would reach into the inner one. /_refs/ is @@ -135,6 +138,11 @@ struct SpoolOwner { // locked but not yet written its record reads as an empty host and pid 0. bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner); +// Whether one of THIS process's descriptors holds /.owner.lock, as the +// kernel reports it in /proc/self/fdinfo (falling back to the recorded host +// and pid where /proc cannot be read). What kHeldByCaller requires. +bool SpoolOwnedByThisProcess(const std::string& dir); + // The owner lock of one spool directory: flock(LOCK_EX) on /.owner.lock, // released with the object (or Release()), and by the kernel when the // process dies. The descriptor is close-on-exec; a child forked WITHOUT exec @@ -214,7 +222,7 @@ class Spool { public: // Opens (creating) the root, after the node-local check and, with kTake, // after taking its owner lock (kOwned when another holder has it); - // kHeldByCaller is refused when nothing holds it. Recovery of + // kHeldByCaller is refused unless this process holds it. Recovery of // pre-existing files is explicit via Recover(), matching the Python // constructor + recover() split. static SpoolStatus Open(SpoolConfig config, Spool* out, std::string* error); diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 320b77ebe..86ea7675d 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -4,8 +4,9 @@ // even in one process: flock binds to an open file description, not to // the process. That is why the engine holds one SpoolOwnerLock and both // of its Spools (sink and service) open with held_by_caller. -// 2. held_by_caller opens beside a holder, and is refused when nothing -// holds the lock. +// 2. held_by_caller opens beside a holder in this process, and is +// refused beside another process's holder or when nothing holds the +// lock. // 3. A second process is refused, told the holder's pid and host; the // lock goes with its holder, even one killed with SIGKILL. // 4. Nesting: a directory under, or containing, an owned one is refused. @@ -131,6 +132,46 @@ void TestHeldByCallerOpensBesideTheHolder() { CHECK(Spool::Open({root, 1 << 20}, &rival, &error) == SpoolStatus::kOwned); } +// (2b) held_by_caller is for a Spool in the process that holds the lock. +// Beside ANOTHER process's lock it is refused, naming that holder: were +// "something holds it" enough, any process could open a live writer's +// directory that way, and its Recover would delete the writer's .open file. +void TestHeldByCallerBesideAnotherProcessIsRefused() { + const std::string root = FreshRoot("held-elsewhere") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t child = ::fork(); + if (child == 0) { + ::close(ready[0]); + SpoolOwnerLock lock; + std::string error; + const bool ok = SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk; + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + + SpoolConfig config{root, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool spool; + std::string error; + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kOwned); + CHECK(Contains(error, "held_by_caller")); + CHECK(Contains(error, "pid " + std::to_string(child))); + CHECK(Contains(error, Hostname())); + + ::kill(child, SIGKILL); + int status = 0; + ::waitpid(child, &status, 0); +} + void TestHeldByCallerWithoutAHolderIsRefused() { const std::string root = FreshRoot("unheld") + "/spool"; SpoolConfig config{root, 1 << 20}; @@ -373,6 +414,7 @@ void TestTheDirectoryLayout() { int main() { TestTwoTakesInOneProcessRefuseEachOther(); TestHeldByCallerOpensBesideTheHolder(); + TestHeldByCallerBesideAnotherProcessIsRefused(); TestHeldByCallerWithoutAHolderIsRefused(); TestTheLockGoesWithItsSpool(); TestASecondProcessIsRefusedUntilTheHolderDies(); diff --git a/tests/test_native_capture_storage_live.py b/tests/test_native_capture_storage_live.py index 068241a3d..ad64a7983 100644 --- a/tests/test_native_capture_storage_live.py +++ b/tests/test_native_capture_storage_live.py @@ -23,6 +23,7 @@ import base64 import json import re +import shutil import socket import subprocess import threading @@ -113,14 +114,20 @@ def _storage_config(endpoint, prefix, **overrides): return NativeCaptureStorageConfig(**fields) -# Every spool directory a test points a service or a driver at is owned by -# the harness for the rest of the test: one SpoolOwnerLock held here, and -# each service, sink driver and store driver opens the directory with -# owner_lock="held_by_caller". That is the engine's arrangement -- it holds -# the lock around its sink and its service -- stretched over the processes a -# test uses: the drivers stage and upload from processes of their own while -# a service is up, and a test often builds a second service on a directory -# the first still has open. Each taking the lock would refuse the others. +# Every spool directory a test points a service at is owned by the harness +# for the rest of the test: one SpoolOwnerLock held here, and each service +# opens the directory with owner_lock="held_by_caller". That is the engine's +# arrangement -- it holds the lock around its sink and its service -- and a +# test often builds a second service on a directory the first still has +# open; each taking the lock would refuse the others. +# +# The drivers are processes of their own, so they cannot open a directory +# this process owns: held_by_caller is refused unless the opening process +# holds the lock. They take it, as the plan has standalone callers do. The +# sink driver stages into a scratch directory it owns, and _stage moves its +# sealed packs into the spool, where a sink in this process would have +# staged them; the store driver uploads from a spool before any service of +# the test has opened it. _HARNESS_LOCKS: dict = {} @@ -172,16 +179,22 @@ def _record(index: int): def _stage(spool_root: Path, indexes, *, records_per_pack: int = 2): - """Stage records through the native sink core; return the tensors.""" + """Stage records through the native sink core; return the tensors. + + The driver takes a scratch directory of its own beside spool_root, and + once its sink is closed the sealed packs move into spool_root under the + same relative paths -- one rename each, so a service scanning the spool + meets a whole ready file or none.""" + scratch = spool_root.parent / f".stage-{uuid.uuid4().hex[:8]}" sink = _Driver(SINK_DRIVER) tensors = {} try: assert sink.call( - op="open", root=str(spool_root), max_bytes=1 << 40, + op="open", root=str(scratch), max_bytes=1 << 40, max_queue_records=256, max_queue_bytes=1 << 24, max_pack_bytes=8 << 20, max_pack_records=records_per_pack, max_linger_ns=1_000_000_000, overload="drop_newest", - admission_timeout=-1, owner_lock=_held(spool_root))["ok"] + admission_timeout=-1)["ok"] for index in indexes: metadata, tensor = _record(index) response = sink.call( @@ -195,6 +208,11 @@ def _stage(spool_root: Path, indexes, *, records_per_pack: int = 2): assert snapshot["persisted_records"] == len(tensors), snapshot finally: sink.close() + for ready in sorted(scratch.rglob("*.dmi-pack.ready")): + target = spool_root / ready.relative_to(scratch) + target.parent.mkdir(parents=True, exist_ok=True) + ready.replace(target) + shutil.rmtree(scratch) return tensors @@ -618,8 +636,7 @@ def test_a_pack_uploaded_but_never_indexed_is_reconciled_at_start( region=REGION, access=ACCESS, secret=SECRET, token=None, insecure=True, connect_timeout=5, read_timeout=15, max_attempts=4, store_id="s3", root=str(spool_root), spool_max_bytes=1 << 40, - limit=-1, max_workers=4, max_in_flight_bytes=1 << 30, - owner_lock=_held(spool_root)) + limit=-1, max_workers=4, max_in_flight_bytes=1 << 30) assert uploaded["ok"], uploaded # And a foreign object where packs live, which must not be indexed. foreign = store.call( @@ -1477,8 +1494,7 @@ def test_a_pass_whose_lease_changed_while_it_read_rereads_the_replay_guard( region=REGION, access=ACCESS, secret=SECRET, token=None, insecure=True, connect_timeout=5, read_timeout=15, max_attempts=4, store_id="s3", root=str(spool_root), spool_max_bytes=1 << 40, - limit=-1, max_workers=4, max_in_flight_bytes=1 << 30, - owner_lock=_held(spool_root)) + limit=-1, max_workers=4, max_in_flight_bytes=1 << 30) assert uploaded["ok"], uploaded finally: store.close() @@ -1548,8 +1564,7 @@ def test_a_conflicted_publish_reports_the_conflict_unless_the_lease_was_lost( region=REGION, access=ACCESS, secret=SECRET, token=None, insecure=True, connect_timeout=5, read_timeout=15, max_attempts=4, store_id="s3", root=str(spool_root), spool_max_bytes=1 << 40, - limit=-1, max_workers=4, max_in_flight_bytes=1 << 30, - owner_lock=_held(spool_root)) + limit=-1, max_workers=4, max_in_flight_bytes=1 << 30) assert uploaded["ok"], uploaded finally: store.close() @@ -1730,8 +1745,7 @@ def _upload_behind_the_service(spool_root: Path, endpoint: str) -> list: region=REGION, access=ACCESS, secret=SECRET, token=None, insecure=True, connect_timeout=5, read_timeout=15, max_attempts=4, store_id="s3", root=str(spool_root), spool_max_bytes=1 << 40, - limit=-1, max_workers=4, max_in_flight_bytes=1 << 30, - owner_lock=_held(spool_root)) + limit=-1, max_workers=4, max_in_flight_bytes=1 << 30) assert uploaded["ok"], uploaded finally: store.close() @@ -2411,8 +2425,7 @@ def test_a_64_mib_pack_over_https_with_a_private_ca_hydrates_exactly( max_queue_records=records, max_queue_bytes=2 * records * record_bytes, max_pack_bytes=2 * records * record_bytes, max_pack_records=records, max_linger_ns=60_000_000_000, - overload="drop_newest", admission_timeout=-1, - owner_lock=_held(spool_root))["ok"] + overload="drop_newest", admission_timeout=-1)["ok"] for index in range(records): metadata = CaptureMetadata( capture_id=f"tls-{index:04d}", tenant_id="t", diff --git a/tests/test_native_live_spool.py b/tests/test_native_live_spool.py index e0a5669ec..bf05cc3ec 100644 --- a/tests/test_native_live_spool.py +++ b/tests/test_native_live_spool.py @@ -72,13 +72,13 @@ def _start_writer(binary: Path, spool: Path, packs: int = 1): f"writer never opened its temp file: {seen} {process.stderr.read()}") -def _recover(spool: Path) -> dict: +def _recover(spool: Path, **fields) -> dict: import json proc = subprocess.run( [str(SPOOL_DRIVER)], input=json.dumps({"op": "recover", "root": str(spool), - "max_bytes": 1 << 30}) + "\n", + "max_bytes": 1 << 30, **fields}) + "\n", capture_output=True, text=True, timeout=30) return json.loads(proc.stdout.strip()) @@ -113,6 +113,30 @@ def test_an_upload_scan_is_refused_while_a_writer_holds_the_spool( process.communicate() +def test_held_by_caller_cannot_sweep_a_live_writers_spool(tmp_path, writer_binary): + """owner_lock=held_by_caller opens without a lock of its own, for a + second Spool in the process that holds the directory. From any other + process it is refused, naming the writer: it used to pass on "something + holds the lock", and its Recover then deleted the writer's .open file, + failing the writer's stage with "cannot link ready file".""" + spool = tmp_path / "spool" + process = _start_writer(writer_binary, spool) + try: + result = _recover(spool, owner_lock="held_by_caller") + assert not result["ok"], result + assert "held_by_caller" in result["what"], result + assert f"pid {process.pid}" in result["what"], result + assert len(list(spool.rglob("*.open"))) == 1 + + output, error = process.communicate("continue\n", timeout=10) + assert process.returncode == 0, output + error + assert len(list(spool.rglob("*.dmi-pack.ready"))) == 1 + finally: + if process.poll() is None: + process.kill() + process.communicate() + + def test_a_sigkilled_writers_spool_is_recovered_by_the_next_process( fake_s3, tmp_path, writer_binary): spool = tmp_path / "spool" diff --git a/tests/test_native_spool_owner_lock.py b/tests/test_native_spool_owner_lock.py index 5732e2854..60c1520ef 100644 --- a/tests/test_native_spool_owner_lock.py +++ b/tests/test_native_spool_owner_lock.py @@ -9,8 +9,11 @@ * a directory nested under, or containing, an owned directory: Scan walks recursively, so the outer spool's cleanup would reach into the inner one; -* ``owner_lock="held_by_caller"`` with nothing holding the lock: that mode - opens without a lock of its own, on the caller's word that one is held; +* ``owner_lock="held_by_caller"`` unless THIS process holds the lock: that + mode opens without a lock of its own, for a second Spool in the process + that holds one. Beside another process's lock it is refused, naming the + holder -- else any process could open a live writer's directory that way + and sweep its in-flight ``.open`` files; * ``/_refs/``, where the upload handoff's ref files will live: no scan may sweep, quarantine or list anything under it. @@ -165,14 +168,22 @@ def test_a_stale_lock_file_still_marks_an_owned_directory(tmp_path): assert "nested" in response["what"], response -def test_held_by_caller_opens_beside_the_holder(tmp_path): +def test_held_by_caller_is_refused_beside_another_processs_holder(tmp_path): + """A process that does not hold the lock cannot open held_by_caller: + the holder is another process, so the open is refused, naming it, and + its Recover never runs.""" root = tmp_path / "spool" holder = _Holder(root) try: assert holder.opened["ok"], holder.opened response = _spool(op="recover", root=str(root), owner_lock="held_by_caller") - assert response["ok"], response + assert not response["ok"], response + assert response["status"] == "open", response + what = response["what"] + assert "held_by_caller" in what, what + assert f"pid {holder.pid}" in what, what + assert socket.gethostname() in what, what finally: holder.close() diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index fa3962f61..8cd617f65 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -254,6 +254,49 @@ def test_held_by_caller_with_nothing_held_is_refused(tmp_path): spool_owner_lock="held_by_caller") +class _OtherProcessHolder: + """Another process holding a directory's SpoolOwnerLock until closed.""" + + def __init__(self, directory): + script = ( + "import sys; sys.path.insert(0, sys.argv[1]);" + "import _dmi_native_store as m\n" + "lock = m.SpoolOwnerLock(sys.argv[2])\n" + "print('held', flush=True)\n" + "sys.stdin.read()\n") + self.proc = subprocess.Popen( + [sys.executable, "-c", script, str(BUILD), str(directory)], + stdin=subprocess.PIPE, stdout=subprocess.PIPE, text=True) + assert self.proc.stdout.readline().strip() == "held" + + @property + def pid(self) -> int: + return self.proc.pid + + def close(self) -> None: + self.proc.stdin.close() + self.proc.wait(timeout=30) + + +def test_held_by_caller_beside_another_processs_lock_is_refused(tmp_path): + """held_by_caller is for the process that holds the lock. The service + and the sink opened that way beside ANOTHER process's lock are refused, + naming it, rather than sweeping and uploading from its directory.""" + directory = tmp_path / "spool" + holder = _OtherProcessHolder(directory) + try: + with pytest.raises(RuntimeError, + match=f"held_by_caller.*pid {holder.pid}"): + _service(_config(), directory, spool_owner_lock="held_by_caller") + if SINK_BUILT: + pytest.importorskip("torch") + with pytest.raises(RuntimeError, + match=f"held_by_caller.*pid {holder.pid}"): + _sink(directory, owner_lock="held_by_caller") + finally: + holder.close() + + def test_a_drained_directory_is_removed_with_its_lock(tmp_path): store = _store() directory = tmp_path / "spool" From 0d1a42952f37fd6658a155eb78ae0faaa51f420b Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 03:07:12 -0400 Subject: [PATCH 05/43] Check a spool's nesting after taking its lock, and only against held locks above Two defects in SpoolOwnerLock::Acquire's nesting check, one ordering and one rule. Ordering: CheckNotNested ran before the lock was taken, so an outer directory and one nested in it, taken at the same moment by two processes, each checked before the other's lock file existed, and both won. Two owners of overlapping trees is what the check exists to prevent: the outer one's recursive Scan, UploadPending or Recover reaches the inner live directory, sweeping its .open files and uploading its packs under the wrong keys. The engine really produces the pair -- a sink-only or explicit record_sink run takes spool_root, a default persistent run claims spool_root//r0-. Acquire now takes the lock (LockInPlace or CreateLocked, which renames the new directory into place first) and checks after, so each side publishes before it looks and at least one sees the other. A refused take lets go of its lock and leaves nothing behind: the directory it created, while still empty, is removed, and a lock file it added to an existing directory is unlinked. Rule: an ANCESTOR with a .owner.lock file refused, held or not. A take never deletes its lock file, so after one sink-only run on spool_root, or one explicit-record_sink run whose service took it (the documented rollback), every later default run was refused as "nested under the owned spool directory " until someone deleted the file by hand. An ancestor now refuses only while its lock is held. Safety does not need more: the next take of that ancestor walks its descendants, meets this directory's lock file, and is refused. Descendants still refuse on the lock file alone, held or not -- a dead directory's packs are its successor's to adopt, not an outer spool's to sweep. Tests, each red before: * test_spool_owner_lock.cpp: 200 trials of two forked processes released at once on an outer directory and a rank directory in it -- both won in 195 of 200 before, never now; a stale lock file above no longer refuses a nested take, and the lock file it leaves still refuses the next take of the outer directory; a refused take leaves no directory or lock file; * test_native_spool_owner_lock.py: the stale-lock test asserted the nested refusal, and now asserts the admission and the outer refusal; * test_native_spool_ownership.py: claim_spool_directory under a spool_root that a SpoolOwnerLock once held succeeds (it raised ValueError "nested under the owned spool directory"). --- native/csrc/store/spool.cpp | 52 +++++++++---- native/csrc/store/spool.h | 14 +++- tests/native/test_spool_owner_lock.cpp | 103 ++++++++++++++++++++++++- tests/test_native_spool_owner_lock.py | 22 ++++-- tests/test_native_spool_ownership.py | 21 +++++ 5 files changed, 186 insertions(+), 26 deletions(-) diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 6bd82c3d2..3fad4c6c8 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -355,26 +355,38 @@ std::string CanonicalPath(const std::string& path, std::string* error) { } // A spool directory must not be nested under, or contain, another owned -// directory -- one with a lock file, held or not: Scan walks recursively, -// so the outer spool's Recover would sweep the inner one's .open files and -// upload its packs under the outer spool's keys. -SpoolStatus CheckNotNested(const std::string& dir, bool exists, - std::string* error) { +// directory: Scan walks recursively, so the outer spool's Recover would +// sweep the inner one's .open files and upload its packs under the outer +// spool's keys. Run AFTER `dir`'s own lock is taken, so that of two takes +// racing on an outer directory and one inside it, at least one sees the +// other: each publishes its lock before it looks. +// - An ancestor refuses while its lock is HELD. A lock file nobody holds +// is a spool that was (every take leaves its file behind); whoever +// takes that ancestor next meets this directory's lock file in its own +// descendant walk, and is refused. +// - A descendant refuses when it has a lock file at all, held or not: a +// dead directory's packs are its successor's to adopt, not this +// spool's to sweep and upload under its own keys. +SpoolStatus CheckNotNested(const std::string& dir, std::string* error) { fs::path ancestor(dir); while (ancestor.has_parent_path() && ancestor.parent_path() != ancestor) { ancestor = ancestor.parent_path(); std::error_code ec; - if (fs::exists(ancestor / kOwnerLockFile, ec)) { + SpoolOwner owner; + if (fs::exists(ancestor / kOwnerLockFile, ec) && + ReadSpoolOwner(ancestor.string(), &owner)) { if (error) { - *error = "spool directory " + dir + " is nested under the owned " - "spool directory " + ancestor.string() + " (it has " + - kOwnerLockFile + "), whose recovery would sweep this one; " - "use a directory outside it"; + *error = "spool directory " + dir + " is nested under the spool " + "directory " + ancestor.string() + ", which " + + (owner.pid > 0 ? "pid " + std::to_string(owner.pid) + + " on host " + owner.host + : std::string("another owner")) + + " holds (" + kOwnerLockFile + "), and whose recovery " + "would sweep this one; use a directory outside it"; } return SpoolStatus::kBadArgument; } } - if (!exists) return SpoolStatus::kOk; std::error_code ec; for (auto it = fs::recursive_directory_iterator( dir, fs::directory_options::skip_permission_denied, ec); @@ -681,14 +693,28 @@ SpoolStatus SpoolOwnerLock::Acquire(const std::string& dir, CheckNodeLocal(exists ? canonical : parent, allow_shared_filesystem, error); if (status != SpoolStatus::kOk) return status; - status = CheckNotNested(canonical, exists, error); - if (status != SpoolStatus::kOk) return status; + const std::string lock_file = canonical + "/" + kOwnerLockFile; + const bool had_lock_file = exists && fs::exists(lock_file, ec); int fd = -1; status = exists ? LockInPlace(canonical, &fd, error) : CreateLocked(canonical, &fd, error); if (status != SpoolStatus::kOk) return status; out->fd_ = fd; out->dir_ = canonical; + // Only now, with this lock published: see CheckNotNested. + status = CheckNotNested(canonical, error); + if (status != SpoolStatus::kOk) { + // Leave nothing of this take behind: the directory it created (while + // it is still empty), or the lock file it added to one that existed. + if (!exists) { + std::string ignored; + out->ReleaseAndRemoveIfEmpty(&ignored); + } else { + if (!had_lock_file) ::unlink(lock_file.c_str()); + out->Release(); + } + return status; + } return SpoolStatus::kOk; } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 7ffa6dd35..79734ca7d 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -30,9 +30,14 @@ // naming the holder, beside another process's lock, and kBadArgument // when nothing holds it. Standalone callers -- the drivers, adoption -- // take. -// A directory nested under, or containing, an owned directory (one with a -// .owner.lock file, held or not) is refused: Scan walks recursively, so the -// outer spool's Recover would reach into the inner one. /_refs/ is +// A directory nested under a HELD spool directory, or containing one with +// a .owner.lock file (held or not: a dead directory's packs are for its +// successor to adopt), is refused: Scan walks recursively, so the outer +// spool's Recover would reach into the inner one. An unheld lock file +// ABOVE refuses nothing -- every take leaves its file behind -- since the +// next take of that directory meets this one's lock file below it. The +// check runs after the lock is taken, so of two processes taking an outer +// and a nested directory at once, at least one is refused. /_refs/ is // never scanned: the upload handoff's ref files live there (plan section // 2.4). The spool must be node-local: NFS and Lustre are refused by statfs // f_type unless allow_shared_filesystem is set, since neither guarantees a @@ -163,7 +168,8 @@ class SpoolOwnerLock { // parent ever meets the directory before its owner holds it -- and records // this host and pid in it. kOwned, naming the holder, when another holder // has it; kBadArgument for a shared filesystem or a directory nested - // under, or containing, an owned one. + // under, or containing, an owned one (see above), after letting go of + // the lock and of whatever this call created. static SpoolStatus Acquire(const std::string& dir, bool allow_shared_filesystem, SpoolOwnerLock* out, std::string* error); diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 86ea7675d..2319e4858 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -9,7 +9,8 @@ // lock. // 3. A second process is refused, told the holder's pid and host; the // lock goes with its holder, even one killed with SIGKILL. -// 4. Nesting: a directory under, or containing, an owned one is refused. +// 4. Nesting: a directory under a HELD one, or containing one with a lock +// file, is refused -- also when two processes take the pair at once. // 5. The node-local check refuses NFS and Lustre by statfs f_type, unless // explicitly allowed (a test seam stands in for statfs). // 6. Adoption's try-lock never creates a directory, and a released @@ -276,6 +277,104 @@ void TestNestedDirectoriesAreRefused() { Spool spool; CHECK(Spool::Open({base + "/other", 1 << 20}, &spool, &error) == SpoolStatus::kBadArgument); + // A refused take leaves nothing of its own behind: not the directory it + // created, nor a lock file it added to one that existed. + CHECK(!fs::exists(base + "/outer/inner")); + CHECK(!fs::exists(base + "/other/.owner.lock")); +} + +// (4b) An outer directory and one nested in it, taken at the same moment by +// two processes: each checks the other's lock only after publishing its +// own, so at most one of them wins. Checked first and locked second, both +// won most of the time (275 of 300 in the review's probe). +void TestAnOuterAndANestedTakeRacingNeverBothWin() { + const std::string base = FreshRoot("nest-race"); + int both = 0; + int outer_won = 0; + int inner_won = 0; + for (int trial = 0; trial < 200; ++trial) { + const std::string outer = base + "/t" + std::to_string(trial); + const std::string inner = outer + "/0123456789ab/r0-0a1b2c3d"; + if (trial % 2 == 0) fs::create_directories(outer); + int go[2], report[2], done[2]; + CHECK(::pipe(go) == 0 && ::pipe(report) == 0 && ::pipe(done) == 0); + pid_t children[2]; + for (int side = 0; side < 2; ++side) { + children[side] = ::fork(); + if (children[side] == 0) { + ::close(go[1]); + ::close(report[0]); + ::close(done[1]); + char byte = 0; + (void)!::read(go[0], &byte, 1); // EOF: the parent let both go + SpoolOwnerLock lock; + std::string error; + const bool won = + SpoolOwnerLock::Acquire(side == 0 ? outer : inner, false, &lock, + &error) == SpoolStatus::kOk; + byte = static_cast(side == 0 ? (won ? 'O' : 'o') + : (won ? 'I' : 'i')); + if (::write(report[1], &byte, 1) != 1) ::_exit(3); + (void)!::read(done[0], &byte, 1); // hold it until both reported + ::_exit(0); + } + } + ::close(go[0]); + ::close(report[1]); + ::close(done[0]); + ::close(go[1]); + char results[2] = {0, 0}; + CHECK(::read(report[0], &results[0], 1) == 1); + CHECK(::read(report[0], &results[1], 1) == 1); + const std::string seen(results, 2); + const bool o = seen.find('O') != std::string::npos; + const bool i = seen.find('I') != std::string::npos; + if (o && i) ++both; + if (o) ++outer_won; + if (i) ++inner_won; + ::close(done[1]); + ::close(report[0]); + for (const pid_t child : children) { + int status = 0; + ::waitpid(child, &status, 0); + } + } + if (both != 0) { + std::cerr << "outer and nested both acquired in " << both + << " of 200 trials\n"; + } + CHECK(both == 0); + CHECK(outer_won + inner_won > 0); +} + +// (4c) An ancestor refuses only while its lock is HELD. A lock file nobody +// holds is a spool that was -- every take leaves its file behind -- and +// refusing on it kept a spool_root that a sink-only run once owned from +// ever holding rank directories. Whoever takes the outer directory next +// meets the inner one's lock file below it and is refused, held or not. +void TestAStaleLockFileAboveDoesNotRefuseANestedDirectory() { + const std::string base = FreshRoot("stale-above"); + std::string error; + { + SpoolOwnerLock once; + CHECK(SpoolOwnerLock::Acquire(base + "/root", false, &once, &error) == + SpoolStatus::kOk); + } + CHECK(fs::exists(base + "/root/.owner.lock")); + SpoolOwnerLock inner; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(base + "/root/0123456789ab/r0-0a1b2c3d", + false, &inner, &error) == SpoolStatus::kOk); + CHECK(error.empty()); + SpoolOwnerLock outer; + CHECK(SpoolOwnerLock::Acquire(base + "/root", false, &outer, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "contains")); + inner.Release(); // its lock file stays: still refused + error.clear(); + CHECK(SpoolOwnerLock::Acquire(base + "/root", false, &outer, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "contains")); } // (5) The node-local check, through the test seam. @@ -419,6 +518,8 @@ int main() { TestTheLockGoesWithItsSpool(); TestASecondProcessIsRefusedUntilTheHolderDies(); TestNestedDirectoriesAreRefused(); + TestAnOuterAndANestedTakeRacingNeverBothWin(); + TestAStaleLockFileAboveDoesNotRefuseANestedDirectory(); TestSharedFilesystemsAreRefusedUnlessAllowed(); TestAdoptionLocksOnlyWhatExistsAndIsDead(); TestANewDirectoryAppearsWithItsLockHeld(); diff --git a/tests/test_native_spool_owner_lock.py b/tests/test_native_spool_owner_lock.py index 60c1520ef..6de5ffbe8 100644 --- a/tests/test_native_spool_owner_lock.py +++ b/tests/test_native_spool_owner_lock.py @@ -7,8 +7,9 @@ second process that tries is refused, told who holds it. What that lock must also refuse, and what it must leave alone: -* a directory nested under, or containing, an owned directory: Scan walks - recursively, so the outer spool's cleanup would reach into the inner one; +* a directory nested under a held one, or containing one with a lock file: + Scan walks recursively, so the outer spool's cleanup would reach into the + inner one; * ``owner_lock="held_by_caller"`` unless THIS process holds the lock: that mode opens without a lock of its own, for a second Spool in the process that holds one. Beside another process's lock it is refused, naming the @@ -156,16 +157,21 @@ def test_a_directory_containing_an_owned_one_is_refused(tmp_path): holder.close() -def test_a_stale_lock_file_still_marks_an_owned_directory(tmp_path): - """Nesting is judged by the lock FILE, not by a live holder: an outer - directory some spool once owned is still a spool directory, and the - next process to open it would sweep the inner one.""" +def test_a_stale_lock_file_above_does_not_refuse_a_nested_directory(tmp_path): + """Above, only a HELD lock refuses: every take leaves its lock file + behind, so a spool_root some spool once owned would otherwise refuse + every directory under it for good. Below, the lock FILE refuses, held + or not: a dead directory's packs are for its successor to adopt, not + for the outer spool to sweep and upload under its own keys -- so the + next process to take the outer directory is refused.""" outer = tmp_path / "spool" assert _spool(op="recover", root=str(outer))["ok"] assert (outer / ".owner.lock").exists() - response = _spool(op="recover", root=str(outer / "inner")) + assert _spool(op="recover", root=str(outer / "inner"))["ok"] + assert (outer / "inner" / ".owner.lock").exists() + response = _spool(op="recover", root=str(outer)) assert not response["ok"], response - assert "nested" in response["what"], response + assert "contains" in response["what"], response def test_held_by_caller_is_refused_beside_another_processs_holder(tmp_path): diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 8cd617f65..232a18aa8 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -192,6 +192,27 @@ def test_a_nested_spool_directory_is_refused(tmp_path): _service(_config(), tmp_path / "outer" / "inner") +def test_a_spool_root_a_sink_once_owned_still_takes_rank_directories(tmp_path): + """A sink-only run (or the explicit-record_sink rollback) takes + spool_root itself and leaves its lock file there. The next default run + claims a rank directory under it, which that unheld lock file must not + refuse as nested.""" + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + root = tmp_path / "root" + _store().SpoolOwnerLock(str(root)).release() + assert (root / ".owner.lock").exists() + claim = claim_spool_directory(NativeSinkConfig(spool_root=str(root)), + _config()) + try: + assert claim.held + assert Path(claim.directory).parent.parent == root + finally: + claim.release() + + def test_a_shared_filesystem_is_refused_unless_allowed(tmp_path): store = _store() store._set_spool_filesystem_type_for_testing(NFS_SUPER_MAGIC) From 55ffbd21906d77cdfa943e50e0085533c4657a39 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 03:10:49 -0400 Subject: [PATCH 06/43] Ignore and clear the staging copy a claim killed before its rename leaves SpoolOwnerLock::Acquire creates a new spool directory as a staging copy, /..<8 hex>.creating, with its lock file created and flocked inside, then renames it into place. A claim killed anywhere between the lock file's open and the rename (or a power loss before the parent's fsync) leaves that copy and its lock file behind, and nothing removed it: adoption skipped it, since the name is not a rank directory, and the nesting check's descendant walk counted its lock file. So a later take of or / -- a sink-only run, or the explicit record_sink rollback, both of which take spool_root -- was refused for good as "contains the owned spool directory .../.r0-....creating", by a directory nobody owns that holds no pack. * The nesting check skips a staging copy whose lock nobody holds. One that is held is a claim in progress and still refuses; one not yet flocked publishes after the walk, and its own ancestor check (which runs after its rename) then sees this lock. * Adoption clears the dead ones it meets under the catalog key: a staging copy older than 60 s whose lock it can take is removed with it. A younger one is left alone, since a claim between its mkdir and its flock holds no lock yet and clearing its copy would fail it. Nothing is owed either way. * IsSpoolClaimStagingName names the pattern, and CreateLocked builds it. Tests: test_spool_owner_lock.cpp -- with an unheld staging copy under root/, taking root and then root/ succeed (refused before), and a held one still refuses the outer take; test_native_spool_adoption_live.py -- the successor of a SIGKILLed process removes an hour-old staging copy beside the dead directory and leaves a fresh one in place (the old one survived before). --- native/csrc/catalog/storage_service.cpp | 38 ++++++++++++++++++++-- native/csrc/catalog/storage_service.h | 5 +++ native/csrc/store/spool.cpp | 30 +++++++++++++++-- native/csrc/store/spool.h | 7 ++++ tests/native/test_spool_owner_lock.cpp | 41 ++++++++++++++++++++++++ tests/test_native_spool_adoption_live.py | 17 ++++++++++ 6 files changed, 133 insertions(+), 5 deletions(-) diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index af73f1265..2e11964b7 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -571,6 +571,7 @@ void CaptureStorageService::adopt_siblings() { adoption_owed_ = true; const fs::path own(spool_.root()); std::vector siblings; + std::vector claim_staging; std::error_code ec; for (fs::directory_iterator it(own.parent_path(), ec), end; !ec && it != end; it.increment(ec)) { @@ -578,9 +579,15 @@ void CaptureStorageService::adopt_siblings() { std::string incarnation; std::error_code type_ec; if (it->path() == own || it->is_symlink(type_ec) || - !it->is_directory(type_ec) || - !dmi_store::ParseSpoolRankDirectoryName( - it->path().filename().string(), &rank, &incarnation)) { + !it->is_directory(type_ec)) { + continue; + } + const std::string name = it->path().filename().string(); + if (dmi_store::IsSpoolClaimStagingName(name)) { + claim_staging.push_back(it->path()); + continue; + } + if (!dmi_store::ParseSpoolRankDirectoryName(name, &rank, &incarnation)) { continue; } siblings.push_back(it->path().string()); @@ -590,6 +597,7 @@ void CaptureStorageService::adopt_siblings() { ": " + ec.message()); return; } + clear_dead_claim_staging(claim_staging); std::sort(siblings.begin(), siblings.end()); bool owed = false; for (const std::string& sibling : siblings) { @@ -598,6 +606,30 @@ void CaptureStorageService::adopt_siblings() { adoption_owed_ = owed; } +void CaptureStorageService::clear_dead_claim_staging( + const std::vector& staging) { + namespace fs = std::filesystem; + // A claim builds its directory's staging copy and renames it into place + // within milliseconds, so one this old whose lock nobody holds belongs + // to a claim that died before its rename. Younger ones are left alone: a + // claim between its mkdir and its flock holds no lock yet, and clearing + // its copy would fail it. Nothing is owed either way. + constexpr auto kDeadAfter = std::chrono::seconds(60); + for (const fs::path& path : staging) { + std::error_code ec; + const auto written = fs::last_write_time(path, ec); + if (ec || fs::file_time_type::clock::now() - written < kDeadAfter) { + continue; + } + dmi_store::SpoolOwnerLock lock; + std::string error; + if (dmi_store::SpoolOwnerLock::TryAdopt(path.string(), &lock, &error) == + dmi_store::SpoolStatus::kOk) { + lock.ReleaseAndRemoveIfEmpty(&error); + } + } +} + bool CaptureStorageService::adopt_sibling(const std::string& directory) { dmi_store::SpoolOwnerLock lock; std::string error; diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index 7a29dff3f..b809d2539 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -70,6 +70,7 @@ #include #include #include +#include #include #include #include @@ -292,6 +293,10 @@ class CaptureStorageService { void adopt_siblings(); // Adopts one sibling; false when it is owed another try. bool adopt_sibling(const std::string& directory); + // Removes the staging copies (dmi_store::IsSpoolClaimStagingName) that + // claims killed before their rename left under the catalog key. + void clear_dead_claim_staging( + const std::vector& staging); // Indexes refs that are gone from their spool, keeping whatever does not // index in pending_index_ -- the only record of it in-process. With no // catalog, keeps them all. Requires cycle_mutex_. diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 3fad4c6c8..20043b73a 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -29,6 +29,9 @@ namespace { constexpr const char* kReadySuffix = ".dmi-pack.ready"; constexpr const char* kOpenSuffix = ".open"; constexpr const char* kOwnerLockFile = ".owner.lock"; +// SpoolOwnerLock::Acquire builds a new directory as +// /..<8 hex>.creating and renames it into place. +constexpr const char* kClaimStagingSuffix = ".creating"; // /_refs/: the upload handoff's ref files (plan section 2.4). Not a // legal object-key component (those start with an alphanumeric), so no pack // is ever staged under it. @@ -366,7 +369,11 @@ std::string CanonicalPath(const std::string& path, std::string* error) { // descendant walk, and is refused. // - A descendant refuses when it has a lock file at all, held or not: a // dead directory's packs are its successor's to adopt, not this -// spool's to sweep and upload under its own keys. +// spool's to sweep and upload under its own keys. Except a claim's +// staging copy (IsSpoolClaimStagingName) that nobody holds: a claim +// killed before its rename, which holds nothing but its lock file. (One +// that is held is a claim in progress, and refuses; one not yet locked +// publishes after this walk, so its own ancestor check sees this lock.) SpoolStatus CheckNotNested(const std::string& dir, std::string* error) { fs::path ancestor(dir); while (ancestor.has_parent_path() && ancestor.parent_path() != ancestor) { @@ -394,6 +401,10 @@ SpoolStatus CheckNotNested(const std::string& dir, std::string* error) { if (it->path().filename() != kOwnerLockFile) continue; const fs::path owned = it->path().parent_path(); if (owned == fs::path(dir)) continue; + if (IsSpoolClaimStagingName(owned.filename().string()) && + !ReadSpoolOwner(owned.string(), nullptr)) { + continue; + } if (error) { *error = "spool directory " + dir + " contains the owned spool " "directory " + owned.string() + " (it has " + kOwnerLockFile + @@ -479,6 +490,20 @@ bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner) { return held; } +bool IsSpoolClaimStagingName(const std::string& name) { + // "." + + "." + 8 hex + ".creating", not empty. + const size_t suffix = std::strlen(kClaimStagingSuffix); + if (name.size() < 1 + 1 + 1 + 8 + suffix || name[0] != '.' || + !HasSuffix(name, kClaimStagingSuffix)) { + return false; + } + const size_t dot = name.size() - suffix - 9; + if (name[dot] != '.') return false; + return std::all_of(name.begin() + dot + 1, name.end() - suffix, [](char c) { + return (c >= '0' && c <= '9') || (c >= 'a' && c <= 'f'); + }); +} + bool SpoolOwnedByThisProcess(const std::string& dir) { const std::string file = dir + "/" + kOwnerLockFile; struct stat target{}; @@ -589,8 +614,9 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, char suffix[16]; std::snprintf(suffix, sizeof(suffix), "%08x", static_cast(random())); + // IsSpoolClaimStagingName's pattern. const std::string staging = - parent + "/." + name + "." + suffix + ".creating"; + parent + "/." + name + "." + suffix + kClaimStagingSuffix; if (::mkdir(staging.c_str(), 0755) != 0) { if (errno == EEXIST) continue; if (error) *error = "cannot create " + staging + ": " + Errno(errno); diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 79734ca7d..f41be6597 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -143,6 +143,13 @@ struct SpoolOwner { // locked but not yet written its record reads as an empty host and pid 0. bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner); +// Whether `name` is the staging copy of a directory SpoolOwnerLock::Acquire +// is creating, "..<8 hex>.creating": built beside its target with its +// lock file held, then renamed into place. One nobody holds was left by a +// claim killed before its rename; it holds nothing but its lock file, is +// ignored by the nesting check, and an adopter clears it. +bool IsSpoolClaimStagingName(const std::string& name); + // Whether one of THIS process's descriptors holds /.owner.lock, as the // kernel reports it in /proc/self/fdinfo (falling back to the recorded host // and pid where /proc cannot be read). What kHeldByCaller requires. diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 2319e4858..38a6ea361 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -377,6 +377,46 @@ void TestAStaleLockFileAboveDoesNotRefuseANestedDirectory() { CHECK(Contains(error, "contains")); } +// (4d) A claim killed between creating its directory's staging copy +// (...creating, lock file inside) and renaming it into place +// leaves that copy behind. Nobody holds it, it holds no pack, and it must +// not refuse a take of the directories above it for good. A staging copy +// whose lock IS held is a claim in progress, and still refuses. +void TestAnUnheldClaimStagingDirectoryRefusesNothing() { + const std::string base = FreshRoot("staging"); + const std::string key = base + "/root/0123456789ab"; + const std::string leftover = key + "/.r0-0a1b2c3d.0badf00d.creating"; + CHECK(dmi_store::IsSpoolClaimStagingName(".r0-0a1b2c3d.0badf00d.creating")); + CHECK(!dmi_store::IsSpoolClaimStagingName("r0-0a1b2c3d")); + CHECK(!dmi_store::IsSpoolClaimStagingName(".creating")); + fs::create_directories(leftover); + std::ofstream(leftover + "/.owner.lock") << "host 1\n"; + std::string error; + { + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(base + "/root", false, &lock, &error) == + SpoolStatus::kOk); + CHECK(error.empty()); + } + { + SpoolOwnerLock lock; + CHECK(SpoolOwnerLock::Acquire(key, false, &lock, &error) == + SpoolStatus::kOk); + CHECK(error.empty()); + } + // Held: a claim in the middle of creating its directory. + const std::string live = base + "/other/0123456789ab"; + SpoolOwnerLock claiming; + CHECK(SpoolOwnerLock::Acquire(live + "/.r1-0a1b2c3d.00c0ffee.creating", + false, &claiming, &error) == + SpoolStatus::kOk); + SpoolOwnerLock outer; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(base + "/other", false, &outer, &error) == + SpoolStatus::kBadArgument); + CHECK(Contains(error, "contains")); +} + // (5) The node-local check, through the test seam. void TestSharedFilesystemsAreRefusedUnlessAllowed() { const std::string root = FreshRoot("statfs") + "/spool"; @@ -520,6 +560,7 @@ int main() { TestNestedDirectoriesAreRefused(); TestAnOuterAndANestedTakeRacingNeverBothWin(); TestAStaleLockFileAboveDoesNotRefuseANestedDirectory(); + TestAnUnheldClaimStagingDirectoryRefusesNothing(); TestSharedFilesystemsAreRefusedUnlessAllowed(); TestAdoptionLocksOnlyWhatExistsAndIsDead(); TestANewDirectoryAppearsWithItsLockHeld(); diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 867f293cb..f056edd10 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -173,6 +173,16 @@ def _stale_open_file(directory: Path) -> Path: return stale +def _claim_staging(parent: Path, name: str, *, age_s: float) -> Path: + """The staging copy a claim killed before its rename leaves behind.""" + staging = parent / f".{name}.0badf00d.creating" + staging.mkdir() + (staging / ".owner.lock").write_text("host 1\n") + then = time.time() - age_s + os.utime(staging, (then, then)) + return staging + + def _read_all(config) -> dict: from dmi.storage.native_capture import NativeCaptureReader @@ -210,6 +220,11 @@ def test_a_sigkilled_process_spool_is_adopted_by_its_successor( _sigkill(child) stale = _stale_open_file(dead) assert _store().spool_owner(str(dead)) is None # died with it + # A claim killed before it renamed its directory into place leaves + # the staging copy (lock file inside). An old one is cleared; a + # fresh one may be a claim in progress, and is left alone. + old_claim = _claim_staging(dead.parent, "r0-0badf00d", age_s=3600) + fresh_claim = _claim_staging(dead.parent, "r0-00c0ffee", age_s=0) # Another process's live spool, bound for the same catalog. live = _claim(base, config) @@ -227,6 +242,8 @@ def test_a_sigkilled_process_spool_is_adopted_by_its_successor( assert snapshot["adoption_owed"] is False, snapshot assert not stale.exists() assert not dead.exists() + assert not old_claim.exists() + assert fresh_claim.exists() service.flush(60.0) assert _read_all(config) == _expected() finally: From 6e8724bf99740b8923d1a56e7b214fa6de7ebcbe Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 03:14:22 -0400 Subject: [PATCH 07/43] Let a child forked without exec drop its copy of every spool owner lock flock binds to an open file description, and a child created with fork() and no exec shares its parent's. O_CLOEXEC only acts at exec, so a fork-started worker (a multiprocessing or DataLoader worker under the fork start method, the default on Linux for the venv's Python 3.10) kept the spool's owner lock after its parent was SIGKILLed. The dead parent's directory then read as owned, naming the dead pid; its successor's adopt_sibling took kOwned for "it lives" and never came back to it, so its ready packs waited for the next restart after the worker exited. Reproduced from Python: the owner took SpoolOwnerLock, os.fork()ed a sleeping child and was SIGKILLed; spool_owner still named the dead pid and a new SpoolOwnerLock raised SpoolOwnedError until the child exited. Every held SpoolOwnerLock is now registered (per binary: each extension and driver that compiles spool.cpp keeps its own registry), and a pthread_atfork child handler closes the child's descriptor of each and marks it not held. It closes, never LOCK_UN: an unlock on the shared description would drop the parent's hold as well, and closing one of its descriptors does not. The registry mutex is held across the fork by the prepare handler, so no thread is mid-update when the child copies it. A kTake Spool's lock is a SpoolOwnerLock too, so the sink-only case is covered. posix_spawn and vfork run no atfork handler, and exec closes the descriptor there anyway. Tests, red before: test_spool_owner_lock.cpp forks an owner that takes the lock and forks a worker; the worker sees its own lock as not held while the parent's still is, and once the owner is SIGKILLed the lock is free with the worker alive (it stayed held, naming the dead owner). test_native_spool_ownership.py does the same through Python's os.fork. --- native/csrc/store/spool.cpp | 89 ++++++++++++++++++++++---- native/csrc/store/spool.h | 18 +++++- tests/native/test_spool_owner_lock.cpp | 77 +++++++++++++++++++++- tests/test_native_spool_ownership.py | 40 ++++++++++++ 4 files changed, 208 insertions(+), 16 deletions(-) diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 20043b73a..8ee73f160 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -16,6 +16,7 @@ #include #include +#include #include #include #include @@ -666,27 +667,91 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, } // namespace +namespace { +// Every held SpoolOwnerLock in this binary. Leaked on purpose, so no +// static destructor runs while a lock is still registered. Each binary +// that compiles spool.cpp (the store and sink extensions, the drivers) +// keeps its own registry and its own fork handlers, for its own locks. +std::mutex& LockRegistryMutex() { + static std::mutex* mutex = new std::mutex; + return *mutex; +} +std::unordered_set& LockRegistry() { + static auto* registry = new std::unordered_set; + return *registry; +} +} // namespace + +void SpoolOwnerLock::BeforeFork() { LockRegistryMutex().lock(); } +void SpoolOwnerLock::AfterForkInParent() { LockRegistryMutex().unlock(); } + +void SpoolOwnerLock::AfterForkInChild() { + // Close, never LOCK_UN: an unlock on the shared description would drop + // the parent's hold too, and closing one of its descriptors does not. + for (SpoolOwnerLock* lock : LockRegistry()) { + ::close(lock->fd_); + lock->fd_ = -1; + lock->dir_.clear(); + } + LockRegistry().clear(); + LockRegistryMutex().unlock(); +} + +void SpoolOwnerLock::Track(SpoolOwnerLock* lock) { + static std::once_flag handlers; + std::call_once(handlers, [] { + ::pthread_atfork(&SpoolOwnerLock::BeforeFork, + &SpoolOwnerLock::AfterForkInParent, + &SpoolOwnerLock::AfterForkInChild); + }); + std::lock_guard guard(LockRegistryMutex()); + LockRegistry().insert(lock); +} + +void SpoolOwnerLock::Untrack(SpoolOwnerLock* lock) { + std::lock_guard guard(LockRegistryMutex()); + LockRegistry().erase(lock); +} + +void SpoolOwnerLock::Hold(int fd, std::string dir) { + fd_ = fd; + dir_ = std::move(dir); + Track(this); +} + SpoolOwnerLock::~SpoolOwnerLock() { Release(); } -SpoolOwnerLock::SpoolOwnerLock(SpoolOwnerLock&& other) noexcept - : fd_(other.fd_), dir_(std::move(other.dir_)) { - other.fd_ = -1; - other.dir_.clear(); +SpoolOwnerLock::SpoolOwnerLock(SpoolOwnerLock&& other) noexcept { + if (other.held()) { + const int fd = other.fd_; + std::string dir = std::move(other.dir_); + Untrack(&other); + other.fd_ = -1; + other.dir_.clear(); + Hold(fd, std::move(dir)); + } } SpoolOwnerLock& SpoolOwnerLock::operator=(SpoolOwnerLock&& other) noexcept { if (this != &other) { Release(); - fd_ = other.fd_; - dir_ = std::move(other.dir_); - other.fd_ = -1; - other.dir_.clear(); + if (other.held()) { + const int fd = other.fd_; + std::string dir = std::move(other.dir_); + Untrack(&other); + other.fd_ = -1; + other.dir_.clear(); + Hold(fd, std::move(dir)); + } } return *this; } void SpoolOwnerLock::Release() { - if (fd_ >= 0) ::close(fd_); // closing the last descriptor unlocks + if (fd_ >= 0) { + Untrack(this); + ::close(fd_); // closing the last descriptor unlocks + } fd_ = -1; dir_.clear(); } @@ -725,8 +790,7 @@ SpoolStatus SpoolOwnerLock::Acquire(const std::string& dir, status = exists ? LockInPlace(canonical, &fd, error) : CreateLocked(canonical, &fd, error); if (status != SpoolStatus::kOk) return status; - out->fd_ = fd; - out->dir_ = canonical; + out->Hold(fd, canonical); // Only now, with this lock published: see CheckNotNested. status = CheckNotNested(canonical, error); if (status != SpoolStatus::kOk) { @@ -757,8 +821,7 @@ SpoolStatus SpoolOwnerLock::TryAdopt(const std::string& dir, int fd = -1; const SpoolStatus status = LockInPlace(resolved, &fd, error); if (status != SpoolStatus::kOk) return status; - out->fd_ = fd; - out->dir_ = resolved; + out->Hold(fd, resolved); return SpoolStatus::kOk; } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index f41be6597..685b141da 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -157,8 +157,13 @@ bool SpoolOwnedByThisProcess(const std::string& dir); // The owner lock of one spool directory: flock(LOCK_EX) on /.owner.lock, // released with the object (or Release()), and by the kernel when the -// process dies. The descriptor is close-on-exec; a child forked WITHOUT exec -// shares it, and keeps the lock held for as long as it lives. +// process dies. flock binds to an open file description, which a child +// shares after fork(): the descriptor is close-on-exec, and a child forked +// WITHOUT exec (a fork-started worker) closes its copy of every held lock +// at once (pthread_atfork), so the lock never outlives its owner in a child +// -- the owner's own hold is untouched, and the child's objects read as not +// held. A child the owner spawns through posix_spawn or vfork runs no +// atfork handler, and loses the descriptor at exec. class SpoolOwnerLock { public: static constexpr const char* kFileName = ".owner.lock"; @@ -200,6 +205,15 @@ class SpoolOwnerLock { bool ReleaseAndRemoveIfEmpty(std::string* error); private: + // Sets fd_ and dir_, and registers the lock for the fork handler. + void Hold(int fd, std::string dir); + // The registry of held locks in this binary, and its fork handlers. + static void Track(SpoolOwnerLock* lock); + static void Untrack(SpoolOwnerLock* lock); + static void BeforeFork(); + static void AfterForkInParent(); + static void AfterForkInChild(); + int fd_ = -1; std::string dir_; }; diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 38a6ea361..794ac7ffa 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -8,7 +8,8 @@ // refused beside another process's holder or when nothing holds the // lock. // 3. A second process is refused, told the holder's pid and host; the -// lock goes with its holder, even one killed with SIGKILL. +// lock goes with its holder, even one killed with SIGKILL, and a child +// it forked without exec does not keep it. // 4. Nesting: a directory under a HELD one, or containing one with a lock // file, is refused -- also when two processes take the pair at once. // 5. The node-local check refuses NFS and Lustre by statfs f_type, unless @@ -252,6 +253,79 @@ void TestASecondProcessIsRefusedUntilTheHolderDies() { CHECK(owner.pid == ::getpid()); } +// (3b) flock binds to an open file description, which a child forked +// without exec shares. Such a child -- a fork-started worker -- used to keep +// the lock after its parent was SIGKILLed, so the dead parent's directory +// read as owned (naming the dead pid) and was never adopted. A forked child +// now lets go of its copies of every held lock at once, and the parent's +// hold is untouched. +void TestAForkedChildDoesNotKeepTheLockPastItsParent() { + const std::string root = FreshRoot("fork-child") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t owner = ::fork(); + if (owner == 0) { + ::close(ready[0]); + SpoolOwnerLock lock; + std::string error; + if (SpoolOwnerLock::Acquire(root, false, &lock, &error) != + SpoolStatus::kOk) { + ::_exit(3); + } + const pid_t worker = ::fork(); // no exec + if (worker == 0) { + // The child's own view: it holds nothing, and the lock is still held + // (by the parent). + const char mine = lock.held() ? 'H' : 'h'; + const char held = dmi_store::ReadSpoolOwner(root, nullptr) ? 'P' : 'p'; + if (::write(ready[1], &mine, 1) != 1 || ::write(ready[1], &held, 1) != 1) { + ::_exit(3); + } + ::pause(); // outlives its parent until killed + ::_exit(0); + } + char pid_text[32]; + const int n = std::snprintf(pid_text, sizeof(pid_text), "%d\n", worker); + if (::write(ready[1], pid_text, n) != n) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + // The worker's two bytes and the owner's "\n", in any order. + std::string seen; + char byte = 0; + int flags = 0; + while (seen.find('\n') == std::string::npos || flags < 2) { + if (::read(ready[0], &byte, 1) != 1) break; + seen.push_back(byte); + if (byte == 'h' || byte == 'H' || byte == 'p' || byte == 'P') ++flags; + } + ::close(ready[0]); + std::string digits; + for (const char c : seen) { + if (c >= '0' && c <= '9') digits.push_back(c); + } + const pid_t worker = static_cast(std::atoi(digits.c_str())); + CHECK(worker > 0); + CHECK(seen.find('h') != std::string::npos); // the child holds nothing + CHECK(seen.find('P') != std::string::npos); // the parent still does + dmi_store::SpoolOwner owner_record; + CHECK(dmi_store::ReadSpoolOwner(root, &owner_record)); + CHECK(owner_record.pid == owner); + + ::kill(owner, SIGKILL); + int status = 0; + ::waitpid(owner, &status, 0); + // The worker lives on, and the lock went with its parent. + CHECK(::kill(worker, 0) == 0); + CHECK(!dmi_store::ReadSpoolOwner(root, &owner_record)); + SpoolOwnerLock successor; + std::string error; + CHECK(SpoolOwnerLock::TryAdopt(root, &successor, &error) == + SpoolStatus::kOk); + ::kill(worker, SIGKILL); +} + // (4) Nesting, both ways. void TestNestedDirectoriesAreRefused() { const std::string base = FreshRoot("nested"); @@ -557,6 +631,7 @@ int main() { TestHeldByCallerWithoutAHolderIsRefused(); TestTheLockGoesWithItsSpool(); TestASecondProcessIsRefusedUntilTheHolderDies(); + TestAForkedChildDoesNotKeepTheLockPastItsParent(); TestNestedDirectoriesAreRefused(); TestAnOuterAndANestedTakeRacingNeverBothWin(); TestAStaleLockFileAboveDoesNotRefuseANestedDirectory(); diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 232a18aa8..10e213186 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -19,6 +19,7 @@ import hashlib import json import os +import signal import socket import subprocess import sys @@ -162,6 +163,45 @@ def test_a_second_process_is_refused_naming_the_holder(tmp_path): result.stdout), result.stdout +def test_a_forked_worker_does_not_keep_a_dead_owners_lock(tmp_path): + """A child forked without exec -- a fork-started DataLoader or + multiprocessing worker -- shares the lock's open file description. It + used to keep the lock after its parent was SIGKILLed, so the dead + parent's directory read as owned, naming the dead pid, and was never + adopted.""" + directory = tmp_path / "spool" + script = ( + "import os, sys, time; sys.path.insert(0, sys.argv[1]);" + "import _dmi_native_store as m\n" + "lock = m.SpoolOwnerLock(sys.argv[2])\n" + "worker = os.fork()\n" + "if worker == 0:\n" + " time.sleep(3600)\n" + " os._exit(0)\n" + "print(worker, flush=True)\n" + "time.sleep(3600)\n") + owner = subprocess.Popen( + [sys.executable, "-c", script, str(BUILD), str(directory)], + stdout=subprocess.PIPE, text=True) + worker = int(owner.stdout.readline()) + try: + assert _store().spool_owner(str(directory))["pid"] == owner.pid + os.kill(owner.pid, signal.SIGKILL) + owner.wait(timeout=30) + os.kill(worker, 0) # the worker outlives its parent + assert _store().spool_owner(str(directory)) is None + with _store().SpoolOwnerLock(str(directory)): + pass + finally: + try: + os.kill(worker, signal.SIGKILL) + except ProcessLookupError: + pass + if owner.poll() is None: + owner.kill() + owner.wait(timeout=30) + + def test_spool_owner_names_the_holder_while_it_holds(tmp_path): store = _store() directory = str(tmp_path / "spool") From a56a163eb44f27e7f160fc4ded0cc1986a8da29e Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 03:18:23 -0400 Subject: [PATCH 08/43] Look again for siblings that were alive when the storage service started adopt_sibling took a sibling whose lock was held for "it lives", which is not owed, so the loop never came back to it: a sibling owned at start() was left for the process's whole lifetime, possibly days for a serving process. The previous commit removes the fork-child case, where the owner was dead but the lock still read as held. Two remain. A sibling can be briefly held at start() -- a predecessor still inside its close() in a rolling restart, which leaves the directory with packs if it cannot drain them -- and a rank on the node can die while this one runs. Either way its packs waited for the next restart on the node. While the last pass found a live sibling, run_cycle now passes over the siblings again every adoption_recheck_interval_ns (30 s by default, 0 for never; a native config knob, not on NativeCaptureStorageConfig), under the same rules as an owed pass: with the lease, nothing owed, no upload of its own failing. A live sibling is still not owed, so flush() is not held up by another process's directory. adopt_sibling probes the lock without blocking first, so a live sibling costs one probe per pass rather than TryAdopt's few milliseconds of retries. The snapshot's live_siblings says how many the last pass left alone. Test (test_native_spool_adoption_live.py, red before): a sibling this process holds, with three packs staged by the real sink, is left alone at start (live_siblings 1, nothing adopted or owed); once its lock is released the service adopts it within the recheck interval, removes the directory, and every capture reads back from the catalog. --- native/csrc/catalog/bindings_store.cpp | 3 + native/csrc/catalog/storage_service.cpp | 39 ++++++++--- native/csrc/catalog/storage_service.h | 21 +++++- tests/test_native_spool_adoption_live.py | 82 +++++++++++++++++++++++- 4 files changed, 133 insertions(+), 12 deletions(-) diff --git a/native/csrc/catalog/bindings_store.cpp b/native/csrc/catalog/bindings_store.cpp index 50dcaf05f..ac0a949df 100644 --- a/native/csrc/catalog/bindings_store.cpp +++ b/native/csrc/catalog/bindings_store.cpp @@ -100,6 +100,8 @@ dc::StorageServiceConfig service_config(const py::dict& d) { d, "spool_allow_shared_filesystem", c.spool_allow_shared_filesystem); c.adopt_sibling_spools = get(d, "adopt_sibling_spools", c.adopt_sibling_spools); + c.adoption_recheck_interval_ns = get( + d, "adoption_recheck_interval_ns", c.adoption_recheck_interval_ns); c.s3 = s3_config(d); c.uploader.store_id = get(d, "store_id", c.uploader.store_id); c.uploader.max_workers = get(d, "uploader_max_workers", c.uploader.max_workers); @@ -152,6 +154,7 @@ py::dict snapshot_dict(const dc::StorageServiceSnapshot& s) { out["adopted_spools"] = s.adopted_spools; out["adopted_packs"] = s.adopted_packs; out["adoption_owed"] = s.adoption_owed; + out["live_siblings"] = s.live_siblings; out["failed"] = s.failed; out["lease_state"] = s.lease_state; // Seconds on the monotonic clock, comparable with time.monotonic() (both diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index 2e11964b7..f597750ae 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -482,11 +482,16 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { // 3. Index them. index_or_owe(std::move(to_index), catalog); - // 3a. A dead sibling an earlier adoption pass left undrained, under the - // same rule as the uploads above: only with the lease, nothing owed - // and nothing of our own failing. - if (catalog && adoption_owed_ && pending_index_.empty() && - upload_failures == 0) { + // 3a. A dead sibling an earlier adoption pass left undrained, or -- on + // the recheck interval -- a sibling that was alive then and may have + // died since, under the same rule as the uploads above: only with + // the lease, nothing owed and nothing of our own failing. + const bool recheck_due = + live_siblings_ && config_.adoption_recheck_interval_ns > 0 && + steady_ns() - last_adoption_ns_ >= + config_.adoption_recheck_interval_ns; + if (catalog && (adoption_owed_ || recheck_due) && + pending_index_.empty() && upload_failures == 0) { adopt_siblings(); } @@ -600,10 +605,17 @@ void CaptureStorageService::adopt_siblings() { clear_dead_claim_staging(claim_staging); std::sort(siblings.begin(), siblings.end()); bool owed = false; + uint64_t live = 0; for (const std::string& sibling : siblings) { - if (!adopt_sibling(sibling)) owed = true; + bool alive = false; + if (!adopt_sibling(sibling, &alive)) owed = true; + if (alive) ++live; } adoption_owed_ = owed; + live_siblings_ = live != 0; + last_adoption_ns_ = steady_ns(); + std::lock_guard state(state_mutex_); + state_.live_siblings = live; } void CaptureStorageService::clear_dead_claim_staging( @@ -630,12 +642,23 @@ void CaptureStorageService::clear_dead_claim_staging( } } -bool CaptureStorageService::adopt_sibling(const std::string& directory) { +bool CaptureStorageService::adopt_sibling(const std::string& directory, + bool* live) { + *live = false; + // A live owner answers a non-blocking probe at once; TryAdopt would retry + // for a few milliseconds first, on every recheck. + if (dmi_store::ReadSpoolOwner(directory, nullptr)) { + *live = true; + return true; + } dmi_store::SpoolOwnerLock lock; std::string error; const dmi_store::SpoolStatus locked = dmi_store::SpoolOwnerLock::TryAdopt(directory, &lock, &error); - if (locked == dmi_store::SpoolStatus::kOwned) return true; // it lives + if (locked == dmi_store::SpoolStatus::kOwned) { // it lives + *live = true; + return true; + } if (locked != dmi_store::SpoolStatus::kOk) { // Another adopter drained and removed it meanwhile: nothing is owed. if (!std::filesystem::exists(directory)) return true; diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index b809d2539..fcf92ba90 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -112,8 +112,15 @@ struct StorageServiceConfig { // could not be drained (no lease, an upload that failed, a pack still // owed) is retried by the loop's cycles, and flush() does not report // drained until it has been. Live siblings -- another rank or job on this - // node -- are left alone. + // node, a predecessor still closing -- are left alone, and are not owed. bool adopt_sibling_spools = false; + // While a pass found a live sibling, the loop passes over the siblings + // again this often, so one whose owner dies later -- a predecessor that + // was still inside close() when this service started, a rank that + // crashes while this one runs -- is adopted then, not at the next + // restart on the node. A live sibling costs one non-blocking lock probe + // per pass. 0 never looks again after start(). + uint64_t adoption_recheck_interval_ns = 30'000'000'000ull; dmi_store::S3Config s3; dmi_store::UploaderConfig uploader; // uploader.store_id names the store @@ -203,6 +210,8 @@ struct StorageServiceSnapshot { uint64_t adopted_spools = 0; uint64_t adopted_packs = 0; bool adoption_owed = false; + // Siblings whose owner was alive at the last adoption pass. + uint64_t live_siblings = 0; // A foreign lease outlived 2 x TTL: the service stopped for good. bool failed = false; // "none" before start, "held", "quarantined" (an unknown outcome set the @@ -291,8 +300,9 @@ class CaptureStorageService { // adoption_owed_ to whether one was left undrained. Requires // cycle_mutex_. Only a lost lease propagates. void adopt_siblings(); - // Adopts one sibling; false when it is owed another try. - bool adopt_sibling(const std::string& directory); + // Adopts one sibling; false when it is owed another try. Sets *live + // when its owner is alive (not owed: it is its owner's). + bool adopt_sibling(const std::string& directory, bool* live); // Removes the staging copies (dmi_store::IsSpoolClaimStagingName) that // claims killed before their rename left under the catalog key. void clear_dead_claim_staging( @@ -380,6 +390,11 @@ class CaptureStorageService { bool reconcile_owed_ = false; // An adoption pass left a dead sibling undrained. Guarded by cycle_mutex_. bool adoption_owed_ = false; + // The last adoption pass found a live sibling, and when it ran: the loop + // passes again every adoption_recheck_interval_ns. Guarded by + // cycle_mutex_. + bool live_siblings_ = false; + uint64_t last_adoption_ns_ = 0; int failure_streak_ = 0; // consecutive failed cycles, for the backoff // Uploaded, so gone from the spool, but not yet in the catalog. std::vector pending_index_; diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index f056edd10..fd82a4ed2 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -15,7 +15,8 @@ When the object store is down at start, the dead directory stays as it was (its packs are durable there), flush() does not report drained, and -the loop adopts it once the store is back. +the loop adopts it once the store is back. A sibling whose owner is still +alive at start and dies later is adopted by a later pass. Needs ClickHouse on 127.0.0.1:8123/9000 and the native sink and store modules: make -C native build/_dmi_native_sink build/_dmi_native_store @@ -292,5 +293,84 @@ def test_a_dead_spool_waits_in_place_while_the_object_store_is_down( lock.release_and_remove_if_empty() +def _stage_into(directory: str, indexes) -> None: + """Stage records into a directory this process holds, as its own sink + would: the REAL native pack sink, held_by_caller.""" + import torch # noqa: F401 -- the sink extension links against it + + sys.path.insert(0, str(BUILD)) + try: + import _dmi_native_sink + finally: + sys.path.remove(str(BUILD)) + sink = _dmi_native_sink.NativePackSink( + spool_root=directory, layout=LAYOUT, + max_pack_records=RECORDS_PER_PACK, max_linger_ns=600 * 10**9, + owner_lock="held_by_caller") + lease = sink.attach() + envelope = _envelope(indexes) + sink.submit_envelope(LAYOUT, envelope.rows, envelope.payload()) + assert sink.flush_and_wait(60.0) + sink.rethrow_if_failed() + del lease, sink + + +def test_a_sibling_whose_owner_dies_after_start_is_adopted_by_a_recheck( + fake_s3, tmp_path): + """A sibling still owned when the service starts -- a predecessor still + inside its close(), another rank that dies later -- is left alone then. + The service looks again every adoption_recheck_interval_ns while it + has such a sibling, and adopts it once its owner is gone, rather than + leaving its packs for the next restart on the node.""" + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix) + sibling = _claim(base, config) + sibling_directory = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + staged = sorted(sibling_directory.rglob("*.dmi-pack.ready")) + assert len(staged) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + + lock = _claim(base, config) + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="recheck-test", poll_interval_ns=50_000_000, + sweep_spool_on_start=True, spool_owner_lock="held_by_caller", + adopt_sibling_spools=True, + adoption_recheck_interval_ns=200_000_000, + **config._lease_native()) + service = _store().StorageService(native) + service.start() + try: + snapshot = service.snapshot() + assert snapshot["live_siblings"] == 1, snapshot + assert snapshot["adopted_spools"] == 0, snapshot + assert snapshot["adoption_owed"] is False, snapshot + assert sorted(sibling_directory.rglob( + "*.dmi-pack.ready")) == staged + + sibling.release() # its owner is gone + deadline = time.monotonic() + 30 + while ((service.snapshot()["adopted_spools"] == 0 + or service.snapshot()["live_siblings"] != 0) + and time.monotonic() < deadline): + time.sleep(0.05) + snapshot = service.snapshot() + assert snapshot["adopted_spools"] == 1, snapshot + assert snapshot["adopted_packs"] == len(staged), snapshot + assert snapshot["live_siblings"] == 0, snapshot + assert not sibling_directory.exists() + assert service.flush(60.0) + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in _envelope( + STAGED_BY_THE_DEAD).expected.items()} + assert _read_all(config) == expected + finally: + service.stop() + lock.release_and_remove_if_empty() + + if __name__ == "__main__": _dead_capture_process(*sys.argv[1:4]) From 055ce112d5f0a93205b7b0afde0f2904db7d0f04 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 03:20:31 -0400 Subject: [PATCH 09/43] Keep the spool directory owned while a sink that did not seal may still stage close() and enable_ring_transport's replacement released the engine's spool claim in a finally once the service stopped, whether or not the sink had sealed: _seal_capture_sink swallows a flush timeout (close_flush_timeout_s is shared with the storage drain) and only logs it. The NativePackSink outlives that release. RingEngine keeps record_sink_ after stop(), NativePackSink::on_engine_release does nothing, and the user's RecordRuntime keeps the RingEngine through its RingTransport. So its stagers carry on and ~PackSink stages whatever it still holds, into a directory nobody owns any more, or into a removed one that Spool::Stage recreates with no .owner.lock. A successor's AdoptSiblings (another process on the node, or this process's next engine) then takes the directory and its Recover deletes the old sink's in-flight .open file, losing those records; or it adopts a directory still being written. The review reproduced it with the real sink and lock: one record submitted and not flushed, claim released and the directory removed; destroying the sink recreated it with a ready pack and no lock file. With a pack flushed and one record still queued, a conformance_spool recover (take) from another process succeeded on the directory while the sink was alive. Dropping an engine without close() had the same effect: the SpoolClaim went with it, and ~SpoolOwnerLock unlocked while the globally activated ring and its sink kept capturing. * _seal_capture_sink returns whether the sink sealed. The claim is let go only then; otherwise the engine keeps the directory owned until the process exits, logging that it does and which directory, and the next process on the node adopts it. * A SpoolClaim is registered in native_capture._HELD_SPOOL_CLAIMS until release(), so garbage collection never unlocks it; only release() or the process's exit does. * A create_record_runtime that fails in attach still releases at once: no record reached the sink. * The v1 contract says so. Tests, red before (test_native_capture_storage_wiring.py, fakes): a close whose sink flush times out stops the service but leaves the lock held and registered, and warns naming the directory (the lock was released); the same for replacing the record ring. test_native_spool_ownership.py: a real claim dropped and garbage-collected still reads as owned by this process (it read as unowned). --- docs/integration-api-v1.md | 6 ++- src/dmi/engine.py | 54 ++++++++++++++++----- src/dmi/storage/native_capture.py | 17 ++++++- tests/test_native_capture_storage_wiring.py | 48 +++++++++++++++--- tests/test_native_spool_ownership.py | 26 ++++++++++ 5 files changed, 131 insertions(+), 20 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index daceeecf5..346b9fa8f 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -192,7 +192,11 @@ hex digits of sha256 of `database/table_prefix/store_id`, the rank torchrun's `RANK`, 0 when unset, and the incarnation fresh on every `create_record_runtime`), and owns it: an flock on its `.owner.lock`, taken before the service starts and let go after the sink and the service are done, -when a drained directory is removed. A second process on a directory is +when a drained directory is removed. If the sink did not seal within +`close_flush_timeout_s`, it may still be staging, so the directory stays owned +by the process until it exits (a warning names it) and the next process on the +node adopts it; so does the directory of an engine dropped without `close()`. +A second process on a directory is refused, naming the holder's pid and host. At start the service adopts the directories under the same catalog key whose owners have died: their stale `.open` files are swept, their ready packs uploaded and indexed, and the diff --git a/src/dmi/engine.py b/src/dmi/engine.py index 9dd85d9d7..eb6649504 100644 --- a/src/dmi/engine.py +++ b/src/dmi/engine.py @@ -824,8 +824,8 @@ def enable_ring_transport( drain_deadline = None if storage is None else ( time.monotonic() + self._capture_storage_config.close_flush_timeout_s) - if storage is not None: - self._seal_capture_sink(drain_deadline) + sealed = storage is not None and self._seal_capture_sink( + drain_deadline) if old_record_mode: self._report_capture_failure() try: @@ -846,7 +846,8 @@ def enable_ring_transport( self._record_mode = False self._record_sink = None if storage is not None: - self._retire_capture_storage(storage, drain_deadline) + self._retire_capture_storage(storage, drain_deadline, + sink_sealed=sealed) # Pass the DMXHostEngine C++ object directly; RingEngine builds a # SubmitFn that calls submit_direct without touching Python/GIL. @@ -886,8 +887,8 @@ def next_auto_group_id(self) -> int: self._auto_batch_group_id += 1 return gid - def _seal_capture_sink(self, deadline: float) -> None: - """Flush the record sink before its ring stops. + def _seal_capture_sink(self, deadline: float) -> bool: + """Flush the record sink before its ring stops; whether it sealed. Stopping the ring releases the sink WITHOUT flushing it, so the records of its open pack would still be in memory while the service @@ -898,8 +899,11 @@ def _seal_capture_sink(self, deadline: float) -> None: max(0.0, deadline - time.monotonic())) except Exception as exc: _LOG.warning("capture sink did not flush: %s", exc) + return False + return True - def _retire_capture_storage(self, storage: Any, deadline: float) -> None: + def _retire_capture_storage(self, storage: Any, deadline: float, *, + sink_sealed: bool) -> None: """Drain the storage service until ``deadline``, then stop it.""" if self._capture_storage is not storage: return @@ -916,9 +920,33 @@ def _retire_capture_storage(self, storage: Any, deadline: float) -> None: try: storage.stop() finally: - # Last: the ring is stopped and the sink sealed by now, and - # the service has stopped touching the directory. - self._release_spool_claim() + # Last: the ring is stopped by now, and the service has + # stopped touching the directory. The sink is let go of only + # if it sealed. + if sink_sealed: + self._release_spool_claim() + else: + self._keep_spool_claim_held() + + def _keep_spool_claim_held(self) -> None: + """Leave the spool directory owned by this process until it exits. + + A sink that did not seal may still be staging: it outlives the ring + (the user's RecordRuntime keeps it), its stagers carry on, and its + destructor stages what it still holds. Let go of now, the directory + is another process's to adopt (or this process's next engine's), + and that adoption would sweep a stage in flight; removed, it would + be recreated by the next stage with no lock at all. The claim stays + in ``native_capture``'s registry, so nothing lets go of it until the + kernel does, at exit; the next process on the node adopts it then. + """ + claim, self._spool_claim = self._spool_claim, None + if claim is not None: + _LOG.warning( + "capture sink did not seal, so its spool directory %s stays " + "owned by this process until it exits; the next process on " + "the node for this catalog adopts what it holds", + claim.directory) def close(self) -> None: """Tear down backend resources.""" @@ -928,6 +956,9 @@ def close(self) -> None: # pack, then getting everything staged into the catalog. drain_deadline = None if storage is None else ( time.monotonic() + self._capture_storage_config.close_flush_timeout_s) + # Whether nothing can still stage into the spool: no record sink to + # seal, or one whose seal went through. + sealed = True if self._ring_transport is not None: record_mode = self._record_mode stopped = False @@ -940,7 +971,7 @@ def close(self) -> None: except Exception: pass if record_mode and storage is not None: - self._seal_capture_sink(drain_deadline) + sealed = self._seal_capture_sink(drain_deadline) if record_mode: self._report_capture_failure() try: @@ -965,7 +996,8 @@ def close(self) -> None: self._record_sink = None if storage is not None: - self._retire_capture_storage(storage, drain_deadline) + self._retire_capture_storage(storage, drain_deadline, + sink_sealed=sealed) if self._host_engine is not None: try: diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index d42147bde..a1ae61918 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -585,18 +585,29 @@ def _spool_producer_rank() -> int: return int(text) if text.isdigit() else 0 +# Every SpoolClaim not yet released. Dropping a claim must not let go of its +# directory: an engine dropped without close() drops its claim while the +# ring and the sink it activated may still be capturing into the directory, +# and another process's adoption would then sweep it from under them. So a +# claim lives until release(), or until the process exits and the kernel +# drops its lock. +_HELD_SPOOL_CLAIMS: set["SpoolClaim"] = set() + + class SpoolClaim: """This process's own spool directory, owned through its lock. From :func:`claim_spool_directory`. Hold it for as long as anything in the process writes or reads the directory -- the engine holds it from before its storage service starts until the sink and the service are - done -- then :meth:`release` it. + done -- then :meth:`release` it. Only ``release()`` lets go: a claim + that is merely dropped stays held until the process exits. """ def __init__(self, lock: Any) -> None: self._lock = lock self.directory: str = lock.directory + _HELD_SPOOL_CLAIMS.add(self) @property def held(self) -> bool: @@ -606,7 +617,9 @@ def release(self) -> bool: """Let go of the directory, removing it if nothing but its lock file is left. Whatever did not drain stays, and the next process on the node for this catalog adopts it. Returns whether it was - removed.""" + removed. Call it only once nothing in the process can still write + the directory.""" + _HELD_SPOOL_CLAIMS.discard(self) return bool(self._lock.release_and_remove_if_empty()) diff --git a/tests/test_native_capture_storage_wiring.py b/tests/test_native_capture_storage_wiring.py index de34cb4b8..e4c6306be 100644 --- a/tests/test_native_capture_storage_wiring.py +++ b/tests/test_native_capture_storage_wiring.py @@ -783,22 +783,58 @@ def test_close_flushes_the_sink_before_the_ring_stops(monkeypatch, tmp_path): assert engine._capture_storage is None -def test_close_still_stops_when_the_sink_flush_fails(monkeypatch, tmp_path): - engine, events, _services = _capture_engine(monkeypatch, tmp_path) - engine.create_record_runtime(_record_format()) - +def _fail_the_sink_flush(engine, events): def _failing_flush(timeout_s): events.append(("sink", "flush", timeout_s)) raise TimeoutError("timed out waiting for durable record completion") engine._ring_transport.flush_records_and_wait = _failing_flush + + +def test_close_still_stops_when_the_sink_flush_fails(monkeypatch, tmp_path, + caplog): + """The service still stops. The spool lock does not go: a sink that did + not seal may still be staging -- it outlives close() through the user's + RecordRuntime, and its stagers and destructor write into the directory + -- so another process's adoption (or this one's next engine) must not + take the directory from under it. The kernel lets go at exit.""" + from dmi.storage import native_capture + + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + _fail_the_sink_flush(engine, events) + (lock,) = engine._test_locks events.clear() - engine.close() + with caplog.at_level("WARNING", logger="dmi.engine"): + engine.close() assert [event[:2] for event in events] == [ ("sink", "flush"), ("ring", "stop"), - ("service", "flush"), ("service", "stop"), ("lock", "release")] + ("service", "flush"), ("service", "stop")] + assert lock.held + assert engine._spool_claim is None + # Kept alive for the process, so no garbage collection lets go of it. + assert any(claim._lock is lock + for claim in native_capture._HELD_SPOOL_CLAIMS) + assert lock.directory in caplog.text + assert "stays owned" in caplog.text + + +def test_replacing_a_record_ring_keeps_the_lock_when_the_sink_did_not_seal( + monkeypatch, tmp_path): + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + _fail_the_sink_flush(engine, events) + (lock,) = engine._test_locks + events.clear() + + engine.enable_ring_transport(object()) + + assert [event[:2] for event in events] == [ + ("sink", "flush"), ("ring", "stop"), + ("service", "flush"), ("service", "stop"), ("ring", "create")] + assert lock.held def test_close_releases_the_lease_even_when_the_drain_fails(monkeypatch, tmp_path): diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 10e213186..777d7e6ae 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -253,6 +253,32 @@ def test_a_spool_root_a_sink_once_owned_still_takes_rank_directories(tmp_path): claim.release() +def test_a_dropped_spool_claim_keeps_its_directory_owned(tmp_path): + """Only release() lets go of a claim. An engine dropped without close() + drops its claim, while the ring and the sink it activated may still be + capturing into the directory: garbage collection must not unlock it for + another process to adopt. The kernel lets go when the process exits.""" + import gc + + from dmi.storage import native_capture + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + claim = claim_spool_directory( + NativeSinkConfig(spool_root=str(tmp_path / "root")), _config()) + directory = claim.directory + del claim + gc.collect() + try: + assert _store().spool_owner(directory)["pid"] == os.getpid() + finally: + for kept in list(native_capture._HELD_SPOOL_CLAIMS): + if kept.directory == directory: + kept.release() + assert _store().spool_owner(directory) is None + + def test_a_shared_filesystem_is_refused_unless_allowed(tmp_path): store = _store() store._set_spool_filesystem_type_for_testing(NFS_SUPER_MAGIC) From 553b01b8d46d95426fea69f11a743e43e6560d7f Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 03:25:25 -0400 Subject: [PATCH 10/43] Key a spool's catalog directory by the servers its packs go to, not only the names The adoption domain, //, was keyed by sha256(database/table_prefix/store_id)[:12] -- the plan's section 2.3 definition -- which names no server. With the defaults (default, dmi, s3) every deployment on a node that shares a spool_root got the same key, so the first to start after another one's crash adopted its dead directories: their packs went up through its own S3 client to its own bucket and were indexed into its own ClickHouse, and never reached their own. Nothing on the adoption path checks a pack's destination to stop that. A site-wide spool path with staging and production under one service account is enough. The key now hashes where the packs go: //\n clickhouse :\n s3 / SpoolCatalogKey and SpoolRankDirectory take a SpoolDestination; the store module's spool_catalog_key / spool_rank_directory take its fields as a dict (every one required); NativeCaptureStorageConfig._spool_destination builds it for claim_spool_directory; and the service's adopt_sibling_spools check computes the same key from its ClickHouse connection, writer, S3 config and store id. The servers are hashed as spelled, so a successor that reaches the same store through a different spelling does not see the dead directory as its sibling: it waits for a process that spells it the same way. That is the safe side of the trade (packs delayed, not misdirected), and spool.h and the v1 contract say so. This departs from the plan's key definition, which needs amending to match. Tests: test_spool_owner_lock.cpp -- the key is that sha256, and changing any one of the seven fields changes it; test_native_spool_ownership.py -- the layout through the binding, a destination missing a field is a KeyError, a staging and a production config differing only in their servers claim directories under different keys, and a staging service refuses to adopt from under production's (they shared one before). The adoption live test's store-down case now has the dead process reach the store through the same switch URL as its successor. --- docs/integration-api-v1.md | 4 +- native/csrc/catalog/bindings_store.cpp | 41 +++++++++--- native/csrc/catalog/storage_service.cpp | 15 +++-- native/csrc/catalog/storage_service.h | 5 +- native/csrc/store/spool.cpp | 18 +++--- native/csrc/store/spool.h | 42 +++++++++---- src/dmi/storage/native_capture.py | 24 +++++-- tests/native/test_spool_owner_lock.cpp | 40 ++++++++++-- tests/test_native_capture_storage_wiring.py | 7 +-- tests/test_native_spool_adoption_live.py | 9 ++- tests/test_native_spool_ownership.py | 70 ++++++++++++++++++--- 11 files changed, 211 insertions(+), 64 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index 346b9fa8f..61709828d 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -188,7 +188,9 @@ same catalog is refused at `create_record_runtime`. The engine spools into a directory of its own under `capture_sink_config.spool_root`, `//r-/` (the key is the first 12 -hex digits of sha256 of `database/table_prefix/store_id`, the rank torchrun's +hex digits of a sha256 of where the packs go -- the ClickHouse host and port, +`database`, `table_prefix`, the S3 endpoint and bucket, and `store_id`, as +spelled in the config -- the rank torchrun's `RANK`, 0 when unset, and the incarnation fresh on every `create_record_runtime`), and owns it: an flock on its `.owner.lock`, taken before the service starts and let go after the sink and the service are done, diff --git a/native/csrc/catalog/bindings_store.cpp b/native/csrc/catalog/bindings_store.cpp index ac0a949df..c1c75e05b 100644 --- a/native/csrc/catalog/bindings_store.cpp +++ b/native/csrc/catalog/bindings_store.cpp @@ -90,6 +90,27 @@ dmi_store::OwnerLock owner_lock(const std::string& text) { return mode; } +// Where a spool's packs go (store/spool.h); every field is required. +dmi_store::SpoolDestination spool_destination(const py::dict& d) { + for (const char* key : {"clickhouse_host", "clickhouse_port", "database", + "table_prefix", "s3_endpoint", "s3_bucket", + "store_id"}) { + if (!d.contains(key)) { + throw py::key_error(std::string("spool destination needs '") + key + + "'"); + } + } + dmi_store::SpoolDestination out; + out.clickhouse_host = d["clickhouse_host"].cast(); + out.clickhouse_port = d["clickhouse_port"].cast(); + out.database = d["database"].cast(); + out.table_prefix = d["table_prefix"].cast(); + out.s3_endpoint = d["s3_endpoint"].cast(); + out.s3_bucket = d["s3_bucket"].cast(); + out.store_id = d["store_id"].cast(); + return out; +} + dc::StorageServiceConfig service_config(const py::dict& d) { dc::StorageServiceConfig c; c.spool_root = get(d, "spool_root", ""); @@ -388,11 +409,16 @@ PYBIND11_MODULE(_dmi_native_store, m) { py::arg("directory"), "Who holds a spool directory's owner lock, or None when nothing " "does."); - m.def("spool_catalog_key", &dmi_store::SpoolCatalogKey, py::arg("database"), - py::arg("table_prefix"), py::arg("store_id")); + m.def("spool_catalog_key", + [](const py::dict& destination) { + return dmi_store::SpoolCatalogKey(spool_destination(destination)); + }, + py::arg("destination"), + "The catalog key of a destination dict (clickhouse_host, " + "clickhouse_port, database, table_prefix, s3_endpoint, s3_bucket, " + "store_id)."); m.def("spool_rank_directory", - [](const std::string& base, const std::string& database, - const std::string& table_prefix, const std::string& store_id, + [](const std::string& base, const py::dict& destination, uint64_t producer_rank, std::optional incarnation) { const std::string fresh = incarnation ? *incarnation : dmi_store::NewSpoolIncarnation(); @@ -404,11 +430,10 @@ PYBIND11_MODULE(_dmi_native_store, m) { throw py::value_error("incarnation must be 8 lowercase hex " "digits"); } - return dmi_store::SpoolRankDirectory(base, database, table_prefix, - store_id, producer_rank, fresh); + return dmi_store::SpoolRankDirectory( + base, spool_destination(destination), producer_rank, fresh); }, - py::arg("base"), py::arg("database"), py::arg("table_prefix"), - py::arg("store_id"), py::arg("producer_rank"), + py::arg("base"), py::arg("destination"), py::arg("producer_rank"), py::arg("incarnation") = py::none(), "//r-, the section " "2.3 spool layout; a fresh incarnation when none is given."); diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index f597750ae..7c6007a91 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -169,9 +169,15 @@ CaptureStorageService::CaptureStorageService(StorageServiceConfig config) // under this catalog's key: a directory under another catalog's key // would index that catalog's packs here. const std::filesystem::path own(spool_.root()); - const std::string key = dmi_store::SpoolCatalogKey( - config_.writer.database, config_.writer.table_prefix, - config_.uploader.store_id); + dmi_store::SpoolDestination destination; + destination.clickhouse_host = config_.clickhouse.host; + destination.clickhouse_port = config_.clickhouse.port; + destination.database = config_.writer.database; + destination.table_prefix = config_.writer.table_prefix; + destination.s3_endpoint = config_.s3.endpoint; + destination.s3_bucket = config_.s3.bucket; + destination.store_id = config_.uploader.store_id; + const std::string key = dmi_store::SpoolCatalogKey(destination); uint64_t rank = 0; std::string incarnation; if (!dmi_store::ParseSpoolRankDirectoryName(own.filename().string(), @@ -180,7 +186,8 @@ CaptureStorageService::CaptureStorageService(StorageServiceConfig config) throw std::invalid_argument( "storage service: adopt_sibling_spools needs spool_root to be a " "rank directory /" + key + "/r- (this " - "catalog's key for database, table_prefix and store_id), got " + + "catalog's key for its ClickHouse host and port, database, " + "table_prefix, S3 endpoint, bucket and store_id), got " + spool_.root()); } } diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index fcf92ba90..dd8a82664 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -103,8 +103,9 @@ struct StorageServiceConfig { // Adopt the spools of dead processes bound for this catalog. spool_root // must then be a rank directory of the plan's section 2.3 layout, // //r-/ - // under THIS catalog's key (SpoolCatalogKey of writer.database, - // writer.table_prefix and uploader.store_id), or construction throws. + // under THIS catalog's key (SpoolCatalogKey of clickhouse.host and + // .port, writer.database and .table_prefix, s3.endpoint and .bucket, and + // uploader.store_id), or construction throws. // start(), after the lease and the sweep of its own directory, tries the // owner lock of every sibling rank directory; each one whose owner is // gone has its .open files swept, its .ready packs uploaded and indexed, diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 8ee73f160..c4134b37a 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -868,10 +868,12 @@ bool SpoolOwnerLock::ReleaseAndRemoveIfEmpty(std::string* error) { return removed; } -std::string SpoolCatalogKey(const std::string& database, - const std::string& table_prefix, - const std::string& store_id) { - const std::string text = database + "/" + table_prefix + "/" + store_id; +std::string SpoolCatalogKey(const SpoolDestination& destination) { + const std::string text = + destination.database + "/" + destination.table_prefix + "/" + + destination.store_id + "\nclickhouse " + destination.clickhouse_host + + ":" + std::to_string(destination.clickhouse_port) + "\ns3 " + + destination.s3_endpoint + "/" + destination.s3_bucket; unsigned char digest[SHA256_DIGEST_LENGTH]; SHA256(reinterpret_cast(text.data()), text.size(), digest); @@ -927,15 +929,13 @@ std::string NewSpoolIncarnation() { } std::string SpoolRankDirectory(const std::string& base, - const std::string& database, - const std::string& table_prefix, - const std::string& store_id, + const SpoolDestination& destination, uint64_t producer_rank, const std::string& incarnation) { std::string root = base; while (root.size() > 1 && root.back() == '/') root.pop_back(); - return root + "/" + SpoolCatalogKey(database, table_prefix, store_id) + - "/" + SpoolRankDirectoryName(producer_rank, incarnation); + return root + "/" + SpoolCatalogKey(destination) + "/" + + SpoolRankDirectoryName(producer_rank, incarnation); } SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 685b141da..263035300 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -218,18 +218,36 @@ class SpoolOwnerLock { std::string dir_; }; +// Where a spool's packs go: the catalog -- its ClickHouse server, database +// and table prefix -- and the store -- its S3 endpoint, bucket and store id. +struct SpoolDestination { + std::string clickhouse_host; + uint64_t clickhouse_port = 0; + std::string database; + std::string table_prefix; + std::string s3_endpoint; + std::string s3_bucket; + std::string store_id; +}; + // The spool layout of the plan's section 2.3. A capture process spools into // //r-/ -// where catalog_key is the first 12 hex digits of -// sha256("//"), and incarnation is 8 hex -// digits fresh for every process start, so no two processes -- two jobs on -// one node, or a restart of the same rank -- ever share a directory. The -// directories under one catalog key are siblings: packs bound for one -// catalog and store, which a successor on the node adopts once their owner -// has died (CaptureStorageService, adopt_sibling_spools). -std::string SpoolCatalogKey(const std::string& database, - const std::string& table_prefix, - const std::string& store_id); +// where catalog_key is the first 12 hex digits of the sha256 of +// "//\n" +// "clickhouse :\n" +// "s3 /" +// and incarnation is 8 hex digits fresh for every process start, so no two +// processes -- two jobs on one node, or a restart of the same rank -- ever +// share a directory. The directories under one catalog key are siblings: +// packs bound for one catalog and store, which a successor on the node +// adopts once their owner has died (CaptureStorageService, +// adopt_sibling_spools). The servers are in the key, not only the names: +// with the plan's sha256(database/table_prefix/store_id), two deployments +// sharing a spool_root and the default names but not a server adopted each +// other's dead directories into the wrong catalog and bucket. The servers +// are hashed as spelled, so spell them alike on every process of a +// deployment, or a dead directory waits for a process that does. +std::string SpoolCatalogKey(const SpoolDestination& destination); bool IsSpoolCatalogKey(const std::string& name); std::string SpoolRankDirectoryName(uint64_t producer_rank, const std::string& incarnation); @@ -239,9 +257,7 @@ bool ParseSpoolRankDirectoryName(const std::string& name, std::string* incarnation); std::string NewSpoolIncarnation(); std::string SpoolRankDirectory(const std::string& base, - const std::string& database, - const std::string& table_prefix, - const std::string& store_id, + const SpoolDestination& destination, uint64_t producer_rank, const std::string& incarnation); diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index a1ae61918..308d8d2f6 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -508,6 +508,20 @@ def _native_dict(self) -> dict[str, Any]: "uploader_max_in_flight_bytes": self.uploader_max_in_flight_bytes, } + def _spool_destination(self) -> dict[str, Any]: + """Where the packs go, as the spool's catalog key hashes it: the + catalog's server and names, and the store's endpoint, bucket and + id (native/csrc/store/spool.h, SpoolDestination).""" + return { + "clickhouse_host": self.clickhouse_host, + "clickhouse_port": self.clickhouse_port, + "database": self.database, + "table_prefix": self.table_prefix, + "s3_endpoint": self.s3_endpoint, + "s3_bucket": self.s3_bucket, + "store_id": self.store_id, + } + def _native_reader_dict(self) -> dict[str, Any]: """The reader's native config: the reader account, when one is set.""" native = self._native_dict() @@ -630,8 +644,11 @@ def claim_spool_directory( """Create and lock this process's spool directory. ``//r-/``: the catalog key is - the first 12 hex digits of sha256 of ``database/table_prefix/store_id``, - so every directory under it holds packs for this catalog and store; the + the first 12 hex digits of a sha256 of where the packs go -- the + ClickHouse host and port, ``database``, ``table_prefix``, the S3 endpoint + and bucket, and ``store_id`` -- so every directory under it holds packs + for this catalog and store, and two deployments that share the default + names but not a server never adopt each other's directories; the incarnation is fresh for every call, so no two processes -- two jobs on one node, or a restart -- share a directory; the rank is torchrun's ``RANK`` (0 when unset). The directory is created with its owner lock @@ -641,8 +658,7 @@ def claim_spool_directory( """ module = _load_native_store_extension() directory = module.spool_rank_directory( - sink_config.spool_root, storage_config.database, - storage_config.table_prefix, storage_config.store_id, + sink_config.spool_root, storage_config._spool_destination(), _spool_producer_rank()) return SpoolClaim(module.SpoolOwnerLock( directory, diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 794ac7ffa..d9f63af5b 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -589,14 +589,43 @@ void TestANewDirectoryAppearsWithItsLockHeld() { } // (7) The layout: //r-/. +dmi_store::SpoolDestination Destination() { + dmi_store::SpoolDestination destination; + destination.clickhouse_host = "ch"; + destination.clickhouse_port = 8123; + destination.database = "db"; + destination.table_prefix = "prefix"; + destination.s3_endpoint = "http://s3:9000"; + destination.s3_bucket = "bucket"; + destination.store_id = "s3"; + return destination; +} + void TestTheDirectoryLayout() { - const std::string key = dmi_store::SpoolCatalogKey("db", "prefix", "s3"); - CHECK(key == Sha256Hex("db/prefix/s3").substr(0, 12)); + const std::string key = dmi_store::SpoolCatalogKey(Destination()); + CHECK(key == Sha256Hex("db/prefix/s3\nclickhouse ch:8123\n" + "s3 http://s3:9000/bucket").substr(0, 12)); CHECK(dmi_store::IsSpoolCatalogKey(key)); CHECK(!dmi_store::IsSpoolCatalogKey("0123456789aB")); CHECK(!dmi_store::IsSpoolCatalogKey("0123456789a")); - CHECK(dmi_store::SpoolCatalogKey("db", "prefix", "s3") != - dmi_store::SpoolCatalogKey("db", "prefix", "s4")); + // Every part of where the packs go is in the key: two deployments that + // share a spool_root and the default names, but not a server, never + // adopt each other's directories. + std::set keys{key}; + for (int field = 0; field < 7; ++field) { + dmi_store::SpoolDestination other = Destination(); + switch (field) { + case 0: other.clickhouse_host = "ch2"; break; + case 1: other.clickhouse_port = 8124; break; + case 2: other.database = "db2"; break; + case 3: other.table_prefix = "prefix2"; break; + case 4: other.s3_endpoint = "http://s3b:9000"; break; + case 5: other.s3_bucket = "bucket2"; break; + case 6: other.store_id = "s4"; break; + } + keys.insert(dmi_store::SpoolCatalogKey(other)); + } + CHECK(keys.size() == 8); CHECK(dmi_store::SpoolRankDirectoryName(3, "0a1b2c3d") == "r3-0a1b2c3d"); uint64_t rank = 0; @@ -617,8 +646,7 @@ void TestTheDirectoryLayout() { seen.insert(fresh); } CHECK(seen.size() == 64); - CHECK(dmi_store::SpoolRankDirectory("/b", "db", "prefix", "s3", 2, - "0a1b2c3d") == + CHECK(dmi_store::SpoolRankDirectory("/b", Destination(), 2, "0a1b2c3d") == "/b/" + key + "/r2-0a1b2c3d"); } diff --git a/tests/test_native_capture_storage_wiring.py b/tests/test_native_capture_storage_wiring.py index e4c6306be..3ad1bf000 100644 --- a/tests/test_native_capture_storage_wiring.py +++ b/tests/test_native_capture_storage_wiring.py @@ -489,8 +489,8 @@ def _lock(directory, allow_shared_filesystem=False): locks.append(lock) return lock - def _rank_directory(base, database, table_prefix, store_id, rank): - events.append(("layout", database, table_prefix, store_id, rank)) + def _rank_directory(base, destination, rank): + events.append(("layout", destination, rank)) return RANK_DIRECTORY.format(base=base, rank=rank) def _load_named_extension(name): @@ -580,8 +580,7 @@ def test_the_service_starts_before_the_sink_opens_the_spool(monkeypatch, tmp_pat directory = RANK_DIRECTORY.format(base=tmp_path / "spool", rank=0) storage = _storage_config() assert events[:5] == [ - ("layout", storage.database, storage.table_prefix, storage.store_id, - 0), + ("layout", storage._spool_destination(), 0), ("lock", "acquire", directory), ("service", "construct"), ("service", "start"), diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index fd82a4ed2..03b9a4cdd 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -90,7 +90,7 @@ def _claim(base: Path, config): """What the engine does first: a fresh rank directory, locked.""" store = _store() directory = store.spool_rank_directory( - str(base), config.database, config.table_prefix, config.store_id, 0) + str(base), config._spool_destination(), 0) return store.SpoolOwnerLock(directory) @@ -261,12 +261,15 @@ def test_a_dead_spool_waits_in_place_while_the_object_store_is_down( base = tmp_path / "spool" with _catalog() as prefix: - child, dead = _spawn_dead_process(base, fake_s3, prefix) + # Both processes reach the store through the switch: the endpoint is + # part of the catalog key, so a successor spelling it differently + # would not see the dead directory as its sibling. + switch = _Switch.to_url(fake_s3) + child, dead = _spawn_dead_process(base, switch.url, prefix) _sigkill(child) staged = sorted(dead.rglob("*.dmi-pack.ready")) assert staged - switch = _Switch.to_url(fake_s3) switch.cut() config = _storage_config(switch.url, prefix) lock = _claim(base, config) diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 777d7e6ae..023cef2f5 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -63,8 +63,14 @@ def _config(**overrides): def _rank_directory(base: Path, config, rank: int = 0) -> str: return _store().spool_rank_directory( - str(base), config.database, config.table_prefix, config.store_id, - rank) + str(base), config._spool_destination(), rank) + + +# Where a spool's packs go: the catalog's server and names, the store's. +DESTINATION = dict( + clickhouse_host="ch", clickhouse_port=8123, database="db", + table_prefix="prefix", s3_endpoint="http://s3:9000", s3_bucket="bucket", + store_id="s3") def _service(config, spool_root, **options): @@ -300,18 +306,62 @@ def test_a_shared_filesystem_is_refused_unless_allowed(tmp_path): def test_the_rank_directory_layout(tmp_path): store = _store() - key = hashlib.sha256(b"db/prefix/s3").hexdigest()[:12] - assert store.spool_catalog_key("db", "prefix", "s3") == key + key = hashlib.sha256( + b"db/prefix/s3\nclickhouse ch:8123\ns3 http://s3:9000/bucket" + ).hexdigest()[:12] + assert store.spool_catalog_key(DESTINATION) == key assert store.spool_rank_directory( - "/base", "db", "prefix", "s3", 3, "0a1b2c3d") == ( - f"/base/{key}/r3-0a1b2c3d") - fresh = {store.spool_rank_directory("/base", "db", "prefix", "s3", 0) + "/base", DESTINATION, 3, "0a1b2c3d") == f"/base/{key}/r3-0a1b2c3d" + fresh = {store.spool_rank_directory("/base", DESTINATION, 0) for _ in range(16)} assert len(fresh) == 16 # a fresh incarnation each time for path in fresh: assert path.startswith(f"/base/{key}/r0-") with pytest.raises(ValueError, match="incarnation"): - store.spool_rank_directory("/base", "db", "prefix", "s3", 0, "XYZ") + store.spool_rank_directory("/base", DESTINATION, 0, "XYZ") + incomplete = dict(DESTINATION) + del incomplete["s3_bucket"] + with pytest.raises(KeyError, match="s3_bucket"): + store.spool_catalog_key(incomplete) + assert _config()._spool_destination() == { + name: getattr(_config(), name) for name in DESTINATION} + + +def test_the_catalog_key_names_the_servers_not_only_the_names(tmp_path): + """The key was sha256(database/table_prefix/store_id), so two + deployments on one node with the default names (default, dmi, s3) but + different ClickHouse servers and buckets shared a key: whichever started + first adopted the other's dead directories, uploading its packs to its + own bucket and indexing them into its own catalog. Every server and + name the packs go to is in the key now.""" + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + store = _store() + key = store.spool_catalog_key(DESTINATION) + for field, value in (("clickhouse_host", "ch2"), ("clickhouse_port", 8124), + ("s3_endpoint", "http://s3b:9000"), + ("s3_bucket", "bucket2")): + assert store.spool_catalog_key({**DESTINATION, field: value}) != key + + staging = _config(clickhouse_host="ch-staging", s3_bucket="staging") + production = _config(clickhouse_host="ch-production", + s3_bucket="production") + sink = NativeSinkConfig(spool_root=str(tmp_path / "spool")) + staged = claim_spool_directory(sink, staging) + produced = claim_spool_directory(sink, production) + try: + assert (Path(staged.directory).parent + != Path(produced.directory).parent) + # A staging service cannot adopt from under production's key. + with pytest.raises(ValueError, match="adopt_sibling_spools"): + _service(staging, produced.directory, + spool_owner_lock="held_by_caller", + adopt_sibling_spools=True) + finally: + staged.release() + produced.release() def test_adoption_needs_a_rank_directory_under_this_catalogs_key(tmp_path): @@ -321,8 +371,8 @@ def test_adoption_needs_a_rank_directory_under_this_catalogs_key(tmp_path): # A rank directory, but under another catalog's key: adopting its # siblings would index that catalog's packs into this one. other = _store().spool_rank_directory( - str(tmp_path / "base"), config.database, "another_prefix", - config.store_id, 0) + str(tmp_path / "base"), + {**config._spool_destination(), "table_prefix": "another_prefix"}, 0) with pytest.raises(ValueError, match="adopt_sibling_spools"): _service(config, other, adopt_sibling_spools=True) mine = _rank_directory(tmp_path / "base", config) From db8728032ad2200275b310c2eeeffe417fd9d6bd Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 03:27:53 -0400 Subject: [PATCH 11/43] Refuse BeeGFS, CIFS/SMB2 and FUSE spool roots, not only NFS and Lustre The node-local check refused two statfs f_type values, NFS (0x6969) and Lustre (0x0BD00BD0). Others give a flock that does not keep out a process on another node just as well, and the owner lock is then no lock at all across the nodes that mount the directory: * BeeGFS (0x19830326), common for HPC scratch, keeps flock client-local unless tuneUseGlobalFileLocks is set, and that is off by default; * CIFS (0xFF534D42) and SMB2 (0xFE534D42); * FUSE (0x65735546), which covers sshfs, s3fs, gcsfuse and the GlusterFS client -- and local filesystems too (fuse-overlayfs, ntfs-3g), which f_type cannot tell apart. Refusing a local one is loud and has the override; admitting a network one is silent, so FUSE is refused, and its refusal names the local case the override is for. allow_shared_filesystem admits any of them, as before. The list still needs maintenance, as the plan's risk section says. Tests, red before: test_spool_owner_lock.cpp names all six and refuses a lock and a Spool on each through the statfs seam, with the override still admitting them, while ext4, xfs, overlayfs and tmpfs pass; test_native_spool_ownership.py refuses each through the binding, naming it. --- docs/integration-api-v1.md | 5 +++-- native/csrc/store/spool.cpp | 21 +++++++++++++++++++-- native/csrc/store/spool.h | 14 ++++++++------ src/dmi/storage/native_capture.py | 7 ++++--- tests/native/test_spool_owner_lock.cpp | 26 +++++++++++++++++++------- tests/test_native_spool_ownership.py | 16 ++++++++++++++++ 6 files changed, 69 insertions(+), 20 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index 61709828d..414f7de6d 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -204,8 +204,9 @@ directories under the same catalog key whose owners have died: their stale `.open` files are swept, their ready packs uploaded and indexed, and the directory removed, so a crashed process's packs reach the catalog through the next one on the node, whatever run it belongs to. The spool root must be -node-local: NFS and Lustre are refused unless -`NativeSinkConfig.spool_allow_shared_filesystem`. Without +node-local: NFS, Lustre, BeeGFS, CIFS/SMB2 and FUSE are refused (by statfs +`f_type`) unless `NativeSinkConfig.spool_allow_shared_filesystem`, which a FUSE +filesystem that is itself local, such as fuse-overlayfs, needs too. Without `capture_storage_config` the sink owns `spool_root` itself. With an explicit `record_sink`, the service drains `spool_root` as that sink writes it, unswept and adopting nothing. diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index c4134b37a..2b1f96929 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -38,9 +38,18 @@ constexpr const char* kClaimStagingSuffix = ".creating"; // is ever staged under it. constexpr const char* kRefsDirectory = "_refs"; -// statfs(2) f_type values (linux/magic.h has NFS; Lustre's is its own). +// statfs(2) f_type values of filesystems whose flock does not keep out a +// process on another node (linux/magic.h has NFS, SMB2, CIFS and FUSE; +// Lustre's and BeeGFS's are their own). BeeGFS keeps flock client-local +// unless tuneUseGlobalFileLocks is set. FUSE covers network filesystems +// (sshfs, s3fs, gcsfuse, GlusterFS) and local ones alike, and f_type cannot +// tell them apart, so a local one needs the override. constexpr uint32_t kNfsSuperMagic = 0x6969; constexpr uint32_t kLustreSuperMagic = 0x0BD00BD0; +constexpr uint32_t kBeeGfsSuperMagic = 0x19830326; +constexpr uint32_t kCifsSuperMagic = 0xFF534D42; +constexpr uint32_t kSmb2SuperMagic = 0xFE534D42; +constexpr uint32_t kFuseSuperMagic = 0x65735546; std::atomic g_filesystem_type_for_testing{-1}; @@ -439,6 +448,10 @@ const char* SharedFilesystemName(int64_t f_type) { switch (static_cast(f_type)) { case kNfsSuperMagic: return "NFS"; case kLustreSuperMagic: return "Lustre"; + case kBeeGfsSuperMagic: return "BeeGFS"; + case kCifsSuperMagic: return "CIFS"; + case kSmb2SuperMagic: return "SMB2"; + case kFuseSuperMagic: return "FUSE"; default: return nullptr; } } @@ -469,7 +482,11 @@ SpoolStatus CheckNodeLocal(const std::string& dir, "since its owner lock (flock) does not keep out a process on " "another node there. Use a local disk, or set " "allow_shared_filesystem if no process on another node can " - "reach this directory"; + "reach this directory" + + (std::string(shared) == "FUSE" + ? " (a FUSE filesystem that is itself local, such as " + "fuse-overlayfs or ntfs-3g, is one)" + : std::string()); } return SpoolStatus::kBadArgument; } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 263035300..77dc705cf 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -39,9 +39,9 @@ // check runs after the lock is taken, so of two processes taking an outer // and a nested directory at once, at least one is refused. /_refs/ is // never scanned: the upload handoff's ref files live there (plan section -// 2.4). The spool must be node-local: NFS and Lustre are refused by statfs -// f_type unless allow_shared_filesystem is set, since neither guarantees a -// flock that excludes a process on another node. +// 2.4). The spool must be node-local: NFS, Lustre, BeeGFS, CIFS/SMB2 and +// FUSE are refused by statfs f_type unless allow_shared_filesystem is set, +// since none guarantees a flock that excludes a process on another node. // // The Python DurablePackSpool (spool.py) takes no lock, and its recover() // deletes every .open file under its root; the C++ spool is deliberately @@ -73,8 +73,9 @@ struct SpoolConfig { std::string root; uint64_t max_bytes = 0; OwnerLock owner_lock = OwnerLock::kTake; - // Admit a root on NFS or Lustre. Only safe when every process that could - // open the directory runs on this node. + // Admit a root on a shared filesystem (SharedFilesystemName), or on a + // FUSE filesystem that is local after all. Only safe when every process + // that could open the directory runs on this node. bool allow_shared_filesystem = false; }; @@ -119,7 +120,8 @@ inline const char* SpoolStatusName(SpoolStatus s) { } // The statfs f_type names of the shared filesystems a spool refuses: "NFS", -// "Lustre", or nullptr for any other. The list needs maintenance as +// "Lustre", "BeeGFS", "CIFS", "SMB2", "FUSE", or nullptr for any other. +// The list needs maintenance as // deployments meet new ones. const char* SharedFilesystemName(int64_t f_type); diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index 308d8d2f6..75bff9b45 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -137,9 +137,10 @@ class NativeSinkConfig: under it, ``//r-/`` (see :func:`claim_spool_directory`), and its storage service adopts the directories of dead processes beside it. Without one, the sink owns - ``spool_root`` itself. A root on NFS or Lustre is refused unless - ``spool_allow_shared_filesystem``: flock there does not keep out a - process on another node. + ``spool_root`` itself. A root on NFS, Lustre, BeeGFS, CIFS/SMB2 or FUSE + is refused unless ``spool_allow_shared_filesystem``: flock there does + not keep out a process on another node (a FUSE filesystem that is local, + such as fuse-overlayfs, needs the override too). """ spool_root: str diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index d9f63af5b..55af67186 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -12,8 +12,9 @@ // it forked without exec does not keep it. // 4. Nesting: a directory under a HELD one, or containing one with a lock // file, is refused -- also when two processes take the pair at once. -// 5. The node-local check refuses NFS and Lustre by statfs f_type, unless -// explicitly allowed (a test seam stands in for statfs). +// 5. The node-local check refuses NFS, Lustre, BeeGFS, CIFS/SMB2 and FUSE +// by statfs f_type, unless explicitly allowed (a test seam stands in +// for statfs). // 6. Adoption's try-lock never creates a directory, and a released // directory that holds nothing but its lock file can be removed. // 7. The directory layout of the plan's section 2.3. @@ -494,20 +495,31 @@ void TestAnUnheldClaimStagingDirectoryRefusesNothing() { // (5) The node-local check, through the test seam. void TestSharedFilesystemsAreRefusedUnlessAllowed() { const std::string root = FreshRoot("statfs") + "/spool"; - CHECK(std::string(dmi_store::SharedFilesystemName(0x6969)) == "NFS"); - CHECK(std::string(dmi_store::SharedFilesystemName(0x0BD00BD0)) == - "Lustre"); + // Each one's flock does not keep out a process on another node: NFS and + // Lustre (the plan's two), BeeGFS (client-local unless + // tuneUseGlobalFileLocks), CIFS/SMB2, and FUSE, which cannot tell sshfs, + // s3fs, gcsfuse or GlusterFS from a local filesystem. + const std::vector> shared = { + {0x6969, "NFS"}, {0x0BD00BD0, "Lustre"}, + {0x19830326, "BeeGFS"}, {0xFF534D42, "CIFS"}, + {0xFE534D42, "SMB2"}, {0x65735546, "FUSE"}}; + for (const auto& [magic, name] : shared) { + const char* named = dmi_store::SharedFilesystemName(magic); + CHECK(named != nullptr && std::string(named) == name); + } CHECK(dmi_store::SharedFilesystemName(0xEF53) == nullptr); // ext4 CHECK(dmi_store::SharedFilesystemName(0x58465342) == nullptr); // xfs + CHECK(dmi_store::SharedFilesystemName(0x794C7630) == nullptr); // overlayfs + CHECK(dmi_store::SharedFilesystemName(0x01021994) == nullptr); // tmpfs std::string error; - for (const int64_t magic : {int64_t{0x6969}, int64_t{0x0BD00BD0}}) { + for (const auto& [magic, name] : shared) { dmi_store::SetFilesystemTypeForTesting(magic); SpoolOwnerLock lock; error.clear(); CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == SpoolStatus::kBadArgument); - CHECK(Contains(error, magic == 0x6969 ? "NFS" : "Lustre")); + CHECK(Contains(error, " is on " + name + " ")); CHECK(Contains(error, "node-local")); CHECK(!lock.held()); Spool spool; diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 023cef2f5..0ccf4df76 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -285,6 +285,22 @@ def test_a_dropped_spool_claim_keeps_its_directory_owned(tmp_path): assert _store().spool_owner(directory) is None +@pytest.mark.parametrize("f_type, name", [ + (0x6969, "NFS"), (0x0BD00BD0, "Lustre"), (0x19830326, "BeeGFS"), + (0xFF534D42, "CIFS"), (0xFE534D42, "SMB2"), (0x65735546, "FUSE")]) +def test_each_shared_filesystem_is_refused_by_name(tmp_path, f_type, name): + store = _store() + store._set_spool_filesystem_type_for_testing(f_type) + try: + with pytest.raises(ValueError, match=f"is on {name} .*node-local"): + store.SpoolOwnerLock(str(tmp_path / "shared")) + with store.SpoolOwnerLock(str(tmp_path / "shared"), + allow_shared_filesystem=True): + pass + finally: + store._set_spool_filesystem_type_for_testing(None) + + def test_a_shared_filesystem_is_refused_unless_allowed(tmp_path): store = _store() store._set_spool_filesystem_type_for_testing(NFS_SUPER_MAGIC) From bc7b56f00f50b55129e7e9abe776c68e05c7d7cc Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 04:10:51 -0400 Subject: [PATCH 12/43] Give the adoption test's live sibling real packs and a stage in flight The one service-level check that adoption leaves a live sibling alone wrote a file named "marker" into it and asserted it still read "live", plus live.held -- the test's own Python object, always true. Recover() unlinks only *.open files and validates only *.ready ones, so a service that swept a live sibling passed: the review's mutation, making the live branch open the sibling held_by_caller and Recover() it, left every CPU and live adoption test green. The live sibling now holds what a live writer's directory does: ready packs its real sink staged, and a stage in flight (a .open temp file). After the successor has adopted the dead directory and flushed, both are exactly where they were, none of the live packs' ids is in the object store, live_siblings is 1, and the catalog holds only the dead process's captures. The sibling's lock is held by this process, which is the case that matters: an adopter that took it for dead would get past the held_by_caller check, and only the liveness probe keeps it out. Checked: with the review's mutation applied to adopt_sibling's live branch, the test fails on the swept .open file (FileNotFoundError); as the code stands it passes. --- tests/test_native_spool_adoption_live.py | 27 ++++++++++++++++++++---- 1 file changed, 23 insertions(+), 4 deletions(-) diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 03b9a4cdd..9f096c0a0 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -37,7 +37,7 @@ # Module-level so the fake-S3 fixture registers in this module. from tests.test_native_s3_client import ( # noqa: E402 - ACCESS, BUCKET, REGION, SECRET, fake_s3, + ACCESS, BUCKET, REGION, SECRET, STATE, fake_s3, ) REPO = Path(__file__).resolve().parents[1] @@ -64,6 +64,7 @@ LEASE = dict(lease_ttl_s=3.0, publish_timeout_s=1.0) INDEXED_BY_THE_DEAD = range(0, 4) # flushed to the catalog before the kill STAGED_BY_THE_DEAD = range(4, 10) # only in its spool when it dies +STAGED_BY_THE_LIVE = range(20, 24) # a live sibling's, never adopted RECORDS_PER_PACK = 2 @@ -227,9 +228,18 @@ def test_a_sigkilled_process_spool_is_adopted_by_its_successor( old_claim = _claim_staging(dead.parent, "r0-0badf00d", age_s=3600) fresh_claim = _claim_staging(dead.parent, "r0-00c0ffee", age_s=0) - # Another process's live spool, bound for the same catalog. + # A live sibling bound for the same catalog: ready packs its sink + # staged, and a stage it has in flight. Its lock is held by this + # process, so an adopter that took it for dead would get past the + # held_by_caller check and sweep it; only its liveness saves it. live = _claim(base, config) - (Path(live.directory) / "marker").write_text("live") + live_directory = Path(live.directory) + _stage_into(live.directory, STAGED_BY_THE_LIVE) + live_packs = sorted(live_directory.rglob("*.dmi-pack.ready")) + assert len(live_packs) == len(STAGED_BY_THE_LIVE) // RECORDS_PER_PACK + live_open = live_packs[0].parent / ( + ".018f0000-0000-7000-8000-00000000beef.0badf00d.open") + live_open.write_bytes(b"half a pack") lock = _claim(base, config) service = _service(config, lock.directory) @@ -245,11 +255,20 @@ def test_a_sigkilled_process_spool_is_adopted_by_its_successor( assert not dead.exists() assert not old_claim.exists() assert fresh_claim.exists() + assert snapshot["live_siblings"] == 1, snapshot service.flush(60.0) + # Only the dead process's captures reach the catalog. assert _read_all(config) == _expected() finally: service.stop() - assert (Path(live.directory) / "marker").read_text() == "live" + # The live sibling is as it was: nothing swept, nothing uploaded. + assert sorted(live_directory.rglob("*.dmi-pack.ready")) == live_packs + assert live_open.read_bytes() == b"half a pack" + live_ids = {path.name.split(".")[0] for path in live_packs} + with STATE.lock: + uploaded = list(STATE.objects) + assert not [key for key in uploaded + if any(pack_id in key for pack_id in live_ids)] assert live.held assert lock.release_and_remove_if_empty() live.release() From efc42b0754939059e4fd31adbe3a680cbf5e8b75 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 04:11:50 -0400 Subject: [PATCH 13/43] Pin that no dead spool is uploaded while an adopted pack is owed adopt_sibling rechecks, for each sibling, that the lease is held and nothing is owed to the catalog (pending_index_ empty) before it uploads: the owner's decision that packs stay in the durable spool while the service cannot index them, applied to adoption. Nothing tested it. Deleting the line passed every CPU test and both live adoption tests; the only outage test cut the object store, not the catalog, and had one dead sibling. The new live test has two dead siblings, each staged by the real sink into its own rank directory. A switch in front of the fake S3 cuts a second switch, in front of ClickHouse, at the first upload: the first sibling's packs reach the object store, their index fails and they are owed in memory. The test asserts that the second sibling's ready packs are all still in its spool, adoption is owed, and flush() times out; then restores the catalog and asserts that both siblings are adopted and removed and every capture of both reads back from the catalog. Checked: with `|| !pending_index_.empty()` removed from adopt_sibling, the test fails on the second sibling's packs having left its spool; as the code stands it passes. --- tests/test_native_spool_adoption_live.py | 82 ++++++++++++++++++++++++ 1 file changed, 82 insertions(+) diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 9f096c0a0..7af4e507f 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -30,6 +30,7 @@ import signal import subprocess import sys +import threading import time from pathlib import Path @@ -185,6 +186,14 @@ def _claim_staging(parent: Path, name: str, *, age_s: float) -> Path: return staging +def _wait_for(predicate, timeout_s: float = 30.0) -> None: + deadline = time.monotonic() + timeout_s + while not predicate(): + if time.monotonic() >= deadline: + raise AssertionError("timed out waiting for the service") + time.sleep(0.05) + + def _read_all(config) -> dict: from dmi.storage.native_capture import NativeCaptureReader @@ -337,6 +346,79 @@ def _stage_into(directory: str, indexes) -> None: del lease, sink +def test_no_dead_spool_is_uploaded_while_an_adopted_pack_is_owed( + fake_s3, tmp_path): + """Owner decision 7, inside adoption. Two dead siblings; the catalog + goes away just after the first one's packs reach the object store, so + they cannot be indexed and are owed in memory. The second sibling's + packs must then stay in its spool, where a crash cannot lose them -- + not be uploaded into a list only this process remembers. Once the + catalog is back, both are indexed.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + # The servers are part of the catalog key, so the siblings are + # claimed through the same switch URLs as the service reaches. + clickhouse = _Switch(CLICKHOUSE_HOST, CLICKHOUSE_HTTP_PORT) + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, + clickhouse_host="127.0.0.1", + clickhouse_port=clickhouse.port) + halves = (range(0, 4), range(4, 10)) + siblings = [] + for indexes in halves: + sibling = _claim(base, config) + siblings.append(Path(sibling.directory)) + _stage_into(sibling.directory, indexes) + sibling.release() # its owner is gone + first, second = sorted(siblings) # adopted in this order + second_packs = sorted(second.rglob("*.dmi-pack.ready")) + assert second_packs + + cut = threading.Event() + + def _cut_the_catalog_at_the_first_upload(request: bytes) -> float: + if request.startswith(b"PUT ") and not cut.is_set(): + cut.set() + clickhouse.cut() + return 0.0 + + store.delay_requests(_cut_the_catalog_at_the_first_upload) + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + _wait_for(lambda: service.snapshot()["pending_index"] > 0) + snapshot = service.snapshot() + assert cut.is_set() + assert not sorted(first.rglob("*.dmi-pack.ready")), snapshot + # Nothing of the second left its spool while the first's packs + # were owed. + assert sorted(second.rglob("*.dmi-pack.ready")) == second_packs + assert snapshot["adoption_owed"] is True, snapshot + with pytest.raises(TimeoutError): + service.flush(2.0) + assert sorted(second.rglob("*.dmi-pack.ready")) == second_packs + + clickhouse.restore() + service.flush(60.0) + _wait_for(lambda: service.snapshot()["adopted_spools"] == 2, 60.0) + snapshot = service.snapshot() + assert snapshot["adoption_owed"] is False, snapshot + assert not first.exists() and not second.exists() + service.flush(60.0) + expected = {} + for indexes in halves: + for capture_id, tensor in _envelope(indexes).expected.items(): + expected[capture_id] = ( + tensor.contiguous().view(-1).numpy().tobytes()) + assert _read_all(config) == expected + finally: + service.stop() + lock.release_and_remove_if_empty() + + def test_a_sibling_whose_owner_dies_after_start_is_adopted_by_a_recheck( fake_s3, tmp_path): """A sibling still owned when the service starts -- a predecessor still From 99321954e25e7d9f47f6e52fe04003f9b0f772e0 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 04:14:50 -0400 Subject: [PATCH 14/43] Cover the owner-lock paths the review's surviving mutations showed untested Three mutations passed every test, and one test checked less than its name says: * SpoolClaim.release() replaced by a plain release(): every drained engine directory and its lock file kept. The wiring fake's lock always reports removal, and nothing else called release() on a real claim. test_releasing_a_claim_removes_its_directory_once_drained claims through claim_spool_directory, and asserts that a drained directory is removed (release() is True) and one holding a pack is kept, unowned, with the claim out of the held registry either way. * Spool::Open(held_by_caller) without its node-local check. test_held_by_caller_still_refuses_a_shared_filesystem holds a real lock and opens the service held_by_caller with the statfs seam reporting NFS: refused, and admitted with the override. * LockInPlace without its re-check that the locked file is still the one at the path (a remover unlinks the lock file before it removes a drained directory). This one is race-only, so the store gains a test seam, SetLockOpenHookForTesting, called between the open and the flock. test_spool_owner_lock.cpp unlinks the file there once: the take must open and lock again (two calls), and what it holds must be the file at the path -- ReadSpoolOwner sees it held and a rival TryAdopt is refused. * test_held_by_caller_with_nothing_held_is_refused passed a directory that does not exist, so it exercised realpath() failing rather than "exists, but nothing holds its lock". It now creates the directory, and checks again with a lock file nobody holds. Each new check was run against its mutation and failed (release() is False; DID NOT RAISE; calls == 1 and the lock not seen at the path). --- native/csrc/store/spool.cpp | 9 ++++ native/csrc/store/spool.h | 5 +++ tests/native/test_spool_owner_lock.cpp | 24 +++++++++++ tests/test_native_spool_ownership.py | 60 ++++++++++++++++++++++++-- 4 files changed, 95 insertions(+), 3 deletions(-) diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 2b1f96929..e30b48c40 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -52,6 +52,10 @@ constexpr uint32_t kSmb2SuperMagic = 0xFE534D42; constexpr uint32_t kFuseSuperMagic = 0x65735546; std::atomic g_filesystem_type_for_testing{-1}; +std::function& LockOpenHookForTesting() { + static auto* hook = new std::function; + return *hook; +} // Whether a recursive walk of a spool root is at /_refs, which no scan // enters. @@ -460,6 +464,10 @@ void SetFilesystemTypeForTesting(int64_t f_type) { g_filesystem_type_for_testing.store(f_type); } +void SetLockOpenHookForTesting(std::function hook) { + LockOpenHookForTesting() = std::move(hook); +} + SpoolStatus CheckNodeLocal(const std::string& dir, bool allow_shared_filesystem, std::string* error) { int64_t f_type = g_filesystem_type_for_testing.load(); @@ -584,6 +592,7 @@ SpoolStatus LockInPlace(const std::string& dir, int* fd_out, if (error) *error = "cannot open " + file + ": " + Errno(errno); return SpoolStatus::kIo; } + if (LockOpenHookForTesting()) LockOpenHookForTesting()(file); if (::flock(fd, LOCK_EX | LOCK_NB) != 0) { const int failure = errno; SpoolOwner owner; diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 77dc705cf..2de6b8bce 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -134,6 +134,11 @@ SpoolStatus CheckNodeLocal(const std::string& dir, // calling statfs(2). A negative value restores statfs. void SetFilesystemTypeForTesting(int64_t f_type); +// Test seam: taking an existing directory's lock calls `hook` with the lock +// file's path after opening the file and before locking it -- the window in +// which a remover can unlink it. An empty function removes the hook. +void SetLockOpenHookForTesting(std::function hook); + // The holder recorded in a directory's owner lock file. struct SpoolOwner { std::string host; diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 55af67186..f9680ff8e 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -583,6 +583,29 @@ void TestAdoptionLocksOnlyWhatExistsAndIsDead() { CHECK(fs::exists(base + "/kept/.owner.lock")); } +// (6b) A lock taken on a file a remover unlinked meanwhile guards nothing +// (ReleaseAndRemoveIfEmpty unlinks the lock file, then removes the +// directory), so it is let go and the file at the path is locked instead. +void TestALockOnAnUnlinkedFileIsTakenAgain() { + const std::string dir = FreshRoot("unlinked") + "/spool"; + fs::create_directories(dir); + std::ofstream(dir + "/.owner.lock") << ""; + int calls = 0; + dmi_store::SetLockOpenHookForTesting([&calls](const std::string& file) { + if (calls++ == 0) ::unlink(file.c_str()); // the remover's unlink + }); + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::TryAdopt(dir, &lock, &error) == SpoolStatus::kOk); + dmi_store::SetLockOpenHookForTesting(nullptr); + CHECK(calls == 2); + CHECK(lock.held()); + // What is held is the lock file at the path, which anyone else meets. + CHECK(dmi_store::ReadSpoolOwner(dir, nullptr)); + SpoolOwnerLock rival; + CHECK(SpoolOwnerLock::TryAdopt(dir, &rival, &error) == SpoolStatus::kOwned); +} + void TestANewDirectoryAppearsWithItsLockHeld() { // Created beside its lock file and renamed into place, so no scan of the // parent can meet the directory before its owner holds it. @@ -678,6 +701,7 @@ int main() { TestAnUnheldClaimStagingDirectoryRefusesNothing(); TestSharedFilesystemsAreRefusedUnlessAllowed(); TestAdoptionLocksOnlyWhatExistsAndIsDead(); + TestALockOnAnUnlinkedFileIsTakenAgain(); TestANewDirectoryAppearsWithItsLockHeld(); TestTheDirectoryLayout(); if (g_failures != 0) { diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 0ccf4df76..1a7442e64 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -285,6 +285,34 @@ def test_a_dropped_spool_claim_keeps_its_directory_owned(tmp_path): assert _store().spool_owner(directory) is None +def test_releasing_a_claim_removes_its_directory_once_drained(tmp_path): + """SpoolClaim.release() is how the engine lets go of its directory: a + drained one is removed with its lock file, one still holding a pack + stays for the next process on the node to adopt.""" + from dmi.storage import native_capture + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + sink = NativeSinkConfig(spool_root=str(tmp_path / "root")) + claim = claim_spool_directory(sink, _config()) + drained = Path(claim.directory) + (drained / "v1" / "tenant=t").mkdir(parents=True) + assert claim.release() is True + assert not drained.exists() + assert not claim.held + assert claim not in native_capture._HELD_SPOOL_CLAIMS + + claim = claim_spool_directory(sink, _config()) + kept = Path(claim.directory) + (kept / "v1").mkdir() + (kept / "v1" / "left.dmi-pack.ready").write_bytes(b"pack") + assert claim.release() is False + assert (kept / "v1" / "left.dmi-pack.ready").exists() + assert _store().spool_owner(str(kept)) is None + assert claim not in native_capture._HELD_SPOOL_CLAIMS + + @pytest.mark.parametrize("f_type, name", [ (0x6969, "NFS"), (0x0BD00BD0, "Lustre"), (0x19830326, "BeeGFS"), (0xFF534D42, "CIFS"), (0xFE534D42, "SMB2"), (0x65735546, "FUSE")]) @@ -402,9 +430,35 @@ def test_an_unknown_owner_lock_mode_is_refused(tmp_path): def test_held_by_caller_with_nothing_held_is_refused(tmp_path): - with pytest.raises(RuntimeError, match="held_by_caller"): - _service(_config(), tmp_path / "spool", - spool_owner_lock="held_by_caller") + """The directory exists -- so this is the check that nothing holds its + lock, not a failure to resolve a path -- first with no lock file, then + with one nobody holds.""" + directory = tmp_path / "spool" + directory.mkdir() + with pytest.raises(RuntimeError, match="held_by_caller, but nothing holds"): + _service(_config(), directory, spool_owner_lock="held_by_caller") + _store().SpoolOwnerLock(str(directory)).release() + assert (directory / ".owner.lock").exists() + with pytest.raises(RuntimeError, match="held_by_caller, but nothing holds"): + _service(_config(), directory, spool_owner_lock="held_by_caller") + + +def test_held_by_caller_still_refuses_a_shared_filesystem(tmp_path): + """The node-local check applies to a Spool opened held_by_caller too: + the caller's lock was taken on a local disk, but what the Spool opens + is judged again.""" + store = _store() + directory = tmp_path / "spool" + with store.SpoolOwnerLock(str(directory)): + store._set_spool_filesystem_type_for_testing(NFS_SUPER_MAGIC) + try: + with pytest.raises(RuntimeError, match="is on NFS .*node-local"): + _service(_config(), directory, + spool_owner_lock="held_by_caller") + _service(_config(), directory, spool_owner_lock="held_by_caller", + spool_allow_shared_filesystem=True) + finally: + store._set_spool_filesystem_type_for_testing(None) class _OtherProcessHolder: From ec5779687cd194b18efad029cad11dbb94de2e3c Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 04:16:51 -0400 Subject: [PATCH 15/43] Refuse GPFS, 9p, AFS and OrangeFS spool roots as well The previous commit brought the node-local refusal to NFS, Lustre, BeeGFS, CIFS/SMB2 and FUSE. The review named more network filesystems that clusters mount where a spool root could land, each of which lets a successor on another node take a live rank directory for a dead one -- adoption then sweeps its .open files, the race B6 removes: * GPFS / IBM Storage Scale (0x47504653): its fcntl locks are cluster-wide, but flock is node-local (NCAR's experience on its GPFS home filesystem: a flock on one login node does not exclude another); * 9p (0x01021997), whose flock the client only mirrors to a server that may not enforce it (QEMU's does not); * AFS, OpenAFS's (0x5346414F) and kAFS's (0x6B414653); * OrangeFS (0x20030528), which has no cross-client flock. The rule the plan states is that a spool root is node-local, which none of these is; allow_shared_filesystem still admits any of them. CephFS, whose flock is coherent across clients, is left alone, as are the local ones (ext4, xfs, btrfs, zfs, tmpfs, overlayfs). Tests, red before: test_spool_owner_lock.cpp names and refuses each new magic through the statfs seam (the override admitting it), and btrfs and zfs pass; test_native_spool_ownership.py refuses each through the binding, by name. --- docs/integration-api-v1.md | 4 ++-- native/csrc/store/spool.cpp | 24 ++++++++++++++++++------ native/csrc/store/spool.h | 11 ++++++----- src/dmi/storage/native_capture.py | 9 +++++---- tests/native/test_spool_owner_lock.cpp | 19 +++++++++++++------ tests/test_native_spool_ownership.py | 4 +++- 6 files changed, 47 insertions(+), 24 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index 414f7de6d..e805c50c1 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -204,8 +204,8 @@ directories under the same catalog key whose owners have died: their stale `.open` files are swept, their ready packs uploaded and indexed, and the directory removed, so a crashed process's packs reach the catalog through the next one on the node, whatever run it belongs to. The spool root must be -node-local: NFS, Lustre, BeeGFS, CIFS/SMB2 and FUSE are refused (by statfs -`f_type`) unless `NativeSinkConfig.spool_allow_shared_filesystem`, which a FUSE +node-local: NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, GPFS, 9p, AFS and OrangeFS +are refused (by statfs `f_type`) unless `NativeSinkConfig.spool_allow_shared_filesystem`, which a FUSE filesystem that is itself local, such as fuse-overlayfs, needs too. Without `capture_storage_config` the sink owns `spool_root` itself. With an explicit `record_sink`, the service drains `spool_root` as that sink writes it, diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index e30b48c40..74d8c1209 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -38,18 +38,25 @@ constexpr const char* kClaimStagingSuffix = ".creating"; // is ever staged under it. constexpr const char* kRefsDirectory = "_refs"; -// statfs(2) f_type values of filesystems whose flock does not keep out a -// process on another node (linux/magic.h has NFS, SMB2, CIFS and FUSE; -// Lustre's and BeeGFS's are their own). BeeGFS keeps flock client-local -// unless tuneUseGlobalFileLocks is set. FUSE covers network filesystems -// (sshfs, s3fs, gcsfuse, GlusterFS) and local ones alike, and f_type cannot -// tell them apart, so a local one needs the override. +// statfs(2) f_type values of network filesystems, whose flock does not +// keep out a process on another node -- or is not a place a node-local +// spool can be (linux/magic.h has NFS, SMB2, CIFS, FUSE, 9p and both AFS +// values; the others are their own). BeeGFS keeps flock client-local +// unless tuneUseGlobalFileLocks is set, and GPFS (IBM Storage Scale) keeps +// it node-local. FUSE covers network filesystems (sshfs, s3fs, gcsfuse, +// GlusterFS) and local ones alike, and f_type cannot tell them apart, so a +// local one needs the override. AFS is OpenAFS's and kAFS's. constexpr uint32_t kNfsSuperMagic = 0x6969; constexpr uint32_t kLustreSuperMagic = 0x0BD00BD0; constexpr uint32_t kBeeGfsSuperMagic = 0x19830326; constexpr uint32_t kCifsSuperMagic = 0xFF534D42; constexpr uint32_t kSmb2SuperMagic = 0xFE534D42; constexpr uint32_t kFuseSuperMagic = 0x65735546; +constexpr uint32_t kGpfsSuperMagic = 0x47504653; +constexpr uint32_t kV9fsMagic = 0x01021997; +constexpr uint32_t kAfsSuperMagic = 0x5346414F; +constexpr uint32_t kAfsFsMagic = 0x6B414653; +constexpr uint32_t kOrangeFsSuperMagic = 0x20030528; std::atomic g_filesystem_type_for_testing{-1}; std::function& LockOpenHookForTesting() { @@ -456,6 +463,11 @@ const char* SharedFilesystemName(int64_t f_type) { case kCifsSuperMagic: return "CIFS"; case kSmb2SuperMagic: return "SMB2"; case kFuseSuperMagic: return "FUSE"; + case kGpfsSuperMagic: return "GPFS"; + case kV9fsMagic: return "9p"; + case kAfsSuperMagic: return "AFS"; + case kAfsFsMagic: return "AFS"; + case kOrangeFsSuperMagic: return "OrangeFS"; default: return nullptr; } } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 2de6b8bce..c6e90cbec 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -39,9 +39,10 @@ // check runs after the lock is taken, so of two processes taking an outer // and a nested directory at once, at least one is refused. /_refs/ is // never scanned: the upload handoff's ref files live there (plan section -// 2.4). The spool must be node-local: NFS, Lustre, BeeGFS, CIFS/SMB2 and -// FUSE are refused by statfs f_type unless allow_shared_filesystem is set, -// since none guarantees a flock that excludes a process on another node. +// 2.4). The spool must be node-local: NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, +// GPFS, 9p, AFS and OrangeFS are refused by statfs f_type unless +// allow_shared_filesystem is set, since none guarantees a flock that +// excludes a process on another node. // // The Python DurablePackSpool (spool.py) takes no lock, and its recover() // deletes every .open file under its root; the C++ spool is deliberately @@ -120,8 +121,8 @@ inline const char* SpoolStatusName(SpoolStatus s) { } // The statfs f_type names of the shared filesystems a spool refuses: "NFS", -// "Lustre", "BeeGFS", "CIFS", "SMB2", "FUSE", or nullptr for any other. -// The list needs maintenance as +// "Lustre", "BeeGFS", "CIFS", "SMB2", "FUSE", "GPFS", "9p", "AFS", +// "OrangeFS", or nullptr for any other. The list needs maintenance as // deployments meet new ones. const char* SharedFilesystemName(int64_t f_type); diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index 75bff9b45..bbec90521 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -137,10 +137,11 @@ class NativeSinkConfig: under it, ``//r-/`` (see :func:`claim_spool_directory`), and its storage service adopts the directories of dead processes beside it. Without one, the sink owns - ``spool_root`` itself. A root on NFS, Lustre, BeeGFS, CIFS/SMB2 or FUSE - is refused unless ``spool_allow_shared_filesystem``: flock there does - not keep out a process on another node (a FUSE filesystem that is local, - such as fuse-overlayfs, needs the override too). + ``spool_root`` itself. A root on NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, + GPFS, 9p, AFS or OrangeFS is refused unless + ``spool_allow_shared_filesystem``: flock there does not keep out a + process on another node (a FUSE filesystem that is local, such as + fuse-overlayfs, needs the override too). """ spool_root: str diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index f9680ff8e..4c390b825 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -12,9 +12,9 @@ // it forked without exec does not keep it. // 4. Nesting: a directory under a HELD one, or containing one with a lock // file, is refused -- also when two processes take the pair at once. -// 5. The node-local check refuses NFS, Lustre, BeeGFS, CIFS/SMB2 and FUSE -// by statfs f_type, unless explicitly allowed (a test seam stands in -// for statfs). +// 5. The node-local check refuses NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, +// GPFS, 9p, AFS and OrangeFS by statfs f_type, unless explicitly +// allowed (a test seam stands in for statfs). // 6. Adoption's try-lock never creates a directory, and a released // directory that holds nothing but its lock file can be removed. // 7. The directory layout of the plan's section 2.3. @@ -497,12 +497,17 @@ void TestSharedFilesystemsAreRefusedUnlessAllowed() { const std::string root = FreshRoot("statfs") + "/spool"; // Each one's flock does not keep out a process on another node: NFS and // Lustre (the plan's two), BeeGFS (client-local unless - // tuneUseGlobalFileLocks), CIFS/SMB2, and FUSE, which cannot tell sshfs, - // s3fs, gcsfuse or GlusterFS from a local filesystem. + // tuneUseGlobalFileLocks), CIFS/SMB2, FUSE, which cannot tell sshfs, + // s3fs, gcsfuse or GlusterFS from a local filesystem, GPFS (IBM Storage + // Scale, whose flock is node-local), 9p, AFS (OpenAFS and kAFS) and + // OrangeFS -- network filesystems all, where a spool is never node-local. const std::vector> shared = { {0x6969, "NFS"}, {0x0BD00BD0, "Lustre"}, {0x19830326, "BeeGFS"}, {0xFF534D42, "CIFS"}, - {0xFE534D42, "SMB2"}, {0x65735546, "FUSE"}}; + {0xFE534D42, "SMB2"}, {0x65735546, "FUSE"}, + {0x47504653, "GPFS"}, {0x01021997, "9p"}, + {0x5346414F, "AFS"}, {0x6B414653, "AFS"}, + {0x20030528, "OrangeFS"}}; for (const auto& [magic, name] : shared) { const char* named = dmi_store::SharedFilesystemName(magic); CHECK(named != nullptr && std::string(named) == name); @@ -511,6 +516,8 @@ void TestSharedFilesystemsAreRefusedUnlessAllowed() { CHECK(dmi_store::SharedFilesystemName(0x58465342) == nullptr); // xfs CHECK(dmi_store::SharedFilesystemName(0x794C7630) == nullptr); // overlayfs CHECK(dmi_store::SharedFilesystemName(0x01021994) == nullptr); // tmpfs + CHECK(dmi_store::SharedFilesystemName(0x9123683E) == nullptr); // btrfs + CHECK(dmi_store::SharedFilesystemName(0x2FC12FC1) == nullptr); // zfs std::string error; for (const auto& [magic, name] : shared) { diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 1a7442e64..c1a2520e0 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -315,7 +315,9 @@ def test_releasing_a_claim_removes_its_directory_once_drained(tmp_path): @pytest.mark.parametrize("f_type, name", [ (0x6969, "NFS"), (0x0BD00BD0, "Lustre"), (0x19830326, "BeeGFS"), - (0xFF534D42, "CIFS"), (0xFE534D42, "SMB2"), (0x65735546, "FUSE")]) + (0xFF534D42, "CIFS"), (0xFE534D42, "SMB2"), (0x65735546, "FUSE"), + (0x47504653, "GPFS"), (0x01021997, "9p"), (0x5346414F, "AFS"), + (0x6B414653, "AFS"), (0x20030528, "OrangeFS")]) def test_each_shared_filesystem_is_refused_by_name(tmp_path, f_type, name): store = _store() store._set_spool_filesystem_type_for_testing(f_type) From a1ded469b4dbe06c0e2fd1adf8e60b57d32ce1c3 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 04:19:18 -0400 Subject: [PATCH 16/43] Keep spool.cpp portable to macOS: guard statfs and the no-replace rename spool.cpp was portable POSIX at d3fe7c2. The owner lock added , which only Linux has (macOS declares statfs in ), read statfs's Linux-only f_type magic, and fell back to plain rename() where RENAME_NOREPLACE is missing -- which on macOS silently replaces an empty target directory. Three CPU tests compile spool.cpp directly and carry Homebrew OpenSSL paths for macOS (test_native_spool_reservations.py, test_native_pack_sink_timeout.py, test_native_spool_owner_lock_unit.py); on a Mac they would now fail to compile, and CI, Ubuntu-only, would not notice. * only on Linux; elsewhere and . * The node-local check reads f_type on Linux, as before, and elsewhere f_fstypename, refusing nfs, smbfs, afpfs, webdav, lustre, afs and the FUSE mounts (macfuse, osxfuse, FreeBSD's fusefs[.]) under the same override. The statfs test seam still takes a Linux magic. * The new directory's rename uses renameat2(RENAME_NOREPLACE) where it exists (Linux, unchanged), renamex_np(RENAME_EXCL) on macOS, and otherwise refuses a target that exists before a plain rename(), so an existing directory still goes to the lock-in-place path. /proc/self/fdinfo, which held_by_caller reads, is already optional: where /proc cannot be read it falls back to the owner record. Not compiled on macOS here. Each non-Linux branch was syntax-checked on Linux with __linux__ undefined against a stand-in (and the plain-rename branch forced), and the Linux build and the spool CPU tests (66) pass unchanged. --- native/csrc/store/spool.cpp | 76 +++++++++++++++++++++++++++++++------ 1 file changed, 65 insertions(+), 11 deletions(-) diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 74d8c1209..096d8ec44 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -19,7 +19,13 @@ #include #include #include +#if defined(__linux__) #include +#else +// statfs(2) with f_fstypename: macOS and the BSDs. +#include +#include +#endif #include namespace dmi_store { @@ -59,6 +65,26 @@ constexpr uint32_t kAfsFsMagic = 0x6B414653; constexpr uint32_t kOrangeFsSuperMagic = 0x20030528; std::atomic g_filesystem_type_for_testing{-1}; + +#if !defined(__linux__) +// Where statfs names the filesystem rather than giving Linux's magic: the +// same refusal, by f_fstypename (FreeBSD spells a FUSE mount +// "fusefs."). +const char* SharedFilesystemTypeName(const char* name) { + static const char* const kShared[][2] = { + {"nfs", "NFS"}, {"smbfs", "SMB"}, {"afpfs", "AFP"}, + {"webdav", "WebDAV"}, {"lustre", "Lustre"}, {"macfuse", "FUSE"}, + {"osxfuse", "FUSE"}, {"fusefs", "FUSE"}, {"afs", "AFS"}}; + for (const auto& entry : kShared) { + const size_t n = std::strlen(entry[0]); + if (std::strncmp(name, entry[0], n) == 0 && + (name[n] == '\0' || name[n] == '.')) { + return entry[1]; + } + } + return nullptr; +} +#endif std::function& LockOpenHookForTesting() { static auto* hook = new std::function; return *hook; @@ -482,23 +508,38 @@ void SetLockOpenHookForTesting(std::function hook) { SpoolStatus CheckNodeLocal(const std::string& dir, bool allow_shared_filesystem, std::string* error) { - int64_t f_type = g_filesystem_type_for_testing.load(); - if (f_type < 0) { + const int64_t f_type = g_filesystem_type_for_testing.load(); + const char* shared = nullptr; + std::string seen; // what statfs said, for the refusal + const auto magic = [](int64_t value) { + char text[48]; + std::snprintf(text, sizeof(text), "statfs f_type 0x%llx", + static_cast(value)); + return std::string(text); + }; + if (f_type >= 0) { + shared = SharedFilesystemName(f_type); + seen = magic(f_type); + } else { struct statfs info{}; if (::statfs(dir.c_str(), &info) != 0) { if (error) *error = "cannot statfs " + dir + ": " + Errno(errno); return SpoolStatus::kIo; } - f_type = static_cast(static_cast(info.f_type)); +#if defined(__linux__) + const int64_t type = + static_cast(static_cast(info.f_type)); + shared = SharedFilesystemName(type); + seen = magic(type); +#else + shared = SharedFilesystemTypeName(info.f_fstypename); + seen = std::string("statfs f_fstypename ") + info.f_fstypename; +#endif } - const char* shared = SharedFilesystemName(f_type); if (shared == nullptr || allow_shared_filesystem) return SpoolStatus::kOk; if (error) { - char magic[32]; - std::snprintf(magic, sizeof(magic), "0x%llx", - static_cast(f_type)); - *error = "spool directory " + dir + " is on " + shared + - " (statfs f_type " + magic + "): a spool must be node-local, " + *error = "spool directory " + dir + " is on " + shared + " (" + seen + + "): a spool must be node-local, " "since its owner lock (flock) does not keep out a process on " "another node there. Use a local disk, or set " "allow_shared_filesystem if no process on another node can " @@ -675,11 +716,24 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, WriteOwnerRecord(fd); ::fsync(fd); FsyncDir(staging, nullptr); -#ifdef RENAME_NOREPLACE + // A rename that refuses an existing target: plain rename() silently + // replaces an empty directory. +#if defined(RENAME_NOREPLACE) const int renamed = ::renameat2(AT_FDCWD, staging.c_str(), AT_FDCWD, dir.c_str(), RENAME_NOREPLACE); +#elif defined(__APPLE__) && defined(RENAME_EXCL) + const int renamed = + ::renamex_np(staging.c_str(), dir.c_str(), RENAME_EXCL); #else - const int renamed = ::rename(staging.c_str(), dir.c_str()); + // Neither: refuse a target that exists before renaming. The window + // left is between the check and the rename, and only another claim of + // the same fresh incarnation could fall into it. + int renamed = -1; + if (::access(dir.c_str(), F_OK) == 0) { + errno = EEXIST; + } else { + renamed = ::rename(staging.c_str(), dir.c_str()); + } #endif if (renamed != 0) { const int failure = errno; From 1cf61b58ee29d5d1cd214f287861cc1c6d522aa6 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 04:23:35 -0400 Subject: [PATCH 17/43] Charge what dead incarnations left beside a spool against its budget Before the section 2.3 layout every restart reused one spool_root, and Spool::Open counted what was already there against spool_max_bytes, so the node's spool stayed within one budget however often a job restarted. Now every create_record_runtime claims a fresh r- directory, and its Spool counts only that directory. So while uploads are blocked (the object store down, the catalog up, so the successor still starts and its adoption is merely owed) each crash-restart adds a whole spool_max_bytes: the sink fills its directory, fails the worker ("spool stage failed"), HF's on_failure=raise ends the job, torchelastic / Slurm --requeue / a k8s restartPolicy starts the next incarnation, and so on until the local disk is full. The review measured four incarnations under one base with spool_max_bytes=65536 holding 63748 bytes each, 254992 in all. The plan's per-node budget line was not implemented alongside the layout that causes this. SpoolConfig.charge_dead_siblings: the root is a rank directory, and what its sibling rank directories hold (ready packs and temp files, _refs/ aside) counts against max_bytes too. A sibling another live process holds is left out -- that is its own budget -- while one nobody holds (a dead incarnation waiting to be adopted) or one THIS process holds (its service adopting it) counts. The charge is refreshed wherever the committed account is: at Open, and in the reconcile a stage runs before it refuses, so capacity comes back as adoption drains them. The kFull message says how much of the total is the dead directories', and the snapshot carries sibling_bytes. The engine's sink sets it (SinkConfig.spool_charge_dead_siblings, NativePackSink's charge_dead_siblings, create_native_pack_sink's) for the directory it claims; the sink-only and explicit-sink modes, which have no layout, do not. A deployment that runs several live ranks on a node still gives each its own spool_max_bytes; dividing a node budget among them is the plan's per-node line, still to do. Tests, red before (the fields did not exist): * test_spool_owner_lock.cpp: beside a dead sibling holding 300 bytes, a 450-byte spool admits one 100-byte pack and refuses the next, naming the dead bytes; once the sibling's packs are gone the stage goes through; without the charge the same budget admits 400; a sibling a live child process holds is not charged; one this process holds is; * test_native_spool_ownership.py: the real sink in a claimed directory beside a dead one holding 1 MiB, with 1 MiB + 256 bytes of budget, refuses its first pack with charge_dead_siblings and stages it without; * test_native_capture_storage_wiring.py: the engine passes charge_dead_siblings=True to its sink with a claim, False without. --- docs/integration-api-v1.md | 6 +- native/csrc/sink/bindings_sink.cpp | 10 +- native/csrc/sink/pack_sink.cpp | 1 + native/csrc/sink/pack_sink.h | 5 + native/csrc/store/spool.cpp | 84 ++++++++++++--- native/csrc/store/spool.h | 23 +++++ src/dmi/engine.py | 5 +- src/dmi/storage/capture/native_sink.py | 9 +- src/dmi/storage/native_capture.py | 7 +- tests/native/test_spool_owner_lock.cpp | 108 ++++++++++++++++++++ tests/test_native_capture_storage_wiring.py | 3 + tests/test_native_spool_ownership.py | 37 +++++++ 12 files changed, 278 insertions(+), 20 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index e805c50c1..b2c410778 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -203,7 +203,11 @@ refused, naming the holder's pid and host. At start the service adopts the directories under the same catalog key whose owners have died: their stale `.open` files are swept, their ready packs uploaded and indexed, and the directory removed, so a crashed process's packs reach the catalog through the -next one on the node, whatever run it belongs to. The spool root must be +next one on the node, whatever run it belongs to. `spool_max_bytes` bounds the +directory together with what those dead directories still hold (a live +process's directory is its own budget), so restarts while uploads are blocked +cannot each add a whole budget; the room comes back as they are adopted. The +spool root must be node-local: NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, GPFS, 9p, AFS and OrangeFS are refused (by statfs `f_type`) unless `NativeSinkConfig.spool_allow_shared_filesystem`, which a FUSE filesystem that is itself local, such as fuse-overlayfs, needs too. Without diff --git a/native/csrc/sink/bindings_sink.cpp b/native/csrc/sink/bindings_sink.cpp index 915fc499e..e341b8f75 100644 --- a/native/csrc/sink/bindings_sink.cpp +++ b/native/csrc/sink/bindings_sink.cpp @@ -159,7 +159,8 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { const std::string& overload, std::optional admission_timeout_s, const std::string& owner_lock, - bool allow_shared_filesystem) { + bool allow_shared_filesystem, + bool charge_dead_siblings) { dmi_sink::SinkConfig config; if (!dmi_store::ParseOwnerLock(owner_lock, &config.spool_owner_lock)) { @@ -168,6 +169,7 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { owner_lock + "'"); } config.spool_allow_shared_filesystem = allow_shared_filesystem; + config.spool_charge_dead_siblings = charge_dead_siblings; config.overload = ParseOverload(overload); config.admission_timeout_s = ParseAdmissionTimeout(admission_timeout_s); @@ -200,7 +202,11 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { // a SpoolOwnerLock on it, as the engine does around its sink and // storage service. py::arg("owner_lock") = "take", - py::arg("allow_shared_filesystem") = false) + py::arg("allow_shared_filesystem") = false, + // spool_root is a rank directory of the spool layout, and the + // dead incarnations' packs beside it count against + // spool_max_bytes, as the engine's claimed directory does. + py::arg("charge_dead_siblings") = false) .def("attach", [](std::shared_ptr self) { // Simulates engine ownership for tests (the real engine takes diff --git a/native/csrc/sink/pack_sink.cpp b/native/csrc/sink/pack_sink.cpp index 6398bdb6c..42aacd4fa 100644 --- a/native/csrc/sink/pack_sink.cpp +++ b/native/csrc/sink/pack_sink.cpp @@ -98,6 +98,7 @@ std::string PackSink::Start(std::string* spool_error) { spool_config.max_bytes = config_.spool_max_bytes; spool_config.owner_lock = config_.spool_owner_lock; spool_config.allow_shared_filesystem = config_.spool_allow_shared_filesystem; + spool_config.charge_dead_siblings = config_.spool_charge_dead_siblings; std::string error; const dmi_store::SpoolStatus st = dmi_store::Spool::Open(spool_config, &spool_, &error); diff --git a/native/csrc/sink/pack_sink.h b/native/csrc/sink/pack_sink.h index ab57613e6..2161e75df 100644 --- a/native/csrc/sink/pack_sink.h +++ b/native/csrc/sink/pack_sink.h @@ -73,6 +73,11 @@ struct SinkConfig { // holds one SpoolOwnerLock and opens both with kHeldByCaller. dmi_store::OwnerLock spool_owner_lock = dmi_store::OwnerLock::kTake; bool spool_allow_shared_filesystem = false; + // spool_root is a rank directory of the section 2.3 layout, and what the + // dead incarnations beside it still hold counts against spool_max_bytes + // (SpoolConfig::charge_dead_siblings). The engine sets it for the + // directory it claims. + bool spool_charge_dead_siblings = false; // Pack assembler workers. Records route by scope hash // (tenant, session, producer_rank), so one scope always lands on one // worker: per-scope ordering and single-scope packs are preserved at any diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 096d8ec44..dada9c105 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -1094,6 +1094,8 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { } out->root_ = resolved; out->max_bytes_ = config.max_bytes; + out->charge_dead_siblings_ = config.charge_dead_siblings; + out->sibling_bytes_ = 0; out->committed_bytes_ = 0; out->committed_entries_ = 0; out->reserved_bytes_ = 0; @@ -1125,9 +1127,53 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { } } out->peak_bytes_ = out->committed_bytes_; + if (out->charge_dead_siblings_) { + out->sibling_bytes_ = out->ChargedSiblingBytes(); + } return SpoolStatus::kOk; } +uint64_t Spool::ChargedSiblingBytes() const { + const fs::path own(root_); + uint64_t bytes = 0; + std::error_code ec; + for (fs::directory_iterator it(own.parent_path(), ec), end; + !ec && it != end; it.increment(ec)) { + uint64_t rank = 0; + std::string incarnation; + std::error_code type_ec; + if (it->path() == own || it->is_symlink(type_ec) || + !it->is_directory(type_ec) || + !ParseSpoolRankDirectoryName(it->path().filename().string(), &rank, + &incarnation)) { + continue; + } + const std::string sibling = it->path().string(); + // Another live process's directory is its own budget. + if (ReadSpoolOwner(sibling, nullptr) && + !SpoolOwnedByThisProcess(sibling)) { + continue; + } + std::error_code walk_ec; + for (fs::recursive_directory_iterator walk(sibling, walk_ec), last; + !walk_ec && walk != last; walk.increment(walk_ec)) { + if (AtRefsDirectory(walk)) { + walk.disable_recursion_pending(); + continue; + } + std::error_code entry_ec; + if (!walk->is_regular_file(entry_ec)) continue; + const std::string name = walk->path().filename().string(); + if (!HasSuffix(name, kReadySuffix) && !HasSuffix(name, kOpenSuffix)) { + continue; + } + const uint64_t size = walk->file_size(entry_ec); + if (!entry_ec) bytes += size; + } + } + return bytes; +} + bool Spool::AccountReadyLocked(const std::string& path, uint64_t object_bytes) { if (accounted_ready_.count(path) != 0) return false; @@ -1194,11 +1240,27 @@ void Spool::ReconcileCommittedLocked() { // The path ledger is rebuilt with the aggregate it describes, so the two // never disagree about which files the committed account holds. accounted_ready_ = std::move(seen_ready); + // And the dead siblings' charge with it, so what adoption has drained + // since is capacity again. + if (charge_dead_siblings_) sibling_bytes_ = ChargedSiblingBytes(); // The scan can raise the committed total (files another object wrote), and // peak_bytes_ must never read below what the account holds right now. peak_bytes_ = std::max(peak_bytes_, committed_bytes_ + reserved_bytes_); } +std::string Spool::FullMessage(uint64_t n) const { + std::string message = + "spool byte limit exceeded: " + + std::to_string(committed_bytes_ + reserved_bytes_ + sibling_bytes_ + n) + + " > " + std::to_string(max_bytes_); + if (sibling_bytes_ > 0) { + message += " (" + std::to_string(sibling_bytes_) + + " bytes of it in the dead spool directories beside this one, " + "still to be adopted)"; + } + return message; +} + void Spool::SetStageHookForTesting(std::function hook) { std::lock_guard lock(mutex_); stage_hook_for_testing_ = std::move(hook); @@ -1277,14 +1339,12 @@ SpoolStatus Spool::Stage(const std::string& pack_id, uint64_t created_at_ns, return SpoolStatus::kConflict; } } - if (committed_bytes_ + reserved_bytes_ + n > max_bytes_) { + if (committed_bytes_ + reserved_bytes_ + sibling_bytes_ + n > + max_bytes_) { ReconcileCommittedLocked(); - if (committed_bytes_ + reserved_bytes_ + n > max_bytes_) { - if (error) { - *error = "spool byte limit exceeded: " + - std::to_string(committed_bytes_ + reserved_bytes_ + n) + - " > " + std::to_string(max_bytes_); - } + if (committed_bytes_ + reserved_bytes_ + sibling_bytes_ + n > + max_bytes_) { + if (error) *error = FullMessage(n); return SpoolStatus::kFull; } } @@ -1347,12 +1407,9 @@ SpoolStatus Spool::Stage(const std::string& pack_id, uint64_t created_at_ns, // ready file -- a state Python cannot reach at all, since it holds its // lock across the whole of stage()), while a serial retry under a lowered // cap is admitted the way the reference admits it. - if (reserved_bytes_ > 0 && committed_bytes_ + reserved_bytes_ > max_bytes_) { - if (error) { - *error = "spool byte limit exceeded: " + - std::to_string(committed_bytes_ + reserved_bytes_) + " > " + - std::to_string(max_bytes_); - } + if (reserved_bytes_ > 0 && + committed_bytes_ + reserved_bytes_ + sibling_bytes_ > max_bytes_) { + if (error) *error = FullMessage(0); return SpoolStatus::kFull; } if (AccountReadyLocked(ready, n)) ++generation_; @@ -1624,6 +1681,7 @@ SpoolSnapshot Spool::Snapshot() const { snapshot.bytes = committed_bytes_ + reserved_bytes_; snapshot.peak_bytes = peak_bytes_; snapshot.max_bytes = max_bytes_; + snapshot.sibling_bytes = sibling_bytes_; return snapshot; } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index c6e90cbec..8dc652482 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -78,6 +78,18 @@ struct SpoolConfig { // FUSE filesystem that is local after all. Only safe when every process // that could open the directory runs on this node. bool allow_shared_filesystem = false; + // The root is a rank directory of the section 2.3 layout, and what its + // SIBLING rank directories hold (ready packs and temp files) counts + // against max_bytes as well -- except a sibling another live process + // holds, which is that process's own budget. A sibling nobody holds is a + // dead incarnation's, waiting to be adopted; one THIS process holds is + // being adopted by its service. Every process start gets a fresh rank + // directory, so without this each crash-restart while uploads are + // blocked would add a whole max_bytes to the node's spool; before the + // layout every restart reused one directory and one budget. The charge is + // refreshed wherever the committed account is (Open, and before a stage + // is refused), so the capacity comes back as adoption drains them. + bool charge_dead_siblings = false; }; struct StagedPack { @@ -95,6 +107,9 @@ struct SpoolSnapshot { uint64_t bytes = 0; uint64_t peak_bytes = 0; uint64_t max_bytes = 0; + // charge_dead_siblings: what the sibling directories were charged, as of + // the last refresh; capacity is judged against bytes plus this. + uint64_t sibling_bytes = 0; }; enum class SpoolStatus { @@ -334,9 +349,17 @@ class Spool { // stale-high. The retry/EEXIST-loser paths run it too, so a charge that // would exceed the cap is judged against the same durable truth. void ReconcileCommittedLocked(); + // charge_dead_siblings: the bytes of ready and temp files in the sibling + // rank directories no other live process holds. + uint64_t ChargedSiblingBytes() const; + // The kFull refusal of a stage of `n` bytes. `mutex_` must be held. + std::string FullMessage(uint64_t n) const; std::string root_; uint64_t max_bytes_ = 0; + bool charge_dead_siblings_ = false; + // What ChargedSiblingBytes() found last. Under `mutex_`. + uint64_t sibling_bytes_ = 0; // Held for the object's life under OwnerLock::kTake; empty otherwise. SpoolOwnerLock owner_lock_; mutable std::mutex mutex_; diff --git a/src/dmi/engine.py b/src/dmi/engine.py index eb6649504..b36316839 100644 --- a/src/dmi/engine.py +++ b/src/dmi/engine.py @@ -585,12 +585,15 @@ def _attach_record_runtime( from .storage.capture.native_sink import create_native_pack_sink # Into the directory the engine claimed for the service, under - # its lock; without a service the sink owns the spool root. + # its lock, with what dead incarnations left beside it charged + # against its budget; without a service the sink owns the spool + # root. claim = self._spool_claim record_sink = create_native_pack_sink( self._capture_sink_config, spool_root=None if claim is None else claim.directory, owner_lock="take" if claim is None else "held_by_caller", + charge_dead_siblings=claim is not None, ).native_sink _native_engine = _native_module() diff --git a/src/dmi/storage/capture/native_sink.py b/src/dmi/storage/capture/native_sink.py index 3482cc4a3..78d4d6889 100644 --- a/src/dmi/storage/capture/native_sink.py +++ b/src/dmi/storage/capture/native_sink.py @@ -94,6 +94,9 @@ class NativePackSinkHandle: sink's life, ``"held_by_caller"`` when the caller holds a ``SpoolOwnerLock`` on it -- as the engine does around its sink and its storage service, which would otherwise refuse each other. + ``charge_dead_siblings`` says the spool root is a rank directory of the + spool layout, and what the dead incarnations beside it still hold + counts against ``spool_max_bytes`` (the engine's claimed directory). """ def __init__( @@ -102,12 +105,14 @@ def __init__( *, spool_root: str | None = None, owner_lock: str = "take", + charge_dead_siblings: bool = False, ) -> None: module = _load_native_sink_extension() self._native_sink = module.NativePackSink( spool_root=config.spool_root if spool_root is None else spool_root, owner_lock=owner_lock, allow_shared_filesystem=config.spool_allow_shared_filesystem, + charge_dead_siblings=charge_dead_siblings, layout=LAYOUT_NAME, num_workers=config.num_workers, max_queue_records=config.max_queue_records, @@ -141,13 +146,15 @@ def create_native_pack_sink( *, spool_root: str | None = None, owner_lock: str = "take", + charge_dead_siblings: bool = False, ) -> NativePackSinkHandle: """Select the native capture writer for one record runtime.""" if not isinstance(config, NativeSinkConfig): raise TypeError("config must be a NativeSinkConfig") return NativePackSinkHandle(config, spool_root=spool_root, - owner_lock=owner_lock) + owner_lock=owner_lock, + charge_dead_siblings=charge_dead_siblings) __all__ = [ diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index bbec90521..2d129d6f0 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -136,8 +136,11 @@ class NativeSinkConfig: ``capture_storage_config`` the engine spools into a directory of its own under it, ``//r-/`` (see :func:`claim_spool_directory`), and its storage service adopts the - directories of dead processes beside it. Without one, the sink owns - ``spool_root`` itself. A root on NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, + directories of dead processes beside it. ``spool_max_bytes`` then + bounds this directory together with what the dead incarnations beside + it still hold, so crash-restarts while uploads are blocked cannot each + add a whole budget. Without one, the sink owns ``spool_root`` itself. + A root on NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, GPFS, 9p, AFS or OrangeFS is refused unless ``spool_allow_shared_filesystem``: flock there does not keep out a process on another node (a FUSE filesystem that is local, such as diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 4c390b825..beb5f012e 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -18,6 +18,7 @@ // 6. Adoption's try-lock never creates a directory, and a released // directory that holds nothing but its lock file can be removed. // 7. The directory layout of the plan's section 2.3. +// 8. The spool budget charges what dead sibling directories hold. // // Built and run by tests/test_native_spool_owner_lock_unit.py. @@ -630,6 +631,112 @@ void TestANewDirectoryAppearsWithItsLockHeld() { CHECK(fs::exists(base + "/r0-0a1b2c3d/.owner.lock")); } +// (8) The budget across incarnations. Every process start gets a fresh +// rank directory, so a spool that charged only its own directory let each +// crash-restart add a full max_bytes while uploads were blocked. With +// charge_dead_siblings, what the sibling rank directories hold counts +// against max_bytes too -- unless another live process holds one (that is +// its own budget); one this process holds, as its service does while it +// adopts it, still counts. +void TestDeadSiblingsCountAgainstTheBudget() { + const std::string base = FreshRoot("budget"); + std::string error; + const auto stage_packs = [&](const std::string& dir, int first, int n) { + Spool spool; + CHECK(Spool::Open({dir, 1 << 20}, &spool, &error) == SpoolStatus::kOk); + for (int i = first; i < first + n; ++i) { + CHECK(StageOne(spool, i, &error) == SpoolStatus::kOk); + } + }; // the Spool, and with it its lock, goes: a dead incarnation + + // A dead incarnation left three 100-byte packs beside the new one. + const std::string key = base + "/0123456789ab"; + stage_packs(key + "/r0-0000000a", 1, 3); + SpoolConfig config{key + "/r0-0000000b", 450}; + config.charge_dead_siblings = true; + Spool spool; + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kOk); + CHECK(spool.Snapshot().sibling_bytes == 300); + CHECK(StageOne(spool, 4, &error) == SpoolStatus::kOk); // 300 + 100 + error.clear(); + CHECK(StageOne(spool, 5, &error) == SpoolStatus::kFull); // 300 + 200 + CHECK(Contains(error, "300 bytes")); + CHECK(Contains(error, "dead")); + // As adoption drains the dead one, the capacity comes back. + for (const auto& entry : + fs::recursive_directory_iterator(key + "/r0-0000000a")) { + if (entry.path().extension() == ".ready") fs::remove(entry.path()); + } + CHECK(StageOne(spool, 5, &error) == SpoolStatus::kOk); + CHECK(spool.Snapshot().sibling_bytes == 0); + + // Without the charge each incarnation had the whole budget to itself. + const std::string uncharged = base + "/ba9876543210"; + stage_packs(uncharged + "/r0-0000000a", 1, 3); + Spool alone; + CHECK(Spool::Open({uncharged + "/r0-0000000b", 450}, &alone, &error) == + SpoolStatus::kOk); + for (int i = 4; i < 8; ++i) { + CHECK(StageOne(alone, i, &error) == SpoolStatus::kOk); + } + + // A sibling another live process holds is its own budget. + const std::string shared = base + "/aaaaaaaaaaaa"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t child = ::fork(); + if (child == 0) { + ::close(ready[0]); + Spool live; + std::string child_error; + bool ok = Spool::Open({shared + "/r1-0000000d", 1 << 20}, &live, + &child_error) == SpoolStatus::kOk; + for (int i = 1; ok && i <= 3; ++i) { + ok = StageOne(live, i, &child_error) == SpoolStatus::kOk; + } + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + SpoolConfig beside{shared + "/r0-0000000e", 450}; + beside.charge_dead_siblings = true; + Spool next; + CHECK(Spool::Open(beside, &next, &error) == SpoolStatus::kOk); + CHECK(next.Snapshot().sibling_bytes == 0); + for (int i = 4; i < 8; ++i) { + CHECK(StageOne(next, i, &error) == SpoolStatus::kOk); + } + ::kill(child, SIGKILL); + int status = 0; + ::waitpid(child, &status, 0); + + // One THIS process holds -- its service adopting it -- still counts. + const std::string adopting = base + "/bbbbbbbbbbbb"; + SpoolOwnerLock held; + CHECK(SpoolOwnerLock::Acquire(adopting + "/r0-0000000f", false, &held, + &error) == SpoolStatus::kOk); + { + SpoolConfig adopted{adopting + "/r0-0000000f", 1 << 20}; + adopted.owner_lock = OwnerLock::kHeldByCaller; + Spool writer; + CHECK(Spool::Open(adopted, &writer, &error) == SpoolStatus::kOk); + for (int i = 1; i <= 3; ++i) { + CHECK(StageOne(writer, i, &error) == SpoolStatus::kOk); + } + } + SpoolConfig own{adopting + "/r0-00000010", 450}; + own.charge_dead_siblings = true; + Spool mine; + CHECK(Spool::Open(own, &mine, &error) == SpoolStatus::kOk); + CHECK(mine.Snapshot().sibling_bytes == 300); +} + // (7) The layout: //r-/. dmi_store::SpoolDestination Destination() { dmi_store::SpoolDestination destination; @@ -711,6 +818,7 @@ int main() { TestALockOnAnUnlinkedFileIsTakenAgain(); TestANewDirectoryAppearsWithItsLockHeld(); TestTheDirectoryLayout(); + TestDeadSiblingsCountAgainstTheBudget(); if (g_failures != 0) { std::cerr << g_failures << " check(s) failed\n"; return 1; diff --git a/tests/test_native_capture_storage_wiring.py b/tests/test_native_capture_storage_wiring.py index 3ad1bf000..60400e6a5 100644 --- a/tests/test_native_capture_storage_wiring.py +++ b/tests/test_native_capture_storage_wiring.py @@ -596,6 +596,8 @@ def test_the_service_starts_before_the_sink_opens_the_spool(monkeypatch, tmp_pat assert config["holder"] # a generated lease holder, never empty (sink,) = engine._test_sinks assert sink["owner_lock"] == "held_by_caller" + # Dead incarnations' packs beside it count against its budget. + assert sink["charge_dead_siblings"] is True assert engine._test_locks[0].held @@ -713,6 +715,7 @@ def test_a_sink_without_a_service_owns_its_spool_itself(monkeypatch, tmp_path): (sink,) = engine._test_sinks assert sink["spool_root"] == str(tmp_path / "spool") assert sink["owner_lock"] == "take" + assert sink["charge_dead_siblings"] is False # no layout, no siblings assert not any(event[0] in ("lock", "layout") for event in events) diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index c1a2520e0..20117d1a2 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -259,6 +259,43 @@ def test_a_spool_root_a_sink_once_owned_still_takes_rank_directories(tmp_path): claim.release() +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_a_dead_siblings_bytes_count_against_the_sinks_budget(tmp_path): + """Each process start claims a fresh rank directory. A sink that + charged only its own let every crash-restart add a whole + spool_max_bytes while uploads were blocked; charge_dead_siblings (what + the engine passes) counts what the dead incarnations beside it still + hold, as the one directory every restart reused did before.""" + pytest.importorskip("torch") + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + sink_config = NativeSinkConfig(spool_root=str(tmp_path / "root")) + dead = claim_spool_directory(sink_config, _config()) + dead_directory = Path(dead.directory) + (dead_directory / "v1").mkdir() + (dead_directory / "v1" / "left.dmi-pack.ready").write_bytes( + b"\0" * (1 << 20)) + dead._lock.release() # killed: the kernel let go, the packs stay + budget = (1 << 20) + 256 # room for the dead bytes, not for a pack + + claim = claim_spool_directory(sink_config, _config()) + try: + sink = _sink(claim.directory, owner_lock="held_by_caller", + spool_max_bytes=budget, max_pack_bytes=budget, + charge_dead_siblings=True) + with pytest.raises(RuntimeError, match="spool byte limit exceeded"): + _stage_one(sink) + del sink + uncharged = _sink(claim.directory, owner_lock="held_by_caller", + spool_max_bytes=budget, max_pack_bytes=budget) + _stage_one(uncharged) # the whole budget, as if nothing were left + del uncharged + finally: + claim.release() + + def test_a_dropped_spool_claim_keeps_its_directory_owned(tmp_path): """Only release() lets go of a claim. An engine dropped without close() drops its claim, while the ring and the sink it activated may still be From 42106b075e65e829c6c4a21972ed882cdf018dfb Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 04:38:28 -0400 Subject: [PATCH 18/43] Adopt dead spools in the loop, a slice per cycle, not in start() or flush() Adoption uploaded a dead sibling's whole backlog synchronously, with no deadline: at start(), so create_record_runtime waited for it, and in step 3a of every flush() cycle, so flush(timeout) did too, even with nothing of this process's own pending. With the object store refusing connections every dead pack ran its whole retry chain (four uploader attempts of four client attempts each) first: the review measured start() at 16.4 s and flush(0.5) at 14.7 s for eight packs, linear in the backlog (about 33 minutes for 1000), and far worse against a store that accepts and never answers. Ctrl-C cannot interrupt the native call. Before B6 start() uploaded nothing; the loop did. And every retry swept and re-hashed the whole dead directory twice (Recover, then the uploader's ListPending). * start() only owes a look at the siblings -- still after the lease and the sweep of its own directory, the plan's order -- and kicks the loop, whose first cycle runs at once. * The loop's cycles adopt; flush()'s never do (run_cycle(adopt)). A cycle probes the siblings' locks, queues the dead ones, and works through them: it takes one's lock, sweeps and validates it once (Recover), then uploads and indexes its packs a round of uploader.max_workers at a time (SpoolUploader::UploadEntries, the uploader's upload of entries already listed), checking the lease and that nothing is owed to the catalog before every round, until adoption_slice_ns (1 s) has passed or stop() is asking. The lock, the Spool and the remaining packs are kept between cycles, so a backlog is hashed once, and neither a flush() nor the service's own uploads wait behind all of it; stop() lets go of a half-adopted sibling, which keeps what is left for the next process. * flush() is about this process's records (plan section 2.6): a sibling still to adopt no longer keeps it from reporting drained, though adopted packs uploaded and not yet indexed are owed in pending_index_ like its own. Only an adoption upload that failed counts towards the loop's backoff. NativeCaptureStorage.flush's TimeoutError no longer mentions adoption, and the storage part of capture_status() is where it shows (adopted_spools, adoption_owed). * A sibling that cannot be locked or opened is recorded, and looked at again after the backoff. Tests (test_native_spool_adoption_live.py): the store-down case, red before -- with the dead lease lapsed and the store cut, start() returned after 11.8 s; it must return within 3 s, and flush(0.5) within 2 s, while the dead spool stays whole and adoption owed; once the store is back the loop adopts it. The other adoption tests wait for the loop's adoption rather than expect it done when start() returns, and the owed-pack test's mutation (no pending-index check before a round) still fails it. The wiring test for the old flush message goes. The storage live suite, the chain suite and the wiring tests pass. --- docs/integration-api-v1.md | 15 +- native/csrc/catalog/bindings_store.cpp | 2 + native/csrc/catalog/storage_service.cpp | 283 ++++++++++++-------- native/csrc/catalog/storage_service.h | 101 ++++--- native/csrc/store/uploader.cpp | 6 + native/csrc/store/uploader.h | 6 + src/dmi/storage/native_capture.py | 19 +- tests/test_native_capture_storage_wiring.py | 13 - tests/test_native_spool_adoption_live.py | 39 ++- 9 files changed, 312 insertions(+), 172 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index b2c410778..9a959b133 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -199,11 +199,16 @@ when a drained directory is removed. If the sink did not seal within by the process until it exits (a warning names it) and the next process on the node adopts it; so does the directory of an engine dropped without `close()`. A second process on a directory is -refused, naming the holder's pid and host. At start the service adopts the -directories under the same catalog key whose owners have died: their stale -`.open` files are swept, their ready packs uploaded and indexed, and the -directory removed, so a crashed process's packs reach the catalog through the -next one on the node, whatever run it belongs to. `spool_max_bytes` bounds the +refused, naming the holder's pid and host. Once started, the service's +background loop adopts the directories under the same catalog key whose owners +have died: their stale `.open` files are swept, their ready packs uploaded and +indexed, a round at a time and a slice of each cycle, and the directory +removed, so a crashed process's packs reach the catalog through the next one on +the node, whatever run it belongs to. Neither `create_record_runtime` nor +`flush_and_wait` waits for that: a flush covers this process's records (an +adopted pack uploaded and not yet indexed is waited for like its own), and the +storage part of `capture_status()` reports the adoption (`adopted_spools`, +`adopted_packs`, `adoption_owed`, `live_siblings`). `spool_max_bytes` bounds the directory together with what those dead directories still hold (a live process's directory is its own budget), so restarts while uploads are blocked cannot each add a whole budget; the room comes back as they are adopted. The diff --git a/native/csrc/catalog/bindings_store.cpp b/native/csrc/catalog/bindings_store.cpp index c1c75e05b..e67044c26 100644 --- a/native/csrc/catalog/bindings_store.cpp +++ b/native/csrc/catalog/bindings_store.cpp @@ -123,6 +123,8 @@ dc::StorageServiceConfig service_config(const py::dict& d) { get(d, "adopt_sibling_spools", c.adopt_sibling_spools); c.adoption_recheck_interval_ns = get( d, "adoption_recheck_interval_ns", c.adoption_recheck_interval_ns); + c.adoption_slice_ns = + get(d, "adoption_slice_ns", c.adoption_slice_ns); c.s3 = s3_config(d); c.uploader.store_id = get(d, "store_id", c.uploader.store_id); c.uploader.max_workers = get(d, "uploader_max_workers", c.uploader.max_workers); diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index 7c6007a91..1f6eef4d2 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -5,6 +5,7 @@ #include #include #include +#include #include #include #include @@ -119,6 +120,17 @@ class CaptureStorageService::LeaseScope { uint64_t counted_on_entry_ = 0; }; +// The dead sibling a cycle is adopting, kept across cycles: its owner lock, +// a Spool on it (held_by_caller, the lock being this service's), and the +// packs its Recover() swept and validated that are not uploaded yet -- so a +// backlog is hashed once, however many cycles its upload takes. +struct CaptureStorageService::Adoption { + std::string directory; + dmi_store::SpoolOwnerLock lock; + dmi_store::Spool spool; + std::deque remaining; +}; + CaptureStorageService::CaptureStorageService(StorageServiceConfig config) : config_(std::move(config)), s3_(config_.s3), @@ -241,8 +253,10 @@ void CaptureStorageService::start() { last_reconcile_ns_ = steady_ns(); { + // With siblings to look at, the loop's first cycle runs at once: it, + // not start(), adopts them. std::lock_guard lock(wake_mutex_); - kick_ = false; + kick_ = adoption_scan_owed_; } { LeaseScope lease(this); // publishes the lease state @@ -271,25 +285,14 @@ void CaptureStorageService::sweep_and_reconcile_at_start() { state_.swept_on_start = recovered.size(); } - // Dead siblings next, after this directory's own sweep and before the - // reconcile, which then finds their packs committed. Nothing here fails - // start(): a sibling left undrained is owed, and the loop retries it. + // Dead siblings are the loop's, from its first cycle on: after the lease + // and this directory's own sweep, as the plan orders them, but not here, + // where a backlog the object store refuses would hold start() -- and + // create_record_runtime -- through every pack's retry chain. if (config_.adopt_sibling_spools) { - try { - adopt_siblings(); - } catch (const CatalogError& exc) { - adoption_owed_ = true; - record_error(std::string("adopting dead spools at start ") + - (is_lease_refusal(exc) ? "lost the publisher lease: " - : "failed: ") + - exc.what()); - } catch (const std::exception& exc) { - adoption_owed_ = true; - record_error(std::string("adopting dead spools at start failed: ") + - exc.what()); - } + adoption_scan_owed_ = true; std::lock_guard lock(state_mutex_); - state_.adoption_owed = adoption_owed_; + state_.adoption_owed = true; } // A failed pass is not fatal -- the bucket is still there next time. Nor @@ -328,6 +331,10 @@ void CaptureStorageService::stop() { // lease has to keep renewing until that is done. stop_lease_thread(); std::lock_guard cycle(cycle_mutex_); + // A sibling half adopted keeps what is left of it, for the next process + // on the node; what was uploaded from it was indexed, or is owed. + adopting_.reset(); + adoption_queue_.clear(); if (started_) { started_ = false; // A quarantined writer holds no lease, so it writes no tombstone: the @@ -372,7 +379,7 @@ bool CaptureStorageService::flush(double timeout_s) { if (!cycle.try_lock_until(deadline)) return false; if (!started_) throw std::logic_error("storage service: not started"); rethrow_if_failed(); - const bool drained = run_cycle().drained; + const bool drained = run_cycle(false).drained; if (!rejected_unreported_.empty()) { std::string message = "storage service: " + std::to_string(rejected_unreported_.size()) + @@ -417,7 +424,7 @@ void CaptureStorageService::loop() { if (failure_) return; // another publisher holds the catalog } std::lock_guard cycle(cycle_mutex_); - run_cycle(); + run_cycle(true); // poll_interval * 2^streak, capped: flush() shares the streak, so an // outage it saw also slows the loop, and a success from either resets it. wait_ns = config_.poll_interval_ns; @@ -429,7 +436,8 @@ void CaptureStorageService::loop() { } } -CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { +CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle( + bool adopt) { CycleOutcome outcome; // The catalog phase needs the lease. Without one -- quarantined after an // unknown outcome, or refused by another holder -- the cycle uploads @@ -489,17 +497,14 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { // 3. Index them. index_or_owe(std::move(to_index), catalog); - // 3a. A dead sibling an earlier adoption pass left undrained, or -- on - // the recheck interval -- a sibling that was alive then and may have - // died since, under the same rule as the uploads above: only with - // the lease, nothing owed and nothing of our own failing. - const bool recheck_due = - live_siblings_ && config_.adoption_recheck_interval_ns > 0 && - steady_ns() - last_adoption_ns_ >= - config_.adoption_recheck_interval_ns; - if (catalog && (adoption_owed_ || recheck_due) && + // 3a. The loop's cycles adopt dead siblings, a slice at a time, under + // the same rule as the uploads above: only with the lease, nothing + // owed and nothing of our own failing. flush()'s cycles do not: a + // dead backlog is not this process's records. + bool adoption_failed = false; + if (adopt && config_.adopt_sibling_spools && catalog && pending_index_.empty() && upload_failures == 0) { - adopt_siblings(); + adoption_failed = !adopt_step(); } // 4. Reconcile on its interval, or when the pass at start() lost the @@ -523,8 +528,10 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { // would be re-hashed on every cycle of an outage. // Without the lease nothing can be confirmed in the catalog, so the // cycle is not drained, and it counts towards the backoff. + // A dead sibling still to adopt is not a failure, nor undrained: only + // an adoption upload that failed backs the loop off. outcome.failed = !catalog || upload_failures != 0 || - !pending_index_.empty() || adoption_owed_; + !pending_index_.empty() || adoption_failed; bool nothing_pending = !batch.refs.empty(); if (batch.refs.empty() && !outcome.failed) { std::vector pending; @@ -553,7 +560,7 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle() { { std::lock_guard lock(state_mutex_); state_.pending_index = pending_index_.size(); - state_.adoption_owed = adoption_owed_; + state_.adoption_owed = adoption_owed(); } failure_streak_ = outcome.failed ? std::min(failure_streak_ + 1, 32) : 0; return outcome; @@ -577,10 +584,61 @@ void CaptureStorageService::index_or_owe(std::vector refs, unindexed.end()); } -void CaptureStorageService::adopt_siblings() { +bool CaptureStorageService::adoption_owed() const { + return adoption_scan_owed_ || adopting_ != nullptr || + !adoption_queue_.empty(); +} + +bool CaptureStorageService::stop_requested() { + std::lock_guard lock(wake_mutex_); + return stop_requested_; +} + +bool CaptureStorageService::adopt_step() { + const uint64_t started = steady_ns(); + if (adopting_ == nullptr && adoption_queue_.empty()) { + // Look at the siblings when that is owed (from start() on) or, while + // the last look found a live one, again on the recheck interval. + const bool recheck_due = + live_siblings_ && config_.adoption_recheck_interval_ns > 0 && + started - last_adoption_scan_ns_ >= + config_.adoption_recheck_interval_ns; + if (!adoption_scan_owed_ && !recheck_due) return true; + if (!scan_siblings()) return false; + } + bool ok = true; + while (!stop_requested()) { + if (adopting_ == nullptr) { + if (adoption_queue_.empty()) break; + const std::string next = adoption_queue_.front(); + adoption_queue_.pop_front(); + if (!begin_adoption(next)) ok = false; + continue; + } + if (adopting_->remaining.empty()) { + finish_adoption(); + continue; + } + // The rules of the cycle's uploads, before every round: none without + // the lease, and none while an uploaded pack is still owed to the + // catalog -- the dead spool is where the rest are durable. + bool catalog = false; + { + LeaseScope lease(this); + catalog = writer_.held_lease() != nullptr; + } + if (!catalog || !pending_index_.empty()) break; + if (!upload_adopted_round()) { + ok = false; // the failed packs stay in the dead spool + break; + } + if (steady_ns() - started >= config_.adoption_slice_ns) break; + } + return ok; +} + +bool CaptureStorageService::scan_siblings() { namespace fs = std::filesystem; - // Owed until the pass completes: a lost lease propagates from the middle. - adoption_owed_ = true; const fs::path own(spool_.root()); std::vector siblings; std::vector claim_staging; @@ -607,106 +665,91 @@ void CaptureStorageService::adopt_siblings() { if (ec) { record_error("adoption: cannot list " + own.parent_path().string() + ": " + ec.message()); - return; + return false; } clear_dead_claim_staging(claim_staging); std::sort(siblings.begin(), siblings.end()); - bool owed = false; uint64_t live = 0; for (const std::string& sibling : siblings) { - bool alive = false; - if (!adopt_sibling(sibling, &alive)) owed = true; - if (alive) ++live; + // A live owner answers a non-blocking probe at once; TryAdopt would + // retry for a few milliseconds first, on every recheck. + if (dmi_store::ReadSpoolOwner(sibling, nullptr)) { + ++live; + } else { + adoption_queue_.push_back(sibling); + } } - adoption_owed_ = owed; + adoption_scan_owed_ = false; live_siblings_ = live != 0; - last_adoption_ns_ = steady_ns(); + last_adoption_scan_ns_ = steady_ns(); std::lock_guard state(state_mutex_); state_.live_siblings = live; + return true; } -void CaptureStorageService::clear_dead_claim_staging( - const std::vector& staging) { - namespace fs = std::filesystem; - // A claim builds its directory's staging copy and renames it into place - // within milliseconds, so one this old whose lock nobody holds belongs - // to a claim that died before its rename. Younger ones are left alone: a - // claim between its mkdir and its flock holds no lock yet, and clearing - // its copy would fail it. Nothing is owed either way. - constexpr auto kDeadAfter = std::chrono::seconds(60); - for (const fs::path& path : staging) { - std::error_code ec; - const auto written = fs::last_write_time(path, ec); - if (ec || fs::file_time_type::clock::now() - written < kDeadAfter) { - continue; - } - dmi_store::SpoolOwnerLock lock; - std::string error; - if (dmi_store::SpoolOwnerLock::TryAdopt(path.string(), &lock, &error) == - dmi_store::SpoolStatus::kOk) { - lock.ReleaseAndRemoveIfEmpty(&error); - } - } -} - -bool CaptureStorageService::adopt_sibling(const std::string& directory, - bool* live) { - *live = false; - // A live owner answers a non-blocking probe at once; TryAdopt would retry - // for a few milliseconds first, on every recheck. - if (dmi_store::ReadSpoolOwner(directory, nullptr)) { - *live = true; - return true; - } - dmi_store::SpoolOwnerLock lock; +bool CaptureStorageService::begin_adoption(const std::string& directory) { + auto adoption = std::make_unique(); + adoption->directory = directory; std::string error; const dmi_store::SpoolStatus locked = - dmi_store::SpoolOwnerLock::TryAdopt(directory, &lock, &error); - if (locked == dmi_store::SpoolStatus::kOwned) { // it lives - *live = true; + dmi_store::SpoolOwnerLock::TryAdopt(directory, &adoption->lock, &error); + if (locked == dmi_store::SpoolStatus::kOwned) { + // Alive after all (taken since the look): its owner's, and not owed. + live_siblings_ = true; + std::lock_guard state(state_mutex_); + ++state_.live_siblings; return true; } if (locked != dmi_store::SpoolStatus::kOk) { // Another adopter drained and removed it meanwhile: nothing is owed. if (!std::filesystem::exists(directory)) return true; record_error("adopting dead spool " + directory + ": " + error); + adoption_scan_owed_ = true; // look again, after the backoff return false; } - // The rules of the cycle's uploads: none without the lease, and none while - // an uploaded pack is still owed to the catalog. - bool catalog = false; - { - LeaseScope lease(this); - catalog = writer_.held_lease() != nullptr; - } - if (!catalog || !pending_index_.empty()) return false; - dmi_store::SpoolConfig config{directory, config_.spool_max_bytes}; - config.owner_lock = dmi_store::OwnerLock::kHeldByCaller; // `lock` + config.owner_lock = dmi_store::OwnerLock::kHeldByCaller; // adoption->lock config.allow_shared_filesystem = config_.spool_allow_shared_filesystem; - dmi_store::Spool spool; std::vector ready; - if (dmi_store::Spool::Open(config, &spool, &error) != + if (dmi_store::Spool::Open(config, &adoption->spool, &error) != dmi_store::SpoolStatus::kOk || - spool.Recover(&ready, &error) != dmi_store::SpoolStatus::kOk) { + adoption->spool.Recover(&ready, &error) != + dmi_store::SpoolStatus::kOk) { record_error("adopting dead spool " + directory + ": " + error); + adoption_scan_owed_ = true; return false; } // Each pack's identity and object key come from the pack and its path in // the dead directory, exactly as its owner would have uploaded it. - dmi_store::SpoolUploader uploader(&spool, &s3_, config_.uploader); - const dmi_store::UploadBatchResult batch = uploader.UploadPending(-1); + adoption->remaining.assign(ready.begin(), ready.end()); + adopting_ = std::move(adoption); + return true; +} + +bool CaptureStorageService::upload_adopted_round() { + Adoption& adoption = *adopting_; + const size_t round = + static_cast(std::max(1, config_.uploader.max_workers)); + std::vector entries; + while (!adoption.remaining.empty() && entries.size() < round) { + entries.push_back(std::move(adoption.remaining.front())); + adoption.remaining.pop_front(); + } + dmi_store::SpoolUploader uploader(&adoption.spool, &s3_, config_.uploader); + const dmi_store::UploadBatchResult batch = uploader.UploadEntries(entries); std::vector to_index; uint64_t uploaded_bytes = 0; size_t failures = 0; for (size_t i = 0; i < batch.refs.size(); ++i) { const dmi_store::PackRef& ref = batch.refs[i]; if (ref.pack_id.empty()) { + // Still in the dead spool: a later cycle retries it. ++failures; + adoption.remaining.push_back(entries[i]); if (i < batch.failures.size()) { - record_error("adopting dead spool " + directory + ": upload failed " - "for " + batch.failures[i].object_key + ": " + - batch.failures[i].error); + record_error("adopting dead spool " + adoption.directory + + ": upload failed for " + batch.failures[i].object_key + + ": " + batch.failures[i].error); } continue; } @@ -724,24 +767,48 @@ bool CaptureStorageService::adopt_sibling(const std::string& directory, // Uploaded, so gone from the dead spool: indexed now, or owed in // pending_index_ like any pack of this service's own. index_or_owe(std::move(to_index), true); - if (failures != 0) return false; // they stay in the dead spool - std::vector left; - if (spool.ListPending(&left, &error) != dmi_store::SpoolStatus::kOk || - !left.empty()) { - return false; - } + return failures == 0; +} + +void CaptureStorageService::finish_adoption() { + std::unique_ptr adoption = std::move(adopting_); { std::lock_guard state(state_mutex_); ++state_.adopted_spools; } - if (!lock.ReleaseAndRemoveIfEmpty(&error)) { + std::string error; + if (!adoption->lock.ReleaseAndRemoveIfEmpty(&error)) { // Nothing to upload is left, only files that are not packs (a // quarantined one, say): the directory stays for someone to look at, // and is not owed. - record_error("adopted dead spool " + directory + " was drained but " - "still holds files that are not packs; left in place"); + record_error("adopted dead spool " + adoption->directory + + " was drained but still holds files that are not packs; " + "left in place"); + } +} + +void CaptureStorageService::clear_dead_claim_staging( + const std::vector& staging) { + namespace fs = std::filesystem; + // A claim builds its directory's staging copy and renames it into place + // within milliseconds, so one this old whose lock nobody holds belongs + // to a claim that died before its rename. Younger ones are left alone: a + // claim between its mkdir and its flock holds no lock yet, and clearing + // its copy would fail it. Nothing is owed either way. + constexpr auto kDeadAfter = std::chrono::seconds(60); + for (const fs::path& path : staging) { + std::error_code ec; + const auto written = fs::last_write_time(path, ec); + if (ec || fs::file_time_type::clock::now() - written < kDeadAfter) { + continue; + } + dmi_store::SpoolOwnerLock lock; + std::string error; + if (dmi_store::SpoolOwnerLock::TryAdopt(path.string(), &lock, &error) == + dmi_store::SpoolStatus::kOk) { + lock.ReleaseAndRemoveIfEmpty(&error); + } } - return true; } void CaptureStorageService::index_bounded(std::vector refs, diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index dd8a82664..0f394b803 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -24,9 +24,9 @@ // directory's owner lock (store/spool.h), so a second process on it is // refused at construction, naming the holder. With adopt_sibling_spools // the service's directory is one rank directory of the plan's section 2.3 -// layout, and at start() it adopts the siblings whose owner has died: a -// crashed process's spool is recovered by the next process on the node for -// the same catalog, whatever run it belongs to. +// layout, and once started its loop adopts the siblings whose owner has +// died: a crashed process's spool is recovered by the next process on the +// node for the same catalog, whatever run it belongs to. // // Deployment shape: the service holds the catalog's single publisher lease, so // run ONE service per (database, table_prefix). A second one waits up to @@ -69,6 +69,7 @@ #include #include #include +#include #include #include #include @@ -106,13 +107,24 @@ struct StorageServiceConfig { // under THIS catalog's key (SpoolCatalogKey of clickhouse.host and // .port, writer.database and .table_prefix, s3.endpoint and .bucket, and // uploader.store_id), or construction throws. - // start(), after the lease and the sweep of its own directory, tries the - // owner lock of every sibling rank directory; each one whose owner is - // gone has its .open files swept, its .ready packs uploaded and indexed, - // and is removed once nothing but its lock file is left. A sibling that - // could not be drained (no lease, an upload that failed, a pack still - // owed) is retried by the loop's cycles, and flush() does not report - // drained until it has been. Live siblings -- another rank or job on this + // The loop's cycles adopt, starting with the first, which start() kicks + // off once it holds the lease and has swept its own directory; start() + // itself uploads nothing of a dead backlog, and neither does flush(). A + // cycle probes the owner lock of every sibling rank directory, then works + // through those whose owner is gone: it takes one's lock, sweeps its + // .open files and validates its .ready packs once, and uploads and + // indexes them a round (uploader.max_workers packs) at a time -- under + // the cycle's upload rules, the lease and nothing owed to the catalog + // checked before every round -- until adoption_slice_ns has passed or a + // stop is requested. The sibling's lock and its remaining packs are kept + // between cycles, so a large backlog is hashed once, not on every cycle, + // and neither a flush() nor the service's own uploads wait behind all of + // it. A drained sibling is removed once nothing but its lock file is + // left. An upload that failed stays in the dead spool and is retried by + // a later cycle, after the backoff. flush() covers this process's + // records: a sibling still to adopt does not keep it from reporting + // drained, though adopted packs uploaded and not yet indexed are owed + // like the service's own. Live siblings -- another rank or job on this // node, a predecessor still closing -- are left alone, and are not owed. bool adopt_sibling_spools = false; // While a pass found a live sibling, the loop passes over the siblings @@ -120,8 +132,12 @@ struct StorageServiceConfig { // was still inside close() when this service started, a rank that // crashes while this one runs -- is adopted then, not at the next // restart on the node. A live sibling costs one non-blocking lock probe - // per pass. 0 never looks again after start(). + // per pass. 0 never looks again after the first pass. uint64_t adoption_recheck_interval_ns = 30'000'000'000ull; + // How long one cycle may spend adopting before it lets go of the cycle -- + // to a flush(), the service's own uploads, stop() -- and carries on in + // the next. Checked between rounds, so a cycle can outrun it by one. + uint64_t adoption_slice_ns = 1'000'000'000ull; dmi_store::S3Config s3; dmi_store::UploaderConfig uploader; // uploader.store_id names the store @@ -207,7 +223,8 @@ struct StorageServiceSnapshot { uint64_t pending_index = 0; // uploaded packs awaiting a retried index uint64_t rejected_packs = 0; // set aside: cannot be indexed (see flush) // adopt_sibling_spools: dead siblings drained, the ready packs of theirs - // that were uploaded, and whether one is still owed a retry. + // that were uploaded, and whether one is still to adopt (or a look at + // the siblings is due). uint64_t adopted_spools = 0; uint64_t adopted_packs = 0; bool adoption_owed = false; @@ -244,9 +261,10 @@ class CaptureStorageService { // Ensure the catalog schema, take the publisher lease (waiting up to // start_lease_wait_ns for another holder's to expire, or for a claim that // timed out to go through -- past it, once, to wait out the quarantine - // such a claim left), sweep the spool, adopt dead siblings - // (adopt_sibling_spools), reconcile once, then start the background - // cycle. The lease renews from the moment it is taken. + // such a claim left), sweep the spool, reconcile once, then start the + // background cycle -- whose first cycle, at once, starts adopting dead + // siblings (adopt_sibling_spools). The lease renews from the moment it is + // taken. // Throws if the lease is still held by another publisher when the wait // ends, or its claim still times out. A lease lost while the reconcile // runs does not fail start(): the loop takes a fresh one, as it would @@ -258,6 +276,8 @@ class CaptureStorageService { // everything it will stage is already staged. Returns false on timeout, // including while a cycle already in flight outlives the deadline; it can // overrun only by its own last cycle, whose requests are all bounded. + // Its cycles adopt nothing, and a dead sibling still to adopt does not + // keep it from returning true (adopt_sibling_spools). // Throws, once, if packs were set aside since the last flush: they are in // the object store but can never reach the catalog. bool flush(double timeout_s); @@ -293,17 +313,32 @@ class CaptureStorageService { // start() waits out like any holder's. class LeaseScope; + // The dead sibling being adopted (storage_service.cpp). + struct Adoption; + void loop(); - // start()'s spool sweep, adoption and reconcile, with the lease held and - // the lease thread renewing it. Requires cycle_mutex_. + // start()'s spool sweep and reconcile, with the lease held and the lease + // thread renewing it. Requires cycle_mutex_. void sweep_and_reconcile_at_start(); - // One adoption pass over the sibling rank directories; sets - // adoption_owed_ to whether one was left undrained. Requires - // cycle_mutex_. Only a lost lease propagates. - void adopt_siblings(); - // Adopts one sibling; false when it is owed another try. Sets *live - // when its owner is alive (not owed: it is its owner's). - bool adopt_sibling(const std::string& directory, bool* live); + // One cycle's share of adoption (see adopt_sibling_spools): looks at the + // siblings when that is due, then adopts until the slice ends. False when + // an upload failed, so the cycle backs off. Requires cycle_mutex_. Only a + // lost lease propagates. + bool adopt_step(); + // Probes every sibling rank directory's owner lock and queues the dead + // ones. False when the directory cannot be listed. Requires cycle_mutex_. + bool scan_siblings(); + // Takes a queued sibling's lock, sweeps it and lists its packs into + // adopting_; leaves adopting_ empty for a live or vanished one. False + // when it could not be locked or opened. + bool begin_adoption(const std::string& directory); + // Uploads and indexes one round of adopting_'s packs; false when an + // upload failed. + bool upload_adopted_round(); + // adopting_ holds no pack any more: removes the directory. + void finish_adoption(); + bool adoption_owed() const; // requires cycle_mutex_ + bool stop_requested(); // Removes the staging copies (dmi_store::IsSpoolClaimStagingName) that // claims killed before their rename left under the catalog key. void clear_dead_claim_staging( @@ -314,7 +349,9 @@ class CaptureStorageService { void index_or_owe(std::vector refs, bool catalog); // Stops the lease thread and waits for it. void stop_lease_thread(); - CycleOutcome run_cycle(); // requires cycle_mutex_ + // Requires cycle_mutex_. `adopt`: the loop's cycles adopt, flush()'s do + // not. + CycleOutcome run_cycle(bool adopt); // Indexes refs in bounded batches, appending every ref that did not index // to *unindexed. Only a lost lease propagates; other failures are recorded. void index_bounded(std::vector refs, @@ -389,13 +426,15 @@ class CaptureStorageService { // The reconcile at start() lost the lease before it finished; the loop // runs one once it holds a lease again. Guarded by cycle_mutex_. bool reconcile_owed_ = false; - // An adoption pass left a dead sibling undrained. Guarded by cycle_mutex_. - bool adoption_owed_ = false; - // The last adoption pass found a live sibling, and when it ran: the loop - // passes again every adoption_recheck_interval_ns. Guarded by - // cycle_mutex_. + // Adoption's state, guarded by cycle_mutex_: whether a look at the + // siblings is due (from start() on), the dead ones the last look found, + // the one being adopted, whether the last look found a live one and when + // it ran (the loop looks again every adoption_recheck_interval_ns). + bool adoption_scan_owed_ = false; + std::deque adoption_queue_; + std::unique_ptr adopting_; bool live_siblings_ = false; - uint64_t last_adoption_ns_ = 0; + uint64_t last_adoption_scan_ns_ = 0; int failure_streak_ = 0; // consecutive failed cycles, for the backoff // Uploaded, so gone from the spool, but not yet in the catalog. std::vector pending_index_; diff --git a/native/csrc/store/uploader.cpp b/native/csrc/store/uploader.cpp index 8f8de47b4..bd1be5980 100644 --- a/native/csrc/store/uploader.cpp +++ b/native/csrc/store/uploader.cpp @@ -235,6 +235,12 @@ UploadBatchResult SpoolUploader::UploadPending(int limit) { if (limit != -1 && static_cast(limit) < pending.size()) { pending.resize(static_cast(limit)); } + return UploadEntries(std::move(pending)); +} + +UploadBatchResult SpoolUploader::UploadEntries( + std::vector pending) { + UploadBatchResult result; // Both vectors are positional from the start: sized to the recover() // order up front, oversized refusals written into their own slot, and // workers below fill the rest by index. diff --git a/native/csrc/store/uploader.h b/native/csrc/store/uploader.h index a0d8e6ed3..09d29abc7 100644 --- a/native/csrc/store/uploader.h +++ b/native/csrc/store/uploader.h @@ -91,6 +91,12 @@ class SpoolUploader { // uploads nothing — see the comment at the byte gate in uploader.cpp. UploadBatchResult UploadPending(int limit = -1); + // Upload these entries, as UploadPending uploads the ones it lists, in + // this order: for a caller that has listed the spool already (adoption + // lists a dead spool once, through Recover, and uploads it a round at a + // time) and must not pay for another hash of everything it holds. + UploadBatchResult UploadEntries(std::vector entries); + // Upload one staged entry with retry. Public for tests. bool UploadOne(const StagedPack& staged, PackRef* ref, int* attempts_out, std::string* error); diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index 2d129d6f0..2d5c21465 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -693,8 +693,9 @@ class NativeCaptureStorage: refuse each other. ``adopt_sibling_spools`` needs ``spool_root`` to be a rank directory of - the spool layout (``spool_rank_directory``); ``start`` then drains the - sibling directories whose owners have died into this catalog. + the spool layout (``spool_rank_directory``); once started, the service's + background loop drains the sibling directories whose owners have died + into this catalog, a slice per cycle. """ def __init__( @@ -757,18 +758,18 @@ def flush(self, timeout_s: float) -> None: Call after the sink's own flush. Raises TimeoutError, carrying the last upload or index error, if the spool has not drained in time. + The dead processes' spools the service adopts + (``adopt_sibling_spools``) are not part of it -- they are the + background loop's, and ``snapshot()`` reports them + (``adopted_spools``, ``adoption_owed``) -- though an adopted pack + uploaded and not yet indexed is waited for like this process's own. """ if not self._service.flush(float(timeout_s)): snapshot = self._service.snapshot() - # A dead process's spool this service has still to adopt keeps - # it undrained too (adopt_sibling_spools). - adopting = (", and a dead process's spool still to adopt" - if snapshot.get("adoption_owed") else "") raise TimeoutError( "timed out waiting for staged packs to reach the catalog " - f"({snapshot['pending_index']} uploaded but unindexed" - f"{adopting}); last error: " - f"{snapshot['last_error'] or 'none'}") + f"({snapshot['pending_index']} uploaded but unindexed); last " + f"error: {snapshot['last_error'] or 'none'}") def stop(self) -> None: """Stop the background thread and release the lease. No flush.""" diff --git a/tests/test_native_capture_storage_wiring.py b/tests/test_native_capture_storage_wiring.py index 60400e6a5..a037fd873 100644 --- a/tests/test_native_capture_storage_wiring.py +++ b/tests/test_native_capture_storage_wiring.py @@ -752,19 +752,6 @@ def test_flush_reports_packs_that_did_not_reach_the_catalog(monkeypatch, tmp_pat engine.flush_and_wait(1.0) -def test_flush_says_when_a_dead_spool_is_still_to_adopt(monkeypatch, tmp_path): - engine, _events, services = _capture_engine(monkeypatch, tmp_path) - engine.create_record_runtime(_record_format()) - services[0].flush_results = [False] - services[0].snapshot = lambda: { - "pending_index": 0, "adoption_owed": True, - "last_error": "adopting dead spool /x: upload failed"} - - with pytest.raises(TimeoutError, - match="dead process's spool still to adopt.*upload"): - engine.flush_and_wait(1.0) - - def test_close_flushes_the_sink_before_the_ring_stops(monkeypatch, tmp_path): """The sink's open pack is in memory until a flush seals it, and stopping the ring releases the sink without one. So close() flushes the sink diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 7af4e507f..4b3172509 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -194,6 +194,16 @@ def _wait_for(predicate, timeout_s: float = 30.0) -> None: time.sleep(0.05) +def _adopted(service, spools: int): + """Whether `spools` dead siblings have been adopted and nothing more is + to adopt, as of the last cycle's end.""" + def done(): + snapshot = service.snapshot() + return (snapshot["adopted_spools"] == spools + and not snapshot["adoption_owed"]) + return done + + def _read_all(config) -> dict: from dmi.storage.native_capture import NativeCaptureReader @@ -253,8 +263,9 @@ def test_a_sigkilled_process_spool_is_adopted_by_its_successor( lock = _claim(base, config) service = _service(config, lock.directory) started = time.monotonic() - service.start() # waits out the dead lease, then adopts + service.start() # waits out the dead lease; the loop adopts try: + _wait_for(_adopted(service, 1)) snapshot = service.snapshot() assert time.monotonic() - started < 30 assert snapshot["adopted_spools"] == 1, snapshot @@ -298,26 +309,41 @@ def test_a_dead_spool_waits_in_place_while_the_object_store_is_down( staged = sorted(dead.rglob("*.dmi-pack.ready")) assert staged + # Let the dead process's lease lapse, so start() has none to wait + # out and its time is its own. + time.sleep(LEASE["lease_ttl_s"] + 1.0) switch.cut() config = _storage_config(switch.url, prefix) lock = _claim(base, config) service = _service(config, lock.directory) + started = time.monotonic() service.start() try: + # start() leaves the dead backlog to the loop: it does not sit + # through every dead pack's retry chain against a store that + # refuses them (about 8 s a round of four packs). + assert time.monotonic() - started < 3.0 + # Nor does flush(), which is about this process's own records: + # it returns within its deadline, drained or not. + flushed = time.monotonic() + try: + service.flush(0.5) + except TimeoutError: + pass # a loop cycle still in an upload round holds the cycle + assert time.monotonic() - flushed < 2.0 + _wait_for(lambda: service.snapshot()["upload_failures"] > 0) snapshot = service.snapshot() assert snapshot["adoption_owed"] is True, snapshot assert snapshot["adopted_spools"] == 0, snapshot # Nothing left the dead spool: its packs are durable there. assert sorted(dead.rglob("*.dmi-pack.ready")) == staged - with pytest.raises(TimeoutError): - service.flush(2.0) switch.restore() - service.flush(60.0) + _wait_for(_adopted(service, 1), 90.0) snapshot = service.snapshot() assert snapshot["adoption_owed"] is False, snapshot - assert snapshot["adopted_spools"] == 1, snapshot assert not dead.exists() + service.flush(60.0) assert _read_all(config) == _expected() finally: service.stop() @@ -403,7 +429,7 @@ def _cut_the_catalog_at_the_first_upload(request: bytes) -> float: clickhouse.restore() service.flush(60.0) - _wait_for(lambda: service.snapshot()["adopted_spools"] == 2, 60.0) + _wait_for(_adopted(service, 2), 60.0) snapshot = service.snapshot() assert snapshot["adoption_owed"] is False, snapshot assert not first.exists() and not second.exists() @@ -447,6 +473,7 @@ def test_a_sibling_whose_owner_dies_after_start_is_adopted_by_a_recheck( service = _store().StorageService(native) service.start() try: + _wait_for(lambda: not service.snapshot()["adoption_owed"]) snapshot = service.snapshot() assert snapshot["live_siblings"] == 1, snapshot assert snapshot["adopted_spools"] == 0, snapshot From 9cacc53daf4c97b7e0482078ed9ced8cb59305b4 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 04:44:36 -0400 Subject: [PATCH 19/43] Leave a dead spool this service can never upload, and say so once When a dead sibling held a pack that can never be uploaded, adoption failed on it every pass and never finished: adoption_owed stayed set, every cycle reported failed, the loop sat at its maximum backoff (the review saw 1 cycle in 20 s where the poll interval gives about 400), and each retry re-hashed the dead pack. With the adoption of the previous commit it also kept every sibling queued after it waiting. A pack larger than this process's uploader_max_in_flight_bytes (a crashed run with a larger bound left it), a different object already at its key, and staged bytes that no longer match are all like that; so is a dead directory whose lock cannot be taken at all (EACCES on another user's lock file). * dmi_store::UploadFailure gains `retryable`: false for the oversized pack, the conflict and the corrupt staged bytes -- the three exits the uploader already treats as final -- and true otherwise (transport, a 5xx, a failed HEAD). The store driver reports it. * An adoption round re-queues only retryable failures, and only they back the loop off. A non-retryable one blocks its directory: the rest of its packs still go up, then the directory is left in place with its lock let go, for a process that can adopt it (or a person), recorded once in last_error and listed in the snapshot's blocked_siblings. It is never looked at again by this service, and is not owed. A sibling that cannot be locked or opened, and one drained of packs that still holds other files, are blocked the same way. Tests, red before: test_native_spool_adoption_live.py -- two dead siblings, one holding packs over a 4096-byte in-flight bound: the other is adopted and removed, the first is listed in blocked_siblings with its packs where they were and the in-flight limit in last_error, no more uploads fail after that, the loop keeps its poll interval (at least 5 cycles a second), adoption is not owed, a flush returns at once, and the catalog holds exactly the adoptable sibling's captures (before: the adoption never settled). test_native_uploader.py: the oversized pack and the conflict are reported not retryable, a TLS failure retryable. --- docs/integration-api-v1.md | 7 ++- native/csrc/catalog/bindings_store.cpp | 1 + native/csrc/catalog/storage_service.cpp | 65 +++++++++++++++++------- native/csrc/catalog/storage_service.h | 26 ++++++++-- native/csrc/store/conformance_store.cpp | 6 ++- native/csrc/store/uploader.cpp | 13 +++-- native/csrc/store/uploader.h | 10 +++- tests/test_native_spool_adoption_live.py | 64 +++++++++++++++++++++-- tests/test_native_uploader.py | 8 ++- 9 files changed, 164 insertions(+), 36 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index 9a959b133..52ffd11b4 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -208,7 +208,12 @@ the node, whatever run it belongs to. Neither `create_record_runtime` nor `flush_and_wait` waits for that: a flush covers this process's records (an adopted pack uploaded and not yet indexed is waited for like its own), and the storage part of `capture_status()` reports the adoption (`adopted_spools`, -`adopted_packs`, `adoption_owed`, `live_siblings`). `spool_max_bytes` bounds the +`adopted_packs`, `adoption_owed`, `live_siblings`). A dead directory the +service can never adopt -- one holding a pack it can never upload, such as one +larger than its `uploader_max_in_flight_bytes` or one whose key already holds a +different object, or one it cannot lock -- is left in place once the rest of +its packs are up, listed in `blocked_siblings`, reported once in `last_error`, +and neither retried nor owed. `spool_max_bytes` bounds the directory together with what those dead directories still hold (a live process's directory is its own budget), so restarts while uploads are blocked cannot each add a whole budget; the room comes back as they are adopted. The diff --git a/native/csrc/catalog/bindings_store.cpp b/native/csrc/catalog/bindings_store.cpp index e67044c26..29735f04b 100644 --- a/native/csrc/catalog/bindings_store.cpp +++ b/native/csrc/catalog/bindings_store.cpp @@ -178,6 +178,7 @@ py::dict snapshot_dict(const dc::StorageServiceSnapshot& s) { out["adopted_packs"] = s.adopted_packs; out["adoption_owed"] = s.adoption_owed; out["live_siblings"] = s.live_siblings; + out["blocked_siblings"] = s.blocked_siblings; out["failed"] = s.failed; out["lease_state"] = s.lease_state; // Seconds on the monotonic clock, comparable with time.monotonic() (both diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index 1f6eef4d2..f00f20640 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -129,6 +129,9 @@ struct CaptureStorageService::Adoption { dmi_store::SpoolOwnerLock lock; dmi_store::Spool spool; std::deque remaining; + // The first pack this service can never upload (UploadFailure:: + // retryable false); the directory is left, with it, once the rest are up. + std::string blocked; }; CaptureStorageService::CaptureStorageService(StorageServiceConfig config) @@ -657,7 +660,8 @@ bool CaptureStorageService::scan_siblings() { claim_staging.push_back(it->path()); continue; } - if (!dmi_store::ParseSpoolRankDirectoryName(name, &rank, &incarnation)) { + if (!dmi_store::ParseSpoolRankDirectoryName(name, &rank, &incarnation) || + blocked_siblings_.count(it->path().string()) != 0) { continue; } siblings.push_back(it->path().string()); @@ -703,9 +707,9 @@ bool CaptureStorageService::begin_adoption(const std::string& directory) { if (locked != dmi_store::SpoolStatus::kOk) { // Another adopter drained and removed it meanwhile: nothing is owed. if (!std::filesystem::exists(directory)) return true; - record_error("adopting dead spool " + directory + ": " + error); - adoption_scan_owed_ = true; // look again, after the backoff - return false; + // Its lock cannot be taken at all -- another user's lock file, say. + block_sibling(directory, "cannot lock it: " + error); + return true; } dmi_store::SpoolConfig config{directory, config_.spool_max_bytes}; config.owner_lock = dmi_store::OwnerLock::kHeldByCaller; // adoption->lock @@ -715,9 +719,9 @@ bool CaptureStorageService::begin_adoption(const std::string& directory) { dmi_store::SpoolStatus::kOk || adoption->spool.Recover(&ready, &error) != dmi_store::SpoolStatus::kOk) { - record_error("adopting dead spool " + directory + ": " + error); - adoption_scan_owed_ = true; - return false; + adoption.reset(); // lets go of its lock + block_sibling(directory, "cannot open it: " + error); + return true; } // Each pack's identity and object key come from the pack and its path in // the dead directory, exactly as its owner would have uploaded it. @@ -740,16 +744,26 @@ bool CaptureStorageService::upload_adopted_round() { std::vector to_index; uint64_t uploaded_bytes = 0; size_t failures = 0; + size_t retryable = 0; for (size_t i = 0; i < batch.refs.size(); ++i) { const dmi_store::PackRef& ref = batch.refs[i]; if (ref.pack_id.empty()) { - // Still in the dead spool: a later cycle retries it. + // Still in the dead spool either way. One a later try could upload + // is retried by a later cycle; one no try by this service can is + // not, and blocks the directory once the rest are up. ++failures; - adoption.remaining.push_back(entries[i]); - if (i < batch.failures.size()) { + const dmi_store::UploadFailure failure = + i < batch.failures.size() ? batch.failures[i] + : dmi_store::UploadFailure{}; + const std::string what = failure.object_key + ": " + failure.error; + if (failure.retryable) { + ++retryable; + adoption.remaining.push_back(entries[i]); record_error("adopting dead spool " + adoption.directory + - ": upload failed for " + batch.failures[i].object_key + - ": " + batch.failures[i].error); + ": upload failed for " + what); + } else if (adoption.blocked.empty()) { + adoption.blocked = "it holds a pack this service can never upload, " + + what; } continue; } @@ -767,11 +781,18 @@ bool CaptureStorageService::upload_adopted_round() { // Uploaded, so gone from the dead spool: indexed now, or owed in // pending_index_ like any pack of this service's own. index_or_owe(std::move(to_index), true); - return failures == 0; + return retryable == 0; } void CaptureStorageService::finish_adoption() { std::unique_ptr adoption = std::move(adopting_); + const std::string directory = adoption->directory; + if (!adoption->blocked.empty()) { + const std::string reason = adoption->blocked; + adoption.reset(); // lets go of its lock; the directory stays + block_sibling(directory, reason); + return; + } { std::lock_guard state(state_mutex_); ++state_.adopted_spools; @@ -779,14 +800,22 @@ void CaptureStorageService::finish_adoption() { std::string error; if (!adoption->lock.ReleaseAndRemoveIfEmpty(&error)) { // Nothing to upload is left, only files that are not packs (a - // quarantined one, say): the directory stays for someone to look at, - // and is not owed. - record_error("adopted dead spool " + adoption->directory + - " was drained but still holds files that are not packs; " - "left in place"); + // quarantined one, say): the directory stays for someone to look at. + block_sibling(directory, "it was drained, but still holds files that " + "are not packs"); } } +void CaptureStorageService::block_sibling(const std::string& directory, + const std::string& reason) { + blocked_siblings_.insert(directory); + record_error("dead spool " + directory + " is left in place, not to be " + "adopted by this service: " + reason); + std::lock_guard state(state_mutex_); + state_.blocked_siblings.assign(blocked_siblings_.begin(), + blocked_siblings_.end()); +} + void CaptureStorageService::clear_dead_claim_staging( const std::vector& staging) { namespace fs = std::filesystem; diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index 0f394b803..f81ae17be 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -75,6 +75,7 @@ #include #include #include +#include #include #include #include @@ -121,10 +122,18 @@ struct StorageServiceConfig { // and neither a flush() nor the service's own uploads wait behind all of // it. A drained sibling is removed once nothing but its lock file is // left. An upload that failed stays in the dead spool and is retried by - // a later cycle, after the backoff. flush() covers this process's - // records: a sibling still to adopt does not keep it from reporting - // drained, though adopted packs uploaded and not yet indexed are owed - // like the service's own. Live siblings -- another rank or job on this + // a later cycle, after the backoff -- unless no retry by this service can + // ever succeed (dmi_store::UploadFailure::retryable: a pack over + // uploader.max_in_flight_bytes, a different object at its key, bytes that + // no longer match), or the sibling cannot be locked or opened at all. + // Such a sibling is blocked: once the rest of its packs are up it is left + // in place, with its lock let go, for a process that can adopt it (or a + // person); it is reported once (last_error, snapshot blocked_siblings), + // never retried or re-hashed by this service, and not owed. So is one + // drained of packs that still holds other files. flush() covers this + // process's records: a sibling still to adopt does not keep it from + // reporting drained, though adopted packs uploaded and not yet indexed + // are owed like the service's own. Live siblings -- another rank or job on this // node, a predecessor still closing -- are left alone, and are not owed. bool adopt_sibling_spools = false; // While a pass found a live sibling, the loop passes over the siblings @@ -230,6 +239,9 @@ struct StorageServiceSnapshot { bool adoption_owed = false; // Siblings whose owner was alive at the last adoption pass. uint64_t live_siblings = 0; + // Dead siblings this service will not adopt, left in place (see + // adopt_sibling_spools), by directory. + std::vector blocked_siblings; // A foreign lease outlived 2 x TTL: the service stopped for good. bool failed = false; // "none" before start, "held", "quarantined" (an unknown outcome set the @@ -335,8 +347,11 @@ class CaptureStorageService { // Uploads and indexes one round of adopting_'s packs; false when an // upload failed. bool upload_adopted_round(); - // adopting_ holds no pack any more: removes the directory. + // adopting_ holds no pack any more: removes the directory, or leaves a + // blocked one. void finish_adoption(); + // Leaves a dead sibling in place for good, reporting why. + void block_sibling(const std::string& directory, const std::string& reason); bool adoption_owed() const; // requires cycle_mutex_ bool stop_requested(); // Removes the staging copies (dmi_store::IsSpoolClaimStagingName) that @@ -433,6 +448,7 @@ class CaptureStorageService { bool adoption_scan_owed_ = false; std::deque adoption_queue_; std::unique_ptr adopting_; + std::set blocked_siblings_; bool live_siblings_ = false; uint64_t last_adoption_scan_ns_ = 0; int failure_streak_ = 0; // consecutive failed cycles, for the backoff diff --git a/native/csrc/store/conformance_store.cpp b/native/csrc/store/conformance_store.cpp index 952856486..b3891788f 100644 --- a/native/csrc/store/conformance_store.cpp +++ b/native/csrc/store/conformance_store.cpp @@ -17,7 +17,9 @@ // "objects":[{"key":"...","size":N,"etag":"..."}...],"attempts":N} // {"op":"upload_one"|"upload_pending",...,"root":"...", // "owner_lock":"take"|"held_by_caller" (optional, take by default)} -// open the spool at root, owner lock included, for the one op. +// open the spool at root, owner lock included, for the one op; +// upload_pending's failures carry pack_id, object_key, attempts, error +// and retryable. // Errors: {"ok":false,"what":"..."}. #include "s3_client.h" @@ -385,6 +387,8 @@ int main() { out += ",\"attempts\":" + std::to_string(failure.attempts); out += ",\"error\":"; jc::EscapeJson(failure.error, &out); + out += std::string(",\"retryable\":") + + (failure.retryable ? "true" : "false"); out += "}"; first = false; } diff --git a/native/csrc/store/uploader.cpp b/native/csrc/store/uploader.cpp index bd1be5980..2914d20c6 100644 --- a/native/csrc/store/uploader.cpp +++ b/native/csrc/store/uploader.cpp @@ -94,7 +94,9 @@ SpoolUploader::SpoolUploader(Spool* spool, S3Client* client, : spool_(spool), client_(client), config_(std::move(config)) {} bool SpoolUploader::UploadOne(const StagedPack& staged, PackRef* ref, - int* attempts_out, std::string* error) { + int* attempts_out, std::string* error, + bool* retryable_out) { + if (retryable_out) *retryable_out = true; const std::string& key = staged.object_key; std::mt19937_64 rng( static_cast(std::hash{}(staged.pack_id))); @@ -164,6 +166,7 @@ bool SpoolUploader::UploadOne(const StagedPack& staged, PackRef* ref, // the retry-exhausted exit carries one. if (attempts_out) *attempts_out = attempts; if (error) *error = last_error; + if (retryable_out) *retryable_out = false; return false; // NOT retryable } } @@ -181,6 +184,7 @@ bool SpoolUploader::UploadOne(const StagedPack& staged, PackRef* ref, // Corrupt staged bytes: no retry can fix local corruption, but report // it as the failure rather than uploading garbage. last_error = "staged bytes do not match the staged checksum"; + if (retryable_out) *retryable_out = false; break; } std::string etag; @@ -266,7 +270,7 @@ UploadBatchResult SpoolUploader::UploadEntries( // test_mixed_batch_reports_oversized_pack_at_its_position), and // docs/benchmarks.md records the accounting decision behind it. result.failures[i] = {pending[i].pack_id, pending[i].object_key, 0, - "pack exceeds the in-flight byte limit"}; + "pack exceeds the in-flight byte limit", false}; ++result.snapshot.attempted_packs; ++result.snapshot.failed_packs; } @@ -339,7 +343,8 @@ UploadBatchResult SpoolUploader::UploadEntries( PackRef ref; std::string error; int attempts = 0; - const bool ok = UploadOne(*staged, &ref, &attempts, &error); + bool retryable = true; + const bool ok = UploadOne(*staged, &ref, &attempts, &error, &retryable); const int64_t elapsed = NowNs() - started; { std::lock_guard lock(mutex); @@ -361,7 +366,7 @@ UploadBatchResult SpoolUploader::UploadEntries( } else { result.refs[slot.index] = PackRef{}; result.failures[slot.index] = {staged->pack_id, staged->object_key, - attempts, error}; + attempts, error, retryable}; ++result.snapshot.failed_packs; if (attempts > 1) { result.snapshot.retries += static_cast(attempts - 1); diff --git a/native/csrc/store/uploader.h b/native/csrc/store/uploader.h index 09d29abc7..d69b59e63 100644 --- a/native/csrc/store/uploader.h +++ b/native/csrc/store/uploader.h @@ -54,6 +54,11 @@ struct UploadFailure { std::string object_key; int attempts = 0; std::string error; + // False when no retry by this uploader can ever succeed: a pack over + // max_in_flight_bytes, a different object already at its key (a + // conflict), staged bytes that no longer match their checksum. The pack + // stays staged either way. + bool retryable = true; }; struct UploadSnapshot { @@ -97,9 +102,10 @@ class SpoolUploader { // time) and must not pay for another hash of everything it holds. UploadBatchResult UploadEntries(std::vector entries); - // Upload one staged entry with retry. Public for tests. + // Upload one staged entry with retry. Public for tests. *retryable_out + // says, on failure, whether a later try could succeed (UploadFailure). bool UploadOne(const StagedPack& staged, PackRef* ref, int* attempts_out, - std::string* error); + std::string* error, bool* retryable_out = nullptr); private: Spool* spool_; diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 4b3172509..d1983bd9c 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -105,14 +105,14 @@ def _service(config, directory: str, **options): adopt_sibling_spools=True, **options) -def _envelope(indexes): +def _envelope(indexes, width: int = 6): from tests.test_native_capture_chain_live import _Envelope import torch envelope = _Envelope() for index in indexes: - envelope.add(index, torch.arange(6, dtype=torch.float16) + index) + envelope.add(index, torch.arange(width, dtype=torch.float16) + index) return envelope @@ -350,7 +350,7 @@ def test_a_dead_spool_waits_in_place_while_the_object_store_is_down( lock.release_and_remove_if_empty() -def _stage_into(directory: str, indexes) -> None: +def _stage_into(directory: str, indexes, width: int = 6) -> None: """Stage records into a directory this process holds, as its own sink would: the REAL native pack sink, held_by_caller.""" import torch # noqa: F401 -- the sink extension links against it @@ -365,7 +365,7 @@ def _stage_into(directory: str, indexes) -> None: max_pack_records=RECORDS_PER_PACK, max_linger_ns=600 * 10**9, owner_lock="held_by_caller") lease = sink.attach() - envelope = _envelope(indexes) + envelope = _envelope(indexes, width) sink.submit_envelope(LAYOUT, envelope.rows, envelope.payload()) assert sink.flush_and_wait(60.0) sink.rethrow_if_failed() @@ -445,6 +445,62 @@ def _cut_the_catalog_at_the_first_upload(request: bytes) -> float: lock.release_and_remove_if_empty() +def test_a_dead_spool_this_service_can_never_upload_is_left_and_reported( + fake_s3, tmp_path): + """A pack this service can never upload -- here larger than its + uploader_max_in_flight_bytes, which a crashed run with a larger bound + left behind -- blocks its dead directory for good: retrying it only + re-hashes it and keeps the loop backing off. It is reported once and + left in place for a process that can (or a person), the siblings after + it are adopted, and nothing is owed or retried.""" + base = tmp_path / "spool" + with _catalog() as prefix: + # A 4096-byte bound: the wide captures' packs exceed it, the narrow + # ones' fit. + config = _storage_config(fake_s3, prefix, + uploader_max_in_flight_bytes=4096) + directories = [] + for indexes, width in ((range(0, 4), 4096), (range(4, 10), 6)): + sibling = _claim(base, config) + directories.append(Path(sibling.directory)) + _stage_into(sibling.directory, indexes, width) + sibling.release() # its owner is gone + blocked, adoptable = directories + blocked_packs = sorted(blocked.rglob("*.dmi-pack.ready")) + assert all(path.stat().st_size > 4096 for path in blocked_packs) + + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + _wait_for(_adopted(service, 1)) + snapshot = service.snapshot() + assert snapshot["blocked_siblings"] == [str(blocked)], snapshot + assert "in-flight byte limit" in snapshot["last_error"], snapshot + assert sorted(blocked.rglob("*.dmi-pack.ready")) == blocked_packs + assert not adoptable.exists() + # Not retried: no more failed uploads, and no backoff -- the + # loop keeps its poll interval. + failures, cycles = (snapshot["upload_failures"], + snapshot["cycles"]) + time.sleep(1.0) + snapshot = service.snapshot() + assert snapshot["upload_failures"] == failures, snapshot + assert snapshot["cycles"] >= cycles + 5, snapshot + assert snapshot["adoption_owed"] is False, snapshot + started = time.monotonic() + service.flush(10.0) + assert time.monotonic() - started < 2.0 + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in _envelope( + range(4, 10)).expected.items()} + assert _read_all(config) == expected + finally: + service.stop() + lock.release_and_remove_if_empty() + + def test_a_sibling_whose_owner_dies_after_start_is_adopted_by_a_recheck( fake_s3, tmp_path): """A sibling still owned when the service starts -- a predecessor still diff --git a/tests/test_native_uploader.py b/tests/test_native_uploader.py index 36ccebfb5..d2e0e9328 100644 --- a/tests/test_native_uploader.py +++ b/tests/test_native_uploader.py @@ -320,6 +320,8 @@ def test_pack_over_the_byte_gate_fails_fast(fake_s3, tmp_path): assert result["snapshot"]["failed_packs"] == 1 assert result["failures"][0]["attempts"] == 0 assert "in-flight" in result["failures"][0]["error"] + # No retry by this uploader can admit it. + assert result["failures"][0]["retryable"] is False finally: sink.close() store.close() @@ -416,8 +418,10 @@ def test_a_pack_conflict_reports_why_it_will_not_overwrite(fake_s3, tmp_path): assert staged["object_key"] in failure["error"], failure assert "do not overwrite" in failure["error"], failure # Counted like every other attempted failure: the HEAD that found - # the conflict was an attempt, and it is not retried. + # the conflict was an attempt, and it is not retried -- now or by a + # later batch. assert failure["attempts"] == 1, failure + assert failure["retryable"] is False, failure # The refusal is real: nothing was written over the existing object, # and the staged pack is still there to inspect. @@ -673,6 +677,8 @@ def test_a_64_mib_pack_uploads_multipart_over_https_with_a_private_ca( assert untrusted["refs"][0]["pack_id"] == "", untrusted assert failure["pack_id"] == staged["pack_id"], failure assert "certificate" in failure["error"].lower(), failure + # A transport failure: a later try (with the CA, here) can succeed. + assert failure["retryable"] is True, failure assert STATE.calls == [] and STATE.objects == {} assert Path(staged["path"]).exists() From dfff3c36760472753ce7ff054b440d60398e3df9 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 04:51:17 -0400 Subject: [PATCH 20/43] Warn of packs left in spool_root outside the layout, and say how to hand them over The engine before the layout spooled into /v1/... and started its service there with the sweep on, so the next start uploaded what a crashed run had left. The engine now claims //r-/, and adoption looks only at rank directories beside it: packs a pre-layout crash left under /v1 are never swept, uploaded or reported, and the next run starts cleanly without a word. The same holds for what a sink-only or explicit-record_sink run leaves in spool_root itself. Adopting them in place is not open to the engine -- spool_root contains owned rank directories, which the nesting rule refuses to let anyone take -- but moving them into a rank directory nobody owns is: the next start adopts it like a dead process's. claim_spool_directory now walks spool_root by name only, skipping the rank directories (adoption's), and when it finds ready packs outside the layout logs a warning with their count, one of them, and the directory to move them into, //r0-00000000, paths below spool_root kept, if they are bound for this catalog and store. The v1 contract describes the upgrade. Tests: test_native_spool_ownership.py, red before (no warning) -- a pack inside the layout warns of nothing; two flat packs under spool_root/v1 (and a temp file, not counted) give one warning naming 2, their directory and the suggested rank directory. test_native_spool_adoption_live.py -- the real sink, sink-only, stages into spool_root; the claim's warning names the directory; moved there, the packs are adopted, the directory removed, and every capture reads back from the catalog. --- docs/integration-api-v1.md | 10 +++- src/dmi/storage/native_capture.py | 51 ++++++++++++++++++- tests/test_native_spool_adoption_live.py | 62 ++++++++++++++++++++++++ tests/test_native_spool_ownership.py | 42 ++++++++++++++++ 4 files changed, 163 insertions(+), 2 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index 52ffd11b4..9f697d85c 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -223,7 +223,15 @@ are refused (by statfs `f_type`) unless `NativeSinkConfig.spool_allow_shared_fil filesystem that is itself local, such as fuse-overlayfs, needs too. Without `capture_storage_config` the sink owns `spool_root` itself. With an explicit `record_sink`, the service drains `spool_root` as that sink writes it, -unswept and adopting nothing. +unswept and adopting nothing. Upgrading from an engine without this layout: it +spooled into `/v1/...` and its next start uploaded what a crashed +run left there, but nothing adopts packs outside the layout now -- nor those a +sink-only or explicit-`record_sink` run leaves in `spool_root`. Each +`create_record_runtime` logs a warning while any are there, with their count, +one of them, and a directory of the layout nobody owns +(`//r0-00000000/`); moved into it with their paths +below `spool_root` kept (`v1/...`), packs bound for this catalog and store are +adopted by the next start on the node. To reach a secured catalog, set `clickhouse_scheme="https"` (and the server's TLS HTTP port, usually 8443) on `NativeCaptureStorageConfig`. The client always verifies the server's certificate and name, against libcurl's built-in CA diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index 2d5c21465..065fc74fc 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -32,6 +32,7 @@ from __future__ import annotations import json +import logging import math import os import re @@ -592,9 +593,36 @@ def validate_capture_bounds( "uploader_max_in_flight_bytes") +_LOG = logging.getLogger(__name__) + # The spool directory's owner-lock modes (native/csrc/store/spool.h). SPOOL_OWNER_LOCKS = ("take", "held_by_caller") +# The spool layout's directory names (native/csrc/store/spool.h). +_CATALOG_KEY = re.compile(r"[0-9a-f]{12}") +_RANK_DIRECTORY = re.compile(r"r(?:0|[1-9][0-9]{0,18})-[0-9a-f]{8}") + + +def _packs_outside_the_layout(spool_root: str) -> tuple[int, Optional[str]]: + """Ready packs under ``spool_root`` that no rank directory of the layout + holds -- left by an engine from before the layout, or by a sink-only or + explicit-record_sink run, which write into ``spool_root`` itself -- and + the first one met. Only names are read, and the rank directories, which + adoption drains, are not walked.""" + count, example = 0, None + for directory, subdirectories, files in os.walk(spool_root): + relative = os.path.relpath(directory, spool_root) + depth = 0 if relative == "." else relative.count(os.sep) + 1 + if depth == 1 and _CATALOG_KEY.fullmatch(os.path.basename(directory)): + subdirectories[:] = [name for name in subdirectories + if not _RANK_DIRECTORY.fullmatch(name)] + for name in files: + if name.endswith(".dmi-pack.ready"): + count += 1 + if example is None: + example = os.path.join(directory, name) + return count, example + def _spool_producer_rank() -> int: """The rank a spool directory is named for: torchrun's global ``RANK``, @@ -660,14 +688,35 @@ def claim_spool_directory( already held. Raises ``SpoolOwnedError`` (a ``RuntimeError``) if another process holds it, and ``ValueError`` for a shared filesystem or a directory nested in another spool. + + Ready packs under ``spool_root`` outside the layout -- an engine from + before it spooled into ``/v1/...``, and a sink-only or + explicit-``record_sink`` run still does -- are adopted by nothing, so a + claim logs a warning naming how many there are, one of them, and a + directory of the layout to move them into for adoption. """ module = _load_native_store_extension() directory = module.spool_rank_directory( sink_config.spool_root, storage_config._spool_destination(), _spool_producer_rank()) - return SpoolClaim(module.SpoolOwnerLock( + claim = SpoolClaim(module.SpoolOwnerLock( directory, allow_shared_filesystem=sink_config.spool_allow_shared_filesystem)) + count, example = _packs_outside_the_layout(sink_config.spool_root) + if count: + orphanage = os.path.join(sink_config.spool_root, + os.path.basename(os.path.dirname(directory)), + "r0-00000000") + _LOG.warning( + "spool_root %s holds %d ready pack(s) outside the per-process " + "layout (for example %s), which no engine adopts: left by an " + "engine from before the layout, or by a sink-only or " + "explicit-record_sink run. If they are bound for this catalog " + "and store, move them, keeping their paths below spool_root " + "(v1/...), into a directory of the layout nobody owns, such as " + "%s, and the next start on this node adopts them", + sink_config.spool_root, count, example, orphanage) + return claim def spool_owner_lock_beside(spool_root: str) -> str: diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index d1983bd9c..8340ceef9 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -501,6 +501,68 @@ def test_a_dead_spool_this_service_can_never_upload_is_left_and_reported( lock.release_and_remove_if_empty() +def test_packs_left_outside_the_layout_are_adopted_once_moved_as_told( + fake_s3, tmp_path, caplog): + """A sink-only run (or an engine from before the layout) leaves its + packs in spool_root itself, where nothing adopts them. The claim's + warning names a directory of the layout to move them into; moved there + with their paths kept, they are adopted like a dead process's.""" + import logging + import shutil + + import torch # noqa: F401 -- the sink extension links against it + + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix) + sys.path.insert(0, str(BUILD)) + try: + import _dmi_native_sink + finally: + sys.path.remove(str(BUILD)) + sink = _dmi_native_sink.NativePackSink( + spool_root=str(base), layout=LAYOUT, + max_pack_records=RECORDS_PER_PACK, max_linger_ns=600 * 10**9) + lease = sink.attach() + envelope = _envelope(STAGED_BY_THE_DEAD) + sink.submit_envelope(LAYOUT, envelope.rows, envelope.payload()) + assert sink.flush_and_wait(60.0) + del lease, sink + flat = sorted((base / "v1").rglob("*.dmi-pack.ready")) + assert len(flat) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + + with caplog.at_level(logging.WARNING, + logger="dmi.storage.native_capture"): + claim = claim_spool_directory( + NativeSinkConfig(spool_root=str(base)), config) + (warning,) = [record.getMessage() for record in caplog.records + if "outside" in record.getMessage()] + key = Path(claim.directory).parent.name + target = base / key / "r0-00000000" + assert f"{len(flat)} ready pack(s)" in warning + assert str(target) in warning + + target.mkdir() + shutil.move(str(base / "v1"), str(target / "v1")) + service = _service(config, claim.directory) + service.start() + try: + _wait_for(_adopted(service, 1)) + assert not target.exists() + service.flush(60.0) + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in envelope.expected.items()} + assert _read_all(config) == expected + finally: + service.stop() + claim.release() + + def test_a_sibling_whose_owner_dies_after_start_is_adopted_by_a_recheck( fake_s3, tmp_path): """A sibling still owned when the service starts -- a predecessor still diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 20117d1a2..4cbf83729 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -296,6 +296,48 @@ def test_a_dead_siblings_bytes_count_against_the_sinks_budget(tmp_path): claim.release() +def test_a_claim_warns_of_packs_left_outside_the_layout(tmp_path, caplog): + """Before the per-process layout the engine spooled into + /v1/... and its next start swept and uploaded whatever a + crashed run left there; a sink-only run still writes there. Nothing + adopts packs outside //r-/ now, so a claim + says so -- how many, where, and how to hand them to adoption -- rather + than leave them silently.""" + import logging + + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + root = tmp_path / "root" + sink = NativeSinkConfig(spool_root=str(root)) + # Packs inside the layout are adoption's, and no reason to warn. + first = claim_spool_directory(sink, _config()) + (Path(first.directory) / "v1").mkdir() + (Path(first.directory) / "v1" / "in.dmi-pack.ready").write_bytes(b"x") + first._lock.release() + with caplog.at_level(logging.WARNING, logger="dmi.storage.native_capture"): + claim = claim_spool_directory(sink, _config()) + claim.release() + assert not [r for r in caplog.records if "outside" in r.getMessage()] + + flat = root / "v1" / "tenant=t" / "date=2026-09-01" + flat.mkdir(parents=True) + for name in ("a", "b"): + (flat / f"{name}.dmi-pack.ready").write_bytes(b"pack") + (flat / ".c.0badf00d.open").write_bytes(b"half") # not a pack + caplog.clear() + with caplog.at_level(logging.WARNING, logger="dmi.storage.native_capture"): + claim = claim_spool_directory(sink, _config()) + key = Path(claim.directory).parent.name + claim.release() + (warning,) = [r.getMessage() for r in caplog.records + if "outside" in r.getMessage()] + assert "2 ready pack(s)" in warning + assert str(flat) in warning + assert f"{root}/{key}/r0-00000000" in warning + + def test_a_dropped_spool_claim_keeps_its_directory_owned(tmp_path): """Only release() lets go of a claim. An engine dropped without close() drops its claim, while the ring and the sink it activated may still be From 01d29a9cd203230ea644406f507c71275bc3298d Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:17:19 -0400 Subject: [PATCH 21/43] Track a spool lock's descriptor for the fork handler from its open() A child forked without exec closes its copies of the owner locks, so a fork-started worker never keeps a dead parent's directory looking live. But the handler only knew locks already held: a take registered its descriptor after LockInPlace or CreateLocked had opened, locked and recorded it, and a release unregistered it before the close. A fork from another thread inside either window -- the storage service's loop takes and lets go of dead siblings' locks without the GIL, while the main thread forks DataLoader workers -- gave the child a descriptor the handler did not close. One copied before the flock shares the description the flock then locks, so the child kept the lock after its owner died. The handler now tracks descriptors, not held objects. Every descriptor on a lock file is added to the set by the open() that makes it and removed by the close() that ends it, each under the mutex the fork's prepare handler takes; the child closes every one in the set. A SpoolOwnerLock records the fork generation it was taken in, which a child bumps, so the child's copies of the objects read as not held and never close a descriptor number reused since. Tests: test_spool_owner_lock.cpp, red before -- a fork inside a take, between the lock file's open() and its flock (the lock-open test seam), left the worker holding the lock after its owner was SIGKILLed; now the lock goes with the owner and the worker lives on. --- native/csrc/store/spool.cpp | 153 ++++++++++++++----------- native/csrc/store/spool.h | 23 ++-- tests/native/test_spool_owner_lock.cpp | 62 ++++++++++ 3 files changed, 157 insertions(+), 81 deletions(-) diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index dada9c105..35fea7091 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -633,6 +633,61 @@ bool SpoolOwnedByThisProcess(const std::string& dir) { namespace { +// Every descriptor this binary has open on a spool owner lock file: held, +// or between its open() and its flock, or on its way to close(). The fork +// handlers close the child's copies of all of them. Tracking starts at the +// open() and ends at the close(), each under the mutex that BeforeFork +// takes, so no fork -- from any thread, at any point of a take or a +// release -- hands a child a copy the handler does not know of: a copy +// made before the flock shares the description the flock then locks. +// Leaked on purpose, so no static destructor runs while one is open. Each +// binary that compiles spool.cpp (the store and sink extensions, the +// drivers) keeps its own set and its own handlers, for its own descriptors. +std::mutex& LockDescriptorsMutex() { + static std::mutex* mutex = new std::mutex; + return *mutex; +} +std::unordered_set& LockDescriptors() { + static auto* descriptors = new std::unordered_set; + return *descriptors; +} +// Bumped in each forked child, where every SpoolOwnerLock taken before the +// fork then reads as released (SpoolOwnerLock::held). +std::atomic g_fork_generation{0}; + +void BeforeForkLockDescriptors() { LockDescriptorsMutex().lock(); } +void AfterForkLockDescriptorsInParent() { LockDescriptorsMutex().unlock(); } +void AfterForkLockDescriptorsInChild() { + // Close, never LOCK_UN: an unlock on the shared description would drop + // the parent's hold too, and closing one of its descriptors does not. + for (const int fd : LockDescriptors()) ::close(fd); + LockDescriptors().clear(); + g_fork_generation.fetch_add(1, std::memory_order_relaxed); + LockDescriptorsMutex().unlock(); +} + +int OpenLockDescriptor(const char* path, int flags, mode_t mode) { + static std::once_flag handlers; + std::call_once(handlers, [] { + ::pthread_atfork(&BeforeForkLockDescriptors, + &AfterForkLockDescriptorsInParent, + &AfterForkLockDescriptorsInChild); + }); + std::lock_guard guard(LockDescriptorsMutex()); + const int fd = ::open(path, flags, mode); + const int failure = errno; + if (fd >= 0) LockDescriptors().insert(fd); + errno = failure; + return fd; +} + +void CloseLockDescriptor(int fd) { + if (fd < 0) return; + std::lock_guard guard(LockDescriptorsMutex()); + LockDescriptors().erase(fd); + ::close(fd); // closing the last descriptor unlocks +} + // Locks the lock file of an existing directory, creating the file if it has // none. Retries a lock lost to a remover's unlink, and a refusal as brief as // another process's ReadSpoolOwner probe. @@ -640,7 +695,8 @@ SpoolStatus LockInPlace(const std::string& dir, int* fd_out, std::string* error) { const std::string file = dir + "/" + kOwnerLockFile; for (int attempt = 0; attempt < 8; ++attempt) { - const int fd = ::open(file.c_str(), O_RDWR | O_CREAT | O_CLOEXEC, 0644); + const int fd = + OpenLockDescriptor(file.c_str(), O_RDWR | O_CREAT | O_CLOEXEC, 0644); if (fd < 0) { if (error) *error = "cannot open " + file + ": " + Errno(errno); return SpoolStatus::kIo; @@ -650,7 +706,7 @@ SpoolStatus LockInPlace(const std::string& dir, int* fd_out, const int failure = errno; SpoolOwner owner; ReadOwnerRecord(fd, &owner); - ::close(fd); + CloseLockDescriptor(fd); if (failure != EWOULDBLOCK) { if (error) *error = "cannot lock " + file + ": " + Errno(failure); return SpoolStatus::kIo; @@ -663,7 +719,7 @@ SpoolStatus LockInPlace(const std::string& dir, int* fd_out, return SpoolStatus::kOwned; } if (!IsFileAt(fd, file)) { - ::close(fd); + CloseLockDescriptor(fd); if (!fs::is_directory(dir)) { if (error) *error = "spool directory " + dir + " was removed while " "it was being locked"; @@ -703,11 +759,11 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, return SpoolStatus::kIo; } const std::string file = staging + "/" + kOwnerLockFile; - const int fd = - ::open(file.c_str(), O_RDWR | O_CREAT | O_EXCL | O_CLOEXEC, 0644); + const int fd = OpenLockDescriptor( + file.c_str(), O_RDWR | O_CREAT | O_EXCL | O_CLOEXEC, 0644); if (fd < 0 || ::flock(fd, LOCK_EX | LOCK_NB) != 0) { const int failure = errno; - if (fd >= 0) ::close(fd); + CloseLockDescriptor(fd); ::unlink(file.c_str()); ::rmdir(staging.c_str()); if (error) *error = "cannot lock " + file + ": " + Errno(failure); @@ -737,7 +793,7 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, #endif if (renamed != 0) { const int failure = errno; - ::close(fd); + CloseLockDescriptor(fd); ::unlink(file.c_str()); ::rmdir(staging.c_str()); if (failure == EEXIST || failure == ENOTEMPTY) { @@ -759,91 +815,48 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, } // namespace -namespace { -// Every held SpoolOwnerLock in this binary. Leaked on purpose, so no -// static destructor runs while a lock is still registered. Each binary -// that compiles spool.cpp (the store and sink extensions, the drivers) -// keeps its own registry and its own fork handlers, for its own locks. -std::mutex& LockRegistryMutex() { - static std::mutex* mutex = new std::mutex; - return *mutex; -} -std::unordered_set& LockRegistry() { - static auto* registry = new std::unordered_set; - return *registry; -} -} // namespace - -void SpoolOwnerLock::BeforeFork() { LockRegistryMutex().lock(); } -void SpoolOwnerLock::AfterForkInParent() { LockRegistryMutex().unlock(); } - -void SpoolOwnerLock::AfterForkInChild() { - // Close, never LOCK_UN: an unlock on the shared description would drop - // the parent's hold too, and closing one of its descriptors does not. - for (SpoolOwnerLock* lock : LockRegistry()) { - ::close(lock->fd_); - lock->fd_ = -1; - lock->dir_.clear(); - } - LockRegistry().clear(); - LockRegistryMutex().unlock(); -} - -void SpoolOwnerLock::Track(SpoolOwnerLock* lock) { - static std::once_flag handlers; - std::call_once(handlers, [] { - ::pthread_atfork(&SpoolOwnerLock::BeforeFork, - &SpoolOwnerLock::AfterForkInParent, - &SpoolOwnerLock::AfterForkInChild); - }); - std::lock_guard guard(LockRegistryMutex()); - LockRegistry().insert(lock); -} - -void SpoolOwnerLock::Untrack(SpoolOwnerLock* lock) { - std::lock_guard guard(LockRegistryMutex()); - LockRegistry().erase(lock); -} - void SpoolOwnerLock::Hold(int fd, std::string dir) { fd_ = fd; dir_ = std::move(dir); - Track(this); + generation_ = g_fork_generation.load(std::memory_order_relaxed); +} + +bool SpoolOwnerLock::held() const { + return fd_ >= 0 && + generation_ == g_fork_generation.load(std::memory_order_relaxed); } SpoolOwnerLock::~SpoolOwnerLock() { Release(); } SpoolOwnerLock::SpoolOwnerLock(SpoolOwnerLock&& other) noexcept { if (other.held()) { - const int fd = other.fd_; - std::string dir = std::move(other.dir_); - Untrack(&other); - other.fd_ = -1; - other.dir_.clear(); - Hold(fd, std::move(dir)); + fd_ = other.fd_; + dir_ = std::move(other.dir_); + generation_ = other.generation_; } + other.fd_ = -1; + other.dir_.clear(); } SpoolOwnerLock& SpoolOwnerLock::operator=(SpoolOwnerLock&& other) noexcept { if (this != &other) { Release(); if (other.held()) { - const int fd = other.fd_; - std::string dir = std::move(other.dir_); - Untrack(&other); - other.fd_ = -1; - other.dir_.clear(); - Hold(fd, std::move(dir)); + fd_ = other.fd_; + dir_ = std::move(other.dir_); + generation_ = other.generation_; } + other.fd_ = -1; + other.dir_.clear(); } return *this; } void SpoolOwnerLock::Release() { - if (fd_ >= 0) { - Untrack(this); - ::close(fd_); // closing the last descriptor unlocks - } + // Untracked and closed in one step (CloseLockDescriptor). In a forked + // child the fork handler closed the descriptor already, and its number + // may be another file's by now: nothing to close. + if (held()) CloseLockDescriptor(fd_); fd_ = -1; dir_.clear(); } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 8dc652482..92962f2c0 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -182,9 +182,11 @@ bool SpoolOwnedByThisProcess(const std::string& dir); // released with the object (or Release()), and by the kernel when the // process dies. flock binds to an open file description, which a child // shares after fork(): the descriptor is close-on-exec, and a child forked -// WITHOUT exec (a fork-started worker) closes its copy of every held lock -// at once (pthread_atfork), so the lock never outlives its owner in a child -// -- the owner's own hold is untouched, and the child's objects read as not +// WITHOUT exec (a fork-started worker) closes its copy of every lock +// descriptor at once (pthread_atfork) -- held, or opened and not yet locked +// or not yet closed, since a fork from another thread can land anywhere in +// a take or a release -- so the lock never outlives its owner in a child. +// The owner's own hold is untouched, and the child's objects read as not // held. A child the owner spawns through posix_spawn or vfork runs no // atfork handler, and loses the descriptor at exec. class SpoolOwnerLock { @@ -215,7 +217,7 @@ class SpoolOwnerLock { static SpoolStatus TryAdopt(const std::string& dir, SpoolOwnerLock* out, std::string* error); - bool held() const { return fd_ >= 0; } + bool held() const; // The canonical path of the directory, while held. const std::string& directory() const { return dir_; } @@ -228,17 +230,16 @@ class SpoolOwnerLock { bool ReleaseAndRemoveIfEmpty(std::string* error); private: - // Sets fd_ and dir_, and registers the lock for the fork handler. + // Takes over `fd`, which the fork handler has tracked since its open(). void Hold(int fd, std::string dir); - // The registry of held locks in this binary, and its fork handlers. - static void Track(SpoolOwnerLock* lock); - static void Untrack(SpoolOwnerLock* lock); - static void BeforeFork(); - static void AfterForkInParent(); - static void AfterForkInChild(); int fd_ = -1; std::string dir_; + // The fork generation the lock was taken in (spool.cpp): in a forked + // child, whose copies of the descriptors the fork handler closed, the + // object reads as not held, and never closes a descriptor number the + // child may have reused since. + uint64_t generation_ = 0; }; // Where a spool's packs go: the catalog -- its ClickHouse server, database diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index beb5f012e..c768df1c8 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -328,6 +328,67 @@ void TestAForkedChildDoesNotKeepTheLockPastItsParent() { ::kill(worker, SIGKILL); } +// (3c) A fork from another thread while a lock is being TAKEN: the child +// gets a copy of the lock file's descriptor between its open() and the +// flock -- the flock then locks the description both share. Registered +// with the fork handler only once the lock was held, that copy was never +// closed in the child, which kept the directory looking live after its +// owner died. Here the fork runs in that window, through the lock-open +// test seam. +void TestAForkWhileALockIsTakenLeavesTheChildNothing() { + const std::string root = FreshRoot("fork-window") + "/spool"; + fs::create_directories(root); + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t owner = ::fork(); + if (owner == 0) { + ::close(ready[0]); + pid_t worker = -1; + dmi_store::SetLockOpenHookForTesting([&](const std::string&) { + if (worker >= 0) return; + worker = ::fork(); // no exec + if (worker == 0) { + ::pause(); // outlives its parent until killed + ::_exit(0); + } + }); + SpoolOwnerLock lock; + std::string error; + const bool ok = + SpoolOwnerLock::TryAdopt(root, &lock, &error) == SpoolStatus::kOk; + dmi_store::SetLockOpenHookForTesting(nullptr); + char text[32]; + const int n = std::snprintf(text, sizeof(text), "%c%d\n", ok ? 'k' : 'x', + static_cast(worker)); + if (::write(ready[1], text, n) != n) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + std::string seen; + char byte = 0; + while (seen.find('\n') == std::string::npos && + ::read(ready[0], &byte, 1) == 1) { + seen.push_back(byte); + } + ::close(ready[0]); + CHECK(!seen.empty() && seen[0] == 'k'); + const pid_t worker = + static_cast(std::atoi(seen.empty() ? "" : seen.c_str() + 1)); + CHECK(worker > 0); + dmi_store::SpoolOwner record; + CHECK(dmi_store::ReadSpoolOwner(root, &record)); + CHECK(record.pid == owner); + + ::kill(owner, SIGKILL); + int status = 0; + ::waitpid(owner, &status, 0); + CHECK(worker > 0 && ::kill(worker, 0) == 0); // the worker lives on + // And the lock went with its owner. + CHECK(!dmi_store::ReadSpoolOwner(root, nullptr)); + if (worker > 0) ::kill(worker, SIGKILL); +} + // (4) Nesting, both ways. void TestNestedDirectoriesAreRefused() { const std::string base = FreshRoot("nested"); @@ -809,6 +870,7 @@ int main() { TestTheLockGoesWithItsSpool(); TestASecondProcessIsRefusedUntilTheHolderDies(); TestAForkedChildDoesNotKeepTheLockPastItsParent(); + TestAForkWhileALockIsTakenLeavesTheChildNothing(); TestNestedDirectoriesAreRefused(); TestAnOuterAndANestedTakeRacingNeverBothWin(); TestAStaleLockFileAboveDoesNotRefuseANestedDirectory(); From a285818318d97529e54ca159806bb35596504408 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:22:21 -0400 Subject: [PATCH 22/43] Lock the spool directory itself, not only its lock file Whether a spool directory's owner lives was judged by whichever inode sat at /.owner.lock at that moment. The owner locks a file it never touches again, and nothing checked that it stayed in place: once something other than its holder removed it -- systemd-tmpfiles aging /tmp (`q /tmp 1777 root root 10d` on RHEL and Fedora) takes a stale file and leaves fresh packs, or a person -- the live directory read as dead. TryAdopt then created a new lock file, took it, and Recover swept the owner's in-flight .open files: the race B6 is there to remove. A take now also flocks the directory itself (O_RDONLY|O_DIRECTORY), after its lock file, and CreateLocked locks the staging copy before the rename, so a new directory still appears with both held. A directory cannot be unlinked while it holds anything, so a lock file replaced behind a live owner's back no longer lets anyone in: the next take meets the directory's lock and is refused (kOwned, saying the lock file was replaced and so cannot name the holder). ReadSpoolOwner probes the directory's lock when the file's is free, and SpoolOwnedByThisProcess accepts this process's lock on either inode, so held_by_caller keeps working for the owner. systemd-tmpfiles, for its part, skips a flocked directory and everything below it, which keeps it from aging a live spool at all. The v1 contract says so, and to keep spool_root out of what other cleaners age. Tests: test_spool_owner_lock.cpp, red before -- with a live owner's lock file removed, the directory reads as held, TryAdopt and a take are refused, and once the owner dies it is adopted; in the owner's own process, SpoolOwnedByThisProcess and a held_by_caller Spool still see it held. test_native_spool_ownership.py -- the same from another process through the bound spool_owner and SpoolOwnerLock. --- docs/integration-api-v1.md | 9 +- native/csrc/store/spool.cpp | 195 +++++++++++++++++++------ native/csrc/store/spool.h | 40 +++-- tests/native/test_spool_owner_lock.cpp | 60 ++++++++ tests/test_native_spool_ownership.py | 29 ++++ 5 files changed, 273 insertions(+), 60 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index 9f697d85c..193d8ebca 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -192,9 +192,12 @@ hex digits of a sha256 of where the packs go -- the ClickHouse host and port, `database`, `table_prefix`, the S3 endpoint and bucket, and `store_id`, as spelled in the config -- the rank torchrun's `RANK`, 0 when unset, and the incarnation fresh on every -`create_record_runtime`), and owns it: an flock on its `.owner.lock`, taken -before the service starts and let go after the sink and the service are done, -when a drained directory is removed. If the sink did not seal within +`create_record_runtime`), and owns it: an flock on its `.owner.lock` and on the +directory itself, taken before the service starts and let go after the sink and +the service are done, when a drained directory is removed. The directory's own +lock keeps it owned should its `.owner.lock` be removed from under it, and +systemd-tmpfiles skips a flocked directory when it ages `/tmp`; other cleaners +may not, so keep `spool_root` out of what they age. If the sink did not seal within `close_flush_timeout_s`, it may still be staging, so the directory stays owned by the process until it exits (a warning names it) and the next process on the node adopts it; so does the directory of an engine dropped without `close()`. diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 35fea7091..6aa5135c2 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -552,20 +552,41 @@ SpoolStatus CheckNodeLocal(const std::string& dir, return SpoolStatus::kBadArgument; } -bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner) { - const std::string file = dir + "/" + kOwnerLockFile; - const int fd = ::open(file.c_str(), O_RDONLY | O_CLOEXEC); - if (fd < 0) return false; - // A probe: if the lock can be taken nobody holds it, and it is let go at - // once. (A take racing the probe retries, see LockInPlace.) +namespace { +// A probe of one flock: if it can be taken nobody holds it, and it is let +// go at once. (A take racing the probe retries, see LockInPlace.) An +// unlock on the probe's own description, so a child forked meanwhile +// keeps nothing either. +bool FlockHeld(int fd) { if (::flock(fd, LOCK_EX | LOCK_NB) == 0) { ::flock(fd, LOCK_UN); - ::close(fd); return false; } - const bool held = errno == EWOULDBLOCK; - if (held && owner != nullptr) ReadOwnerRecord(fd, owner); - ::close(fd); + return errno == EWOULDBLOCK; +} +} // namespace + +bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner) { + const std::string file = dir + "/" + kOwnerLockFile; + const int fd = ::open(file.c_str(), O_RDONLY | O_CLOEXEC); + bool held = fd >= 0 && FlockHeld(fd); + if (!held) { + // The directory's own lock, which its owner keeps however its lock + // file is replaced (SpoolOwnerLock). + const int dir_fd = ::open(dir.c_str(), O_RDONLY | O_DIRECTORY | O_CLOEXEC); + if (dir_fd >= 0) { + held = FlockHeld(dir_fd); + ::close(dir_fd); + } + } + if (held && owner != nullptr) { + if (fd >= 0) { + ReadOwnerRecord(fd, owner); + } else { + *owner = SpoolOwner{}; + } + } + if (fd >= 0) ::close(fd); return held; } @@ -585,8 +606,17 @@ bool IsSpoolClaimStagingName(const std::string& name) { bool SpoolOwnedByThisProcess(const std::string& dir) { const std::string file = dir + "/" + kOwnerLockFile; - struct stat target{}; - if (::stat(file.c_str(), &target) != 0) return false; + // The lock file, and the directory itself, which its owner locks too: a + // lock file replaced behind the owner's back is no longer the one it + // holds, but the directory is. + struct stat targets[2]{}; + size_t n_targets = 0; + if (::stat(file.c_str(), &targets[n_targets]) == 0) ++n_targets; + if (::stat(dir.c_str(), &targets[n_targets]) == 0 && + S_ISDIR(targets[n_targets].st_mode)) { + ++n_targets; + } + if (n_targets == 0) return false; // /proc/self/fdinfo/ lists the flocks each open file description // holds ("lock: 1: FLOCK ADVISORY WRITE ..."), so the kernel says // whether one of this process's descriptors on the file holds the lock -- @@ -608,10 +638,13 @@ bool SpoolOwnedByThisProcess(const std::string& dir) { const long fd = std::strtol(entry->d_name, &end, 10); if (end == entry->d_name || *end != '\0' || fd == listing) continue; struct stat by_fd{}; - if (::fstat(static_cast(fd), &by_fd) != 0 || - by_fd.st_dev != target.st_dev || by_fd.st_ino != target.st_ino) { - continue; + if (::fstat(static_cast(fd), &by_fd) != 0) continue; + bool on_target = false; + for (size_t i = 0; i < n_targets; ++i) { + on_target = on_target || (by_fd.st_dev == targets[i].st_dev && + by_fd.st_ino == targets[i].st_ino); } + if (!on_target) continue; const std::string info = std::string("/proc/self/fdinfo/") + entry->d_name; std::FILE* in = std::fopen(info.c_str(), "re"); @@ -688,15 +721,66 @@ void CloseLockDescriptor(int fd) { ::close(fd); // closing the last descriptor unlocks } -// Locks the lock file of an existing directory, creating the file if it has -// none. Retries a lock lost to a remover's unlink, and a refusal as brief as -// another process's ReadSpoolOwner probe. -SpoolStatus LockInPlace(const std::string& dir, int* fd_out, +// The two locks of a held spool directory: its lock file's, which records +// the holder, and the directory's own. The directory cannot be unlinked +// while it holds anything, so a lock file removed behind a live owner's +// back -- by an age-based cleaner such as systemd-tmpfiles, which also +// skips a directory that is flocked, or by a person -- leaves the +// directory owned: the next take meets its lock and is refused, where by +// the new lock file alone it took the live directory for a dead one. +struct HeldLock { + int file_fd = -1; + int dir_fd = -1; +}; + +void CloseHeldLock(HeldLock* lock) { + CloseLockDescriptor(lock->dir_fd); + CloseLockDescriptor(lock->file_fd); + *lock = HeldLock{}; +} + +// Takes `dir`'s own lock beside its lock file's, which `lock` holds. kOwned +// while another holder has it: an owner whose lock file was replaced (or +// is being probed, which a retry outlasts). +SpoolStatus LockDirectory(const std::string& dir, HeldLock* lock, + std::string* error) { + lock->dir_fd = OpenLockDescriptor(dir.c_str(), + O_RDONLY | O_DIRECTORY | O_CLOEXEC, 0); + if (lock->dir_fd < 0) { + if (error) *error = "cannot open spool directory " + dir + ": " + + Errno(errno); + return SpoolStatus::kIo; + } + if (::flock(lock->dir_fd, LOCK_EX | LOCK_NB) == 0) return SpoolStatus::kOk; + const int failure = errno; + CloseLockDescriptor(lock->dir_fd); + lock->dir_fd = -1; + if (failure != EWOULDBLOCK) { + if (error) *error = "cannot lock spool directory " + dir + ": " + + Errno(failure); + return SpoolStatus::kIo; + } + if (error) { + *error = "spool directory " + dir + " is owned by another process, " + "which holds the directory's own lock: its " + kOwnerLockFile + + " was replaced since, so that file does not name it; a spool " + "directory has one owner process"; + } + return SpoolStatus::kOwned; +} + +// Locks an existing directory -- its lock file, creating the file if it +// has none, and the directory itself. Retries a lock lost to a remover's +// unlink, and a refusal as brief as another process's ReadSpoolOwner +// probe. +SpoolStatus LockInPlace(const std::string& dir, HeldLock* out, std::string* error) { const std::string file = dir + "/" + kOwnerLockFile; for (int attempt = 0; attempt < 8; ++attempt) { - const int fd = + HeldLock lock; + lock.file_fd = OpenLockDescriptor(file.c_str(), O_RDWR | O_CREAT | O_CLOEXEC, 0644); + const int fd = lock.file_fd; if (fd < 0) { if (error) *error = "cannot open " + file + ": " + Errno(errno); return SpoolStatus::kIo; @@ -706,7 +790,7 @@ SpoolStatus LockInPlace(const std::string& dir, int* fd_out, const int failure = errno; SpoolOwner owner; ReadOwnerRecord(fd, &owner); - CloseLockDescriptor(fd); + CloseHeldLock(&lock); if (failure != EWOULDBLOCK) { if (error) *error = "cannot lock " + file + ": " + Errno(failure); return SpoolStatus::kIo; @@ -719,7 +803,7 @@ SpoolStatus LockInPlace(const std::string& dir, int* fd_out, return SpoolStatus::kOwned; } if (!IsFileAt(fd, file)) { - CloseLockDescriptor(fd); + CloseHeldLock(&lock); if (!fs::is_directory(dir)) { if (error) *error = "spool directory " + dir + " was removed while " "it was being locked"; @@ -727,20 +811,29 @@ SpoolStatus LockInPlace(const std::string& dir, int* fd_out, } continue; } + const SpoolStatus directory = LockDirectory(dir, &lock, error); + if (directory != SpoolStatus::kOk) { + CloseHeldLock(&lock); + if (directory == SpoolStatus::kOwned && attempt < 2) { + std::this_thread::sleep_for(std::chrono::milliseconds(2)); + continue; + } + return directory; + } WriteOwnerRecord(fd); - *fd_out = fd; + *out = lock; return SpoolStatus::kOk; } if (error) *error = "cannot lock " + file + ": it keeps being replaced"; return SpoolStatus::kIo; } -// Creates `dir` with its lock file already held: built under a hidden name +// Creates `dir` with its locks already held: built under a hidden name // beside it, then renamed into place, so a scan of the parent never meets // the directory unowned (an adopter would otherwise take a brand-new // sibling for a dead one). Falls back to LockInPlace if `dir` appears // meanwhile. -SpoolStatus CreateLocked(const std::string& dir, int* fd_out, +SpoolStatus CreateLocked(const std::string& dir, HeldLock* out, std::string* error) { const fs::path target(dir); const std::string parent = target.parent_path().string(); @@ -759,16 +852,25 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, return SpoolStatus::kIo; } const std::string file = staging + "/" + kOwnerLockFile; - const int fd = OpenLockDescriptor( + HeldLock lock; + lock.file_fd = OpenLockDescriptor( file.c_str(), O_RDWR | O_CREAT | O_EXCL | O_CLOEXEC, 0644); + const int fd = lock.file_fd; if (fd < 0 || ::flock(fd, LOCK_EX | LOCK_NB) != 0) { const int failure = errno; - CloseLockDescriptor(fd); + CloseHeldLock(&lock); ::unlink(file.c_str()); ::rmdir(staging.c_str()); if (error) *error = "cannot lock " + file + ": " + Errno(failure); return SpoolStatus::kIo; } + // Nobody else knows the staging copy yet, so its own lock is free. + if (LockDirectory(staging, &lock, error) != SpoolStatus::kOk) { + CloseHeldLock(&lock); + ::unlink(file.c_str()); + ::rmdir(staging.c_str()); + return SpoolStatus::kIo; + } WriteOwnerRecord(fd); ::fsync(fd); FsyncDir(staging, nullptr); @@ -793,11 +895,11 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, #endif if (renamed != 0) { const int failure = errno; - CloseLockDescriptor(fd); + CloseHeldLock(&lock); ::unlink(file.c_str()); ::rmdir(staging.c_str()); if (failure == EEXIST || failure == ENOTEMPTY) { - return LockInPlace(dir, fd_out, error); + return LockInPlace(dir, out, error); } if (error) { *error = "cannot create spool directory " + dir + ": " + @@ -806,7 +908,7 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, return SpoolStatus::kIo; } FsyncDir(parent, nullptr); - *fd_out = fd; + *out = lock; // the directory's lock went with the rename return SpoolStatus::kOk; } if (error) *error = "cannot create spool directory " + dir; @@ -815,8 +917,9 @@ SpoolStatus CreateLocked(const std::string& dir, int* fd_out, } // namespace -void SpoolOwnerLock::Hold(int fd, std::string dir) { +void SpoolOwnerLock::Hold(int fd, int dir_fd, std::string dir) { fd_ = fd; + dir_fd_ = dir_fd; dir_ = std::move(dir); generation_ = g_fork_generation.load(std::memory_order_relaxed); } @@ -831,10 +934,12 @@ SpoolOwnerLock::~SpoolOwnerLock() { Release(); } SpoolOwnerLock::SpoolOwnerLock(SpoolOwnerLock&& other) noexcept { if (other.held()) { fd_ = other.fd_; + dir_fd_ = other.dir_fd_; dir_ = std::move(other.dir_); generation_ = other.generation_; } other.fd_ = -1; + other.dir_fd_ = -1; other.dir_.clear(); } @@ -843,21 +948,27 @@ SpoolOwnerLock& SpoolOwnerLock::operator=(SpoolOwnerLock&& other) noexcept { Release(); if (other.held()) { fd_ = other.fd_; + dir_fd_ = other.dir_fd_; dir_ = std::move(other.dir_); generation_ = other.generation_; } other.fd_ = -1; + other.dir_fd_ = -1; other.dir_.clear(); } return *this; } void SpoolOwnerLock::Release() { - // Untracked and closed in one step (CloseLockDescriptor). In a forked - // child the fork handler closed the descriptor already, and its number - // may be another file's by now: nothing to close. - if (held()) CloseLockDescriptor(fd_); + // Each untracked and closed in one step (CloseLockDescriptor). In a + // forked child the fork handler closed them already, and their numbers + // may be other files' by now: nothing to close. + if (held()) { + CloseLockDescriptor(dir_fd_); + CloseLockDescriptor(fd_); + } fd_ = -1; + dir_fd_ = -1; dir_.clear(); } @@ -891,11 +1002,11 @@ SpoolStatus SpoolOwnerLock::Acquire(const std::string& dir, if (status != SpoolStatus::kOk) return status; const std::string lock_file = canonical + "/" + kOwnerLockFile; const bool had_lock_file = exists && fs::exists(lock_file, ec); - int fd = -1; - status = exists ? LockInPlace(canonical, &fd, error) - : CreateLocked(canonical, &fd, error); + HeldLock lock; + status = exists ? LockInPlace(canonical, &lock, error) + : CreateLocked(canonical, &lock, error); if (status != SpoolStatus::kOk) return status; - out->Hold(fd, canonical); + out->Hold(lock.file_fd, lock.dir_fd, canonical); // Only now, with this lock published: see CheckNotNested. status = CheckNotNested(canonical, error); if (status != SpoolStatus::kOk) { @@ -923,10 +1034,10 @@ SpoolStatus SpoolOwnerLock::TryAdopt(const std::string& dir, if (error) *error = "no spool directory to adopt at " + dir; return SpoolStatus::kBadArgument; } - int fd = -1; - const SpoolStatus status = LockInPlace(resolved, &fd, error); + HeldLock lock; + const SpoolStatus status = LockInPlace(resolved, &lock, error); if (status != SpoolStatus::kOk) return status; - out->Hold(fd, resolved); + out->Hold(lock.file_fd, lock.dir_fd, resolved); return SpoolStatus::kOk; } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 92962f2c0..acbd2984f 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -15,9 +15,10 @@ // One owner per directory (B6). Recover() deletes every .open file this // object is not writing, so it is only safe while no other PROCESS writes // there. A spool directory therefore has an owner lock -- flock(LOCK_EX) on -// /.owner.lock, whose content records the holder's host and pid -- -// and a second process that tries to take it is refused, told who holds -// it. The lock goes with its holder, even one killed with SIGKILL. +// /.owner.lock, whose content records the holder's host and pid, and +// on itself -- and a second process that tries to take it is +// refused, told who holds it. The lock goes with its holder, even one +// killed with SIGKILL. // - owner_lock=kTake (the default) takes it in Open(), before anything // reads the directory, and holds it for the Spool object's life. // - owner_lock=kHeldByCaller takes none: the calling process holds a @@ -161,9 +162,11 @@ struct SpoolOwner { int64_t pid = 0; }; -// Whether /.owner.lock is held right now (by any process, this one -// included), and if so who recorded themselves in it. A holder that has -// locked but not yet written its record reads as an empty host and pid 0. +// Whether 's owner lock is held right now (by any process, this one +// included) -- its lock file's, or the directory's own -- and if so who +// recorded themselves in the lock file. A holder that has locked but not +// yet written its record, or whose lock file was replaced, reads as an +// empty host and pid 0. bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner); // Whether `name` is the staging copy of a directory SpoolOwnerLock::Acquire @@ -173,14 +176,19 @@ bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner); // ignored by the nesting check, and an adopter clears it. bool IsSpoolClaimStagingName(const std::string& name); -// Whether one of THIS process's descriptors holds /.owner.lock, as the -// kernel reports it in /proc/self/fdinfo (falling back to the recorded host -// and pid where /proc cannot be read). What kHeldByCaller requires. +// Whether one of THIS process's descriptors holds 's owner lock (its +// lock file's, or the directory's own), as the kernel reports it in +// /proc/self/fdinfo (falling back to the recorded host and pid where /proc +// cannot be read). What kHeldByCaller requires. bool SpoolOwnedByThisProcess(const std::string& dir); -// The owner lock of one spool directory: flock(LOCK_EX) on /.owner.lock, -// released with the object (or Release()), and by the kernel when the -// process dies. flock binds to an open file description, which a child +// The owner lock of one spool directory: flock(LOCK_EX) on /.owner.lock +// and on itself, released with the object (or Release()), and by the +// kernel when the process dies. The directory's lock keeps a live owner's +// directory owned should its lock file be removed from under it -- by an +// age-based cleaner (systemd-tmpfiles, which also leaves a flocked +// directory and everything below it alone) or by a person: the next take +// meets the directory's lock, not only a new lock file nobody holds. flock binds to an open file description, which a child // shares after fork(): the descriptor is close-on-exec, and a child forked // WITHOUT exec (a fork-started worker) closes its copy of every lock // descriptor at once (pthread_atfork) -- held, or opened and not yet locked @@ -230,10 +238,12 @@ class SpoolOwnerLock { bool ReleaseAndRemoveIfEmpty(std::string* error); private: - // Takes over `fd`, which the fork handler has tracked since its open(). - void Hold(int fd, std::string dir); + // Takes over the two descriptors, which the fork handler has tracked + // since their open(). + void Hold(int fd, int dir_fd, std::string dir); - int fd_ = -1; + int fd_ = -1; // on /.owner.lock + int dir_fd_ = -1; // on itself std::string dir_; // The fork generation the lock was taken in (spool.cpp): in a forked // child, whose copies of the descriptors the fork handler closed, the diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index c768df1c8..c9a36e565 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -675,6 +675,65 @@ void TestALockOnAnUnlinkedFileIsTakenAgain() { CHECK(SpoolOwnerLock::TryAdopt(dir, &rival, &error) == SpoolStatus::kOwned); } +// (6c) Liveness is not the lock file's alone. Something other than its +// holder -- an age-based cleaner such as systemd-tmpfiles, a person -- can +// unlink /.owner.lock while the owner lives: the owner's descriptor is +// then on the unlinked file, and whoever opens the path next meets a new +// one that nobody holds. Judged by that file alone the live directory read +// as dead, an adopter took it, and its Recover swept the owner's in-flight +// .open files. The owner also locks the directory itself, which cannot be +// unlinked while it holds anything. +void TestAReplacedLockFileLeavesTheDirectoryOwned() { + const std::string dir = FreshRoot("replaced") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t owner = ::fork(); + if (owner == 0) { + ::close(ready[0]); + SpoolOwnerLock lock; + std::string error; + const bool ok = SpoolOwnerLock::Acquire(dir, false, &lock, &error) == + SpoolStatus::kOk; + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + + CHECK(fs::remove(dir + "/.owner.lock")); + CHECK(dmi_store::ReadSpoolOwner(dir, nullptr)); // still owned + std::string error; + SpoolOwnerLock adopter; + CHECK(SpoolOwnerLock::TryAdopt(dir, &adopter, &error) == + SpoolStatus::kOwned); + CHECK(!adopter.held()); + Spool spool; + CHECK(Spool::Open({dir, 1 << 20}, &spool, &error) == SpoolStatus::kOwned); + + ::kill(owner, SIGKILL); + int status = 0; + ::waitpid(owner, &status, 0); + CHECK(!dmi_store::ReadSpoolOwner(dir, nullptr)); + CHECK(SpoolOwnerLock::TryAdopt(dir, &adopter, &error) == SpoolStatus::kOk); + + // In the owner's own process too: its held_by_caller Spools still open. + adopter.Release(); + SpoolOwnerLock mine; + CHECK(SpoolOwnerLock::Acquire(dir, false, &mine, &error) == + SpoolStatus::kOk); + CHECK(fs::remove(dir + "/.owner.lock")); + CHECK(dmi_store::SpoolOwnedByThisProcess(dir)); + SpoolConfig beside{dir, 1 << 20}; + beside.owner_lock = OwnerLock::kHeldByCaller; + Spool held; + CHECK(Spool::Open(beside, &held, &error) == SpoolStatus::kOk); +} + void TestANewDirectoryAppearsWithItsLockHeld() { // Created beside its lock file and renamed into place, so no scan of the // parent can meet the directory before its owner holds it. @@ -878,6 +937,7 @@ int main() { TestSharedFilesystemsAreRefusedUnlessAllowed(); TestAdoptionLocksOnlyWhatExistsAndIsDead(); TestALockOnAnUnlinkedFileIsTakenAgain(); + TestAReplacedLockFileLeavesTheDirectoryOwned(); TestANewDirectoryAppearsWithItsLockHeld(); TestTheDirectoryLayout(); TestDeadSiblingsCountAgainstTheBudget(); diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 4cbf83729..bed5b6db3 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -208,6 +208,35 @@ def test_a_forked_worker_does_not_keep_a_dead_owners_lock(tmp_path): owner.wait(timeout=30) +def test_a_removed_lock_file_leaves_a_live_directory_owned(tmp_path): + """Something other than the owner -- an age-based cleaner such as + systemd-tmpfiles, a person -- can remove .owner.lock from under a live + owner. Judged by that file alone, another process then saw nobody + holding the directory, took it, and could sweep the owner's stages in + flight. The owner locks the directory itself as well.""" + directory = tmp_path / "spool" + with _store().SpoolOwnerLock(str(directory)) as lock: + (directory / ".owner.lock").unlink() + probe = ( + "import sys; sys.path.insert(0, sys.argv[1]);" + "import _dmi_native_store as m\n" + "print(m.spool_owner(sys.argv[2]) is not None)\n" + "try:\n" + " m.SpoolOwnerLock(sys.argv[2])\n" + " print('taken')\n" + "except m.SpoolOwnedError:\n" + " print('owned')\n") + result = subprocess.run( + [sys.executable, "-c", probe, str(BUILD), str(directory)], + capture_output=True, text=True, timeout=60) + assert result.returncode == 0, result.stderr + assert result.stdout.split() == ["True", "owned"], result.stdout + assert lock.held + assert _store().spool_owner(str(directory)) is None + with _store().SpoolOwnerLock(str(directory)): + pass + + def test_spool_owner_names_the_holder_while_it_holds(tmp_path): store = _store() directory = str(tmp_path / "spool") From 2f2a68b42aa5fd83385793d6ea6ebfa9d7797193 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:23:55 -0400 Subject: [PATCH 23/43] Judge held_by_caller by the owner record where fdinfo lists no flocks held_by_caller -- the engine's service, its sink and every adoption -- asks the kernel whether one of this process's descriptors holds the directory's lock, through the "lock:" lines of /proc/self/fdinfo, and fell back to the owner record only when /proc/self/fd could not be opened. A procfs that lists descriptors but no locks answered "no" for the process that really held the lock: gVisor's fdinfo prints only pos, flags and mnt_id (Modal, GKE Sandbox), and WSL1's lists none either. There every held_by_caller open was refused kOwned, naming the caller's own pid, so the default persistent path could not start at all, and every adoption would have blocked its sibling. Before B6 that configuration worked. Whether this kernel lists flocks is now probed once per process, with a flock on a temporary file; where it does not (or no temporary file can be made), the record decides -- this host and this pid -- as the Python spool_owner_lock_beside already does. SetFdinfoHidesLocksForTesting makes this binary's fdinfo reads see no lock lines. Tests: test_spool_owner_lock.cpp, red before -- with the seam hiding lock lines, the process holding a lock is told so and opens its own held_by_caller Spool; beside another process's lock the open is still refused, naming that holder. --- native/csrc/store/spool.cpp | 64 +++++++++++++++++++------- native/csrc/store/spool.h | 7 ++- tests/native/test_spool_owner_lock.cpp | 56 ++++++++++++++++++++++ 3 files changed, 109 insertions(+), 18 deletions(-) diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 6aa5135c2..288f787ff 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -604,6 +604,48 @@ bool IsSpoolClaimStagingName(const std::string& name) { }); } +namespace { +std::atomic g_fdinfo_hides_locks_for_testing{false}; + +// Whether /proc/self/fdinfo/ lists a write flock on the descriptor's +// open file description ("lock: 1: FLOCK ADVISORY WRITE ..."). +bool FdinfoShowsWriteFlock(const std::string& fd) { + if (g_fdinfo_hides_locks_for_testing.load()) return false; + std::FILE* in = std::fopen(("/proc/self/fdinfo/" + fd).c_str(), "re"); + if (in == nullptr) return false; + bool shown = false; + char line[512]; + while (!shown && std::fgets(line, sizeof(line), in) != nullptr) { + shown = std::strncmp(line, "lock:", 5) == 0 && + std::strstr(line, " FLOCK ") != nullptr && + std::strstr(line, " WRITE ") != nullptr; + } + std::fclose(in); + return shown; +} +// Whether this kernel lists flocks in /proc/self/fdinfo at all. Linux does +// (since 3.8); gVisor's procfs prints only pos, flags and mnt_id, and +// WSL1's lists no locks either. Probed once, on a flock taken on a +// temporary file; no temporary file, no telling, and the record decides. +bool FdinfoListsFlocks() { + if (g_fdinfo_hides_locks_for_testing.load()) return false; + static const bool lists = [] { + std::FILE* temp = std::tmpfile(); + if (temp == nullptr) return false; + const int fd = ::fileno(temp); + const bool shown = ::flock(fd, LOCK_EX | LOCK_NB) == 0 && + FdinfoShowsWriteFlock(std::to_string(fd)); + std::fclose(temp); + return shown; + }(); + return lists; +} +} // namespace + +void SetFdinfoHidesLocksForTesting(bool hide) { + g_fdinfo_hides_locks_for_testing.store(hide); +} + bool SpoolOwnedByThisProcess(const std::string& dir) { const std::string file = dir + "/" + kOwnerLockFile; // The lock file, and the directory itself, which its owner locks too: a @@ -621,9 +663,10 @@ bool SpoolOwnedByThisProcess(const std::string& dir) { // holds ("lock: 1: FLOCK ADVISORY WRITE ..."), so the kernel says // whether one of this process's descriptors on the file holds the lock -- // a SpoolOwnerLock's, or one a Spool took with kTake. The record in the - // file is only a fallback: it is written after the lock is taken, and a - // pid says nothing across pid namespaces. - DIR* fds = ::opendir("/proc/self/fd"); + // file is only a fallback, where /proc cannot be read or lists no flocks + // (FdinfoListsFlocks): it is written after the lock is taken, and a pid + // says nothing across pid namespaces. + DIR* fds = FdinfoListsFlocks() ? ::opendir("/proc/self/fd") : nullptr; if (fds == nullptr) { SpoolOwner owner; return ReadSpoolOwner(dir, &owner) && owner.pid == ::getpid() && @@ -645,20 +688,7 @@ bool SpoolOwnedByThisProcess(const std::string& dir) { by_fd.st_ino == targets[i].st_ino); } if (!on_target) continue; - const std::string info = - std::string("/proc/self/fdinfo/") + entry->d_name; - std::FILE* in = std::fopen(info.c_str(), "re"); - if (in == nullptr) continue; - char line[512]; - while (std::fgets(line, sizeof(line), in) != nullptr) { - if (std::strncmp(line, "lock:", 5) == 0 && - std::strstr(line, " FLOCK ") != nullptr && - std::strstr(line, " WRITE ") != nullptr) { - held = true; - break; - } - } - std::fclose(in); + held = FdinfoShowsWriteFlock(entry->d_name); } ::closedir(fds); return held; diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index acbd2984f..210495ce0 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -151,6 +151,10 @@ SpoolStatus CheckNodeLocal(const std::string& dir, // calling statfs(2). A negative value restores statfs. void SetFilesystemTypeForTesting(int64_t f_type); +// Test seam: every /proc/self/fdinfo read in this binary sees no "lock:" +// lines, as under gVisor or WSL1, whose procfs lists none. +void SetFdinfoHidesLocksForTesting(bool hide); + // Test seam: taking an existing directory's lock calls `hook` with the lock // file's path after opening the file and before locking it -- the window in // which a remover can unlink it. An empty function removes the hook. @@ -179,7 +183,8 @@ bool IsSpoolClaimStagingName(const std::string& name); // Whether one of THIS process's descriptors holds 's owner lock (its // lock file's, or the directory's own), as the kernel reports it in // /proc/self/fdinfo (falling back to the recorded host and pid where /proc -// cannot be read). What kHeldByCaller requires. +// cannot be read, or its fdinfo lists no flocks at all, as under gVisor or +// WSL1). What kHeldByCaller requires. bool SpoolOwnedByThisProcess(const std::string& dir); // The owner lock of one spool directory: flock(LOCK_EX) on /.owner.lock diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index c9a36e565..851fcf0c7 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -176,6 +176,61 @@ void TestHeldByCallerBesideAnotherProcessIsRefused() { ::waitpid(child, &status, 0); } +// (2c) Where the kernel lists no flocks in /proc/self/fdinfo -- gVisor's +// procfs prints only pos/flags/mnt_id, WSL1's none either -- the check +// fell back to the owner record only when /proc/self/fd could not be +// opened, so the process that really held the lock was refused its own +// held_by_caller Spools, naming its own pid, and no default-mode capture +// could start. There the record decides: this host and pid, or not. +void TestHeldByCallerWhereTheKernelListsNoFlocks() { + const std::string root = FreshRoot("no-fdinfo-locks") + "/spool"; + dmi_store::SetFdinfoHidesLocksForTesting(true); + SpoolOwnerLock lock; + std::string error; + CHECK(SpoolOwnerLock::Acquire(root, false, &lock, &error) == + SpoolStatus::kOk); + CHECK(dmi_store::SpoolOwnedByThisProcess(root)); + SpoolConfig config{root, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool spool; + error.clear(); + CHECK(Spool::Open(config, &spool, &error) == SpoolStatus::kOk); + CHECK(error.empty()); + + // Beside another process's lock it is still refused, naming that holder. + const std::string other = FreshRoot("no-fdinfo-locks-other") + "/spool"; + int ready[2]; + CHECK(::pipe(ready) == 0); + const pid_t child = ::fork(); + if (child == 0) { + ::close(ready[0]); + SpoolOwnerLock held; + std::string child_error; + const bool ok = SpoolOwnerLock::Acquire(other, false, &held, + &child_error) == SpoolStatus::kOk; + const char byte = ok ? '1' : '0'; + if (::write(ready[1], &byte, 1) != 1) ::_exit(3); + ::pause(); // until killed + ::_exit(0); + } + ::close(ready[1]); + char byte = 0; + CHECK(::read(ready[0], &byte, 1) == 1); + CHECK(byte == '1'); + ::close(ready[0]); + CHECK(!dmi_store::SpoolOwnedByThisProcess(other)); + SpoolConfig beside{other, 1 << 20}; + beside.owner_lock = OwnerLock::kHeldByCaller; + Spool refused; + error.clear(); + CHECK(Spool::Open(beside, &refused, &error) == SpoolStatus::kOwned); + CHECK(Contains(error, "pid " + std::to_string(child))); + ::kill(child, SIGKILL); + int status = 0; + ::waitpid(child, &status, 0); + dmi_store::SetFdinfoHidesLocksForTesting(false); +} + void TestHeldByCallerWithoutAHolderIsRefused() { const std::string root = FreshRoot("unheld") + "/spool"; SpoolConfig config{root, 1 << 20}; @@ -925,6 +980,7 @@ int main() { TestTwoTakesInOneProcessRefuseEachOther(); TestHeldByCallerOpensBesideTheHolder(); TestHeldByCallerBesideAnotherProcessIsRefused(); + TestHeldByCallerWhereTheKernelListsNoFlocks(); TestHeldByCallerWithoutAHolderIsRefused(); TestTheLockGoesWithItsSpool(); TestASecondProcessIsRefusedUntilTheHolderDies(); From 2cb2337dd206bb5383fb2990d2f3bbf98b314610 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:29:05 -0400 Subject: [PATCH 24/43] Charge a sink only what an adoption can drain beside it charge_dead_siblings exempted only siblings another process held, and charged every one this process held on the grounds that its service was adopting it: SpoolOwnedByThisProcess cannot tell an adoption's lock from a claim's. But two kinds of this-process-held directory are no adoption's -- a claim kept owned because its sink did not seal (by design, until exit), and the claim of an engine dropped without close() -- and the service reads them as live and never adopts them. So after a seal that timed out while the store was down, the same process's next record runtime (a notebook, a server, an adapter re-attach, a ring replacement) had its budget cut by the kept directory's bytes for the rest of the process's life, possibly to nothing, and was refused with "still to be adopted". Siblings the service blocked were charged by every sink on the node for good. The lock file's record now says why its holder took it: TryAdopt, the adopter's take, adds an "adopting" line, and an adopter leaving a directory blocked writes "blocked: " into it (MarkBlocked) before it lets go; the next take rewrites the record. ReadSpoolOwner fills the record whether or not the lock is held. A sink then charges a sibling only while an adoption can drain it: one nobody holds that is not marked blocked and that this process could take and empty (another user's is not), and one this process holds as an adoption. A kept claim, another process's directory, and a blocked one are not charged, and the refusal says the charged bytes are in dead directories this process's storage service adopts. A blocked directory's lock file also says why for a person. The v1 contract says what is charged. Tests: test_spool_owner_lock.cpp, red before -- a sibling this process holds through a take is not charged and the sink gets its whole budget; taken by TryAdopt it is charged; marked blocked and let go it is not (and the record says why), taken again the mark is gone and it is charged; one this process cannot write is not charged. test_native_spool_ownership.py, red before -- a kept claim beside a new one leaves the new sink its budget. test_native_spool_adoption_live.py -- a blocked sibling's lock file carries the reason. --- docs/integration-api-v1.md | 13 +-- native/csrc/catalog/storage_service.cpp | 20 +++-- native/csrc/catalog/storage_service.h | 9 +- native/csrc/store/spool.cpp | 105 +++++++++++++++++------ native/csrc/store/spool.h | 54 ++++++++---- tests/native/test_spool_owner_lock.cpp | 91 ++++++++++++++++---- tests/test_native_spool_adoption_live.py | 8 ++ tests/test_native_spool_ownership.py | 33 +++++++ 8 files changed, 264 insertions(+), 69 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index 193d8ebca..53f5b1cab 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -215,11 +215,14 @@ storage part of `capture_status()` reports the adoption (`adopted_spools`, service can never adopt -- one holding a pack it can never upload, such as one larger than its `uploader_max_in_flight_bytes` or one whose key already holds a different object, or one it cannot lock -- is left in place once the rest of -its packs are up, listed in `blocked_siblings`, reported once in `last_error`, -and neither retried nor owed. `spool_max_bytes` bounds the -directory together with what those dead directories still hold (a live -process's directory is its own budget), so restarts while uploads are blocked -cannot each add a whole budget; the room comes back as they are adopted. The +its packs are up, listed in `blocked_siblings`, reported once in `last_error` +and in its `.owner.lock` (a `blocked: ` line), and neither retried nor +owed. `spool_max_bytes` bounds the directory together with what the dead +directories the service can adopt still hold, so restarts while uploads are +blocked cannot each add a whole budget; the room comes back as they are +adopted. A live process's directory is its own budget, and neither a blocked +directory nor one this process keeps owned itself (an earlier engine's whose +sink did not seal) is charged: no adoption here drains them. The spool root must be node-local: NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, GPFS, 9p, AFS and OrangeFS are refused (by statfs `f_type`) unless `NativeSinkConfig.spool_allow_shared_filesystem`, which a FUSE diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index f00f20640..5aeca5f4b 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -719,8 +719,7 @@ bool CaptureStorageService::begin_adoption(const std::string& directory) { dmi_store::SpoolStatus::kOk || adoption->spool.Recover(&ready, &error) != dmi_store::SpoolStatus::kOk) { - adoption.reset(); // lets go of its lock - block_sibling(directory, "cannot open it: " + error); + block_sibling(directory, "cannot open it: " + error, &adoption->lock); return true; } // Each pack's identity and object key come from the pack and its path in @@ -788,9 +787,8 @@ void CaptureStorageService::finish_adoption() { std::unique_ptr adoption = std::move(adopting_); const std::string directory = adoption->directory; if (!adoption->blocked.empty()) { - const std::string reason = adoption->blocked; - adoption.reset(); // lets go of its lock; the directory stays - block_sibling(directory, reason); + // The directory stays, and its lock goes. + block_sibling(directory, adoption->blocked, &adoption->lock); return; } { @@ -807,7 +805,17 @@ void CaptureStorageService::finish_adoption() { } void CaptureStorageService::block_sibling(const std::string& directory, - const std::string& reason) { + const std::string& reason, + dmi_store::SpoolOwnerLock* lock) { + if (lock != nullptr) { + // Said in its lock file, while this service still holds it, then let + // go of: for a person, and for the sinks on the node, which charge a + // dead directory against their budget only while an adoption can + // drain it (SpoolConfig::charge_dead_siblings). A later take -- by a + // process that can adopt it -- rewrites the mark. + lock->MarkBlocked(reason); + lock->Release(); + } blocked_siblings_.insert(directory); record_error("dead spool " + directory + " is left in place, not to be " "adopted by this service: " + reason); diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index f81ae17be..29daad4c0 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -129,6 +129,8 @@ struct StorageServiceConfig { // Such a sibling is blocked: once the rest of its packs are up it is left // in place, with its lock let go, for a process that can adopt it (or a // person); it is reported once (last_error, snapshot blocked_siblings), + // and in its lock file's record (dmi_store::SpoolOwner::blocked), which + // keeps the sinks on the node from charging it against their budgets; // never retried or re-hashed by this service, and not owed. So is one // drained of packs that still holds other files. flush() covers this // process's records: a sibling still to adopt does not keep it from @@ -350,8 +352,11 @@ class CaptureStorageService { // adopting_ holds no pack any more: removes the directory, or leaves a // blocked one. void finish_adoption(); - // Leaves a dead sibling in place for good, reporting why. - void block_sibling(const std::string& directory, const std::string& reason); + // Leaves a dead sibling in place for good, reporting why -- and, given + // its held lock, marking why in it (SpoolOwnerLock::MarkBlocked) before + // letting go of it. + void block_sibling(const std::string& directory, const std::string& reason, + dmi_store::SpoolOwnerLock* lock = nullptr); bool adoption_owed() const; // requires cycle_mutex_ bool stop_requested(); // Removes the staging copies (dmi_store::IsSpoolClaimStagingName) that diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 288f787ff..95faeaef7 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -329,27 +329,53 @@ void FsyncParent(const std::string& path) { FsyncDir(fs::path(path).parent_path().string(), nullptr); } -// " ": whoever holds the lock records itself, so a refused -// process can say who holds the directory. -void WriteOwnerRecord(int fd) { - const std::string record = - Hostname() + " " + std::to_string(::getpid()) + "\n"; +// The record's roles, on the line after " ". +constexpr const char* kAdoptingRole = "adopting"; +constexpr const char* kBlockedRole = "blocked: "; +constexpr size_t kOwnerRecordBytes = 2048; + +// " \n", then a role line when there is one ("adopting", or +// "blocked: "): whoever holds the lock records itself, so a refused +// process can say who holds the directory, and a sink beside it whether +// an adoption can drain it. +void WriteOwnerRecord(int fd, const std::string& role = "") { + std::string record = Hostname() + " " + std::to_string(::getpid()) + "\n"; + if (!role.empty()) { + std::string line = role.substr(0, kOwnerRecordBytes - record.size() - 2); + std::replace(line.begin(), line.end(), '\n', ' '); + record += line + "\n"; + } if (::ftruncate(fd, 0) == 0) { (void)!::pwrite(fd, record.data(), record.size(), 0); } } void ReadOwnerRecord(int fd, SpoolOwner* owner) { - char buffer[512]; - const ssize_t n = ::pread(fd, buffer, sizeof(buffer) - 1, 0); - owner->host.clear(); - owner->pid = 0; + char buffer[kOwnerRecordBytes]; + const ssize_t n = ::pread(fd, buffer, sizeof(buffer), 0); + *owner = SpoolOwner{}; if (n <= 0) return; std::string record(buffer, static_cast(n)); - while (!record.empty() && std::isspace(static_cast( - record.back()))) { - record.pop_back(); + const auto trim = [](std::string* text) { + while (!text->empty() && std::isspace(static_cast( + text->back()))) { + text->pop_back(); + } + }; + const size_t newline = record.find('\n'); + if (newline != std::string::npos) { + std::string role = record.substr(newline + 1); + record.resize(newline); + trim(&role); + const size_t blocked = std::strlen(kBlockedRole); + if (role == kAdoptingRole) { + owner->adopting = true; + } else if (role.compare(0, blocked, kBlockedRole) == 0) { + owner->blocked = role.substr(blocked); + if (owner->blocked.empty()) owner->blocked = "blocked"; + } } + trim(&record); const size_t space = record.rfind(' '); if (space == std::string::npos) { owner->host = record; @@ -579,7 +605,7 @@ bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner) { ::close(dir_fd); } } - if (held && owner != nullptr) { + if (owner != nullptr) { if (fd >= 0) { ReadOwnerRecord(fd, owner); } else { @@ -803,8 +829,8 @@ SpoolStatus LockDirectory(const std::string& dir, HeldLock* lock, // has none, and the directory itself. Retries a lock lost to a remover's // unlink, and a refusal as brief as another process's ReadSpoolOwner // probe. -SpoolStatus LockInPlace(const std::string& dir, HeldLock* out, - std::string* error) { +SpoolStatus LockInPlace(const std::string& dir, const char* role, + HeldLock* out, std::string* error) { const std::string file = dir + "/" + kOwnerLockFile; for (int attempt = 0; attempt < 8; ++attempt) { HeldLock lock; @@ -850,7 +876,7 @@ SpoolStatus LockInPlace(const std::string& dir, HeldLock* out, } return directory; } - WriteOwnerRecord(fd); + WriteOwnerRecord(fd, role); *out = lock; return SpoolStatus::kOk; } @@ -929,7 +955,7 @@ SpoolStatus CreateLocked(const std::string& dir, HeldLock* out, ::unlink(file.c_str()); ::rmdir(staging.c_str()); if (failure == EEXIST || failure == ENOTEMPTY) { - return LockInPlace(dir, out, error); + return LockInPlace(dir, "", out, error); } if (error) { *error = "cannot create spool directory " + dir + ": " + @@ -1033,7 +1059,7 @@ SpoolStatus SpoolOwnerLock::Acquire(const std::string& dir, const std::string lock_file = canonical + "/" + kOwnerLockFile; const bool had_lock_file = exists && fs::exists(lock_file, ec); HeldLock lock; - status = exists ? LockInPlace(canonical, &lock, error) + status = exists ? LockInPlace(canonical, "", &lock, error) : CreateLocked(canonical, &lock, error); if (status != SpoolStatus::kOk) return status; out->Hold(lock.file_fd, lock.dir_fd, canonical); @@ -1065,12 +1091,20 @@ SpoolStatus SpoolOwnerLock::TryAdopt(const std::string& dir, return SpoolStatus::kBadArgument; } HeldLock lock; - const SpoolStatus status = LockInPlace(resolved, &lock, error); + const SpoolStatus status = + LockInPlace(resolved, kAdoptingRole, &lock, error); if (status != SpoolStatus::kOk) return status; out->Hold(lock.file_fd, lock.dir_fd, resolved); return SpoolStatus::kOk; } +bool SpoolOwnerLock::MarkBlocked(const std::string& reason) { + if (!held()) return false; + WriteOwnerRecord(fd_, kBlockedRole + (reason.empty() ? "blocked" : reason)); + ::fsync(fd_); + return true; +} + bool SpoolOwnerLock::ReleaseAndRemoveIfEmpty(std::string* error) { if (!held()) return false; const std::string dir = dir_; @@ -1287,6 +1321,17 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { return SpoolStatus::kOk; } +namespace { +// Whether this process could take `dir`'s lock and empty it: write its lock +// file (or create one), and unlink in it. +bool CouldAdopt(const std::string& dir) { + const std::string file = dir + "/" + kOwnerLockFile; + if (::access(dir.c_str(), W_OK | X_OK) != 0) return false; + return ::access(file.c_str(), F_OK) != 0 || + ::access(file.c_str(), R_OK | W_OK) == 0; +} +} // namespace + uint64_t Spool::ChargedSiblingBytes() const { const fs::path own(root_); uint64_t bytes = 0; @@ -1303,9 +1348,20 @@ uint64_t Spool::ChargedSiblingBytes() const { continue; } const std::string sibling = it->path().string(); - // Another live process's directory is its own budget. - if (ReadSpoolOwner(sibling, nullptr) && - !SpoolOwnedByThisProcess(sibling)) { + // Only what adoption can drain: a dead directory, or one this process's + // adoption holds. + SpoolOwner owner; + if (ReadSpoolOwner(sibling, &owner)) { + // Another live process's directory is its own budget. One this + // process holds for its own writing -- an earlier engine's claim, + // kept owned while its unsealed sink may still stage -- no adoption + // here drains (its service reads it as live), until the process + // exits and the next one on the node adopts it. + if (!owner.adopting || !SpoolOwnedByThisProcess(sibling)) continue; + } else if (!owner.blocked.empty() || !CouldAdopt(sibling)) { + // Dead, but left for good by an adopter that could never drain it + // (its lock file says why), or not one this process could take and + // empty at all -- another user's, say. continue; } std::error_code walk_ec; @@ -1409,8 +1465,9 @@ std::string Spool::FullMessage(uint64_t n) const { " > " + std::to_string(max_bytes_); if (sibling_bytes_ > 0) { message += " (" + std::to_string(sibling_bytes_) + - " bytes of it in the dead spool directories beside this one, " - "still to be adopted)"; + " bytes of it in dead spool directories beside this one, " + "which this process's storage service adopts: the room " + "comes back as it drains them)"; } return message; } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index 210495ce0..ac74fc717 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -81,15 +81,21 @@ struct SpoolConfig { bool allow_shared_filesystem = false; // The root is a rank directory of the section 2.3 layout, and what its // SIBLING rank directories hold (ready packs and temp files) counts - // against max_bytes as well -- except a sibling another live process - // holds, which is that process's own budget. A sibling nobody holds is a - // dead incarnation's, waiting to be adopted; one THIS process holds is - // being adopted by its service. Every process start gets a fresh rank - // directory, so without this each crash-restart while uploads are - // blocked would add a whole max_bytes to the node's spool; before the - // layout every restart reused one directory and one budget. The charge is - // refreshed wherever the committed account is (Open, and before a stage - // is refused), so the capacity comes back as adoption drains them. + // against max_bytes as well, while an adoption by this process's storage + // service can drain it: a sibling nobody holds (a dead incarnation's, + // waiting to be adopted), and one this process holds through an adoption + // (SpoolOwner::adopting). Not charged: a sibling another live process + // holds, which is that process's own budget; one this process holds for + // its own writing -- an earlier engine's claim kept owned while its + // unsealed sink may still stage -- which no adoption here drains until + // the process exits; and one an adopter left blocked (SpoolOwner:: + // blocked), or that this process could not take and empty at all. Every + // process start gets a fresh rank directory, so without this each + // crash-restart while uploads are blocked would add a whole max_bytes to + // the node's spool; before the layout every restart reused one directory + // and one budget. The charge is refreshed wherever the committed account + // is (Open, and before a stage is refused), so the capacity comes back as + // adoption drains them. bool charge_dead_siblings = false; }; @@ -160,17 +166,24 @@ void SetFdinfoHidesLocksForTesting(bool hide); // which a remover can unlink it. An empty function removes the hook. void SetLockOpenHookForTesting(std::function hook); -// The holder recorded in a directory's owner lock file. +// The record in a directory's owner lock file: its last holder, which +// wrote it on taking the lock, and why it held the directory. Read with +// the lock free, it is the last holder's, whom the kernel has let go of. struct SpoolOwner { std::string host; int64_t pid = 0; + // Taken by an adopter (SpoolOwnerLock::TryAdopt), to drain a dead + // directory, rather than by the process writing it. + bool adopting = false; + // Why an adopter left the directory for good (MarkBlocked), or empty. + std::string blocked; }; // Whether 's owner lock is held right now (by any process, this one -// included) -- its lock file's, or the directory's own -- and if so who -// recorded themselves in the lock file. A holder that has locked but not -// yet written its record, or whose lock file was replaced, reads as an -// empty host and pid 0. +// included) -- its lock file's, or the directory's own. Fills *owner with +// the lock file's record either way (empty without one). A holder that has +// locked but not yet written its record, or whose lock file was replaced, +// reads as an empty host and pid 0. bool ReadSpoolOwner(const std::string& dir, SpoolOwner* owner); // Whether `name` is the staging copy of a directory SpoolOwnerLock::Acquire @@ -225,11 +238,18 @@ class SpoolOwnerLock { SpoolOwnerLock* out, std::string* error); // Adoption's try-lock: takes the lock of an EXISTING directory whose owner - // is gone, creating its lock file if it has none. kOwned while its owner - // lives; never creates the directory. + // is gone, creating its lock file if it has none, and records the take as + // an adoption (SpoolOwner::adopting). kOwned while its owner lives; never + // creates the directory. static SpoolStatus TryAdopt(const std::string& dir, SpoolOwnerLock* out, std::string* error); + // Records in the lock file, while it is held, that the directory is left + // for good and why (SpoolOwner::blocked): an adopter that can never drain + // it lets go of it after this. The next take rewrites the record. Returns + // whether it was written. + bool MarkBlocked(const std::string& reason); + bool held() const; // The canonical path of the directory, while held. const std::string& directory() const { return dir_; } @@ -366,7 +386,7 @@ class Spool { // would exceed the cap is judged against the same durable truth. void ReconcileCommittedLocked(); // charge_dead_siblings: the bytes of ready and temp files in the sibling - // rank directories no other live process holds. + // rank directories an adoption can drain (SpoolConfig). uint64_t ChargedSiblingBytes() const; // The kFull refusal of a stage of `n` bytes. `mutex_` must be held. std::string FullMessage(uint64_t n) const; diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 851fcf0c7..d8da5a84b 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -810,9 +810,10 @@ void TestANewDirectoryAppearsWithItsLockHeld() { // rank directory, so a spool that charged only its own directory let each // crash-restart add a full max_bytes while uploads were blocked. With // charge_dead_siblings, what the sibling rank directories hold counts -// against max_bytes too -- unless another live process holds one (that is -// its own budget); one this process holds, as its service does while it -// adopts it, still counts. +// against max_bytes too, while adoption can drain it: a dead one, or one +// this process's adoption holds -- not one another live process holds +// (that is its own budget), one this process holds for its own writing, +// or one an adopter left blocked. void TestDeadSiblingsCountAgainstTheBudget() { const std::string base = FreshRoot("budget"); std::string error; @@ -891,25 +892,85 @@ void TestDeadSiblingsCountAgainstTheBudget() { int status = 0; ::waitpid(child, &status, 0); - // One THIS process holds -- its service adopting it -- still counts. - const std::string adopting = base + "/bbbbbbbbbbbb"; + // Only what adoption can drain is charged. One THIS process holds for + // its own writing -- an earlier engine's claim, kept owned while its + // unsealed sink might still stage -- is drained by no adoption of this + // process's (its service reads it as live), so charging it left every + // later sink in the process that much less budget until exit. + const std::string mixed = base + "/bbbbbbbbbbbb"; + const std::string sibling = mixed + "/r0-0000000f"; SpoolOwnerLock held; - CHECK(SpoolOwnerLock::Acquire(adopting + "/r0-0000000f", false, &held, - &error) == SpoolStatus::kOk); + CHECK(SpoolOwnerLock::Acquire(sibling, false, &held, &error) == + SpoolStatus::kOk); { - SpoolConfig adopted{adopting + "/r0-0000000f", 1 << 20}; - adopted.owner_lock = OwnerLock::kHeldByCaller; + SpoolConfig kept{sibling, 1 << 20}; + kept.owner_lock = OwnerLock::kHeldByCaller; Spool writer; - CHECK(Spool::Open(adopted, &writer, &error) == SpoolStatus::kOk); + CHECK(Spool::Open(kept, &writer, &error) == SpoolStatus::kOk); for (int i = 1; i <= 3; ++i) { CHECK(StageOne(writer, i, &error) == SpoolStatus::kOk); } } - SpoolConfig own{adopting + "/r0-00000010", 450}; - own.charge_dead_siblings = true; - Spool mine; - CHECK(Spool::Open(own, &mine, &error) == SpoolStatus::kOk); - CHECK(mine.Snapshot().sibling_bytes == 300); + const auto charged = [&]() { + SpoolConfig own{mixed + "/r0-00000010", 450}; + own.charge_dead_siblings = true; + Spool mine; + CHECK(Spool::Open(own, &mine, &error) == SpoolStatus::kOk); + return mine.Snapshot().sibling_bytes; + }; + CHECK(charged() == 0); + { + SpoolConfig own{mixed + "/r0-00000010", 450}; + own.charge_dead_siblings = true; + Spool mine; + CHECK(Spool::Open(own, &mine, &error) == SpoolStatus::kOk); + for (int i = 4; i < 8; ++i) { + CHECK(StageOne(mine, i, &error) == SpoolStatus::kOk); + } + } + fs::remove_all(mixed + "/r0-00000010"); + + // One this process's adoption holds -- a dead one its service drains -- + // still counts: the room comes back as the adoption drains it. + held.Release(); + SpoolOwnerLock adopting; + CHECK(SpoolOwnerLock::TryAdopt(sibling, &adopting, &error) == + SpoolStatus::kOk); + dmi_store::SpoolOwner record; + CHECK(dmi_store::ReadSpoolOwner(sibling, &record)); + CHECK(record.adopting && record.blocked.empty()); + CHECK(record.pid == ::getpid()); + CHECK(charged() == 300); + + // One an adopter left for good -- blocked, its lock let go -- is drained + // by nobody here, and is not charged either; its lock file says why. + CHECK(adopting.MarkBlocked("it holds a pack this service can never upload")); + adopting.Release(); + CHECK(!dmi_store::ReadSpoolOwner(sibling, &record)); + CHECK(record.blocked == "it holds a pack this service can never upload"); + CHECK(!record.adopting); + CHECK(charged() == 0); + // A process that can adopt it after all takes it again, and the mark + // goes with its take. + CHECK(SpoolOwnerLock::TryAdopt(sibling, &adopting, &error) == + SpoolStatus::kOk); + CHECK(dmi_store::ReadSpoolOwner(sibling, &record)); + CHECK(record.adopting && record.blocked.empty()); + CHECK(charged() == 300); + adopting.Release(); + CHECK(charged() == 300); // dead: still to adopt + + // One this process cannot take and empty at all -- another user's -- is + // no adoption's here either. + fs::permissions(sibling + "/.owner.lock", fs::perms::owner_read); + fs::permissions(sibling, fs::perms::owner_read | fs::perms::owner_exec); + if (::access(sibling.c_str(), W_OK) != 0) { // not as root + CHECK(charged() == 0); + } + fs::permissions(sibling, fs::perms::owner_all); + fs::permissions(sibling + "/.owner.lock", + fs::perms::owner_read | fs::perms::owner_write); + CHECK(charged() == 300); } // (7) The layout: //r-/. diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 8340ceef9..423f682ec 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -478,6 +478,14 @@ def test_a_dead_spool_this_service_can_never_upload_is_left_and_reported( assert snapshot["blocked_siblings"] == [str(blocked)], snapshot assert "in-flight byte limit" in snapshot["last_error"], snapshot assert sorted(blocked.rglob("*.dmi-pack.ready")) == blocked_packs + # Let go of, and marked in its lock file -- for a person, and for + # the sinks on the node, which charge a dead directory against + # their budgets only while an adoption can drain it. + assert _store().spool_owner(str(blocked)) is None + record = (blocked / ".owner.lock").read_text() + assert ("\nblocked: it holds a pack this service can never " + "upload") in record, record + assert "in-flight byte limit" in record, record assert not adoptable.exists() # Not retried: no more failed uploads, and no backoff -- the # loop keeps its poll interval. diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index bed5b6db3..3ce1ada89 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -325,6 +325,39 @@ def test_a_dead_siblings_bytes_count_against_the_sinks_budget(tmp_path): claim.release() +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_a_directory_this_process_keeps_is_not_charged_to_its_next_sink( + tmp_path): + """An engine whose sink did not seal keeps its claim held until the + process exits (the sink may still stage), and the same process's next + create_record_runtime claims a fresh directory beside it. No adoption + in this process can drain the kept one -- its service reads it as live + -- so charging it cut every later sink's budget until exit, and its + refusal called the bytes "still to be adopted".""" + pytest.importorskip("torch") + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + sink_config = NativeSinkConfig(spool_root=str(tmp_path / "root")) + kept = claim_spool_directory(sink_config, _config()) + (Path(kept.directory) / "v1").mkdir() + (Path(kept.directory) / "v1" / "left.dmi-pack.ready").write_bytes( + b"\0" * (1 << 20)) + budget = (1 << 20) + 256 # room for the kept bytes, not for a pack + claim = claim_spool_directory(sink_config, _config()) + try: + assert _store().spool_owner(kept.directory)["pid"] == os.getpid() + sink = _sink(claim.directory, owner_lock="held_by_caller", + spool_max_bytes=budget, max_pack_bytes=budget, + charge_dead_siblings=True) + _stage_one(sink) # the whole budget + del sink + finally: + claim.release() + kept.release() + + def test_a_claim_warns_of_packs_left_outside_the_layout(tmp_path, caplog): """Before the per-process layout the engine spooled into /v1/... and its next start swept and uploaded whatever a From 354cd4e7ff373c5b97a34411e9878e42636f00a9 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:29:23 -0400 Subject: [PATCH 25/43] Say that skipping _refs/ is the C++ spool's alone The C++ spool skips /_refs/ in Open's accounting walk, the sibling charge, the committed reconcile and Scan (so Recover and ListPending). The Python reference DurablePackSpool does not: its constructor counts, and its recover() sweeps and quarantines, .open and .ready files there. spool.h's parity paragraph and the tests named only the lock as deliberately stricter, so a parity comparison meeting the difference once ref files live there would not know it was intended. It is, and it is not ported to the reference; spool.h and the driver tests now say so. --- native/csrc/store/spool.h | 5 ++++- tests/test_native_spool_owner_lock.py | 9 +++++++-- 2 files changed, 11 insertions(+), 3 deletions(-) diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index ac74fc717..a73294b6e 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -47,7 +47,10 @@ // // The Python DurablePackSpool (spool.py) takes no lock, and its recover() // deletes every .open file under its root; the C++ spool is deliberately -// stricter, and that is not ported to the reference. +// stricter, and that is not ported to the reference. Nor is skipping +// /_refs/: the reference counts, sweeps and quarantines .open and +// .ready files there like any others, which the C++ spool leaves alone -- +// a C++-only divergence, deliberately not ported. #ifndef DMI_STORE_SPOOL_H_ #define DMI_STORE_SPOOL_H_ diff --git a/tests/test_native_spool_owner_lock.py b/tests/test_native_spool_owner_lock.py index 6de5ffbe8..53192b1ec 100644 --- a/tests/test_native_spool_owner_lock.py +++ b/tests/test_native_spool_owner_lock.py @@ -24,7 +24,10 @@ test_spool_owner_lock.cpp (test_native_spool_owner_lock_unit.py). The Python spool (dmi.storage.capture.spool) takes no lock; the C++ spool is -deliberately stricter, and that is not ported to the reference. +deliberately stricter, and that is not ported to the reference. Skipping +``_refs/`` is C++-only as well: the reference's constructor counts, and its +recover() sweeps and quarantines, ``.open`` and ``.ready`` files there, and +that divergence is deliberate, not ported either. Build: make -C native build/conformance_spool build/conformance_sink """ @@ -217,7 +220,9 @@ def test_an_unknown_owner_lock_mode_is_refused(tmp_path): def test_recovery_leaves_the_refs_directory_alone(tmp_path): """/_refs/ will hold the upload handoff's ref files (E2a). A - sweep that deleted, quarantined or listed them would break it.""" + sweep that deleted, quarantined or listed them would break it. C++ + only: the Python reference spool still walks _refs/, deliberately + unported.""" root = tmp_path / "spool" refs = root / "_refs" refs.mkdir(parents=True) From 72c9929ef1495e57d7d2d6f3c0e50aaa18d93e0a Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:36:47 -0400 Subject: [PATCH 26/43] Pass over nested spool directories, and refuse nesting only while held Two holes in the nesting rule, one in each direction. The descendant half of CheckNotNested refused any lock file below the directory, held or not. A default-mode run that leaves its rank directory under spool_root -- a SIGKILL, a close whose drain left packs, an unsealed close once the process exits, a blocked sibling -- then refused both other modes on that spool_root: sink-only (the sink takes spool_root) and the explicit-record_sink rollback (the service takes it), with "contains the owned spool directory" although nobody held it, until a default-mode start adopted the directory, and for good for a blocked one. The rollback was blocked exactly when it is most likely needed, and nothing said so. The other way, TryAdopt runs no nesting check, although the rule relied on the next take of a directory meeting the lock file below it. An adopter of a dead rank directory holding a live nested spool (a root put there, which an unheld ancestor admits) had its Recover sweep that spool's in-flight .open files, and would have uploaded its packs under the outer directory's keys. The refusal was standing in for what the walks should do. Every walk of a spool -- Open's accounting, the committed reconcile, the sibling charge and Scan (Recover and ListPending) -- now passes over a subdirectory with a lock file of its own: another spool directory, live or dead, is never this one's to count, sweep, list or upload. So a flat spool_root leaves the rank directories under it to adoption, and an adopter leaves a spool nested in the dead directory it drains alone (then keeps the directory, which still holds it, and says so). CheckNotNested refuses a descendant, like an ancestor, only while its lock is held -- the outer walk could meet a live inner directory before its lock file is there -- and names the holder. The two layouts on one spool_root take turns and never run at once. The v1 contract says what switching modes does; skipping nested spools is C++-only, not ported to the Python reference. Tests: test_spool_owner_lock.cpp, red before -- a held rank directory under a root refuses its take naming the holder; once dead, with a pack and a stage in flight, the root is taken, counts and recovers nothing of it, stages its own, and the rank directory cannot be taken while the root is held; an adopter of a dead directory holding a live nested spool recovers only its own pack, leaves the nested .open in place, and keeps the directory. test_native_spool_owner_lock.py, red before -- a dead spool inside another is left alone by the outer recover. test_native_spool_ownership.py, red before -- after a crashed default run, the explicit-sink service and a sink-only sink take spool_root and leave the dead directory as it was; a live claim still refuses both. --- docs/integration-api-v1.md | 8 +- native/csrc/catalog/storage_service.cpp | 8 +- native/csrc/store/spool.cpp | 87 ++++++++++------- native/csrc/store/spool.h | 34 ++++--- src/dmi/engine.py | 4 +- src/dmi/storage/native_capture.py | 3 +- tests/native/test_spool_owner_lock.cpp | 124 +++++++++++++++++++++--- tests/test_native_spool_owner_lock.py | 48 ++++++--- tests/test_native_spool_ownership.py | 46 +++++++++ 9 files changed, 277 insertions(+), 85 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index 53f5b1cab..2b0ec77d2 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -229,7 +229,13 @@ are refused (by statfs `f_type`) unless `NativeSinkConfig.spool_allow_shared_fil filesystem that is itself local, such as fuse-overlayfs, needs too. Without `capture_storage_config` the sink owns `spool_root` itself. With an explicit `record_sink`, the service drains `spool_root` as that sink writes it, -unswept and adopting nothing. Upgrading from an engine without this layout: it +unswept and adopting nothing. Both of those modes pass over the rank +directories under `spool_root` -- what a crashed or undrained default-mode run +left there is the next default-mode start's to adopt -- so switching to them +(the rollback to an explicit `record_sink` included) works after such a run. +They are refused, naming the holder, while a default-mode process on the node +holds a rank directory under that `spool_root`, and a default-mode start is +refused while one of them holds `spool_root`. Upgrading from an engine without this layout: it spooled into `/v1/...` and its next start uploaded what a crashed run left there, but nothing adopts packs outside the layout now -- nor those a sink-only or explicit-`record_sink` run leaves in `spool_root`. Each diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index 5aeca5f4b..79c59d195 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -797,10 +797,12 @@ void CaptureStorageService::finish_adoption() { } std::string error; if (!adoption->lock.ReleaseAndRemoveIfEmpty(&error)) { - // Nothing to upload is left, only files that are not packs (a - // quarantined one, say): the directory stays for someone to look at. + // Nothing to upload is left, only files that are not its packs (a + // quarantined one, a spool directory nested in it): the directory + // stays for someone to look at. block_sibling(directory, "it was drained, but still holds files that " - "are not packs"); + "are not its packs (a quarantined pack, or a " + "spool directory nested in it)"); } } diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 95faeaef7..8b263c826 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -98,6 +98,24 @@ bool AtRefsDirectory(const fs::recursive_directory_iterator& it) { it->is_directory(ec); } +// Whether a recursive walk of a spool is at a subdirectory with a lock +// file of its own: another spool directory nested in this one -- a rank +// directory of the layout under a flat spool_root, a claim's staging copy, +// a root someone put inside a dead rank directory -- live or dead. No walk +// of this spool enters it, for counting, sweeping, listing or uploading: a +// live one's owner is writing it, and a dead one's packs are its +// successor's to adopt, under its own keys, not this spool's. +bool AtNestedSpool(const fs::recursive_directory_iterator& it) { + std::error_code ec; + return it->is_directory(ec) && !it->is_symlink(ec) && + fs::exists(it->path() / kOwnerLockFile, ec); +} + +// Every walk of a spool's own files skips these two. +bool AtSkippedDirectory(const fs::recursive_directory_iterator& it) { + return AtRefsDirectory(it) || AtNestedSpool(it); +} + std::string Hostname() { char host[256] = {0}; if (::gethostname(host, sizeof(host) - 1) != 0 || host[0] == '\0') { @@ -430,24 +448,28 @@ std::string CanonicalPath(const std::string& path, std::string* error) { return out; } -// A spool directory must not be nested under, or contain, another owned -// directory: Scan walks recursively, so the outer spool's Recover would -// sweep the inner one's .open files and upload its packs under the outer -// spool's keys. Run AFTER `dir`'s own lock is taken, so that of two takes -// racing on an outer directory and one inside it, at least one sees the -// other: each publishes its lock before it looks. -// - An ancestor refuses while its lock is HELD. A lock file nobody holds -// is a spool that was (every take leaves its file behind); whoever -// takes that ancestor next meets this directory's lock file in its own -// descendant walk, and is refused. -// - A descendant refuses when it has a lock file at all, held or not: a -// dead directory's packs are its successor's to adopt, not this -// spool's to sweep and upload under its own keys. Except a claim's -// staging copy (IsSpoolClaimStagingName) that nobody holds: a claim -// killed before its rename, which holds nothing but its lock file. (One -// that is held is a claim in progress, and refuses; one not yet locked -// publishes after this walk, so its own ancestor check sees this lock.) +// A spool directory must not be nested under, or contain, another one that +// is owned: Scan walks recursively, and although every walk passes over a +// subdirectory with a lock file of its own (AtNestedSpool), the outer +// spool's walk can reach one before its owner's lock file is there -- a +// take of an existing directory creates it -- and sweep the .open files +// its owner then writes. Run AFTER `dir`'s own lock is taken, so that of +// two takes racing on an outer directory and one inside it, at least one +// sees the other: each publishes its lock before it looks. Only a HELD +// lock refuses, in either direction. One nobody holds is a spool that was: +// every take leaves its file behind, a spool_root a sink-only run once +// owned holds one, and a crashed default-mode run leaves its rank +// directory's under spool_root. Such a directory's packs are left alone by +// every walk of the other (AtNestedSpool) -- a dead one inside is for its +// successor to adopt -- so it refuses nothing: the rank directories under +// a flat spool_root and that spool_root's own modes (sink-only, explicit +// record_sink) take turns, and never run at once. SpoolStatus CheckNotNested(const std::string& dir, std::string* error) { + const auto holder = [](const SpoolOwner& owner) { + return owner.pid > 0 ? "pid " + std::to_string(owner.pid) + " on host " + + owner.host + : std::string("another owner"); + }; fs::path ancestor(dir); while (ancestor.has_parent_path() && ancestor.parent_path() != ancestor) { ancestor = ancestor.parent_path(); @@ -458,11 +480,9 @@ SpoolStatus CheckNotNested(const std::string& dir, std::string* error) { if (error) { *error = "spool directory " + dir + " is nested under the spool " "directory " + ancestor.string() + ", which " + - (owner.pid > 0 ? "pid " + std::to_string(owner.pid) + - " on host " + owner.host - : std::string("another owner")) + - " holds (" + kOwnerLockFile + "), and whose recovery " - "would sweep this one; use a directory outside it"; + holder(owner) + " holds (" + kOwnerLockFile + "), and whose " + "recovery would sweep this one; use a directory outside " + "it, or wait for that owner to end"; } return SpoolStatus::kBadArgument; } @@ -474,15 +494,14 @@ SpoolStatus CheckNotNested(const std::string& dir, std::string* error) { if (it->path().filename() != kOwnerLockFile) continue; const fs::path owned = it->path().parent_path(); if (owned == fs::path(dir)) continue; - if (IsSpoolClaimStagingName(owned.filename().string()) && - !ReadSpoolOwner(owned.string(), nullptr)) { - continue; - } + SpoolOwner owner; + if (!ReadSpoolOwner(owned.string(), &owner)) continue; // dead: left be if (error) { - *error = "spool directory " + dir + " contains the owned spool " - "directory " + owned.string() + " (it has " + kOwnerLockFile + - "), which this one's recovery would sweep; use a directory " - "that does not contain it"; + *error = "spool directory " + dir + " contains the spool directory " + + owned.string() + ", which " + holder(owner) + " holds (" + + kOwnerLockFile + "), and which this one's recovery could " + "sweep while its owner writes it; use a directory that does " + "not contain it, or wait for that owner to end"; } return SpoolStatus::kBadArgument; } @@ -1298,7 +1317,7 @@ SpoolStatus Spool::Open(SpoolConfig config, Spool* out, std::string* error) { // later retry of one of them is recognised as already counted. for (auto it = fs::recursive_directory_iterator(out->root_, ec); it != fs::recursive_directory_iterator(); ++it) { - if (AtRefsDirectory(it)) { + if (AtSkippedDirectory(it)) { it.disable_recursion_pending(); continue; } @@ -1367,7 +1386,7 @@ uint64_t Spool::ChargedSiblingBytes() const { std::error_code walk_ec; for (fs::recursive_directory_iterator walk(sibling, walk_ec), last; !walk_ec && walk != last; walk.increment(walk_ec)) { - if (AtRefsDirectory(walk)) { + if (AtSkippedDirectory(walk)) { walk.disable_recursion_pending(); continue; } @@ -1426,7 +1445,7 @@ void Spool::ReconcileCommittedLocked() { std::error_code walk_ec; for (auto it = fs::recursive_directory_iterator(root_, walk_ec); it != fs::recursive_directory_iterator(); ++it) { - if (AtRefsDirectory(it)) { + if (AtSkippedDirectory(it)) { it.disable_recursion_pending(); continue; } @@ -1749,7 +1768,7 @@ SpoolStatus Spool::Scan(std::vector* out, bool discard_open_files, std::unordered_map seen_ready; for (auto it = fs::recursive_directory_iterator(root_, ec); it != fs::recursive_directory_iterator(); ++it) { - if (AtRefsDirectory(it)) { + if (AtSkippedDirectory(it)) { it.disable_recursion_pending(); continue; } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index a73294b6e..77c26c35c 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -31,16 +31,20 @@ // naming the holder, beside another process's lock, and kBadArgument // when nothing holds it. Standalone callers -- the drivers, adoption -- // take. -// A directory nested under a HELD spool directory, or containing one with -// a .owner.lock file (held or not: a dead directory's packs are for its -// successor to adopt), is refused: Scan walks recursively, so the outer -// spool's Recover would reach into the inner one. An unheld lock file -// ABOVE refuses nothing -- every take leaves its file behind -- since the -// next take of that directory meets this one's lock file below it. The -// check runs after the lock is taken, so of two processes taking an outer -// and a nested directory at once, at least one is refused. /_refs/ is -// never scanned: the upload handoff's ref files live there (plan section -// 2.4). The spool must be node-local: NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, +// A spool never walks into a subdirectory with a .owner.lock of its own -- +// another spool directory nested in it, live or dead -- to count, sweep, +// list or upload what it holds: a dead one's packs are for its successor to +// adopt, under its own keys. So a flat spool_root (the sink-only mode, an +// explicit record_sink) passes over the rank directories a crashed +// default-mode run left under it, and an adopter over a spool someone put +// inside the dead directory it drains. A directory nested under, or +// containing, a HELD spool directory is refused at the take, since the +// outer spool's walk could meet the inner one before its lock file is +// there; one nobody holds refuses nothing (every take leaves its file +// behind). The check runs after the lock is taken, so of two processes +// taking an outer and a nested directory at once, at least one is refused. +// /_refs/ is never scanned: the upload handoff's ref files live there +// (plan section 2.4). The spool must be node-local: NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, // GPFS, 9p, AFS and OrangeFS are refused by statfs f_type unless // allow_shared_filesystem is set, since none guarantees a flock that // excludes a process on another node. @@ -48,9 +52,9 @@ // The Python DurablePackSpool (spool.py) takes no lock, and its recover() // deletes every .open file under its root; the C++ spool is deliberately // stricter, and that is not ported to the reference. Nor is skipping -// /_refs/: the reference counts, sweeps and quarantines .open and -// .ready files there like any others, which the C++ spool leaves alone -- -// a C++-only divergence, deliberately not ported. +// /_refs/, or a nested spool directory: the reference counts, sweeps +// and quarantines .open and .ready files there like any others, which the +// C++ spool leaves alone -- C++-only divergences, deliberately not ported. #ifndef DMI_STORE_SPOOL_H_ #define DMI_STORE_SPOOL_H_ @@ -234,8 +238,8 @@ class SpoolOwnerLock { // parent ever meets the directory before its owner holds it -- and records // this host and pid in it. kOwned, naming the holder, when another holder // has it; kBadArgument for a shared filesystem or a directory nested - // under, or containing, an owned one (see above), after letting go of - // the lock and of whatever this call created. + // under, or containing, a held one (see above), after letting go of the + // lock and of whatever this call created. static SpoolStatus Acquire(const std::string& dir, bool allow_shared_filesystem, SpoolOwnerLock* out, std::string* error); diff --git a/src/dmi/engine.py b/src/dmi/engine.py index b36316839..efca98679 100644 --- a/src/dmi/engine.py +++ b/src/dmi/engine.py @@ -506,7 +506,9 @@ def _start_capture_storage(self, record_sink: Optional[Any]) -> Optional[Any]: An explicit ``record_sink`` writes where it was built to: the service drains ``spool_root`` itself, unswept and with no siblings - to adopt, beside the sink's lock if this process holds one. + to adopt, beside the sink's lock if this process holds one. It + passes over the rank directories under ``spool_root``, which a + default-mode start adopts. """ config = self._capture_storage_config if config is None or self._storage_backend != "persistent": diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index 065fc74fc..99fb0506e 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -687,7 +687,8 @@ def claim_spool_directory( ``RANK`` (0 when unset). The directory is created with its owner lock already held. Raises ``SpoolOwnedError`` (a ``RuntimeError``) if another process holds it, and ``ValueError`` for a shared filesystem or a - directory nested in another spool. + directory nested in a spool that another owner holds (a sink-only or + explicit-``record_sink`` run on ``spool_root``). Ready packs under ``spool_root`` outside the layout -- an engine from before it spooled into ``/v1/...``, and a sink-only or diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index d8da5a84b..a8a58075e 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -10,8 +10,9 @@ // 3. A second process is refused, told the holder's pid and host; the // lock goes with its holder, even one killed with SIGKILL, and a child // it forked without exec does not keep it. -// 4. Nesting: a directory under a HELD one, or containing one with a lock -// file, is refused -- also when two processes take the pair at once. +// 4. Nesting: a directory under, or containing, a HELD one is refused -- +// also when two processes take the pair at once -- and a dead one +// nested in a spool is left alone by all of its walks. // 5. The node-local check refuses NFS, Lustre, BeeGFS, CIFS/SMB2, FUSE, // GPFS, 9p, AFS and OrangeFS by statfs f_type, unless explicitly // allowed (a test seam stands in for statfs). @@ -539,34 +540,126 @@ void TestAnOuterAndANestedTakeRacingNeverBothWin() { CHECK(outer_won + inner_won > 0); } -// (4c) An ancestor refuses only while its lock is HELD. A lock file nobody -// holds is a spool that was -- every take leaves its file behind -- and -// refusing on it kept a spool_root that a sink-only run once owned from -// ever holding rank directories. Whoever takes the outer directory next -// meets the inner one's lock file below it and is refused, held or not. +// (4c) Nesting refuses only while the other directory's lock is HELD. A +// lock file nobody holds is a spool that was -- every take leaves its file +// behind -- and a dead spool directory inside another is never the outer +// one's to sweep, count or upload under its own keys: every walk of a +// spool passes over a subdirectory with its own lock file. So a spool_root +// that a sink-only run once owned still takes rank directories, and a +// spool_root that a crashed default-mode run left a rank directory in +// still takes the sink-only or explicit-record_sink modes (the rollback), +// which leave that directory to adoption. void TestAStaleLockFileAboveDoesNotRefuseANestedDirectory() { const std::string base = FreshRoot("stale-above"); + const std::string root = base + "/root"; + const std::string rank = root + "/0123456789ab/r0-0a1b2c3d"; std::string error; { SpoolOwnerLock once; - CHECK(SpoolOwnerLock::Acquire(base + "/root", false, &once, &error) == + CHECK(SpoolOwnerLock::Acquire(root, false, &once, &error) == SpoolStatus::kOk); } - CHECK(fs::exists(base + "/root/.owner.lock")); + CHECK(fs::exists(root + "/.owner.lock")); SpoolOwnerLock inner; error.clear(); - CHECK(SpoolOwnerLock::Acquire(base + "/root/0123456789ab/r0-0a1b2c3d", - false, &inner, &error) == SpoolStatus::kOk); + CHECK(SpoolOwnerLock::Acquire(rank, false, &inner, &error) == + SpoolStatus::kOk); CHECK(error.empty()); SpoolOwnerLock outer; - CHECK(SpoolOwnerLock::Acquire(base + "/root", false, &outer, &error) == + CHECK(SpoolOwnerLock::Acquire(root, false, &outer, &error) == SpoolStatus::kBadArgument); CHECK(Contains(error, "contains")); - inner.Release(); // its lock file stays: still refused + CHECK(Contains(error, rank)); + CHECK(Contains(error, "pid " + std::to_string(::getpid()))); + // The inner one stages a pack and has a stage in flight, then dies. + { + SpoolConfig config{rank, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool writer; + CHECK(Spool::Open(config, &writer, &error) == SpoolStatus::kOk); + CHECK(StageOne(writer, 1, &error) == SpoolStatus::kOk); + } + const std::string in_flight = + rank + "/v1/.018f0000-0000-7000-8000-00000000beef.0badf00d.open"; + std::ofstream(in_flight) << "half a pack"; + inner.Release(); + // Its lock file stays, and nobody holds it: the outer take goes through, + // and nothing of the outer spool touches the dead one. + Spool flat; error.clear(); - CHECK(SpoolOwnerLock::Acquire(base + "/root", false, &outer, &error) == + CHECK(Spool::Open({root, 150}, &flat, &error) == SpoolStatus::kOk); + CHECK(flat.Snapshot().bytes == 0); + std::vector staged; + CHECK(flat.Recover(&staged, &error) == SpoolStatus::kOk); + CHECK(staged.empty()); + CHECK(fs::exists(in_flight)); + CHECK(StageOne(flat, 2, &error) == SpoolStatus::kOk); // 100 of 150 + CHECK(flat.ListPending(&staged, &error) == SpoolStatus::kOk); + CHECK(staged.size() == 1 && staged[0].object_key.rfind("v1/", 0) == 0); + size_t dead_packs = 0; + for (const auto& entry : fs::recursive_directory_iterator(rank)) { + if (entry.path().extension() == ".ready") ++dead_packs; + } + CHECK(dead_packs == 1); + // And while the outer one holds the root, the rank directory cannot be + // taken: an adopter would first have to wait for it. + SpoolOwnerLock again; + error.clear(); + CHECK(SpoolOwnerLock::Acquire(rank, false, &again, &error) == SpoolStatus::kBadArgument); - CHECK(Contains(error, "contains")); + CHECK(Contains(error, "nested")); +} + +// (4e) An adopter takes a dead directory's lock with TryAdopt, which runs +// no nesting check, then Recovers it. A live spool nested inside that dead +// directory -- a root put there, which the nesting rule admits under an +// unheld lock -- had its in-flight .open files swept by that Recover, and +// its packs listed under the outer directory's keys. Every walk passes +// over a subdirectory with its own lock file, so the adoption drains only +// the dead directory's own packs and leaves the directory in place. +void TestAnAdopterLeavesASpoolNestedInADeadOneAlone() { + const std::string base = FreshRoot("nested-in-dead"); + const std::string dead = base + "/0123456789ab/r0-0000dead"; + const std::string nested = dead + "/inner"; + std::string error; + { + Spool gone; + CHECK(Spool::Open({dead, 1 << 20}, &gone, &error) == SpoolStatus::kOk); + CHECK(StageOne(gone, 1, &error) == SpoolStatus::kOk); + } // its owner died + SpoolOwnerLock live; + CHECK(SpoolOwnerLock::Acquire(nested, false, &live, &error) == + SpoolStatus::kOk); + { + SpoolConfig config{nested, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool writer; + CHECK(Spool::Open(config, &writer, &error) == SpoolStatus::kOk); + CHECK(StageOne(writer, 2, &error) == SpoolStatus::kOk); + } + const std::string in_flight = + nested + "/v1/.018f0000-0000-7000-8000-00000000beef.0badf00d.open"; + std::ofstream(in_flight) << "half a pack"; + + SpoolOwnerLock adopter; + CHECK(SpoolOwnerLock::TryAdopt(dead, &adopter, &error) == SpoolStatus::kOk); + SpoolConfig config{dead, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool adopted; + CHECK(Spool::Open(config, &adopted, &error) == SpoolStatus::kOk); + CHECK(adopted.Snapshot().bytes == 100); // its own pack only + std::vector staged; + CHECK(adopted.Recover(&staged, &error) == SpoolStatus::kOk); + CHECK(fs::exists(in_flight)); + CHECK(staged.size() == 1); + if (staged.size() == 1) { + CHECK(staged[0].object_key.rfind("v1/", 0) == 0); + CHECK(adopted.Remove(staged[0], &error) == SpoolStatus::kOk); + } + // Drained of its own, it still holds the live one: it stays. + CHECK(!adopter.ReleaseAndRemoveIfEmpty(&error)); + CHECK(fs::exists(in_flight)); + CHECK(live.held()); } // (4d) A claim killed between creating its directory's staging copy @@ -1051,6 +1144,7 @@ int main() { TestAnOuterAndANestedTakeRacingNeverBothWin(); TestAStaleLockFileAboveDoesNotRefuseANestedDirectory(); TestAnUnheldClaimStagingDirectoryRefusesNothing(); + TestAnAdopterLeavesASpoolNestedInADeadOneAlone(); TestSharedFilesystemsAreRefusedUnlessAllowed(); TestAdoptionLocksOnlyWhatExistsAndIsDead(); TestALockOnAnUnlinkedFileIsTakenAgain(); diff --git a/tests/test_native_spool_owner_lock.py b/tests/test_native_spool_owner_lock.py index 53192b1ec..8c813ba29 100644 --- a/tests/test_native_spool_owner_lock.py +++ b/tests/test_native_spool_owner_lock.py @@ -7,9 +7,11 @@ second process that tries is refused, told who holds it. What that lock must also refuse, and what it must leave alone: -* a directory nested under a held one, or containing one with a lock file: - Scan walks recursively, so the outer spool's cleanup would reach into the - inner one; +* a directory nested under, or containing, a held one: Scan walks + recursively, so the outer spool's cleanup would reach into the inner one + while its owner writes it -- and every walk passes over a subdirectory + with a lock file of its own, so a dead spool directory inside another is + left for its successor to adopt; * ``owner_lock="held_by_caller"`` unless THIS process holds the lock: that mode opens without a lock of its own, for a second Spool in the process that holds one. Beside another process's lock it is refused, naming the @@ -27,7 +29,8 @@ deliberately stricter, and that is not ported to the reference. Skipping ``_refs/`` is C++-only as well: the reference's constructor counts, and its recover() sweeps and quarantines, ``.open`` and ``.ready`` files there, and -that divergence is deliberate, not ported either. +that divergence is deliberate, not ported either, like passing over +nested spool directories. Build: make -C native build/conformance_spool build/conformance_sink """ @@ -160,21 +163,36 @@ def test_a_directory_containing_an_owned_one_is_refused(tmp_path): holder.close() -def test_a_stale_lock_file_above_does_not_refuse_a_nested_directory(tmp_path): - """Above, only a HELD lock refuses: every take leaves its lock file - behind, so a spool_root some spool once owned would otherwise refuse - every directory under it for good. Below, the lock FILE refuses, held - or not: a dead directory's packs are for its successor to adopt, not - for the outer spool to sweep and upload under its own keys -- so the - next process to take the outer directory is refused.""" +def test_a_dead_spool_directory_inside_another_is_left_alone(tmp_path): + """Nesting refuses only while the other directory's lock is HELD. Above: + every take leaves its lock file behind, so a spool_root some spool once + owned would otherwise refuse every directory under it for good. Below: + a dead directory's packs are for its successor to adopt, not for the + outer spool to sweep and upload under its own keys -- so every walk of + a spool passes over a subdirectory with a lock file of its own, and the + outer take goes through. Refused instead, as it was, a spool_root that + a crashed default-mode run left a rank directory in refused the + sink-only and explicit-record_sink modes, the rollback included.""" outer = tmp_path / "spool" assert _spool(op="recover", root=str(outer))["ok"] assert (outer / ".owner.lock").exists() - assert _spool(op="recover", root=str(outer / "inner"))["ok"] - assert (outer / "inner" / ".owner.lock").exists() + inner = outer / "inner" + assert _spool(op="recover", root=str(inner))["ok"] + assert (inner / ".owner.lock").exists() + (inner / "v1").mkdir() + stale = inner / "v1" / ".018f0000-0000-7000-8000-000000000001.abcd1234.open" + stale.write_bytes(b"in progress") + bogus = inner / "v1" / READY_NAME # the wrong checksum: quarantined if met + bogus.write_bytes(b"not a pack") + response = _spool(op="recover", root=str(outer)) - assert not response["ok"], response - assert "contains" in response["what"], response + + assert response["ok"], response + assert response["staged"] == [] + assert response["snapshot"] == {"entries": 0, "bytes": 0} + assert stale.exists() and bogus.exists() + assert sorted(p.name for p in (inner / "v1").iterdir()) == sorted( + [stale.name, bogus.name]) def test_held_by_caller_is_refused_beside_another_processs_holder(tmp_path): diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 3ce1ada89..7127184f0 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -288,6 +288,52 @@ def test_a_spool_root_a_sink_once_owned_still_takes_rank_directories(tmp_path): claim.release() +@pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") +def test_a_crashed_default_run_leaves_spool_root_to_the_other_modes(tmp_path): + """The reverse of the stale lock above: a default-mode run that dies + (or whose close did not drain) leaves its rank directory, lock file + and all, under spool_root. The sink-only mode (the sink takes + spool_root) and the explicit-record_sink rollback (the service takes + it) were then refused as containing an "owned" directory nobody held, + until a default-mode start adopted it -- the rollback blocked just when + it is most likely needed. Both now take spool_root, pass over the dead + directory, and leave it to adoption; a live one still refuses them.""" + pytest.importorskip("torch") + from dmi.storage.native_capture import ( + NativeSinkConfig, claim_spool_directory, + ) + + root = tmp_path / "root" + sink_config = NativeSinkConfig(spool_root=str(root)) + crashed = claim_spool_directory(sink_config, _config()) + dead = Path(crashed.directory) + (dead / "v1").mkdir() + left = dead / "v1" / "left.dmi-pack.ready" + left.write_bytes(b"pack") + in_flight = dead / "v1" / ".left.0badf00d.open" + in_flight.write_bytes(b"half") + crashed._lock.release() # killed: the kernel let go, the files stay + assert _store().spool_owner(str(dead)) is None + + service = _service(_config(), root, spool_owner_lock="take") # rollback + del service + sink = _sink(root) # sink-only + _stage_one(sink) + del sink + flat = sorted((root / "v1").rglob("*.dmi-pack.ready")) + assert len(flat) == 1 + assert left.read_bytes() == b"pack" and in_flight.exists() + + live = claim_spool_directory(sink_config, _config()) + try: + with pytest.raises(RuntimeError, match="contains the spool directory"): + _service(_config(), root, spool_owner_lock="take") + with pytest.raises(RuntimeError, match=f"pid {os.getpid()}"): + _sink(root) + finally: + live.release() + + @pytest.mark.skipif(not SINK_BUILT, reason="the native sink module is not built") def test_a_dead_siblings_bytes_count_against_the_sinks_budget(tmp_path): """Each process start claims a fresh rank directory. A sink that From e0604aff0355639e1830d2f4f11cd9a19d811cc9 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:43:02 -0400 Subject: [PATCH 27/43] Let go of a half-adopted sibling once the service has latched When a foreign lease outlives 2 x TTL, the service latches and its loop returns -- without letting go of adopting_, which holds the dead sibling's owner lock between cycles. The sibling stayed locked by this process, with nobody working on it, until the engine called stop() at close, while the engine kept capturing, possibly for hours. Meanwhile every process on the node, the one now holding the catalog included, read it as live and could only recheck it every 30 s. A latched loop now lets go of the sibling being adopted and of those queued (what stop() does, now one helper), and looks no more; the snapshot's adoption_owed says so. latch_failure wakes the loop, so this happens at once rather than after its wait, which an outage stretches to max_backoff_ns. Tests: test_native_spool_adoption_live.py, red before -- a service half-way through adopting a dead sibling (the object store cut) loses the catalog to a rival for 2 x TTL and latches; within 5 s the sibling's lock is free, with every pack still in it. --- native/csrc/catalog/storage_service.cpp | 34 ++++++++++++-- native/csrc/catalog/storage_service.h | 4 ++ tests/test_native_spool_adoption_live.py | 58 ++++++++++++++++++++++++ 3 files changed, 91 insertions(+), 5 deletions(-) diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index 79c59d195..4d0b51187 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -334,10 +334,7 @@ void CaptureStorageService::stop() { // lease has to keep renewing until that is done. stop_lease_thread(); std::lock_guard cycle(cycle_mutex_); - // A sibling half adopted keeps what is left of it, for the next process - // on the node; what was uploaded from it was indexed, or is owed. - adopting_.reset(); - adoption_queue_.clear(); + let_go_of_adoption(); if (started_) { started_ = false; // A quarantined writer holds no lease, so it writes no tombstone: the @@ -422,11 +419,22 @@ void CaptureStorageService::loop() { if (stop_requested_) return; kick_ = false; } + bool failed = false; { std::lock_guard lock(state_mutex_); - if (failure_) return; // another publisher holds the catalog + failed = failure_ != nullptr; } std::lock_guard cycle(cycle_mutex_); + if (failed) { + // Another publisher holds the catalog, and no cycle runs again. A + // sibling half adopted would stay locked by this process, with + // nobody working on it, until stop() -- the engine's close(), maybe + // hours of capture later -- and every other process on the node + // would read it as live meanwhile. Nor does it look again. + adoption_scan_owed_ = false; + let_go_of_adoption(); + return; + } run_cycle(true); // poll_interval * 2^streak, capped: flush() shares the streak, so an // outage it saw also slows the loop, and a success from either resets it. @@ -587,6 +595,15 @@ void CaptureStorageService::index_or_owe(std::vector refs, unindexed.end()); } +void CaptureStorageService::let_go_of_adoption() { + // A sibling half adopted keeps what is left of it, for the next process + // on the node; what was uploaded from it was indexed, or is owed. + adopting_.reset(); + adoption_queue_.clear(); + std::lock_guard lock(state_mutex_); + state_.adoption_owed = adoption_owed(); +} + bool CaptureStorageService::adoption_owed() const { return adoption_scan_owed_ || adopting_ != nullptr || !adoption_queue_.empty(); @@ -1420,6 +1437,13 @@ void CaptureStorageService::latch_failure(std::exception_ptr failure, std::fprintf(stderr, "dmi capture storage: indexing stopped: %s\n", line.c_str()); std::fflush(stderr); + // Woken now, not after its wait (up to max_backoff_ns in an outage), so + // the loop lets go of a sibling it was adopting at once. + { + std::lock_guard lock(wake_mutex_); + kick_ = true; + } + wake_.notify_all(); } } // namespace dmi_catalog diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index 29daad4c0..9a0504e1f 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -358,6 +358,10 @@ class CaptureStorageService { void block_sibling(const std::string& directory, const std::string& reason, dmi_store::SpoolOwnerLock* lock = nullptr); bool adoption_owed() const; // requires cycle_mutex_ + // Lets go of the sibling being adopted, and of those queued, leaving what + // is left of them for the next process on the node: at stop(), and once + // the service has latched. Requires cycle_mutex_. + void let_go_of_adoption(); bool stop_requested(); // Removes the staging copies (dmi_store::IsSpoolClaimStagingName) that // claims killed before their rename left under the catalog key. diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 423f682ec..f21ce74d3 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -629,5 +629,63 @@ def test_a_sibling_whose_owner_dies_after_start_is_adopted_by_a_recheck( lock.release_and_remove_if_empty() +def test_a_latched_service_lets_go_of_the_sibling_it_was_adopting( + fake_s3, tmp_path): + """A service whose catalog another publisher keeps for 2 x TTL latches, + and its loop stops for good -- but the dead sibling it was half-way + through adopting stayed locked by this process, with nobody working on + it, until the engine's close(), possibly hours of capture later. Every + other process on the node read it as live meanwhile. The latched loop + lets go of it, at once.""" + from dmi.storage.native_capture import NativeCaptureStorage + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + # The servers are part of the catalog key: the sibling is claimed + # through the same switch URLs as the service reaches. + clickhouse = _Switch(CLICKHOUSE_HOST, CLICKHOUSE_HTTP_PORT) + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, + clickhouse_host="127.0.0.1", + clickhouse_port=clickhouse.port, + reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + sibling.release() # its owner is gone + staged = sorted(dead.rglob("*.dmi-pack.ready")) + store.cut() # so the adoption stays half done + lock = _claim(base, config) + service = _service(config, lock.directory) + rival = NativeCaptureStorage( + _storage_config(fake_s3, prefix, holder="rival-publisher", + start_lease_wait_s=10.0, + reconcile_on_start=False), + spool_root=str(tmp_path / "rival"), spool_max_bytes=1 << 30, + sweep_spool=True) + service.start() + try: + _wait_for(lambda: service.snapshot()["upload_failures"] > 0, 60.0) + owner = _store().spool_owner(str(dead)) + assert owner is not None and owner["pid"] == os.getpid(), owner + + clickhouse.cut() # the service can renew no more + rival.start() # waits out the lease the cut service cannot renew + clickhouse.restore() + _wait_for(lambda: service.snapshot()["failed"], 30.0) + _wait_for(lambda: _store().spool_owner(str(dead)) is None, 5.0) + snapshot = service.snapshot() + assert snapshot["adoption_owed"] is False, snapshot + assert snapshot["adopted_packs"] == 0, snapshot + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + finally: + service.stop() + rival.stop() + clickhouse.close() + store.close() + lock.release_and_remove_if_empty() + + if __name__ == "__main__": _dead_capture_process(*sys.argv[1:4]) From 20191d172a60c9c79ad1502362da0fea3388dc8b Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:43:10 -0400 Subject: [PATCH 28/43] Pin that a flush adopts nothing, with the loop idle The flush half of the fix that moved adoption into the loop -- flush()'s cycles run with adopt=false -- had no test that could fail. The one assertion aimed at it, flush(0.5) returning within 2 s while the store is down, ran while the loop's first cycle held the cycle mutex through its own upload round, so the flush timed out on the lock without running a cycle at all (its own comment says as much). A flush that adopted again survived the whole adoption live suite, and would reopen the overrun for flush() and close()'s drain: a dead pack's retry chain past the deadline, minutes against a stalled store. The new test idles the loop (an hour's poll interval) once its first adoption round has failed with the object store cut, so the flush holds the cycle itself. It must return within 2 s, with upload_failures and adopted_packs unchanged and the dead packs where they were. Tests: test_native_spool_adoption_live.py -- passes; with flush() running run_cycle(true) (the review's surviving mutation) it fails, the flush uploading dead packs and timing out. --- tests/test_native_spool_adoption_live.py | 44 ++++++++++++++++++++++++ 1 file changed, 44 insertions(+) diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index f21ce74d3..36c3e8b69 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -629,6 +629,50 @@ def test_a_sibling_whose_owner_dies_after_start_is_adopted_by_a_recheck( lock.release_and_remove_if_empty() +def test_flush_adopts_nothing_while_the_loop_is_idle(fake_s3, tmp_path): + """flush() covers this process's records, and its cycles never adopt: + a dead backlog is the loop's. With the object store down, a flush that + adopted sat through a dead pack's retry chain, past its deadline by a + round (minutes against a stalled store). Here the loop is idle -- its + first round has failed and its poll interval is an hour -- so the flush + holds the cycle itself and would be seen adopting: uploading (and + failing) dead packs and overrunning its deadline.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, poll_interval_s=3600.0) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + sibling.release() # its owner is gone + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert staged + store.cut() + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + # The loop's first cycle, which start() kicks, begins adopting + # and fails its first round; then it waits out its interval. + _wait_for(lambda: service.snapshot()["upload_failures"] > 0 + and service.snapshot()["cycles"] >= 1, 60.0) + before = service.snapshot() + assert before["adoption_owed"] is True, before + flushed = time.monotonic() + service.flush(0.5) # nothing of its own: drained + assert time.monotonic() - flushed < 2.0 + after = service.snapshot() + assert after["upload_failures"] == before["upload_failures"], after + assert after["adopted_packs"] == 0, after + assert after["adoption_owed"] is True, after + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + finally: + service.stop() + lock.release_and_remove_if_empty() + + def test_a_latched_service_lets_go_of_the_sibling_it_was_adopting( fake_s3, tmp_path): """A service whose catalog another publisher keeps for 2 x TTL latches, From 633f7e0e5a975089257ba767156d0d3cba73f444 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:46:17 -0400 Subject: [PATCH 29/43] Race a scan against new spool directories, not only look at the result TestANewDirectoryAppearsWithItsLockHeld promised that a new directory appears with its lock held -- CreateLocked's staging copy and rename -- but checked only the end state: one entry, its lock file present. Creating the directory and then locking it in place passed every CPU and live suite, although that is exactly the window in which an adopter's scan meets a brand-new sibling unlocked, takes it for dead, and removes it from under its claimer (whose create_record_runtime then fails). The test now races a watcher against the takes: a forked child lists the parent over and over and probes every rank directory it has not yet seen held, while this process creates 100 and holds every one, so their creation is the only moment one could read as unheld. Each take is slowed at the lock-open seam, where a lock file is open and not yet locked, so a directory there to be seen before its lock is seen. Tests: test_spool_owner_lock.cpp -- passes; against a spool.cpp that creates the directory and then locks it (the review's surviving mutation) the watcher meets all 100 unheld, in every run. --- tests/native/test_spool_owner_lock.cpp | 74 ++++++++++++++++++++++++++ 1 file changed, 74 insertions(+) diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index a8a58075e..8dd4b58de 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -23,6 +23,7 @@ // // Built and run by tests/test_native_spool_owner_lock_unit.py. +#include #include #include #include @@ -897,6 +898,79 @@ void TestANewDirectoryAppearsWithItsLockHeld() { CHECK(names == std::set{"r0-0a1b2c3d"}); CHECK(lock.directory() == base + "/r0-0a1b2c3d"); CHECK(fs::exists(base + "/r0-0a1b2c3d/.owner.lock")); + + // The window itself: a watcher -- an adopter's scan, in effect -- lists + // the parent over and over and probes each rank directory it has not yet + // seen held, while this process creates many and keeps holding every + // one, so their creation is the only moment one could read as unheld. + // Each take is slowed where it has opened a lock file and not locked it + // (the lock-open seam), so a directory there to be seen before its lock + // would be seen. One made first and locked after gives such a scan a + // dead-looking sibling, which an adopter would take, and remove from + // under its claimer. + const std::string parent = base + "/race"; + fs::create_directories(parent); + int report[2], stop[2]; + CHECK(::pipe(report) == 0 && ::pipe(stop) == 0); + const pid_t watcher = ::fork(); + if (watcher == 0) { + ::close(report[0]); + ::close(stop[1]); + ::fcntl(stop[0], F_SETFL, O_NONBLOCK); + std::set held, unheld; + char byte = 0; + while (::read(stop[0], &byte, 1) < 0 && errno == EAGAIN) { + std::error_code ec; + for (fs::directory_iterator it(parent, ec), end; !ec && it != end; + it.increment(ec)) { + const std::string name = it->path().filename().string(); + uint64_t rank = 0; + std::string incarnation; + if (held.count(name) != 0 || + !dmi_store::ParseSpoolRankDirectoryName(name, &rank, + &incarnation)) { + continue; + } + if (dmi_store::ReadSpoolOwner(it->path().string(), nullptr)) { + held.insert(name); + } else { + unheld.insert(name); + } + } + } + const int count = static_cast(unheld.size()); + if (::write(report[1], &count, sizeof(count)) != sizeof(count)) { + ::_exit(3); + } + ::_exit(0); + } + ::close(report[1]); + ::close(stop[0]); + dmi_store::SetLockOpenHookForTesting([](const std::string&) { + ::usleep(1000); + }); + std::vector claims(100); + int taken = 0; + for (size_t i = 0; i < claims.size(); ++i) { + char name[32]; + std::snprintf(name, sizeof(name), "/r0-%08zx", i); + if (SpoolOwnerLock::Acquire(parent + name, false, &claims[i], &error) == + SpoolStatus::kOk) { + ++taken; + } + } + dmi_store::SetLockOpenHookForTesting(nullptr); + ::close(stop[1]); // the watcher stops at EOF + int unheld = -1; + CHECK(::read(report[0], &unheld, sizeof(unheld)) == sizeof(unheld)); + ::close(report[0]); + int status = 0; + ::waitpid(watcher, &status, 0); + CHECK(taken == 100); + if (unheld != 0) { + std::cerr << "a scan met " << unheld << " unheld new directories\n"; + } + CHECK(unheld == 0); } // (8) The budget across incarnations. Every process start gets a fresh From 988ee724575480e850b4bd286df602c12c545570 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 06:58:48 -0400 Subject: [PATCH 30/43] Say the fork handler tracks the directory's lock descriptor too The set of lock descriptors the fork handler closes in a child holds the directory's descriptor as well as the lock file's since the directory is locked too; its comment named only the lock file. --- native/csrc/store/spool.cpp | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index 8b263c826..154f9d8d2 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -741,8 +741,9 @@ bool SpoolOwnedByThisProcess(const std::string& dir) { namespace { -// Every descriptor this binary has open on a spool owner lock file: held, -// or between its open() and its flock, or on its way to close(). The fork +// Every descriptor this binary has open on a spool owner lock file, or on +// the directory it locks with it: held, or between its open() and its +// flock, or on its way to close(). The fork // handlers close the child's copies of all of them. Tracking starts at the // open() and ends at the close(), each under the mutex that BeforeFork // takes, so no fork -- from any thread, at any point of a take or a From 10b4fe76b80c70d5adb8ebc66c83749f49afeabb Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 14:09:51 -0400 Subject: [PATCH 31/43] Adopt dead spools through the upload path's chunks, clients and cancels The merge left adoption as B6 wrote it, beside B5's upload path rather than on it. Its rounds uploaded through the index-read client (s3_) with an uploader that had no Cancellation, so stop() cut an adopted transfer only by accident and the uploader went on retrying; a cut pack was booked as an upload failure, logged as "adopting dead spool ... upload failed" and pushed to the back of the queue. Its listing of a dead backlog -- Recover, which hashes every pack -- could not be cut at all, so stop() waited for the whole hash (9 s for 32 sparse 256 MiB packs). - A round is now one chunk of the service's own upload path: upload_chunk (UploadStaged, then the same booking of uploaded, failed and cancelled packs), then index_or_owe before the next chunk. It is at most uploader.max_workers and indexer.max_packs packs, so at most one chunk of a dead spool is out of it and not yet in the catalog. - The adoption uploader writes through upload_s3_ and has upload_cancel_, as the service's own does: stop() cuts its transfers, retries and backoff, and its index reads go through s3_, which read_cancel_ cuts. - A pack a cancel cut goes back to the front of the sibling's remaining packs, is counted in cancelled_uploads, not upload_failures, and makes the cycle cut short rather than failed; what a cancel left owed counts in the cycle's deferred, as the service's own does. - Spool::Recover gains B5's (cancel, cut) form, and begin_adoption lists a dead spool with upload_cancel_: a cut listing lets the directory go, whole and unlocked, and it is looked at again from the start. - A cycle adopts only once every staged pack of its own went up, nothing is owed and no cancel came. stop() still never removes a dead directory: only finish_adoption does, once none of its packs is left. --- native/csrc/catalog/storage_service.cpp | 225 ++++++++++++++--------- native/csrc/catalog/storage_service.h | 60 ++++-- native/csrc/store/spool.cpp | 5 + native/csrc/store/spool.h | 7 + tests/native/test_spool_owner_lock.cpp | 65 ++++++- tests/test_native_spool_adoption_live.py | 130 ++++++++++++- 6 files changed, 387 insertions(+), 105 deletions(-) diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index 2e7d8468f..b7890e2df 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -640,60 +640,41 @@ CaptureStorageService::CycleOutcome CaptureStorageService::run_cycle( } } const size_t end = std::min(staged.size(), next + chunk); - const dmi_store::UploadBatchResult batch = uploader_->UploadStaged( + ChunkOutcome sent = upload_chunk( + uploader_.get(), std::vector(staged.begin() + next, staged.begin() + end)); next = end; - std::vector to_index; - uint64_t uploaded_packs = 0; - uint64_t uploaded_bytes = 0; - size_t failed_uploads = 0; - size_t cancelled_uploads = 0; - for (size_t i = 0; i < batch.refs.size(); ++i) { - const dmi_store::PackRef& ref = batch.refs[i]; - if (!ref.pack_id.empty()) { - to_index.push_back({ref.pack_id, ref.store_id, ref.object_key, - ref.object_bytes, ref.checksum, - ref.record_count}); - ++uploaded_packs; - uploaded_bytes += ref.object_bytes; - } else if (i < batch.failures.size() && - batch.failures[i].cancelled) { - ++cancelled_uploads; // still staged; not the pack's fault - } else { - // A failed upload stays in the spool, so a later cycle - // retries it. - ++failed_uploads; - if (i < batch.failures.size()) { - record_error("upload failed for " + - batch.failures[i].object_key + ": " + - batch.failures[i].error); - } + for (size_t i = 0; i < sent.batch.failures.size(); ++i) { + const dmi_store::UploadFailure& failure = sent.batch.failures[i]; + // A failed upload stays in the spool, so a later cycle retries + // it; a cancelled one is still staged, and not the pack's fault. + if (!sent.batch.refs[i].pack_id.empty() || failure.cancelled) { + continue; } + record_error("upload failed for " + failure.object_key + ": " + + failure.error); } - if (cancelled_uploads != 0) outcome.cut_short = true; - upload_failures += failed_uploads; - { - std::lock_guard lock(state_mutex_); - state_.uploaded_packs += uploaded_packs; - state_.uploaded_bytes += uploaded_bytes; - state_.upload_failures += failed_uploads; - state_.cancelled_uploads += cancelled_uploads; - } - deferred += index_or_owe(std::move(to_index), catalog, deadline_ns); + if (sent.cancelled != 0) outcome.cut_short = true; + upload_failures += sent.failed; + deferred += index_or_owe(std::move(sent.to_index), true, deadline_ns); } uploaded_all = next == staged.size() && upload_failures == 0; } } // 3a. The loop's cycles adopt dead siblings, a slice at a time, under - // the same rule as the uploads above: only with the lease, nothing - // owed and nothing of our own failing. flush()'s cycles do not: a - // dead backlog is not this process's records. + // the same rule as the uploads above: only with the lease, every + // staged pack of this process's up, nothing owed, and no cancel. + // flush()'s cycles do not: a dead backlog is not this process's + // records. Adoption uploads and indexes through the same chunks, + // and the same Cancellations, as the service's own spool. bool adoption_failed = false; - if (adopt && config_.adopt_sibling_spools && catalog && - pending_index_.empty() && upload_failures == 0) { - adoption_failed = !adopt_step(); + if (adopt && config_.adopt_sibling_spools && uploaded_all && + pending_index_.empty() && !upload_cancel_.cancelled()) { + bool adoption_cut = false; + adoption_failed = !adopt_step(deadline_ns, &deferred, &adoption_cut); + if (adoption_cut) outcome.cut_short = true; } // 4. Reconcile on its interval, or when the pass at start() lost the @@ -774,9 +755,39 @@ size_t CaptureStorageService::index_or_owe(std::vector refs, return deferred; } +CaptureStorageService::ChunkOutcome CaptureStorageService::upload_chunk( + dmi_store::SpoolUploader* uploader, + std::vector chunk) { + ChunkOutcome sent; + sent.batch = uploader->UploadStaged(std::move(chunk)); + uint64_t uploaded_bytes = 0; + for (size_t i = 0; i < sent.batch.refs.size(); ++i) { + const dmi_store::PackRef& ref = sent.batch.refs[i]; + if (!ref.pack_id.empty()) { + sent.to_index.push_back({ref.pack_id, ref.store_id, ref.object_key, + ref.object_bytes, ref.checksum, + ref.record_count}); + uploaded_bytes += ref.object_bytes; + } else if (i < sent.batch.failures.size() && + sent.batch.failures[i].cancelled) { + ++sent.cancelled; // still staged; not the pack's fault + } else { + ++sent.failed; // still staged too + } + } + std::lock_guard lock(state_mutex_); + state_.uploaded_packs += sent.to_index.size(); + state_.uploaded_bytes += uploaded_bytes; + state_.upload_failures += sent.failed; + state_.cancelled_uploads += sent.cancelled; + return sent; +} + void CaptureStorageService::let_go_of_adoption() { // A sibling half adopted keeps what is left of it, for the next process - // on the node; what was uploaded from it was indexed, or is owed. + // on the node; what was uploaded from it was indexed, or is owed. Its + // directory is not removed: only finish_adoption() removes one, once no + // pack of it is left. adopting_.reset(); adoption_queue_.clear(); std::lock_guard lock(state_mutex_); @@ -793,7 +804,9 @@ bool CaptureStorageService::stop_requested() { return stop_requested_; } -bool CaptureStorageService::adopt_step() { +bool CaptureStorageService::adopt_step(uint64_t deadline_ns, size_t* deferred, + bool* cut_short) { + *cut_short = false; const uint64_t started = steady_ns(); if (adopting_ == nullptr && adoption_queue_.empty()) { // Look at the siblings when that is owed (from start() on) or, while @@ -807,11 +820,25 @@ bool CaptureStorageService::adopt_step() { } bool ok = true; while (!stop_requested()) { + // The service's uploads' Cancellation is adoption's too: stop() cuts + // it between rounds, and in a round (its listing, its uploads, and, + // through the reads' Cancellation, its index reads). + if (upload_cancel_.cancelled()) { + *cut_short = true; + break; + } if (adopting_ == nullptr) { if (adoption_queue_.empty()) break; const std::string next = adoption_queue_.front(); adoption_queue_.pop_front(); - if (!begin_adoption(next)) ok = false; + bool cut = false; + if (!begin_adoption(next, &cut)) ok = false; + if (cut) { + // Its listing was cut: it is looked at again, from the start. + adoption_queue_.push_front(next); + *cut_short = true; + break; + } continue; } if (adopting_->remaining.empty()) { @@ -827,10 +854,15 @@ bool CaptureStorageService::adopt_step() { catalog = writer_.held_lease() != nullptr; } if (!catalog || !pending_index_.empty()) break; - if (!upload_adopted_round()) { + bool cut = false; + if (!upload_adopted_round(deadline_ns, deferred, &cut)) { ok = false; // the failed packs stay in the dead spool break; } + if (cut) { + *cut_short = true; // the cut packs stay in the dead spool + break; + } if (steady_ns() - started >= config_.adoption_slice_ns) break; } return ok; @@ -887,7 +919,9 @@ bool CaptureStorageService::scan_siblings() { return true; } -bool CaptureStorageService::begin_adoption(const std::string& directory) { +bool CaptureStorageService::begin_adoption(const std::string& directory, + bool* cut) { + *cut = false; auto adoption = std::make_unique(); adoption->directory = directory; std::string error; @@ -913,11 +947,16 @@ bool CaptureStorageService::begin_adoption(const std::string& directory) { std::vector ready; if (dmi_store::Spool::Open(config, &adoption->spool, &error) != dmi_store::SpoolStatus::kOk || - adoption->spool.Recover(&ready, &error) != + adoption->spool.Recover(&ready, &error, &upload_cancel_, cut) != dmi_store::SpoolStatus::kOk) { block_sibling(directory, "cannot open it: " + error, &adoption->lock); return true; } + // Validating hashes every pack of a dead backlog; the uploads' cancel + // (stop()) stops it between packs, as it does the service's own listing. + // The directory is let go of whole, its lock with it: nothing of it was + // uploaded, and nothing is removed. + if (*cut) return true; // Each pack's identity and object key come from the pack and its path in // the dead directory, exactly as its owner would have uploaded it. adoption->remaining.assign(ready.begin(), ready.end()); @@ -925,57 +964,67 @@ bool CaptureStorageService::begin_adoption(const std::string& directory) { return true; } -bool CaptureStorageService::upload_adopted_round() { +bool CaptureStorageService::upload_adopted_round(uint64_t deadline_ns, + size_t* deferred, + bool* cut) { + *cut = false; Adoption& adoption = *adopting_; - const size_t round = - static_cast(std::max(1, config_.uploader.max_workers)); + // A round is a chunk of the service's own upload path (upload_chunk, then + // index_or_owe): uploaded, then indexed before the next is uploaded, so + // at most one chunk (indexer.max_packs) of it is out of the dead spool + // and not yet in the catalog. It is at most uploader.max_workers packs as + // well, since the adoption slice is checked between rounds. + const size_t round = static_cast(std::min( + std::max(1, config_.uploader.max_workers), + std::max(1, config_.indexer.max_packs))); std::vector entries; while (!adoption.remaining.empty() && entries.size() < round) { entries.push_back(std::move(adoption.remaining.front())); adoption.remaining.pop_front(); } - dmi_store::SpoolUploader uploader(&adoption.spool, &s3_, config_.uploader); - const dmi_store::UploadBatchResult batch = uploader.UploadStaged(entries); - std::vector to_index; - uint64_t uploaded_bytes = 0; - size_t failures = 0; + // Through the service's upload client and its Cancellation, so stop() + // cuts an adoption's transfers, retries and backoff as it does the + // service's own. + dmi_store::SpoolUploader uploader(&adoption.spool, &upload_s3_, + config_.uploader); + uploader.set_cancellation(&upload_cancel_); + ChunkOutcome sent = upload_chunk(&uploader, entries); size_t retryable = 0; - for (size_t i = 0; i < batch.refs.size(); ++i) { - const dmi_store::PackRef& ref = batch.refs[i]; - if (ref.pack_id.empty()) { - // Still in the dead spool either way. One a later try could upload - // is retried by a later cycle; one no try by this service can is - // not, and blocks the directory once the rest are up. - ++failures; - const dmi_store::UploadFailure failure = - i < batch.failures.size() ? batch.failures[i] - : dmi_store::UploadFailure{}; - const std::string what = failure.object_key + ": " + failure.error; - if (failure.retryable) { - ++retryable; - adoption.remaining.push_back(entries[i]); - record_error("adopting dead spool " + adoption.directory + - ": upload failed for " + what); - } else if (adoption.blocked.empty()) { - adoption.blocked = "it holds a pack this service can never upload, " + - what; - } - continue; + std::vector cancelled; + for (size_t i = 0; i < sent.batch.refs.size(); ++i) { + if (!sent.batch.refs[i].pack_id.empty()) continue; + // Still in the dead spool either way. One a cancel cut short goes + // back to the front, as it was; one a later try could upload is + // retried by a later cycle; one no try by this service can is not, + // and blocks the directory once the rest are up. + const dmi_store::UploadFailure failure = + i < sent.batch.failures.size() ? sent.batch.failures[i] + : dmi_store::UploadFailure{}; + const std::string what = failure.object_key + ": " + failure.error; + if (failure.cancelled) { + cancelled.push_back(entries[i]); + } else if (failure.retryable) { + ++retryable; + adoption.remaining.push_back(entries[i]); + record_error("adopting dead spool " + adoption.directory + + ": upload failed for " + what); + } else if (adoption.blocked.empty()) { + adoption.blocked = "it holds a pack this service can never upload, " + + what; } - to_index.push_back({ref.pack_id, ref.store_id, ref.object_key, - ref.object_bytes, ref.checksum, ref.record_count}); - uploaded_bytes += ref.object_bytes; } + adoption.remaining.insert(adoption.remaining.begin(), cancelled.begin(), + cancelled.end()); { std::lock_guard state(state_mutex_); - state_.uploaded_packs += to_index.size(); - state_.uploaded_bytes += uploaded_bytes; - state_.upload_failures += failures; - state_.adopted_packs += to_index.size(); - } - // Uploaded, so gone from the dead spool: indexed now, or owed in - // pending_index_ like any pack of this service's own. - index_or_owe(std::move(to_index), true, 0); + state_.adopted_packs += sent.to_index.size(); + } + *cut = !cancelled.empty(); + // Uploaded, so gone from the dead spool: indexed now, a cancel of the + // uploads or not, or owed in pending_index_ like any pack of this + // service's own. After the requeue above, since a lost lease throws out + // of here and the packs still in the dead spool must stay in remaining. + *deferred += index_or_owe(std::move(sent.to_index), true, deadline_ns); return retryable == 0; } diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index 0915d1c59..55993b6f8 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -139,10 +139,16 @@ struct StorageServiceConfig { // cycle probes the owner lock of every sibling rank directory, then works // through those whose owner is gone: it takes one's lock, sweeps its // .open files and validates its .ready packs once, and uploads and - // indexes them a round (uploader.max_workers packs) at a time -- under - // the cycle's upload rules, the lease and nothing owed to the catalog - // checked before every round -- until adoption_slice_ns has passed or a - // stop is requested. The sibling's lock and its remaining packs are kept + // indexes them a round at a time -- a chunk of the service's own upload + // path, at most uploader.max_workers and indexer.max_packs packs, each + // indexed before the next is uploaded -- under the cycle's upload rules, + // the lease and nothing owed to the catalog checked before every round, + // until adoption_slice_ns has passed or a stop is requested. Its + // listing, uploads and index reads go through the service's own clients + // and Cancellations, so stop() cuts an adoption as it cuts the service's + // own work: what it did not upload stays in the dead spool, which is let + // go of and never removed while a pack of it is left, for the next + // process on the node. The sibling's lock and its remaining packs are kept // between cycles, so a large backlog is hashed once, not on every cycle, // and neither a flush() nor the service's own uploads wait behind all of // it. A drained sibling is removed once nothing but its lock file is @@ -355,7 +361,8 @@ class CaptureStorageService { // multipart upload it cut (one attempt, 5 s at most), and the lease // release, each catalog request bounded by the client's request timeout // (under the lease, by the lease deadline); the lease renews until the - // loop is done. + // loop is done. An adoption in flight is cut the same way, its dead + // sibling let go of with what it still holds (adopt_sibling_spools). void stop(); StorageServiceSnapshot snapshot() const; @@ -373,6 +380,16 @@ class CaptureStorageService { bool cut_short = false; }; + // One chunk of the upload path, the service's own spool's or a dead + // sibling's (upload_chunk): the batch, positional as UploadStaged returns + // it, what went up and is to be indexed, and how many were not. + struct ChunkOutcome { + dmi_store::UploadBatchResult batch; + std::vector to_index; + size_t failed = 0; // still staged; counted in upload_failures + size_t cancelled = 0; // still staged: a cancel cut them short + }; + // Holds lease_mutex_ for a stretch of catalog work, and bounds every // request the thread makes meanwhile by the held lease's deadline -- read // afresh per request, so a renewal inside the stretch extends it at once. @@ -397,20 +414,26 @@ class CaptureStorageService { // thread renewing it. Requires cycle_mutex_. void sweep_and_reconcile_at_start(); // One cycle's share of adoption (see adopt_sibling_spools): looks at the - // siblings when that is due, then adopts until the slice ends. False when - // an upload failed, so the cycle backs off. Requires cycle_mutex_. Only a - // lost lease propagates. - bool adopt_step(); + // siblings when that is due, then adopts until the slice ends, a stop is + // requested or the uploads' Cancellation is cancelled -- which also cuts + // a round in flight, and *cut_short then says so. Adds what its index + // passes left owed through a cancel to *deferred (index_bounded). False + // when an upload failed, so the cycle backs off. Requires cycle_mutex_. + // Only a lost lease propagates. + bool adopt_step(uint64_t deadline_ns, size_t* deferred, bool* cut_short); // Probes every sibling rank directory's owner lock and queues the dead // ones. False when the directory cannot be listed. Requires cycle_mutex_. bool scan_siblings(); // Takes a queued sibling's lock, sweeps it and lists its packs into - // adopting_; leaves adopting_ empty for a live or vanished one. False - // when it could not be locked or opened. - bool begin_adoption(const std::string& directory); - // Uploads and indexes one round of adopting_'s packs; false when an - // upload failed. - bool upload_adopted_round(); + // adopting_; leaves adopting_ empty for a live or vanished one, and for + // one whose listing the uploads' Cancellation cut (*cut), whose lock it + // lets go of. False when it could not be locked or opened. + bool begin_adoption(const std::string& directory, bool* cut); + // Uploads and indexes one round of adopting_'s packs, one chunk of the + // service's upload path; *cut when a cancel left some in the dead spool, + // back at the front of adopting_. False when an upload failed. + bool upload_adopted_round(uint64_t deadline_ns, size_t* deferred, + bool* cut); // adopting_ holds no pack any more: removes the directory, or leaves a // blocked one. void finish_adoption(); @@ -429,6 +452,13 @@ class CaptureStorageService { // claims killed before their rename left under the catalog key. void clear_dead_claim_staging( const std::vector& staging); + // Uploads one chunk -- packs a listing returned, in its order -- through + // `uploader`, the service's own or an adoption's over a dead spool, and + // books the outcome. The caller indexes to_index (index_or_owe) before it + // uploads another chunk, so at most one chunk is ever out of a spool and + // not yet in the catalog. Requires cycle_mutex_. + ChunkOutcome upload_chunk(dmi_store::SpoolUploader* uploader, + std::vector chunk); // Indexes refs that are gone from their spool, keeping whatever does not // index in pending_index_ -- the only record of it in-process. With no // catalog, keeps them all. Returns how many of those a cancel or the diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index c20b34c9a..b606b174c 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -1757,6 +1757,11 @@ SpoolStatus Spool::Recover(std::vector* out, std::string* error) { return Scan(out, true, error); } +SpoolStatus Spool::Recover(std::vector* out, std::string* error, + const Cancellation* cancel, bool* cut) { + return Scan(out, true, error, cancel, cut); +} + SpoolStatus Spool::ListPending(std::vector* out, std::string* error) { return Scan(out, false, error); } diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index af0f2cca5..cd1c695d3 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -354,6 +354,13 @@ class Spool { // lock keeps other processes out; writers in this process sharing it // (kHeldByCaller) are the caller's to order. SpoolStatus Recover(std::vector* out, std::string* error); + // The same, but its validation stops between packs once `cancel` is + // cancelled, as ListPending's does: *cut then says so and *out is empty. + // The .open files are swept by then, and the account is left as it was. + // For an adopter's listing of a dead spool, whose backlog it would + // otherwise hash whole before a stop could take effect. + SpoolStatus Recover(std::vector* out, std::string* error, + const Cancellation* cancel, bool* cut); // Validate and list ready packs without deleting in-progress writes. SpoolStatus ListPending(std::vector* out, std::string* error); diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 8dd4b58de..57f1238c2 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -17,7 +17,8 @@ // GPFS, 9p, AFS and OrangeFS by statfs f_type, unless explicitly // allowed (a test seam stands in for statfs). // 6. Adoption's try-lock never creates a directory, and a released -// directory that holds nothing but its lock file can be removed. +// directory that holds nothing but its lock file can be removed; an +// adopter's listing of a dead spool stops between packs on a cancel. // 7. The directory layout of the plan's section 2.3. // 8. The spool budget charges what dead sibling directories hold. // @@ -38,6 +39,7 @@ #include #include +#include "store/cancel.h" #include "store/spool.h" namespace fs = std::filesystem; @@ -973,6 +975,66 @@ void TestANewDirectoryAppearsWithItsLockHeld() { CHECK(unheld == 0); } +// (6e) An adopter lists a dead spool once, through Recover, which hashes +// every pack of what may be a large backlog. The storage service hands it +// the Cancellation stop() cancels, and a cancelled listing stops between +// packs: nothing listed, *cut, every ready pack where it was (nothing is +// lost, and nothing is uploaded from a cut listing), the account as it +// was. The dead owner's stale .open file is swept by then. Not cancelled, +// the same listing lists every pack. +void TestAnAdoptersListingStopsOnACancel() { + const std::string dead = + FreshRoot("adopt-cut") + "/0123456789ab/r0-0000dead"; + std::string error; + { + Spool gone; + CHECK(Spool::Open({dead, 1 << 20}, &gone, &error) == SpoolStatus::kOk); + for (int n = 1; n <= 3; ++n) { + CHECK(StageOne(gone, n, &error) == SpoolStatus::kOk); + } + } // its owner died + const std::string stale = + dead + "/v1/.018f0000-0000-7000-8000-00000000dead.0badf00d.open"; + std::ofstream(stale) << "half a pack"; + const auto readies = [&dead] { + std::set found; + for (const auto& entry : fs::recursive_directory_iterator(dead)) { + const std::string name = entry.path().filename().string(); + if (name.size() > 6 && name.substr(name.size() - 6) == ".ready") { + found.insert(entry.path().string()); + } + } + return found; + }; + const std::set staged_before = readies(); + CHECK(staged_before.size() == 3); + + SpoolOwnerLock adopter; + CHECK(SpoolOwnerLock::TryAdopt(dead, &adopter, &error) == SpoolStatus::kOk); + SpoolConfig config{dead, 1 << 20}; + config.owner_lock = OwnerLock::kHeldByCaller; + Spool adopted; + CHECK(Spool::Open(config, &adopted, &error) == SpoolStatus::kOk); + const uint64_t bytes_before = adopted.Snapshot().bytes; + + dmi_store::Cancellation cancel; + cancel.Cancel(); + std::vector listed; + bool cut = false; + CHECK(adopted.Recover(&listed, &error, &cancel, &cut) == SpoolStatus::kOk); + CHECK(cut); + CHECK(listed.empty()); + CHECK(readies() == staged_before); + CHECK(!fs::exists(stale)); + CHECK(adopted.Snapshot().bytes == bytes_before); + + cancel.Reset(); + CHECK(adopted.Recover(&listed, &error, &cancel, &cut) == SpoolStatus::kOk); + CHECK(!cut); + CHECK(listed.size() == 3); + CHECK(adopted.Snapshot().bytes == 300); +} + // (8) The budget across incarnations. Every process start gets a fresh // rank directory, so a spool that charged only its own directory let each // crash-restart add a full max_bytes while uploads were blocked. With @@ -1224,6 +1286,7 @@ int main() { TestALockOnAnUnlinkedFileIsTakenAgain(); TestAReplacedLockFileLeavesTheDirectoryOwned(); TestANewDirectoryAppearsWithItsLockHeld(); + TestAnAdoptersListingStopsOnACancel(); TestTheDirectoryLayout(); TestDeadSiblingsCountAgainstTheBudget(); if (g_failures != 0) { diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 36c3e8b69..f48726606 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -16,7 +16,9 @@ When the object store is down at start, the dead directory stays as it was (its packs are durable there), flush() does not report drained, and the loop adopts it once the store is back. A sibling whose owner is still -alive at start and dies later is adopted by a later pass. +alive at start and dies later is adopted by a later pass. stop() cuts an +adoption as it cuts the service's own work -- its uploads, and its listing +of a dead backlog -- and leaves the dead directory with every pack in it. Needs ClickHouse on 127.0.0.1:8123/9000 and the native sink and store modules: make -C native build/_dmi_native_sink build/_dmi_native_store @@ -673,6 +675,132 @@ def test_flush_adopts_nothing_while_the_loop_is_idle(fake_s3, tmp_path): lock.release_and_remove_if_empty() +def test_stop_cuts_an_adoption_stalled_on_its_uploads(fake_s3, tmp_path): + """An adoption uploads through the service's own upload client and + Cancellation, so stop() cuts it as it cuts the service's own uploads. + Here every PUT is held open and never answered: before, the adoption's + uploader went round a pack's retries on a client stop() did not cancel, + and stop() sat through read timeouts. Now stop() returns at once, every + pack of the dead spool is still in it (none was uploaded, none lost), + and the directory is left, unlocked, for the next incarnation -- which + adopts it.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, s3_read_timeout_s=30, + reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + sibling.release() # its owner is gone + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert len(staged) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + store.stall_requests(lambda request: request.startswith(b"PUT ")) + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + _wait_for(lambda: store.stalled, 30.0) # an adopted PUT in flight + stopping = time.monotonic() + service.stop() + assert time.monotonic() - stopping < 5.0 + snapshot = service.snapshot() + assert snapshot["cancelled_uploads"] >= 1, snapshot + assert snapshot["upload_failures"] == 0, snapshot + assert snapshot["adopted_packs"] == 0, snapshot + assert "adopting dead spool" not in snapshot["last_error"], snapshot + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + assert _store().spool_owner(str(dead)) is None + finally: + service.stop() + lock.release_and_remove_if_empty() + + store.restore() + successor_lock = _claim(base, config) + successor = _service(config, successor_lock.directory) + successor.start() + try: + _wait_for(_adopted(successor, 1), 60.0) + assert not dead.exists() + successor.flush(60.0) + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in _envelope( + STAGED_BY_THE_DEAD).expected.items()} + assert _read_all(config) == expected + finally: + successor.stop() + store.close() + successor_lock.release_and_remove_if_empty() + + +def _sparse_backlog(directory: Path, count: int, size: int) -> list[Path]: + """`count` ready packs of `size` zero bytes, sparse and named for their + checksum: validating them hashes every byte, which takes seconds, while + they take no disk.""" + import hashlib + import uuid + + digest = hashlib.sha256() + zeros = bytes(1 << 20) + for _ in range(size // len(zeros)): + digest.update(zeros) + packs = directory / "v1" + packs.mkdir(parents=True, exist_ok=True) + backlog = [] + for _ in range(count): + ready = packs / f"{uuid.uuid4()}.1.1.{digest.hexdigest()}.dmi-pack.ready" + with open(ready, "wb") as sparse: + sparse.truncate(size) + backlog.append(ready) + return sorted(backlog) + + +def test_stop_cuts_an_adoption_listing_a_dead_backlog(fake_s3, tmp_path): + """An adoption lists a dead spool once, validating every pack, which + over a backlog takes seconds: 32 sparse packs of 256 MiB, here. That + listing stops between packs once stop() cancels the uploads, as the + service's own listing does, so stop() does not wait for the rest of + it; the directory keeps every pack, and nothing of it was uploaded.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + sibling.release() # its owner is gone + backlog = _sparse_backlog(dead, 32, 256 << 20) + # Were the listing not cut, no pack may reach the fake store's + # memory: every PUT is held back, and stop() cuts those. + store.stall_requests(lambda request: request.startswith(b"PUT ")) + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + # The loop's first cycle, which start() kicks, begins the + # listing; the whole of it takes several seconds. + _wait_for(lambda: _store().spool_owner(str(dead)) is not None, + 10.0) + time.sleep(0.3) + stopping = time.monotonic() + service.stop() + elapsed = time.monotonic() - stopping + assert elapsed < 2.5, elapsed + snapshot = service.snapshot() + assert snapshot["adopted_packs"] == 0, snapshot + assert store.stalled == [], store.stalled + assert sorted(dead.rglob("*.dmi-pack.ready")) == backlog + assert _store().spool_owner(str(dead)) is None + finally: + service.stop() + store.close() + lock.release_and_remove_if_empty() + + def test_a_latched_service_lets_go_of_the_sibling_it_was_adopting( fake_s3, tmp_path): """A service whose catalog another publisher keeps for 2 x TTL latches, From f8ae2d9e409a53c59dab467ea294160b14e9f974 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 14:16:13 -0400 Subject: [PATCH 32/43] Let a flush wait for the adoption step in flight, not the rest of a slice flush() never adopts, and B5 has it return on time; but it takes the cycle lock, and a loop cycle adopting held that for adoption_slice_ns (a second by default) plus a round. Against a slow store a flush of a process with nothing of its own timed out behind a dead backlog: with a second a PUT and a one-minute slice, flush(4.0) returned False. flush() now counts itself in flushes_in_progress_ for its whole call, and a cycle adopting lets go of the cycle at its next step while one is -- after the round, or the sibling's listing, it is in. Each cycle still takes one step at least, so flushes that keep coming slow adoption down but never stop it. --- native/csrc/catalog/storage_service.cpp | 20 ++++++++ native/csrc/catalog/storage_service.h | 13 +++-- tests/test_native_spool_adoption_live.py | 60 +++++++++++++++++++++++- 3 files changed, 89 insertions(+), 4 deletions(-) diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index b7890e2df..c3d8b69f0 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -417,6 +417,15 @@ void CaptureStorageService::stop_lease_thread() { } bool CaptureStorageService::flush(double timeout_s) { + // The loop's adoption gives way to this call at its next step, so the + // cycle lock is not held for the rest of an adoption slice. + struct InProgress { + explicit InProgress(std::atomic* count) : count_(count) { + count_->fetch_add(1, std::memory_order_acq_rel); + } + ~InProgress() { count_->fetch_sub(1, std::memory_order_acq_rel); } + std::atomic* count_; + } in_progress(&flushes_in_progress_); const auto deadline = std::chrono::steady_clock::now() + std::chrono::duration_cast( std::chrono::duration(timeout_s)); @@ -819,6 +828,10 @@ bool CaptureStorageService::adopt_step(uint64_t deadline_ns, size_t* deferred, if (!scan_siblings()) return false; } bool ok = true; + // Whether this call has taken a step yet: a lock taken and a listing, a + // round, or a finish. Each call takes one at least, so flushes that keep + // coming slow adoption down but never stop it. + bool stepped = false; while (!stop_requested()) { // The service's uploads' Cancellation is adoption's too: stop() cuts // it between rounds, and in a round (its listing, its uploads, and, @@ -827,6 +840,13 @@ bool CaptureStorageService::adopt_step(uint64_t deadline_ns, size_t* deferred, *cut_short = true; break; } + // A flush() is waiting for the cycle: the rest of the slice is the + // next cycle's, which the loop runs once the flush is done. + if (stepped && + flushes_in_progress_.load(std::memory_order_acquire) > 0) { + break; + } + stepped = true; if (adopting_ == nullptr) { if (adoption_queue_.empty()) break; const std::string next = adoption_queue_.front(); diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index 55993b6f8..c83ecd05c 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -143,7 +143,9 @@ struct StorageServiceConfig { // path, at most uploader.max_workers and indexer.max_packs packs, each // indexed before the next is uploaded -- under the cycle's upload rules, // the lease and nothing owed to the catalog checked before every round, - // until adoption_slice_ns has passed or a stop is requested. Its + // until adoption_slice_ns has passed, a stop is requested, or a flush() + // is running -- a flush waits for the step the adoption is in (one round, + // or one sibling's listing), not for the rest of a slice. Its // listing, uploads and index reads go through the service's own clients // and Cancellations, so stop() cuts an adoption as it cuts the service's // own work: what it did not upload stays in the dead spool, which is let @@ -177,8 +179,10 @@ struct StorageServiceConfig { // per pass. 0 never looks again after the first pass. uint64_t adoption_recheck_interval_ns = 30'000'000'000ull; // How long one cycle may spend adopting before it lets go of the cycle -- - // to a flush(), the service's own uploads, stop() -- and carries on in - // the next. Checked between rounds, so a cycle can outrun it by one. + // to the service's own uploads, stop() -- and carries on in the next. + // Checked between rounds, so a cycle can outrun it by one. A flush() + // does not wait for it: a cycle that has made one step of adoption lets + // go of the cycle at the next one while a flush is running. uint64_t adoption_slice_ns = 1'000'000'000ull; dmi_store::S3Config s3; @@ -574,6 +578,9 @@ class CaptureStorageService { std::set blocked_siblings_; bool live_siblings_ = false; uint64_t last_adoption_scan_ns_ = 0; + // flush() calls in progress. A cycle adopting lets go of the cycle at + // its next step while one is, so a flush never waits out a slice. + std::atomic flushes_in_progress_{0}; int failure_streak_ = 0; // consecutive failed cycles, for the backoff // steady ns at which the last cycle -- the loop's or a flush's -- ended; // 0 before the first. The loop waits its interval from it. Guarded by diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index f48726606..dcbb4e983 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -18,7 +18,8 @@ the loop adopts it once the store is back. A sibling whose owner is still alive at start and dies later is adopted by a later pass. stop() cuts an adoption as it cuts the service's own work -- its uploads, and its listing -of a dead backlog -- and leaves the dead directory with every pack in it. +of a dead backlog -- and leaves the dead directory with every pack in it; +a flush waits for the adoption step in flight, not for a whole slice. Needs ClickHouse on 127.0.0.1:8123/9000 and the native sink and store modules: make -C native build/_dmi_native_sink build/_dmi_native_store @@ -801,6 +802,63 @@ def test_stop_cuts_an_adoption_listing_a_dead_backlog(fake_s3, tmp_path): lock.release_and_remove_if_empty() +def test_a_flush_waits_for_the_adoption_step_in_flight_not_its_slice( + fake_s3, tmp_path): + """flush() never adopts, and it does not wait out the loop's adoption + either: a cycle adopting lets go of the cycle at its next step while a + flush is running. Here each adopted PUT takes a second and the slice is + a minute, so a cycle would hold the cycle for the whole dead backlog, + five rounds; a flush(4.0) of a process with nothing of its own used to + time out behind it. It now returns drained after the round in flight, + and the adoption goes on afterwards.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + backlog = range(100, 140) # 20 packs: five rounds of four + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, backlog) + sibling.release() # its owner is gone + assert len(sorted(dead.rglob("*.dmi-pack.ready"))) == 20 + store.delay_requests( + lambda request: 1.0 if request.startswith(b"PUT ") else 0.0) + lock = _claim(base, config) + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="flush-yield-test", poll_interval_ns=50_000_000, + sweep_spool_on_start=True, reconcile_on_start=False, + spool_owner_lock="held_by_caller", adopt_sibling_spools=True, + adoption_slice_ns=60_000_000_000, **config._lease_native()) + service = _store().StorageService(native) + service.start() + try: + _wait_for(lambda: service.snapshot()["adopted_packs"] >= 4, 30.0) + flushing = time.monotonic() + assert service.flush(4.0) + elapsed = time.monotonic() - flushing + assert elapsed < 2.5, elapsed + snapshot = service.snapshot() + assert snapshot["adopted_packs"] < 20, snapshot + assert snapshot["adoption_owed"] is True, snapshot + + store.restore() + _wait_for(_adopted(service, 1), 60.0) + assert not dead.exists() + assert service.flush(60.0) + expected = { + capture_id: tensor.contiguous().view(-1).numpy().tobytes() + for capture_id, tensor in _envelope(backlog).expected.items()} + assert _read_all(config) == expected + finally: + service.stop() + store.close() + lock.release_and_remove_if_empty() + + def test_a_latched_service_lets_go_of_the_sibling_it_was_adopting( fake_s3, tmp_path): """A service whose catalog another publisher keeps for 2 x TTL latches, From 87d3a92004e6ddce04d8e65f1af7acabcb5c0324 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 14:19:32 -0400 Subject: [PATCH 33/43] Keep the spool lock until the release backstop has sealed the sink B6 lets go of the engine's spool directory once the sink and the service are done, and judged the sink by close()'s flush before the ring stops. B5 added a second flush after it: stopping the ring drains what it still queued into the sink, and the release backstop (on_engine_release) stages that, bounded by its own timeout. A backstop that timed out left a stage on its way into a directory whose lock close() then gave up -- another process's adoption could sweep it -- and a backstop that sealed a sink whose first flush ran out of budget kept the directory owned until exit. NativePackSink::sealed_on_release() says whether the release's flush went through; a released sink admits nothing, so then no stage is left to come. It is false while attached, after a release flush that failed or timed out, and with the backstop off. close() and a record ring's replacement judge the sink by it once the ring has stopped; a sink that does not report it (not a native pack sink) is still judged by the flush before the stop. --- native/csrc/sink/bindings_sink.cpp | 4 ++ native/csrc/sink/native_pack_sink.cpp | 6 ++- native/csrc/sink/native_pack_sink.h | 20 ++++++- src/dmi/engine.py | 37 +++++++++++-- tests/test_native_capture_storage_wiring.py | 59 +++++++++++++++++++++ tests/test_native_sink_release.py | 45 ++++++++++++++++ 6 files changed, 164 insertions(+), 7 deletions(-) diff --git a/native/csrc/sink/bindings_sink.cpp b/native/csrc/sink/bindings_sink.cpp index fdb758156..bb10b213e 100644 --- a/native/csrc/sink/bindings_sink.cpp +++ b/native/csrc/sink/bindings_sink.cpp @@ -299,6 +299,10 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { [](const dmi_sink::NativePackSink& self) { return std::string(OverloadName(self.sink().config().overload)); }) + // Released by its engine, and the release backstop's flush went + // through: nothing the sink holds can still reach the spool. + .def_property_readonly("sealed_on_release", + &dmi_sink::NativePackSink::sealed_on_release) .def_property_readonly( "release_flush_timeout_s", [](const dmi_sink::NativePackSink& self) { diff --git a/native/csrc/sink/native_pack_sink.cpp b/native/csrc/sink/native_pack_sink.cpp index 5f8d9bcbf..afcc79cb1 100644 --- a/native/csrc/sink/native_pack_sink.cpp +++ b/native/csrc/sink/native_pack_sink.cpp @@ -258,12 +258,16 @@ bool NativePackSink::flush_and_wait(Duration timeout) { } void NativePackSink::on_engine_release() noexcept { + sealed_on_release_.store(false, std::memory_order_release); if (release_flush_timeout_ == Duration::zero()) return; try { std::string error; const bool flushed = sink_->Flush( std::chrono::duration(release_flush_timeout_).count(), &error); - if (flushed) return; + if (flushed) { + sealed_on_release_.store(true, std::memory_order_release); + return; + } // Once released, rethrow_if_failed refuses the sink as not attached, // and a timeout latches nothing: this line is what says what became of // the open pack (a failure also counts in the snapshot's failures). diff --git a/native/csrc/sink/native_pack_sink.h b/native/csrc/sink/native_pack_sink.h index 97b2dfeaa..10b2d19bb 100644 --- a/native/csrc/sink/native_pack_sink.h +++ b/native/csrc/sink/native_pack_sink.h @@ -10,6 +10,7 @@ #ifndef DMI_SINK_NATIVE_PACK_SINK_H_ #define DMI_SINK_NATIVE_PACK_SINK_H_ +#include #include #include #include @@ -66,9 +67,22 @@ class NativePackSink final : public ring::RecordSink { PackSink& sink_for_testing() { return *sink_; } const std::string& layout() const { return layout_; } Duration release_flush_timeout() const { return release_flush_timeout_; } + // Whether the sink can no longer write its spool: its engine released it + // and the release backstop's flush went through. A released sink admits + // nothing, so once that flush has persisted everything admitted, no stage + // is left to come -- not from its stagers, its linger, or its destructor. + // False while attached (and from a new engine's acquire on), after a + // release whose flush failed or timed out, and when the backstop is off + // (release_flush_timeout zero): a stage may then still be on its way. + // The engine lets go of the spool directory's owner lock only on true. + bool sealed_on_release() const { + return sealed_on_release_.load(std::memory_order_acquire); + } protected: - void on_engine_acquire() override {} + void on_engine_acquire() override { + sealed_on_release_.store(false, std::memory_order_release); + } // The backstop for a ring that stops without a flush: RingEngine::stop // drains its record worker into submit() and then releases the sink, and // until a flush seals it the open pack is only in memory -- for up to @@ -80,7 +94,8 @@ class NativePackSink final : public ring::RecordSink { // times out writes one line to stderr, and that line is the only report // of a timeout. A pipeline failure also counts in snapshot()["failures"]; // rethrow_if_failed reports it only once the sink is attached again, - // since a released sink refuses the call as not attached. + // since a released sink refuses the call as not attached. Whether the + // flush went through is sealed_on_release(). void on_engine_release() noexcept override; private: @@ -91,6 +106,7 @@ class NativePackSink final : public ring::RecordSink { std::unique_ptr sink_; const std::string layout_; const Duration release_flush_timeout_; + std::atomic sealed_on_release_{false}; // Counters at construction. Losses are judged against it, as the // reference adapter judges them against its pipeline's baseline. SinkSnapshot baseline_; diff --git a/src/dmi/engine.py b/src/dmi/engine.py index f28e59052..6c487e8bc 100644 --- a/src/dmi/engine.py +++ b/src/dmi/engine.py @@ -837,6 +837,7 @@ def enable_ring_transport( drain_deadline = None if storage is None else ( time.monotonic() + self._capture_storage_config.close_flush_timeout_s) + old_record_sink = self._record_sink sealed = storage is not None and self._seal_capture_sink( drain_deadline) if old_record_mode: @@ -850,6 +851,9 @@ def enable_ring_transport( # A record sink remains leased while its worker may still # call it. Preserve the transport so shutdown can retry. raise + if storage is not None: + sealed = self._sink_sealed_after_release(old_record_sink, + sealed) try: _rt.deactivate() except Exception: @@ -918,6 +922,25 @@ def _seal_capture_sink(self, deadline: float) -> bool: return False return True + @staticmethod + def _sink_sealed_after_release(sink: Any, sealed_before_stop: bool) -> bool: + """Whether nothing can still stage into the spool, the ring stopped. + + Stopping the ring released the sink, and the native pack sink's + release backstop flushed it then; ``sealed_on_release`` says whether + that went through, and a released sink admits nothing more. It + decides either way: it seals a sink whose flush before the stop ran + out of close()'s budget, and a sink it did not get through -- one + wedged since that flush, holding what the stopping ring drained into + it, or with the backstop off -- may still stage, whatever that flush + said. A sink that does not say is judged by the flush before the + stop. + """ + released = getattr(sink, "sealed_on_release", None) + if released is None: + return sealed_before_stop + return bool(released) + def _retire_capture_storage(self, storage: Any, deadline: float, *, sink_sealed: bool) -> None: """Drain the storage service until ``deadline``, then stop it.""" @@ -938,8 +961,9 @@ def _retire_capture_storage(self, storage: Any, deadline: float, *, storage.stop() finally: # Last: the ring is stopped by now, and the service has - # stopped touching the directory. The sink is let go of only - # if it sealed. + # stopped touching the directory. The directory is let go of + # only if the sink can stage no more + # (_sink_sealed_after_release). if sink_sealed: self._release_spool_claim() else: @@ -979,7 +1003,9 @@ def close(self) -> None: ``NativeCaptureStorageConfig.close_flush_timeout_s`` says by how much. What does not drain in time is logged, not raised, and stays where the next start recovers it; ``flush_and_wait`` is the call - that raises. + that raises. The spool directory's owner lock goes last, and only + once the sink can stage no more -- its release backstop went + through; otherwise this process keeps the directory until it exits. """ storage = self._capture_storage @@ -988,10 +1014,11 @@ def close(self) -> None: drain_deadline = None if storage is None else ( time.monotonic() + self._capture_storage_config.close_flush_timeout_s) # Whether nothing can still stage into the spool: no record sink to - # seal, or one whose seal went through. + # seal, or one sealed by its flush or by its release from the ring. sealed = True if self._ring_transport is not None: record_mode = self._record_mode + record_sink = self._record_sink stopped = False # Best-effort reset of the device-global native null flag. This is # needed only after callers explicitly disabled capture; the normal @@ -1016,6 +1043,8 @@ def close(self) -> None: # alive. Leave the state intact so close can be retried. if record_mode and not stopped: return + if record_mode and storage is not None: + sealed = self._sink_sealed_after_release(record_sink, sealed) try: _rt = _ring_module() _rt.deactivate() diff --git a/tests/test_native_capture_storage_wiring.py b/tests/test_native_capture_storage_wiring.py index 91f7ce1c1..e119bc6a9 100644 --- a/tests/test_native_capture_storage_wiring.py +++ b/tests/test_native_capture_storage_wiring.py @@ -834,6 +834,65 @@ def test_close_still_stops_when_the_sink_flush_fails(monkeypatch, tmp_path, assert "stays owned" in caplog.text +def test_close_keeps_the_lock_when_the_release_backstop_did_not_seal( + monkeypatch, tmp_path, caplog): + """The sink's flush went through, but the ring stopping drains what it + still queued into the sink, and the release backstop that stages it + timed out (sealed_on_release false): a stage is still on its way into + the directory, so the lock stays, as for a sink that never sealed.""" + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + engine._record_sink.sealed_on_release = False + (lock,) = engine._test_locks + events.clear() + + with caplog.at_level("WARNING", logger="dmi.engine"): + engine.close() + + assert [event[:2] for event in events] == [ + ("sink", "flush"), ("ring", "stop"), + ("service", "flush"), ("service", "stop")] + assert lock.held + assert "stays owned" in caplog.text + + +def test_close_releases_the_lock_once_the_release_backstop_sealed_the_sink( + monkeypatch, tmp_path): + """The sink's flush ran out of close()'s budget, but its release from + the stopping ring flushed it (sealed_on_release true), and a released + sink admits nothing more: nothing can stage into the directory, so the + lock goes, as for a sink whose flush went through.""" + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + _fail_the_sink_flush(engine, events) + engine._record_sink.sealed_on_release = True + (lock,) = engine._test_locks + events.clear() + + engine.close() + + assert [event[:2] for event in events] == [ + ("sink", "flush"), ("ring", "stop"), + ("service", "flush"), ("service", "stop"), ("lock", "release")] + assert not lock.held + + +def test_replacing_a_record_ring_asks_the_released_sink_whether_it_sealed( + monkeypatch, tmp_path): + engine, events, _services = _capture_engine(monkeypatch, tmp_path) + engine.create_record_runtime(_record_format()) + engine._record_sink.sealed_on_release = False + (lock,) = engine._test_locks + events.clear() + + engine.enable_ring_transport(object()) + + assert [event[:2] for event in events] == [ + ("sink", "flush"), ("ring", "stop"), + ("service", "flush"), ("service", "stop"), ("ring", "create")] + assert lock.held + + def test_replacing_a_record_ring_keeps_the_lock_when_the_sink_did_not_seal( monkeypatch, tmp_path): engine, events, _services = _capture_engine(monkeypatch, tmp_path) diff --git a/tests/test_native_sink_release.py b/tests/test_native_sink_release.py index 84ac52fee..a6a9bc8de 100644 --- a/tests/test_native_sink_release.py +++ b/tests/test_native_sink_release.py @@ -237,6 +237,51 @@ def test_a_zero_release_timeout_leaves_the_open_pack_to_the_pipeline( del lease +def test_sealed_on_release_says_whether_the_sink_can_still_stage( + native_sink_module, tmp_path): + """The engine lets go of its spool directory's owner lock only once + nothing can stage into the directory any more. sealed_on_release says + so: true once a release's flush has persisted everything admitted -- + a released sink admits nothing -- and false while the sink is + attached, after a release whose flush timed out (a stage is still on + its way, and lands later), and with the backstop off.""" + sink = _binding_sink(native_sink_module, tmp_path / "sealed") + assert sink.sealed_on_release is False + lease = sink.attach() + sink.submit_envelope(LAYOUT, *_envelope(0)) + _wait_until_admitted(sink, 1) + assert sink.sealed_on_release is False # attached: it may take more + del lease + assert sink.sealed_on_release is True + assert len(_ready(tmp_path / "sealed")) == 1 + lease = sink.attach() + assert sink.sealed_on_release is False # attached again + del lease + assert sink.sealed_on_release is True # nothing left to flush + + wedged = _binding_sink(native_sink_module, tmp_path / "wedged", + release_flush_timeout_s=0.5) + release_stages = wedged._hold_stages_for_testing(10.0) + lease = wedged.attach() + wedged.submit_envelope(LAYOUT, *_envelope(1)) + _wait_until_admitted(wedged, 1) + del lease + assert wedged.sealed_on_release is False + assert _ready(tmp_path / "wedged") == [] + release_stages() # the stage the release gave up on still lands + deadline = time.monotonic() + 10.0 + while not _ready(tmp_path / "wedged"): + assert time.monotonic() < deadline, wedged.snapshot() + time.sleep(0.01) + assert wedged.sealed_on_release is False + + off = _binding_sink(native_sink_module, tmp_path / "off", + release_flush_timeout_s=0.0) + lease = off.attach() + del lease + assert off.sealed_on_release is False + + @pytest.mark.parametrize("timeout_s", [-1.0, float("nan"), float("inf")]) def test_the_release_timeout_must_be_finite_and_not_negative( native_sink_module, tmp_path, timeout_s): From 1616682b2064fe190e9d1c10f007c19c73b36fb0 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 14:20:14 -0400 Subject: [PATCH 34/43] Say when close() keeps the spool directory, and what a flush waits for of adoption The integration guide judged the sink by close()'s own flush and said a flush does not wait for adoption. Since the merge the directory stays owned only when the release backstop did not seal the sink, a flush waits for the adoption step in flight and no more, and close() cuts an adoption as it cuts the service's own uploads. --- docs/integration-api-v1.md | 15 ++++++++++----- 1 file changed, 10 insertions(+), 5 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index f61a9536b..bcad684e2 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -197,10 +197,11 @@ directory itself, taken before the service starts and let go after the sink and the service are done, when a drained directory is removed. The directory's own lock keeps it owned should its `.owner.lock` be removed from under it, and systemd-tmpfiles skips a flocked directory when it ages `/tmp`; other cleaners -may not, so keep `spool_root` out of what they age. If the sink did not seal within -`close_flush_timeout_s`, it may still be staging, so the directory stays owned -by the process until it exits (a warning names it) and the next process on the -node adopts it; so does the directory of an engine dropped without `close()`. +may not, so keep `spool_root` out of what they age. If the sink did not seal -- +the release backstop's flush, when the stopping ring lets go of it, did not go +through -- it may still be staging, so the directory stays owned by the +process until it exits (a warning names it) and the next process on the node +adopts it; so does the directory of an engine dropped without `close()`. A second process on a directory is refused, naming the holder's pid and host. Once started, the service's background loop adopts the directories under the same catalog key whose owners @@ -209,7 +210,11 @@ indexed, a round at a time and a slice of each cycle, and the directory removed, so a crashed process's packs reach the catalog through the next one on the node, whatever run it belongs to. Neither `create_record_runtime` nor `flush_and_wait` waits for that: a flush covers this process's records (an -adopted pack uploaded and not yet indexed is waited for like its own), and the +adopted pack uploaded and not yet indexed is waited for like its own), waits +for no more of an adoption in progress than the step it is in (one round, or +one directory's listing), and `close()` cuts an adoption as it cuts the +service's own uploads, leaving what it did not upload in the dead directory +for the next process. The storage part of `capture_status()` reports the adoption (`adopted_spools`, `adopted_packs`, `adoption_owed`, `live_siblings`). A dead directory the service can never adopt -- one holding a pack it can never upload, such as one From 9ac51d4828cd40985bcd9123322a4b3c3f18532e Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 14:38:48 -0400 Subject: [PATCH 35/43] Hold the harness's spool lock for B5's live services as for B6's B5's live tests build their services from a raw native dict, which opens the spool with owner_lock take. Under the owner lock that service keeps the directory until it is destroyed, stop() or not, so the successor each test then starts on the same spool (_service, held_by_caller beside the harness's lock) was refused: five tests failed with SpoolOwnedError. They now open it held_by_caller under the harness's lock, as B6 did for every raw service already in the file. --- tests/test_native_capture_storage_live.py | 54 +++++++++++++++-------- 1 file changed, 36 insertions(+), 18 deletions(-) diff --git a/tests/test_native_capture_storage_live.py b/tests/test_native_capture_storage_live.py index dc2e61588..5f67da1c3 100644 --- a/tests/test_native_capture_storage_live.py +++ b/tests/test_native_capture_storage_live.py @@ -897,7 +897,8 @@ def test_a_flush_against_a_black_hole_catalog_overruns_by_one_request( fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="black-hole-test", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="black-hole-test", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, # Every cycle is due a reconcile. @@ -942,7 +943,8 @@ def test_the_stop_after_a_flush_runs_no_further_cycle(fake_s3, tmp_path): fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="stop-after-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stop-after-flush", # Wakes while the flush below holds the cycle lock. poll_interval_ns=3_000_000_000, reconcile_on_start=False, # No renewal falls due while the test runs (a third of the TTL). @@ -1006,7 +1008,8 @@ def test_a_flush_out_of_time_does_not_hash_the_spool(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="flush-out-of-time", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="flush-out-of-time", # The loop sleeps through the test, and nothing uploads. poll_interval_ns=60_000_000_000, reconcile_on_start=False, sweep_spool_on_start=False, @@ -1065,7 +1068,8 @@ def test_stop_cuts_a_listing_that_hashes_a_backlog(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="listing-stop", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="listing-stop", poll_interval_ns=20_000_000, reconcile_on_start=False, sweep_spool_on_start=False, uploader_max_in_flight_bytes=1 << 30) @@ -1101,7 +1105,8 @@ def test_a_flush_cuts_the_listing_its_deadline_passes_in(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="listing-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="listing-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, sweep_spool_on_start=False, @@ -1139,7 +1144,8 @@ def test_a_flush_returns_on_time_while_an_upload_stalls(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-upload-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-upload-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, s3_read_timeout_s=30) @@ -1212,7 +1218,8 @@ def test_a_flush_that_cuts_a_multipart_upload_waits_for_its_abort( with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="multipart-abort-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="multipart-abort-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, clickhouse_request_timeout_s=2.0, @@ -1250,7 +1257,8 @@ def test_stop_returns_promptly_while_an_upload_stalls(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-upload-stop", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-upload-stop", poll_interval_ns=20_000_000, reconcile_on_start=False, s3_read_timeout_s=60) service = _load_native_store_extension().StorageService(native) @@ -1314,7 +1322,8 @@ def test_a_flush_returns_on_time_while_the_index_reads_stall( with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-read-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-read-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, s3_read_timeout_s=3, @@ -1373,7 +1382,8 @@ def test_stop_returns_promptly_while_the_index_reads_stall(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-read-stop", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-read-stop", poll_interval_ns=20_000_000, reconcile_on_start=False, s3_read_timeout_s=3) service = _load_native_store_extension().StorageService(native) @@ -1452,7 +1462,8 @@ def _slow_guard(request: bytes) -> float: s3.url, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="stop-no-replay-guard", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stop-no-replay-guard", poll_interval_ns=20_000_000, reconcile_on_start=False, uploader_max_workers=1, s3_read_timeout_s=30) # Staged first, so the loop's first cycle lists all four. @@ -1518,7 +1529,8 @@ def test_an_object_store_read_outage_sets_no_pack_aside(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="read-outage", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="read-outage", poll_interval_ns=20_000_000, max_backoff_ns=200_000_000, reconcile_on_start=False, s3_read_timeout_s=1, s3_max_attempts=1, max_index_attempts=2) @@ -1575,7 +1587,8 @@ def test_a_flush_against_a_slow_catalog_indexes_one_batch_past_its_deadline( fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="slow-catalog-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="slow-catalog-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, indexer_max_packs=1, @@ -1639,7 +1652,8 @@ def test_a_close_whose_budget_ends_mid_upload_leaves_nothing_owed( with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="close-mid-upload", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="close-mid-upload", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, uploader_max_workers=1, indexer_max_packs=4) @@ -1692,7 +1706,8 @@ def test_the_one_batch_past_a_flushs_deadline_is_a_full_one(fake_s3, fake_s3, catalog.table_prefix, clickhouse_port=switch.port)._native_dict() native.update( - spool_root=str(spool_root), holder="full-batch-flush", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="full-batch-flush", # The loop sleeps through the test, so the flush runs the cycle. poll_interval_ns=60_000_000_000, reconcile_on_start=False, indexer_max_packs=2, clickhouse_request_timeout_s=60.0) @@ -1731,7 +1746,8 @@ def test_only_the_loop_reconciles_never_a_flush(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="flush-no-reconcile", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="flush-no-reconcile", poll_interval_ns=2_000_000_000, reconcile_on_start=False, reconcile_interval_ns=1_000_000) service = _load_native_store_extension().StorageService(native) @@ -1761,7 +1777,8 @@ def test_a_service_started_again_after_stop_uploads_again(fake_s3, tmp_path): spool_root = tmp_path / "spool" with _catalog() as (_client, catalog): native = _storage_config(fake_s3, catalog.table_prefix)._native_dict() - native.update(spool_root=str(spool_root), holder="restarted", + native.update(spool_root=str(spool_root), + spool_owner_lock=_held(spool_root), holder="restarted", reconcile_on_start=False) service = _load_native_store_extension().StorageService(native) service.start() @@ -1793,7 +1810,8 @@ def test_stop_cuts_a_reconcile_whose_listing_stalls(fake_s3, tmp_path): with _catalog() as (_client, catalog): native = _storage_config(s3.url, catalog.table_prefix)._native_dict() native.update( - spool_root=str(spool_root), holder="stalled-reconcile-stop", + spool_root=str(spool_root), spool_owner_lock=_held(spool_root), + holder="stalled-reconcile-stop", poll_interval_ns=20_000_000, reconcile_on_start=False, reconcile_interval_ns=1_000_000, s3_read_timeout_s=30) service = _load_native_store_extension().StorageService(native) From 74bcac5b22404d6ebe4212c1855e6534e8762f25 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 14:58:31 -0400 Subject: [PATCH 36/43] Pin on a real ring that close() lets go of the engine's spool directory The wiring tests fake the sink, so only here does a real ring release a real NativePackSink whose backstop seals it: close() without a flush then drains the tail into the catalog and, the sink sealed on release, lets go of the engine's rank directory and removes it, leaving nothing to adopt. --- tests/test_native_capture_storage_gpu_e2e.py | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/tests/test_native_capture_storage_gpu_e2e.py b/tests/test_native_capture_storage_gpu_e2e.py index b317ea184..7a74a9dc0 100644 --- a/tests/test_native_capture_storage_gpu_e2e.py +++ b/tests/test_native_capture_storage_gpu_e2e.py @@ -237,7 +237,8 @@ def test_close_alone_delivers_the_tail_to_the_catalog(fake_s3, tmp_path): """No flush_and_wait: close() must still seal the sink's open pack and drain it into the catalog. The 60 s linger means nothing but a flush can seal it, and 3 records against max_pack_records=2 leave one record in the - open pack when close() runs.""" + open pack when close() runs. The sink sealed, close() lets go of the + engine's spool directory and removes it.""" from dmi.api.v1 import HookPointV1, HookSpecV1, MonitoringEngine, TransportSpec from dmi.config import MonitoringConfig from dmi.storage.capture import CaptureRecordFormat @@ -283,6 +284,11 @@ def test_close_alone_delivers_the_tail_to_the_catalog(fake_s3, tmp_path): engine.close() # no flush_and_wait + # The ring's release sealed the sink (sealed_on_release), so close() + # let go of the engine's own spool directory, and removed it once + # drained: no rank directory is left for anyone to adopt. + assert list((tmp_path / "spool").glob("*/r*-*")) == [] + reader = NativeCaptureReader(storage) selection = reader.select(tenant_id="tenant-gpu") captures = {capture.descriptor["capture_id"]: capture From 9d5364f7077bcd04de4fbf488697bd1c22f1beb7 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 18:44:43 -0400 Subject: [PATCH 37/43] Wait for a forked worker to have run before killing its lock's owner A child forked without exec closes its copy of the spool lock only in the fork handler, which runs once the child is first scheduled. The two fork tests that killed the owner without hearing from the worker then read the lock at once, and on a loaded host the worker had not run yet: the lock outlived its owner for tens of milliseconds and the check failed (the C++ case every time at load 60-80 on 32 cores). The worker now says it has run, past fork() and its handler, before the owner is killed, as the sibling case already did. --- tests/native/test_spool_owner_lock.cpp | 22 ++++++++++++++++------ tests/test_native_spool_ownership.py | 24 ++++++++++++++++++------ 2 files changed, 34 insertions(+), 12 deletions(-) diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 57f1238c2..129bf3840 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -393,7 +393,10 @@ void TestAForkedChildDoesNotKeepTheLockPastItsParent() { // with the fork handler only once the lock was held, that copy was never // closed in the child, which kept the directory looking live after its // owner died. Here the fork runs in that window, through the lock-open -// test seam. +// test seam. The worker says it has run before its owner is killed: the +// fork handler closes its copy only once the child is first scheduled, +// which on a loaded host can come tens of milliseconds after the owner +// is reaped, and until then the lock outlives its owner. void TestAForkWhileALockIsTakenLeavesTheChildNothing() { const std::string root = FreshRoot("fork-window") + "/spool"; fs::create_directories(root); @@ -407,6 +410,9 @@ void TestAForkWhileALockIsTakenLeavesTheChildNothing() { if (worker >= 0) return; worker = ::fork(); // no exec if (worker == 0) { + // Past fork(), so past the fork handler. + const char ran = 'w'; + if (::write(ready[1], &ran, 1) != 1) ::_exit(3); ::pause(); // outlives its parent until killed ::_exit(0); } @@ -424,16 +430,20 @@ void TestAForkWhileALockIsTakenLeavesTheChildNothing() { ::_exit(0); } ::close(ready[1]); + // The worker's 'w' and the owner's "k\n", in any order. std::string seen; char byte = 0; - while (seen.find('\n') == std::string::npos && - ::read(ready[0], &byte, 1) == 1) { + while (seen.find('\n') == std::string::npos || + seen.find('w') == std::string::npos) { + if (::read(ready[0], &byte, 1) != 1) break; seen.push_back(byte); } ::close(ready[0]); - CHECK(!seen.empty() && seen[0] == 'k'); - const pid_t worker = - static_cast(std::atoi(seen.empty() ? "" : seen.c_str() + 1)); + CHECK(seen.find('w') != std::string::npos); // the worker has run + const size_t outcome = seen.find_first_of("kx"); + CHECK(outcome != std::string::npos && seen[outcome] == 'k'); + const pid_t worker = static_cast(std::atoi( + outcome == std::string::npos ? "" : seen.c_str() + outcome + 1)); CHECK(worker > 0); dmi_store::SpoolOwner record; CHECK(dmi_store::ReadSpoolOwner(root, &record)); diff --git a/tests/test_native_spool_ownership.py b/tests/test_native_spool_ownership.py index 7127184f0..937244512 100644 --- a/tests/test_native_spool_ownership.py +++ b/tests/test_native_spool_ownership.py @@ -174,7 +174,12 @@ def test_a_forked_worker_does_not_keep_a_dead_owners_lock(tmp_path): multiprocessing worker -- shares the lock's open file description. It used to keep the lock after its parent was SIGKILLed, so the dead parent's directory read as owned, naming the dead pid, and was never - adopted.""" + adopted. + + The worker says it has run before its owner is killed: the fork + handler closes its copy of the lock only once the child is first + scheduled, which on a loaded host can come tens of milliseconds after + the owner is reaped, and until then the lock outlives its owner.""" directory = tmp_path / "spool" script = ( "import os, sys, time; sys.path.insert(0, sys.argv[1]);" @@ -182,6 +187,7 @@ def test_a_forked_worker_does_not_keep_a_dead_owners_lock(tmp_path): "lock = m.SpoolOwnerLock(sys.argv[2])\n" "worker = os.fork()\n" "if worker == 0:\n" + " print('worker ran', flush=True)\n" # past fork(), and its handler " time.sleep(3600)\n" " os._exit(0)\n" "print(worker, flush=True)\n" @@ -189,8 +195,13 @@ def test_a_forked_worker_does_not_keep_a_dead_owners_lock(tmp_path): owner = subprocess.Popen( [sys.executable, "-c", script, str(BUILD), str(directory)], stdout=subprocess.PIPE, text=True) - worker = int(owner.stdout.readline()) + worker = None try: + # The worker's line and the owner's, in either order. + lines = {owner.stdout.readline().strip(), + owner.stdout.readline().strip()} + assert "worker ran" in lines, lines + (worker,) = [int(line) for line in lines if line != "worker ran"] assert _store().spool_owner(str(directory))["pid"] == owner.pid os.kill(owner.pid, signal.SIGKILL) owner.wait(timeout=30) @@ -199,10 +210,11 @@ def test_a_forked_worker_does_not_keep_a_dead_owners_lock(tmp_path): with _store().SpoolOwnerLock(str(directory)): pass finally: - try: - os.kill(worker, signal.SIGKILL) - except ProcessLookupError: - pass + if worker is not None: + try: + os.kill(worker, signal.SIGKILL) + except ProcessLookupError: + pass if owner.poll() is None: owner.kill() owner.wait(timeout=30) From 1a10c1eb7365f4bc63ce8895afdebda14adebf03 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 18:44:43 -0400 Subject: [PATCH 38/43] Forward no request its client did not finish, and store no body short of its length The switch forwarded a request whose client went away mid-body -- a flush deadline's cancel landing between libcurl's head and its body, after the switch had answered 100 Continue itself -- and the fake S3, whose signature check takes the payload hash from x-amz-content-sha256, stored the short body. A PUT held back by a delay could so land after the successor had uploaded the same pack, and overwrite it with 0 bytes: test_a_close_whose_budget_ends_mid_upload_leaves_nothing_owed then failed reading the pack's trailer. The switch now drops a request it did not receive whole, and one a delay still held back when close() ran; the fake S3, like S3, stores no body short of its Content-Length or not the one x-amz-content-sha256 hashes. --- tests/test_native_capture_storage_live.py | 96 +++++++++++++++++++- tests/test_native_s3_client.py | 101 +++++++++++++++++++++- 2 files changed, 194 insertions(+), 3 deletions(-) diff --git a/tests/test_native_capture_storage_live.py b/tests/test_native_capture_storage_live.py index 5f67da1c3..f0c3fff83 100644 --- a/tests/test_native_capture_storage_live.py +++ b/tests/test_native_capture_storage_live.py @@ -356,6 +356,9 @@ def __init__(self, host: str, port: int): self.refused: list[float] = [] self._lock = threading.Lock() self._sockets: set[socket.socket] = set() + # close() has run: a request held back by a delay is not forwarded + # once it has. + self._closed = False threading.Thread(target=self._accept, daemon=True).start() @classmethod @@ -421,6 +424,7 @@ def _look_then_route(self, client): what delay_requests() says; anything else goes straight through.""" request = b"" continued = False + whole = False # the head and all Content-Length bytes of body try: client.settimeout(2.0) while True: @@ -428,6 +432,7 @@ def _look_then_route(self, client): if found: length = re.search(rb"(?i)content-length:\s*(\d+)", head) if length is None or len(body) >= int(length.group(1)): + whole = True break # libcurl holds a body over 1 KiB back until the server # says 100 Continue, or for a second; answer for it, so @@ -444,6 +449,12 @@ def _look_then_route(self, client): except OSError: client.close() return + if not whole: + # The client went away mid-request -- a cancel that landed + # after the head, say. Forwarded, the fake S3 stored the short + # body under the key, over what a later upload put there. + client.close() + return if continued: # The server must not answer 100 Continue a second time. head, _, body = request.partition(b"\r\n\r\n") @@ -467,6 +478,10 @@ def _look_then_route(self, client): and b"_publisher_lease" in request) if late: time.sleep(late_by) + if self._closed: + # Held back past close(): the test is done with this server. + client.close() + return try: upstream = socket.create_connection(self._target) upstream.sendall(request) @@ -536,7 +551,9 @@ def close(self): listener does not wake a thread blocked in accept(), which still takes one more queued connection: the stall settings go first, so that connection is refused rather than held open for its client's - whole timeout (a stop()'s lease release after a stall() did).""" + whole timeout (a stop()'s lease release after a stall() did). + A request a delay still holds back is dropped, not forwarded.""" + self._closed = True self._stalled = False self._stall_if = None self._slow_once = None @@ -546,6 +563,83 @@ def close(self): self._listener.close() +class _Recorder: + """A TCP server that keeps every byte it is sent, by connection.""" + + def __init__(self): + self._listener = socket.create_server(("127.0.0.1", 0)) + self.port = self._listener.getsockname()[1] + self.received: list[bytes] = [] + self._lock = threading.Lock() + threading.Thread(target=self._accept, daemon=True).start() + + def _accept(self): + while True: + try: + connection, _ = self._listener.accept() + except OSError: + return + threading.Thread(target=self._keep, args=(connection,), + daemon=True).start() + + def _keep(self, connection): + data = b"" + try: + while chunk := connection.recv(65536): + data += chunk + except OSError: + pass + connection.close() + with self._lock: + self.received.append(data) + + def close(self): + self._listener.close() + + +def test_the_switch_forwards_no_request_its_client_did_not_finish(tmp_path): + """The harness itself. A client that went away after a request's head + -- a flush deadline's cancel landing between libcurl's head and body, + which the switch had answered 100 Continue for -- was forwarded all + the same, its body short: the fake S3, whose signature check trusts + x-amz-content-sha256, then stored an empty object over what a later + upload put at the key. Nor does a request a delay still holds back + when close() runs reach the server after it.""" + recorder = _Recorder() + switch = _Switch("127.0.0.1", recorder.port) + switch.delay_requests(lambda request: 0.3) + head = (b"PUT /b/k HTTP/1.1\r\nHost: x\r\nContent-Length: 2048\r\n" + b"Expect: 100-continue\r\n\r\n") + try: + with socket.create_connection(("127.0.0.1", switch.port), + timeout=10) as client: + client.sendall(head) + assert client.recv(64).startswith(b"HTTP/1.1 100 Continue") + # The control: a whole request, held back and then forwarded. + with socket.create_connection(("127.0.0.1", switch.port), + timeout=10) as client: + client.sendall(head + bytes(2048)) + client.shutdown(socket.SHUT_WR) + time.sleep(1.0) + with recorder._lock: + forwarded = list(recorder.received) + assert len(forwarded) == 1, forwarded + assert forwarded[0].endswith(bytes(2048)), forwarded + + # Held back when close() runs, then dropped. + with socket.create_connection(("127.0.0.1", switch.port), + timeout=10) as client: + client.sendall(head + bytes(2048)) + time.sleep(0.1) + switch.close() + time.sleep(1.0) + with recorder._lock: + assert recorder.received == forwarded, recorder.received + finally: + switch.close() + recorder.close() + + def test_an_index_failure_keeps_the_pack_owed_until_it_lands(fake_s3, tmp_path): spool_root = tmp_path / "spool" switch = _Switch(CLICKHOUSE_HOST, CLICKHOUSE_HTTP_PORT) diff --git a/tests/test_native_s3_client.py b/tests/test_native_s3_client.py index 52178e84c..5d50ec0a3 100644 --- a/tests/test_native_s3_client.py +++ b/tests/test_native_s3_client.py @@ -141,9 +141,22 @@ def _reject_unsigned(self, body: bytes) -> bool: return True return False - def _read_body(self) -> bytes: + def _read_body(self): + """The request's body, or None when it is not the one signed for: + shorter than its Content-Length (the client went away mid-body), or + not what x-amz-content-sha256 hashes. S3 stores neither + (IncompleteBody, XAmzContentSHA256Mismatch), and a body short of + its length passed the signature check here, which trusts that + header -- so a PUT cut off mid-body stored an empty object.""" length = int(self.headers.get("Content-Length", 0)) - return self.rfile.read(length) if length else b"" + body = self.rfile.read(length) if length else b"" + if len(body) < length: + return None + claimed = self.headers.get("x-amz-content-sha256", "") + if (len(claimed) == 64 + and hashlib.sha256(body).hexdigest() != claimed.lower()): + return None + return body def _split(self): from urllib.parse import unquote @@ -216,6 +229,15 @@ def _route(self): return key = segments[1] body = self._read_body() + if body is None: + # Nobody may be listening; the connection goes, and nothing is + # recorded or stored. + self.close_connection = True + try: + self._send(400, {}, b"incomplete or mismatched body") + except OSError: + pass + return self._record(body) if self._reject_unsigned(body): return @@ -987,3 +1009,78 @@ def test_an_uncancelled_request_is_untouched_by_the_cancel_hook(fake_s3): assert put["ok"], put assert put.get("cancelled") is False, put assert STATE.objects["packs/on-time"]["body"] == b"data" + + +def _signed_put_head(endpoint: str, key: str, body_hash: str, + length: int) -> bytes: + """A SigV4-signed PUT's head as the native client sends one, for a body + of `length` bytes that x-amz-content-sha256 says hashes to `body_hash`.""" + import datetime as datetime_module + + host = urlsplit(endpoint).netloc + credentials = botocore.credentials.Credentials(ACCESS, SECRET, None) + request = botocore.awsrequest.AWSRequest( + method="PUT", url=f"http://{host}/{BUCKET}/{key}", data=b"", + headers={"x-amz-content-sha256": body_hash}) + frozen = datetime_module.datetime.now(datetime_module.timezone.utc).replace( + microsecond=0, tzinfo=None) + with mock.patch.object( + botocore.auth, "get_current_datetime", return_value=frozen + ): + botocore.auth.SigV4Auth(credentials, "s3", REGION).add_auth(request) + head = f"PUT /{BUCKET}/{key} HTTP/1.1\r\nHost: {host}\r\n" + for name in ("x-amz-content-sha256", "X-Amz-Date", "Authorization"): + head += f"{name}: {request.headers[name]}\r\n" + head += f"Content-Length: {length}\r\n\r\n" + return head.encode() + + +def _send_raw(endpoint: str, request: bytes) -> bytes: + """Send `request`, say no more (shutdown for writing), and return the + answer, if any.""" + import socket + + parts = urlsplit(endpoint) + with socket.create_connection((parts.hostname, parts.port), + timeout=10) as connection: + connection.sendall(request) + connection.shutdown(socket.SHUT_WR) + answer = b"" + try: + while chunk := connection.recv(65536): + answer += chunk + except OSError: + pass + return answer + + +def test_the_fake_stores_no_put_that_is_not_the_body_it_signed_for(fake_s3): + """The harness itself. Its signature check takes the payload hash from + x-amz-content-sha256, as S3 does, so a PUT whose client went away after + the head -- a cancel landing between libcurl's head and its body -- + passed it with a short body and was stored, empty, over what a later + upload put at the key. S3 stores neither a body short of its + Content-Length nor one its x-amz-content-sha256 does not hash.""" + body = bytes(range(256)) * 8 + digest = hashlib.sha256(body).hexdigest() + + # The control: the same head with its whole body is stored. + answer = _send_raw(fake_s3, _signed_put_head( + fake_s3, "packs/whole", digest, len(body)) + body) + assert answer.startswith(b"HTTP/1.1 200"), answer + assert STATE.objects["packs/whole"]["body"] == body + + # The head, and half the body, then nothing more. + _send_raw(fake_s3, _signed_put_head( + fake_s3, "packs/short", digest, len(body)) + body[:1024]) + # The head alone. + _send_raw(fake_s3, _signed_put_head( + fake_s3, "packs/headless", digest, len(body))) + # A whole body, but not the one the head hashes. + answer = _send_raw(fake_s3, _signed_put_head( + fake_s3, "packs/other", digest, len(body)) + bytes(len(body))) + assert answer.startswith(b"HTTP/1.1 400"), answer + with STATE.lock: + assert set(STATE.objects) == {"packs/whole"}, sorted(STATE.objects) + assert [call["path"] for call in STATE.calls] == [ + f"/{BUCKET}/packs/whole"], STATE.calls From 92295ca226e935fa2e7d2a38a831150667d466e2 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 18:50:33 -0400 Subject: [PATCH 39/43] Give two cut tests the time a loaded runner needs before the cut The multipart-abort flush had 1 s to list and hash its 65 MiB pack, send the preflight HEAD, hash it again and create the upload before its part was stalled; at load 70-80 on 32 cores the deadline came first, cut the upload before the part and left nothing to abort, on main as here. Its budget is 4 s now, the bound moved with it. The listing-cut uploader test ended within one 256 MiB pack's hash of its cancel, which on such a runner took most of its 2 s bound; its backlog is 128 packs of 64 MiB, so one pack is a small part of the bound and the whole listing still far over it. --- tests/test_native_capture_storage_live.py | 12 +++++++++--- tests/test_native_uploader.py | 17 ++++++++++------- 2 files changed, 19 insertions(+), 10 deletions(-) diff --git a/tests/test_native_capture_storage_live.py b/tests/test_native_capture_storage_live.py index f0c3fff83..dc619f6b6 100644 --- a/tests/test_native_capture_storage_live.py +++ b/tests/test_native_capture_storage_live.py @@ -1292,12 +1292,18 @@ def test_a_flush_that_cuts_a_multipart_upload_waits_for_its_abort( outlasts the request timeout. Here the store holds the part and the abort both, so the flush pays the whole bound, and no more. The pack is a sparse file of zeros named for its checksum, over the client's - 64 MiB multipart threshold; it stays staged.""" + 64 MiB multipart threshold; it stays staged. The budget has to see the + part stalled before it ends: the listing and the upload each hash the + pack first, and a HEAD and the CreateMultipartUpload go before the + part, which on a loaded runner outlasted a 1 s budget -- the deadline + then cut the upload before its part was sent, and nothing was left to + abort.""" import hashlib from dmi.storage.native_capture import _load_native_store_extension size = 65 << 20 + budget = 4.0 digest = hashlib.sha256() zeros = bytes(1 << 20) for _ in range(size // len(zeros)): @@ -1323,7 +1329,7 @@ def test_a_flush_that_cuts_a_multipart_upload_waits_for_its_abort( try: s3.stall_requests(_multipart_part_or_abort) started = time.monotonic() - drained = service.flush(1.0) + drained = service.flush(budget) elapsed = time.monotonic() - started snapshot = service.snapshot() finally: @@ -1333,7 +1339,7 @@ def test_a_flush_that_cuts_a_multipart_upload_waits_for_its_abort( assert drained is False assert len(s3.stalled) == 2, s3.stalled # the part, then the abort # The deadline, up to a second for the cut, and the 5 s abort. - assert elapsed < 1.0 + 1.0 + 5.0 + 1.0, (elapsed, snapshot) + assert elapsed < budget + 1.0 + 5.0 + 1.0, (elapsed, snapshot) assert snapshot["cancelled_uploads"] == 1, snapshot assert snapshot["upload_failures"] == 0, snapshot assert _ready(spool_root) == [ready] diff --git a/tests/test_native_uploader.py b/tests/test_native_uploader.py index c3f5c67b1..acb82ac9f 100644 --- a/tests/test_native_uploader.py +++ b/tests/test_native_uploader.py @@ -381,12 +381,15 @@ def test_a_cancel_stops_the_listing_between_packs(fake_s3, tmp_path): """UploadPending lists the spool before it uploads anything, and a listing re-hashes every staged pack: over a backlog, seconds a GiB of it, all before a single worker looked at the cancel. The listing now - stops between packs once cancelled, and nothing is tried. 16 sparse - packs of 256 MiB, zeros named for their checksum, take it about 3 s.""" + stops between packs once cancelled, and nothing is tried. 128 sparse + packs of 64 MiB, zeros named for their checksum, take it about 6 s + idle; one of them takes a small part of the bound below even on a + loaded runner, where 256 MiB packs took it over, the whole listing + slowing with them.""" import hashlib import uuid - size = 256 << 20 + size = 64 << 20 digest = hashlib.sha256() zeros = bytes(1 << 20) for _ in range(size // len(zeros)): @@ -394,7 +397,7 @@ def test_a_cancel_stops_the_listing_between_packs(fake_s3, tmp_path): root = tmp_path / "spool" root.mkdir() backlog = [] - for _ in range(16): + for _ in range(128): ready = root / f"{uuid.uuid4()}.1.1.{digest.hexdigest()}.dmi-pack.ready" with open(ready, "wb") as sparse: sparse.truncate(size) @@ -411,9 +414,9 @@ def test_a_cancel_stops_the_listing_between_packs(fake_s3, tmp_path): assert result["refs"] == [] and result["failures"] == [], result assert result["snapshot"]["cancelled_packs"] == 0, result # Between packs, not after the listing: the cut ends within one pack's - # hash of the cancel, about 0.5 s here, where the whole listing takes - # over 3 s. The bound leaves room for a loaded runner, whose hashing - # slows with it (about 1.1 s at 4x oversubscription). + # hash of the cancel, about 0.1 s here, where the whole listing takes + # about 6 s. The bound leaves room for a loaded runner, whose hashing + # slows with it (a pack about 0.6 s at 2.5x oversubscription). assert elapsed < 2.0, elapsed assert sorted(root.rglob("*.dmi-pack.ready")) == sorted(backlog) assert STATE.calls == [] From 45c5c88f1f66be81b74111e8c413a520bf7f2543 Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 19:21:43 -0400 Subject: [PATCH 40/43] Allow the stalled-upload stop tests eight attempts, so an uncut backoff shows stop() cancels the upload client and the uploader alike. With only the client's cancel an attempt in flight is still aborted and every later one refused at once, but the uploader's own backoff sleeps on: at the default four attempts that is about 1.75 s, which fit under both tests' bounds, so neither the adoption's uploader nor the service's losing its Cancellation turned anything red. At eight attempts it is about 26 s; each mutation now fails its test (27 s, and past the 15 s join). --- tests/test_native_capture_storage_live.py | 7 +++++-- tests/test_native_spool_adoption_live.py | 14 ++++++++++++-- 2 files changed, 17 insertions(+), 4 deletions(-) diff --git a/tests/test_native_capture_storage_live.py b/tests/test_native_capture_storage_live.py index dc619f6b6..9457919e8 100644 --- a/tests/test_native_capture_storage_live.py +++ b/tests/test_native_capture_storage_live.py @@ -1349,7 +1349,10 @@ def test_stop_returns_promptly_while_an_upload_stalls(fake_s3, tmp_path): """stop() joined a loop whose cycle was inside a PUT the store never answers, so it waited out the S3 read timeout on every attempt, with the lease held. It cancels the upload now: the pack stays in the spool, the - lease is released, and the next process uploads the pack.""" + lease is released, and the next process uploads the pack. The uploader + is allowed eight attempts: one whose own retries and backoff stop() did + not cut, on a client stop() did, would sleep out about 26 s of backoff + (four attempts' 1.75 s fit under the bound).""" from dmi.storage.native_capture import _load_native_store_extension spool_root = tmp_path / "spool" @@ -1360,7 +1363,7 @@ def test_stop_returns_promptly_while_an_upload_stalls(fake_s3, tmp_path): spool_root=str(spool_root), spool_owner_lock=_held(spool_root), holder="stalled-upload-stop", poll_interval_ns=20_000_000, reconcile_on_start=False, - s3_read_timeout_s=60) + s3_read_timeout_s=60, uploader_max_attempts=8) service = _load_native_store_extension().StorageService(native) service.start() stopper = None diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index dcbb4e983..199cdec58 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -684,7 +684,10 @@ def test_stop_cuts_an_adoption_stalled_on_its_uploads(fake_s3, tmp_path): and stop() sat through read timeouts. Now stop() returns at once, every pack of the dead spool is still in it (none was uploaded, none lost), and the directory is left, unlocked, for the next incarnation -- which - adopts it.""" + adopts it. The uploader is allowed eight attempts: an adoption uploader + whose own retries and backoff stop() did not cut, on a client stop() + did, would sleep out about 26 s of backoff (four attempts' 1.75 s fit + under the bound).""" from tests.test_native_capture_storage_live import _Switch base = tmp_path / "spool" @@ -700,7 +703,14 @@ def test_stop_cuts_an_adoption_stalled_on_its_uploads(fake_s3, tmp_path): assert len(staged) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK store.stall_requests(lambda request: request.startswith(b"PUT ")) lock = _claim(base, config) - service = _service(config, lock.directory) + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="adoption-stall-stop", poll_interval_ns=50_000_000, + sweep_spool_on_start=True, reconcile_on_start=False, + spool_owner_lock="held_by_caller", adopt_sibling_spools=True, + uploader_max_attempts=8, **config._lease_native()) + service = _store().StorageService(native) service.start() try: _wait_for(lambda: store.stalled, 30.0) # an adopted PUT in flight From 908121198e1933938d56859801973e83d1e040ba Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 19:40:39 -0400 Subject: [PATCH 41/43] Validate an adopted dead spool a pack a step, so a flush waits for one pack An adoption listed a dead spool in one step: Spool::Recover hashed every byte of every ready pack under the cycle lock, and only stop() cut it. A flush, which yields to the adoption only between steps, waited out the whole listing -- as long as the backlog is big, up to the dead spool's 1 TiB budget -- and timed out with its own pack still staged; close() then reported capture storage undrained and left the tail pack for the next process. With 512 sparse packs of 64 MiB, a flush(8.0) of one pack staged during the listing timed out; it now drains in about a second. Spool::BeginRecovery sweeps a dead spool's .open files and lists its ready packs without hashing them, and each ContinueRecovery validates the next, quarantining one that fails; the last rebuilds the account as Recover() does. Adoption validates one pack per step, keeps the listing across cycles, and checks the slice between packs as between rounds, so a flush waits for at most one pack's hash (or one round), the service's own uploads for at most a slice, and stop() still cuts between packs. The cancellable Recover overload, whose only caller this was, is gone. The flush docstring, close_flush_timeout_s and the integration guide say what a flush waits for of an adoption. --- docs/integration-api-v1.md | 9 +- native/csrc/catalog/storage_service.cpp | 77 +++++++------ native/csrc/catalog/storage_service.h | 69 +++++++----- native/csrc/store/spool.cpp | 133 +++++++++++++++-------- native/csrc/store/spool.h | 40 +++++-- src/dmi/storage/native_capture.py | 13 ++- tests/native/test_spool_owner_lock.cpp | 86 ++++++++++----- tests/test_native_spool_adoption_live.py | 69 +++++++++++- 8 files changed, 346 insertions(+), 150 deletions(-) diff --git a/docs/integration-api-v1.md b/docs/integration-api-v1.md index bcad684e2..b952dae58 100644 --- a/docs/integration-api-v1.md +++ b/docs/integration-api-v1.md @@ -211,10 +211,11 @@ removed, so a crashed process's packs reach the catalog through the next one on the node, whatever run it belongs to. Neither `create_record_runtime` nor `flush_and_wait` waits for that: a flush covers this process's records (an adopted pack uploaded and not yet indexed is waited for like its own), waits -for no more of an adoption in progress than the step it is in (one round, or -one directory's listing), and `close()` cuts an adoption as it cuts the -service's own uploads, leaving what it did not upload in the dead directory -for the next process. The +for no more of an adoption in progress than the step it is in (one round of +its uploads, or one of its packs validated -- a dead directory's packs are +hashed one per step, so a large backlog's listing does not hold a flush up), +and `close()` cuts an adoption as it cuts the service's own uploads, leaving +what it did not upload in the dead directory for the next process. The storage part of `capture_status()` reports the adoption (`adopted_spools`, `adopted_packs`, `adoption_owed`, `live_siblings`). A dead directory the service can never adopt -- one holding a pack it can never upload, such as one diff --git a/native/csrc/catalog/storage_service.cpp b/native/csrc/catalog/storage_service.cpp index c3d8b69f0..16c747e88 100644 --- a/native/csrc/catalog/storage_service.cpp +++ b/native/csrc/catalog/storage_service.cpp @@ -5,6 +5,7 @@ #include #include #include +#include #include #include #include @@ -155,13 +156,17 @@ class CaptureStorageService::LeaseScope { }; // The dead sibling a cycle is adopting, kept across cycles: its owner lock, -// a Spool on it (held_by_caller, the lock being this service's), and the -// packs its Recover() swept and validated that are not uploaded yet -- so a -// backlog is hashed once, however many cycles its upload takes. +// a Spool on it (held_by_caller, the lock being this service's), its +// recovery -- the packs listed and those validated so far, a pack a step +// (Spool::BeginRecovery) -- and, once every pack is validated, those not +// uploaded yet. So a backlog is hashed once, however many cycles its +// listing and its upload take. struct CaptureStorageService::Adoption { std::string directory; dmi_store::SpoolOwnerLock lock; dmi_store::Spool spool; + dmi_store::SpoolRecovery recovery; + bool validated = false; // every listed pack; `remaining` holds the valid std::deque remaining; // The first pack this service can never upload (UploadFailure:: // retryable false); the directory is left, with it, once the rest are up. @@ -828,14 +833,16 @@ bool CaptureStorageService::adopt_step(uint64_t deadline_ns, size_t* deferred, if (!scan_siblings()) return false; } bool ok = true; - // Whether this call has taken a step yet: a lock taken and a listing, a - // round, or a finish. Each call takes one at least, so flushes that keep - // coming slow adoption down but never stop it. + // Whether this call has taken a step yet: a lock taken and a sweep, one + // pack of a listing validated, a round, or a finish. Each call takes one + // at least, so flushes that keep coming slow adoption down but never stop + // it. bool stepped = false; while (!stop_requested()) { // The service's uploads' Cancellation is adoption's too: stop() cuts - // it between rounds, and in a round (its listing, its uploads, and, - // through the reads' Cancellation, its index reads). + // it between steps -- between the packs a listing validates too -- and + // in a round (its uploads, and, through the reads' Cancellation, its + // index reads). if (upload_cancel_.cancelled()) { *cut_short = true; break; @@ -851,14 +858,26 @@ bool CaptureStorageService::adopt_step(uint64_t deadline_ns, size_t* deferred, if (adoption_queue_.empty()) break; const std::string next = adoption_queue_.front(); adoption_queue_.pop_front(); - bool cut = false; - if (!begin_adoption(next, &cut)) ok = false; - if (cut) { - // Its listing was cut: it is looked at again, from the start. - adoption_queue_.push_front(next); - *cut_short = true; - break; + begin_adoption(next); + continue; + } + if (!adopting_->validated) { + // One pack of the listing: validating hashes every byte of it, so + // over a dead backlog a whole listing takes as long as the backlog + // is big, and neither a flush nor stop() waits for more than the + // pack in flight. The slice bounds a listing as it does the rounds. + Adoption& adoption = *adopting_; + if (adoption.spool.ContinueRecovery(&adoption.recovery)) { + // Each pack's identity and object key come from the pack and its + // path in the dead directory, exactly as its owner would have + // uploaded it. + adoption.remaining.assign( + std::make_move_iterator(adoption.recovery.valid.begin()), + std::make_move_iterator(adoption.recovery.valid.end())); + adoption.recovery = dmi_store::SpoolRecovery{}; + adoption.validated = true; } + if (steady_ns() - started >= config_.adoption_slice_ns) break; continue; } if (adopting_->remaining.empty()) { @@ -939,9 +958,7 @@ bool CaptureStorageService::scan_siblings() { return true; } -bool CaptureStorageService::begin_adoption(const std::string& directory, - bool* cut) { - *cut = false; +void CaptureStorageService::begin_adoption(const std::string& directory) { auto adoption = std::make_unique(); adoption->directory = directory; std::string error; @@ -952,36 +969,28 @@ bool CaptureStorageService::begin_adoption(const std::string& directory, live_siblings_ = true; std::lock_guard state(state_mutex_); ++state_.live_siblings; - return true; + return; } if (locked != dmi_store::SpoolStatus::kOk) { // Another adopter drained and removed it meanwhile: nothing is owed. - if (!std::filesystem::exists(directory)) return true; + if (!std::filesystem::exists(directory)) return; // Its lock cannot be taken at all -- another user's lock file, say. block_sibling(directory, "cannot lock it: " + error); - return true; + return; } dmi_store::SpoolConfig config{directory, config_.spool_max_bytes}; config.owner_lock = dmi_store::OwnerLock::kHeldByCaller; // adoption->lock config.allow_shared_filesystem = config_.spool_allow_shared_filesystem; - std::vector ready; + // Its .open files swept and its ready packs listed, none hashed yet: + // adopt_step validates them, a pack a step. if (dmi_store::Spool::Open(config, &adoption->spool, &error) != dmi_store::SpoolStatus::kOk || - adoption->spool.Recover(&ready, &error, &upload_cancel_, cut) != + adoption->spool.BeginRecovery(&adoption->recovery, &error) != dmi_store::SpoolStatus::kOk) { block_sibling(directory, "cannot open it: " + error, &adoption->lock); - return true; - } - // Validating hashes every pack of a dead backlog; the uploads' cancel - // (stop()) stops it between packs, as it does the service's own listing. - // The directory is let go of whole, its lock with it: nothing of it was - // uploaded, and nothing is removed. - if (*cut) return true; - // Each pack's identity and object key come from the pack and its path in - // the dead directory, exactly as its owner would have uploaded it. - adoption->remaining.assign(ready.begin(), ready.end()); + return; + } adopting_ = std::move(adoption); - return true; } bool CaptureStorageService::upload_adopted_round(uint64_t deadline_ns, diff --git a/native/csrc/catalog/storage_service.h b/native/csrc/catalog/storage_service.h index c83ecd05c..c1b95a28d 100644 --- a/native/csrc/catalog/storage_service.h +++ b/native/csrc/catalog/storage_service.h @@ -137,24 +137,27 @@ struct StorageServiceConfig { // off once it holds the lease and has swept its own directory; start() // itself uploads nothing of a dead backlog, and neither does flush(). A // cycle probes the owner lock of every sibling rank directory, then works - // through those whose owner is gone: it takes one's lock, sweeps its - // .open files and validates its .ready packs once, and uploads and - // indexes them a round at a time -- a chunk of the service's own upload - // path, at most uploader.max_workers and indexer.max_packs packs, each - // indexed before the next is uploaded -- under the cycle's upload rules, - // the lease and nothing owed to the catalog checked before every round, - // until adoption_slice_ns has passed, a stop is requested, or a flush() - // is running -- a flush waits for the step the adoption is in (one round, - // or one sibling's listing), not for the rest of a slice. Its - // listing, uploads and index reads go through the service's own clients - // and Cancellations, so stop() cuts an adoption as it cuts the service's - // own work: what it did not upload stays in the dead spool, which is let - // go of and never removed while a pack of it is left, for the next - // process on the node. The sibling's lock and its remaining packs are kept - // between cycles, so a large backlog is hashed once, not on every cycle, - // and neither a flush() nor the service's own uploads wait behind all of - // it. A drained sibling is removed once nothing but its lock file is - // left. An upload that failed stays in the dead spool and is retried by + // through those whose owner is gone: it takes one's lock and sweeps its + // .open files, validates its .ready packs once, a pack a step -- which + // hashes every byte of it -- and uploads and indexes them a round at a + // time -- a chunk of the service's own upload path, at most + // uploader.max_workers and indexer.max_packs packs, each indexed before + // the next is uploaded -- under the cycle's upload rules, the lease and + // nothing owed to the catalog checked before every round, until + // adoption_slice_ns has passed, a stop is requested, or a flush() is + // running -- a flush waits for the step the adoption is in (one round, + // one pack's validation, or one sibling's lock and sweep), not for the + // rest of a slice or of a listing. Its uploads and index reads go through + // the service's own clients and Cancellations, and its listing stops + // between packs on the uploads' cancel, so stop() cuts an adoption as it + // cuts the service's own work: what it did not upload stays in the dead + // spool, which is let go of and never removed while a pack of it is + // left, for the next process on the node. The sibling's lock, its + // listing and its remaining packs are kept between cycles, so a large + // backlog is hashed once, not on every cycle, and neither a flush() nor + // the service's own uploads wait behind all of it. A drained sibling is + // removed once nothing but its lock file is left. An upload that failed + // stays in the dead spool and is retried by // a later cycle, after the backoff -- unless no retry by this service can // ever succeed (dmi_store::UploadFailure::retryable: a pack over // uploader.max_in_flight_bytes, a different object at its key, bytes that @@ -180,9 +183,10 @@ struct StorageServiceConfig { uint64_t adoption_recheck_interval_ns = 30'000'000'000ull; // How long one cycle may spend adopting before it lets go of the cycle -- // to the service's own uploads, stop() -- and carries on in the next. - // Checked between rounds, so a cycle can outrun it by one. A flush() - // does not wait for it: a cycle that has made one step of adoption lets - // go of the cycle at the next one while a flush is running. + // Checked between rounds and between the packs a listing validates, so a + // cycle can outrun it by one round, or one pack's hash. A flush() does + // not wait for it: a cycle that has made one step of adoption lets go of + // the cycle at the next one while a flush is running. uint64_t adoption_slice_ns = 1'000'000'000ull; dmi_store::S3Config s3; @@ -348,7 +352,10 @@ class CaptureStorageService { // when that is under about 6 s. At zero it still runs one cycle, which // indexes one batch of what earlier cycles owe and uploads nothing. // Its cycles adopt nothing, and a dead sibling still to adopt does not - // keep it from returning true (adopt_sibling_spools). + // keep it from returning true (adopt_sibling_spools); a loop cycle that + // is adopting when it is called lets go of the cycle after the step it + // is in -- one round of a dead spool's uploads, one of its packs + // validated, or one sibling's lock and sweep. // Throws, once, if packs were set aside since the last flush: they are in // the object store but can never reach the catalog. bool flush(double timeout_s); @@ -418,9 +425,11 @@ class CaptureStorageService { // thread renewing it. Requires cycle_mutex_. void sweep_and_reconcile_at_start(); // One cycle's share of adoption (see adopt_sibling_spools): looks at the - // siblings when that is due, then adopts until the slice ends, a stop is - // requested or the uploads' Cancellation is cancelled -- which also cuts - // a round in flight, and *cut_short then says so. Adds what its index + // siblings when that is due, then adopts a step at a time -- a sibling + // taken and swept, one of its packs validated, a round, a finish -- until + // the slice ends, a flush() is waiting (after one step at least), a stop + // is requested or the uploads' Cancellation is cancelled -- which also + // cuts a round in flight, and *cut_short then says so. Adds what its index // passes left owed through a cancel to *deferred (index_bounded). False // when an upload failed, so the cycle backs off. Requires cycle_mutex_. // Only a lost lease propagates. @@ -428,11 +437,11 @@ class CaptureStorageService { // Probes every sibling rank directory's owner lock and queues the dead // ones. False when the directory cannot be listed. Requires cycle_mutex_. bool scan_siblings(); - // Takes a queued sibling's lock, sweeps it and lists its packs into - // adopting_; leaves adopting_ empty for a live or vanished one, and for - // one whose listing the uploads' Cancellation cut (*cut), whose lock it - // lets go of. False when it could not be locked or opened. - bool begin_adoption(const std::string& directory, bool* cut); + // Takes a queued sibling's lock, sweeps its .open files and lists its + // ready packs into adopting_, hashing none of them (adopt_step validates + // them, a pack a step); leaves adopting_ empty for a live or vanished + // one, and for one it cannot lock or open, which it blocks. + void begin_adoption(const std::string& directory); // Uploads and indexes one round of adopting_'s packs, one chunk of the // service's upload path; *cut when a cancel left some in the dead spool, // back at the front of adopting_. False when an upload failed. diff --git a/native/csrc/store/spool.cpp b/native/csrc/store/spool.cpp index b606b174c..ac392c013 100644 --- a/native/csrc/store/spool.cpp +++ b/native/csrc/store/spool.cpp @@ -1757,9 +1757,27 @@ SpoolStatus Spool::Recover(std::vector* out, std::string* error) { return Scan(out, true, error); } -SpoolStatus Spool::Recover(std::vector* out, std::string* error, - const Cancellation* cancel, bool* cut) { - return Scan(out, true, error, cancel, cut); +SpoolStatus Spool::BeginRecovery(SpoolRecovery* recovery, std::string* error) { + *recovery = SpoolRecovery{}; + std::lock_guard lock(mutex_); + uint64_t open_bytes = 0; // stays 0: every .open file is swept + ListReadyLocked(true, &recovery->listed, &open_bytes); + (void)error; + return SpoolStatus::kOk; +} + +bool Spool::ContinueRecovery(SpoolRecovery* recovery) { + std::lock_guard lock(mutex_); + if (recovery->next < recovery->listed.size()) { + StagedPack staged; + if (ValidateReadyLocked(recovery->listed[recovery->next], &staged)) { + recovery->valid.push_back(std::move(staged)); + } + ++recovery->next; + } + if (recovery->next < recovery->listed.size()) return false; + CommitListingLocked(recovery->valid, 0); + return true; } SpoolStatus Spool::ListPending(std::vector* out, std::string* error) { @@ -1777,10 +1795,29 @@ SpoolStatus Spool::Scan(std::vector* out, bool discard_open_files, out->clear(); if (cut) *cut = false; std::lock_guard lock(mutex_); - std::error_code ec; std::vector readies; - uint64_t bytes = 0; - std::unordered_map seen_ready; + uint64_t open_bytes = 0; + ListReadyLocked(discard_open_files, &readies, &open_bytes); + for (const std::string& path : readies) { + if (cancel != nullptr && cancel->cancelled()) { + // Before the next pack's hash. The account below is rebuilt from a + // whole listing only; quarantines already made stand. + out->clear(); + if (cut) *cut = true; + return SpoolStatus::kOk; + } + StagedPack staged; + if (ValidateReadyLocked(path, &staged)) out->push_back(std::move(staged)); + } + CommitListingLocked(*out, open_bytes); + (void)error; + return SpoolStatus::kOk; +} + +void Spool::ListReadyLocked(bool discard_open_files, + std::vector* readies, + uint64_t* open_bytes) { + std::error_code ec; for (auto it = fs::recursive_directory_iterator(root_, ec); it != fs::recursive_directory_iterator(); ++it) { if (AtSkippedDirectory(it)) { @@ -1799,65 +1836,65 @@ SpoolStatus Spool::Scan(std::vector* out, bool discard_open_files, FsyncDir(entry.path().parent_path().string(), nullptr); } else { const uint64_t size = entry.file_size(ec); - if (!ec) bytes += size; + if (!ec) *open_bytes += size; ec.clear(); // Another writer may have just committed its temp. } continue; } if (HasSuffix(name, kReadySuffix)) { - readies.push_back(path); + readies->push_back(path); } } - std::sort(readies.begin(), readies.end()); - for (const std::string& path : readies) { - if (cancel != nullptr && cancel->cancelled()) { - // Before the next pack's hash. The account below is rebuilt from a - // whole listing only; quarantines already made stand. - out->clear(); - if (cut) *cut = true; - return SpoolStatus::kOk; - } - const std::string name = fs::path(path).filename().string(); - std::string id, sum; - uint64_t created = 0, records = 0; - const uint64_t size = fs::file_size(path, ec); - if (!ParseReadyName(name, &id, &created, &records, &sum) || ec || - Sha256HexFile(path, nullptr) != sum) { - // Quarantine: keep the bytes, drop the .ready suffix. - const std::string target = path.substr(0, path.size() - 6) + - ".quarantined"; - ::rename(path.c_str(), target.c_str()); - FsyncDir(fs::path(path).parent_path().string(), nullptr); - ++generation_; - continue; - } - const std::string rel = fs::relative(path, root_, ec).string(); - const size_t slash = rel.rfind('/'); - const std::string parent = (slash == std::string::npos) ? "" : rel.substr(0, slash); - StagedPack staged; - staged.pack_id = id; - staged.created_at_ns = created; - staged.record_count = records; - staged.checksum = sum; - staged.object_key = (parent.empty() ? "" : parent + "/") + id + ".dmi-pack"; - staged.path = path; - staged.object_bytes = size; - out->push_back(std::move(staged)); - bytes += size; - seen_ready.emplace(path, size); + std::sort(readies->begin(), readies->end()); +} + +bool Spool::ValidateReadyLocked(const std::string& path, StagedPack* out) { + std::error_code ec; + const std::string name = fs::path(path).filename().string(); + std::string id, sum; + uint64_t created = 0, records = 0; + const uint64_t size = fs::file_size(path, ec); + if (!ParseReadyName(name, &id, &created, &records, &sum) || ec || + Sha256HexFile(path, nullptr) != sum) { + // Quarantine: keep the bytes, drop the .ready suffix. + const std::string target = path.substr(0, path.size() - 6) + + ".quarantined"; + ::rename(path.c_str(), target.c_str()); + FsyncDir(fs::path(path).parent_path().string(), nullptr); + ++generation_; + return false; } + const std::string rel = fs::relative(path, root_, ec).string(); + const size_t slash = rel.rfind('/'); + const std::string parent = (slash == std::string::npos) ? "" : rel.substr(0, slash); + out->pack_id = id; + out->created_at_ns = created; + out->record_count = records; + out->checksum = sum; + out->object_key = (parent.empty() ? "" : parent + "/") + id + ".dmi-pack"; + out->path = path; + out->object_bytes = size; + return true; +} + +void Spool::CommitListingLocked(const std::vector& valid, + uint64_t open_bytes) { // Recovery rebuilds the committed account only; a stage in flight on // another thread keeps its reservation. The path ledger is rebuilt with // it (_commit_recovery_locked does the same), so the surviving entries are // exactly the ones a later retry will recognise as already counted, and // the quarantined ones are simply absent. + uint64_t bytes = open_bytes; + std::unordered_map seen_ready; + for (const StagedPack& staged : valid) { + bytes += staged.object_bytes; + seen_ready.emplace(staged.path, staged.object_bytes); + } committed_bytes_ = bytes; - committed_entries_ = out->size(); + committed_entries_ = valid.size(); accounted_ready_ = std::move(seen_ready); peak_bytes_ = std::max(peak_bytes_, committed_bytes_ + reserved_bytes_); ++generation_; - (void)error; - return SpoolStatus::kOk; } SpoolStatus Spool::Remove(const StagedPack& staged, std::string* error) { diff --git a/native/csrc/store/spool.h b/native/csrc/store/spool.h index cd1c695d3..31b4a10db 100644 --- a/native/csrc/store/spool.h +++ b/native/csrc/store/spool.h @@ -118,6 +118,14 @@ struct StagedPack { uint64_t object_bytes = 0; }; +// A recovery in steps (Spool::BeginRecovery): the ready packs it listed, +// sorted, how many of them it has validated, and those that were valid. +struct SpoolRecovery { + std::vector listed; + size_t next = 0; + std::vector valid; +}; + struct SpoolSnapshot { uint64_t entries = 0; uint64_t bytes = 0; @@ -354,13 +362,19 @@ class Spool { // lock keeps other processes out; writers in this process sharing it // (kHeldByCaller) are the caller's to order. SpoolStatus Recover(std::vector* out, std::string* error); - // The same, but its validation stops between packs once `cancel` is - // cancelled, as ListPending's does: *cut then says so and *out is empty. - // The .open files are swept by then, and the account is left as it was. - // For an adopter's listing of a dead spool, whose backlog it would - // otherwise hash whole before a stop could take effect. - SpoolStatus Recover(std::vector* out, std::string* error, - const Cancellation* cancel, bool* cut); + // Recover() a pack at a time, for an adopter's recovery of a dead spool: + // Recover() hashes every byte of a dead backlog in one call, and whatever + // waits for the adopter -- a stop(), a flush() behind its cycle -- waits + // for all of it. BeginRecovery() is Recover()'s sweep of the .open files + // and its listing of the ready packs, and hashes none of them. Each + // ContinueRecovery() then validates the next pack listed, quarantining + // one that fails as Recover() does; once none is left it rebuilds the + // account as Recover() does and returns true, recovery->valid then the + // packs Recover() would have listed. The account is left as it was + // until then. Nothing else may validate or remove this spool's packs + // meanwhile. + SpoolStatus BeginRecovery(SpoolRecovery* recovery, std::string* error); + bool ContinueRecovery(SpoolRecovery* recovery); // Validate and list ready packs without deleting in-progress writes. SpoolStatus ListPending(std::vector* out, std::string* error); @@ -390,6 +404,18 @@ class Spool { SpoolStatus Scan(std::vector* out, bool discard_open_files, std::string* error, const Cancellation* cancel = nullptr, bool* cut = nullptr); + // A scan's parts. The walk: every ready path, sorted, into *readies; + // the .open files this object is not writing deleted (discard_open_files) + // or their bytes added to *open_bytes. The validation of one listed ready + // path: true with *out filled when its name parses and its size and + // sha256 match, and otherwise it is quarantined. And the account rebuilt + // from a whole listing's valid packs. `mutex_` must be held for each. + void ListReadyLocked(bool discard_open_files, + std::vector* readies, + uint64_t* open_bytes); + bool ValidateReadyLocked(const std::string& path, StagedPack* out); + void CommitListingLocked(const std::vector& valid, + uint64_t open_bytes); // Count/uncount one ready path in the committed account, at most once each // -- Python's _account_ready_locked / _unaccount_ready_locked. `mutex_` // must be held. Both return whether they actually changed the account. diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index f2c9902dc..376e51a3d 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -313,7 +313,13 @@ class NativeCaptureStorageConfig: # budget cuts a multipart upload; against a slow catalog that still # answers, by one batch of statements and the release. When the sink # itself is stuck, the flush its release from the ring makes adds up to - # 30 s. close() logs what did not drain; flush_and_wait is what raises. + # 30 s. The budget also pays for the adoption step the background loop + # is in when the drain starts (adopt_sibling_spools, the engine's + # default): the drain's cycle waits for it -- one of a dead process's + # packs validated, which hashes it, or one round of their uploads -- + # but never for a dead backlog's whole listing, and past the budget + # stopping the service cuts it. close() logs what did not drain; + # flush_and_wait is what raises. close_flush_timeout_s: float = 60.0 # Bytes of packs the uploader holds in flight at once. A staged pack # larger than this is never uploaded, so the sink's max_pack_bytes must @@ -863,6 +869,11 @@ def flush(self, timeout_s: float) -> None: background loop's, and ``snapshot()`` reports them (``adopted_spools``, ``adoption_owed``) -- though an adopted pack uploaded and not yet indexed is waited for like this process's own. + Nor does a flush adopt, but it does wait, within ``timeout_s``, for + the adoption step the loop is in when it is called: one of a dead + spool's packs validated, which hashes it, or one round of their + uploads (at most ``uploader_max_workers`` packs) -- not a dead + backlog's whole listing, which goes a pack a step. """ if not self._service.flush(float(timeout_s)): snapshot = self._service.snapshot() diff --git a/tests/native/test_spool_owner_lock.cpp b/tests/native/test_spool_owner_lock.cpp index 129bf3840..cab47b6c8 100644 --- a/tests/native/test_spool_owner_lock.cpp +++ b/tests/native/test_spool_owner_lock.cpp @@ -18,7 +18,7 @@ // allowed (a test seam stands in for statfs). // 6. Adoption's try-lock never creates a directory, and a released // directory that holds nothing but its lock file can be removed; an -// adopter's listing of a dead spool stops between packs on a cancel. +// adopter recovers a dead spool a pack at a time. // 7. The directory layout of the plan's section 2.3. // 8. The spool budget charges what dead sibling directories hold. // @@ -35,11 +35,11 @@ #include #include #include +#include #include #include #include -#include "store/cancel.h" #include "store/spool.h" namespace fs = std::filesystem; @@ -985,16 +985,18 @@ void TestANewDirectoryAppearsWithItsLockHeld() { CHECK(unheld == 0); } -// (6e) An adopter lists a dead spool once, through Recover, which hashes -// every pack of what may be a large backlog. The storage service hands it -// the Cancellation stop() cancels, and a cancelled listing stops between -// packs: nothing listed, *cut, every ready pack where it was (nothing is -// lost, and nothing is uploaded from a cut listing), the account as it -// was. The dead owner's stale .open file is swept by then. Not cancelled, -// the same listing lists every pack. -void TestAnAdoptersListingStopsOnACancel() { +// (6e) An adopter recovers a dead spool a pack at a time. Recover() hashes +// every pack of what may be a large backlog in one call, and whatever +// waited for the adopter -- stop(), a flush() behind its cycle -- waited +// for all of it. BeginRecovery sweeps the dead owner's stale .open file and +// lists the ready packs, hashing none, the account left as it was; each +// ContinueRecovery validates the next one, quarantining a corrupt one at +// its turn, and the last rebuilds the account as Recover() does. Until +// then every other ready pack stays where it was: an adopter let go of +// midway leaves nothing lost. +void TestAnAdopterRecoversADeadSpoolAPackAtATime() { const std::string dead = - FreshRoot("adopt-cut") + "/0123456789ab/r0-0000dead"; + FreshRoot("adopt-steps") + "/0123456789ab/r0-0000dead"; std::string error; { Spool gone; @@ -1018,6 +1020,10 @@ void TestAnAdoptersListingStopsOnACancel() { }; const std::set staged_before = readies(); CHECK(staged_before.size() == 3); + // The second in listing order no longer matches its checksum. + const std::string corrupt = *std::next(staged_before.begin()); + std::ofstream(corrupt, std::ios::binary | std::ios::trunc) + << std::string(100, 'z'); SpoolOwnerLock adopter; CHECK(SpoolOwnerLock::TryAdopt(dead, &adopter, &error) == SpoolStatus::kOk); @@ -1027,22 +1033,52 @@ void TestAnAdoptersListingStopsOnACancel() { CHECK(Spool::Open(config, &adopted, &error) == SpoolStatus::kOk); const uint64_t bytes_before = adopted.Snapshot().bytes; - dmi_store::Cancellation cancel; - cancel.Cancel(); - std::vector listed; - bool cut = false; - CHECK(adopted.Recover(&listed, &error, &cancel, &cut) == SpoolStatus::kOk); - CHECK(cut); - CHECK(listed.empty()); - CHECK(readies() == staged_before); + dmi_store::SpoolRecovery recovery; + CHECK(adopted.BeginRecovery(&recovery, &error) == SpoolStatus::kOk); CHECK(!fs::exists(stale)); + CHECK((std::set(recovery.listed.begin(), + recovery.listed.end()) == staged_before)); + CHECK(recovery.next == 0 && recovery.valid.empty()); + CHECK(readies() == staged_before); CHECK(adopted.Snapshot().bytes == bytes_before); - cancel.Reset(); - CHECK(adopted.Recover(&listed, &error, &cancel, &cut) == SpoolStatus::kOk); - CHECK(!cut); - CHECK(listed.size() == 3); - CHECK(adopted.Snapshot().bytes == 300); + CHECK(!adopted.ContinueRecovery(&recovery)); // the first + CHECK(recovery.next == 1 && recovery.valid.size() == 1); + CHECK(recovery.valid[0].path == *staged_before.begin()); + CHECK(!adopted.ContinueRecovery(&recovery)); // the corrupt one + CHECK(recovery.next == 2 && recovery.valid.size() == 1); + CHECK(!fs::exists(corrupt)); + CHECK(fs::exists(corrupt.substr(0, corrupt.size() - 6) + ".quarantined")); + CHECK(readies().size() == 2); + CHECK(adopted.Snapshot().bytes == bytes_before); // not rebuilt yet + + CHECK(adopted.ContinueRecovery(&recovery)); // the last, and the account + CHECK(recovery.valid.size() == 2); + CHECK(recovery.valid[1].path == *staged_before.rbegin()); + CHECK(recovery.valid[1].object_key == + "v1/018f0000-0000-7000-8000-000000000003.dmi-pack"); + CHECK(adopted.Snapshot().bytes == 200); + CHECK(adopted.Snapshot().entries == 2); + CHECK(adopted.ContinueRecovery(&recovery)); // done stays done + CHECK(recovery.valid.size() == 2); + + // Nothing listed: done at the first step. + const std::string empty = + FreshRoot("adopt-steps-empty") + "/0123456789ab/r0-00000e00"; + { Spool gone; CHECK(Spool::Open({empty, 1 << 20}, &gone, &error) == + SpoolStatus::kOk); } + SpoolOwnerLock empty_adopter; + CHECK(SpoolOwnerLock::TryAdopt(empty, &empty_adopter, &error) == + SpoolStatus::kOk); + SpoolConfig empty_config{empty, 1 << 20}; + empty_config.owner_lock = OwnerLock::kHeldByCaller; + Spool empty_spool; + CHECK(Spool::Open(empty_config, &empty_spool, &error) == SpoolStatus::kOk); + dmi_store::SpoolRecovery nothing; + CHECK(empty_spool.BeginRecovery(¬hing, &error) == SpoolStatus::kOk); + CHECK(nothing.listed.empty()); + CHECK(empty_spool.ContinueRecovery(¬hing)); + CHECK(nothing.valid.empty()); } // (8) The budget across incarnations. Every process start gets a fresh @@ -1296,7 +1332,7 @@ int main() { TestALockOnAnUnlinkedFileIsTakenAgain(); TestAReplacedLockFileLeavesTheDirectoryOwned(); TestANewDirectoryAppearsWithItsLockHeld(); - TestAnAdoptersListingStopsOnACancel(); + TestAnAdopterRecoversADeadSpoolAPackAtATime(); TestTheDirectoryLayout(); TestDeadSiblingsCountAgainstTheBudget(); if (g_failures != 0) { diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 199cdec58..41dc714ea 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -19,7 +19,8 @@ alive at start and dies later is adopted by a later pass. stop() cuts an adoption as it cuts the service's own work -- its uploads, and its listing of a dead backlog -- and leaves the dead directory with every pack in it; -a flush waits for the adoption step in flight, not for a whole slice. +a flush waits for the adoption step in flight, not for a whole slice, nor +for a dead backlog's whole listing. Needs ClickHouse on 127.0.0.1:8123/9000 and the native sink and store modules: make -C native build/_dmi_native_sink build/_dmi_native_store @@ -812,6 +813,72 @@ def test_stop_cuts_an_adoption_listing_a_dead_backlog(fake_s3, tmp_path): lock.release_and_remove_if_empty() +def _no_tenant_in_the_path(request: bytes) -> bool: + """A PUT of a pack outside the capture layout: the sparse backlog's.""" + line = request.split(b"\r\n", 1)[0] + return line.startswith(b"PUT ") and b"tenant" not in line + + +def test_a_flush_does_not_wait_for_an_adoption_to_list_a_dead_backlog( + fake_s3, tmp_path): + """An adoption validates a dead spool's packs before it uploads any, + hashing every byte: over a backlog -- an object-store outage's, up to + the dead spool's own budget -- for as long as the backlog is big. A + flush of this process's own records waited for all of it, since the + listing was one step of the adoption, and timed out with its own pack + still staged: engine.close() then reported capture storage undrained. + The listing now validates one pack per step, and a flush waits for the + one in flight. Here the backlog is 512 sparse packs of 64 MiB, which + take their listing 20 s and more; the flush drains the pack staged + meanwhile well inside its budget, with the listing still going.""" + from tests.test_native_capture_storage_live import _Switch + + base = tmp_path / "spool" + with _catalog() as prefix: + store = _Switch.to_url(fake_s3) + config = _storage_config(store.url, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + sibling.release() # its owner is gone + backlog = _sparse_backlog(dead, 512, 64 << 20) + # Were the listing to end, its packs would go no further than the + # switch: the service's own go through. + store.stall_requests(_no_tenant_in_the_path) + lock = _claim(base, config) + service = _service(config, lock.directory) + service.start() + try: + # The loop's first cycle, which start() kicks, begins listing. + _wait_for(lambda: _store().spool_owner(str(dead)) is not None, + 10.0) + time.sleep(0.3) + _stage_into(lock.directory, range(0, 2)) # as close() would + flushing = time.monotonic() + service.flush(8.0) # raises TimeoutError when not drained + elapsed = time.monotonic() - flushing + snapshot = service.snapshot() + assert snapshot["uploaded_packs"] == 1, snapshot + assert snapshot["indexed_packs"] == 1, snapshot + assert not sorted(Path(lock.directory).rglob("*.dmi-pack.ready")) + # Drained with the listing still going: the flush did not wait + # it out. + assert snapshot["adopted_packs"] == 0, snapshot + assert snapshot["adoption_owed"] is True, snapshot + owner = _store().spool_owner(str(dead)) + assert owner is not None and owner["pid"] == os.getpid(), ( + owner, elapsed) + assert store.stalled == [], store.stalled + stopping = time.monotonic() + service.stop() + assert time.monotonic() - stopping < 5.0 + assert sorted(dead.rglob("*.dmi-pack.ready")) == backlog + assert _store().spool_owner(str(dead)) is None + finally: + service.stop() + store.close() + lock.release_and_remove_if_empty() + + def test_a_flush_waits_for_the_adoption_step_in_flight_not_its_slice( fake_s3, tmp_path): """flush() never adopts, and it does not wait out the loop's adoption From ca392bcad3cdb9a84f3dd9af57bb54a6856e84fb Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 20:02:07 -0400 Subject: [PATCH 42/43] Pin that flushes never stop an adoption, and that a cycle's own packs go first Two scheduling rules of adoption had no test: every adopting cycle takes one step at least, however many flushes wait for it (f8ae2d9), and a cycle adopts only once every pack of its own went up (10b4fe7). Dropping either left the adoption suite green. The first is now pinned by two threads flushing back to back while a sibling dies; each cycle lists a 32 MiB own pack that fails validation and cannot be quarantined, so a flush is always waiting by the time the loop's cycle comes to adopt (Python flushes alone leave gaps a cycle slips through). The second by an own pack the store always refuses: the dead sibling is not even taken while it is there, and is adopted once it is gone. Each rule's mutant fails its test. --- tests/test_native_spool_adoption_live.py | 138 +++++++++++++++++++++++ 1 file changed, 138 insertions(+) diff --git a/tests/test_native_spool_adoption_live.py b/tests/test_native_spool_adoption_live.py index 41dc714ea..fceccff69 100644 --- a/tests/test_native_spool_adoption_live.py +++ b/tests/test_native_spool_adoption_live.py @@ -936,6 +936,144 @@ def test_a_flush_waits_for_the_adoption_step_in_flight_not_its_slice( lock.release_and_remove_if_empty() +def test_flushes_that_keep_coming_slow_an_adoption_down_but_never_stop_it( + fake_s3, tmp_path): + """A cycle adopting lets go of the cycle at its next step while a flush + is waiting for it -- but only after one step at least, or flushes that + keep coming, as a capture loop's do, would stop adoption altogether: + every adopting cycle would find one waiting and take no step. Here two + threads flush back to back while a sibling that was alive at start + dies; it is still adopted, a step a cycle. + + Every cycle -- a flush's and the loop's -- first lists the service's + own spool, which here hashes a 32 MiB pack that fails validation and + cannot be set aside (a directory holds its quarantine name). So by the + time the loop's cycle comes to adopt, a flush is waiting for it. + Python threads flushing on their own leave gaps between their flushes, + and a cycle that took no step while a flush waited still adopted + through those gaps.""" + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert len(staged) == len(STAGED_BY_THE_DEAD) // RECORDS_PER_PACK + lock = _claim(base, config) + (slow,) = _sparse_backlog(Path(lock.directory) / "slow", 1, 32 << 20) + with open(slow, "r+b") as corrupt: + corrupt.write(b"x") # no longer the zeros its name hashes + slow.with_name(slow.name[:-len(".ready")] + ".quarantined").mkdir() + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="flush-storm-test", poll_interval_ns=50_000_000, + sweep_spool_on_start=True, reconcile_on_start=False, + spool_owner_lock="held_by_caller", adopt_sibling_spools=True, + adoption_recheck_interval_ns=100_000_000, + **config._lease_native()) + service = _store().StorageService(native) + done = threading.Event() + flushed = [0, 0] + failures = [] + + def _flush_back_to_back(index: int) -> None: + try: + while not done.is_set(): + service.flush(10.0) + flushed[index] += 1 + # Not straight back for the lock the cycle just let go + # of: the loop, woken for it, takes its turn. + time.sleep(0.005) + except Exception as exc: # noqa: BLE001 -- reported below + failures.append(exc) + + flushers = [threading.Thread(target=_flush_back_to_back, args=(i,), + daemon=True) for i in range(2)] + service.start() + try: + # The first look finds the sibling alive, and leaves it. + _wait_for(lambda: not service.snapshot()["adoption_owed"]) + assert service.snapshot()["live_siblings"] == 1 + for flusher in flushers: + flusher.start() + _wait_for(lambda: min(flushed) >= 3) + sibling.release() # its owner is gone, the flushes still coming + before = list(flushed) + _wait_for(_adopted(service, 1), 60.0) + snapshot = service.snapshot() + # The flushes kept coming all the while. + assert all(now > then for now, then in zip(flushed, before)), ( + flushed, before) + assert not failures, failures + assert snapshot["adopted_packs"] == len(staged), snapshot + assert not dead.exists() + assert slow.exists() # still listed, and still refused + finally: + done.set() + for flusher in flushers: + if flusher.ident is not None: # started + flusher.join(timeout=60.0) + service.stop() + lock.release_and_remove_if_empty() + + +def test_a_cycle_adopts_nothing_while_its_own_packs_are_not_all_up( + fake_s3, tmp_path): + """A cycle adopts only once every pack this process staged has gone + up: its own records come first. Here the service's own spool holds a + pack the store refuses every time (its key is under the fake store's + fault/always-500/), so each cycle's own upload fails; the dead sibling + beside it waits, every pack in place, however many cycles pass. Once + that pack is gone, the next cycle adopts the sibling.""" + import hashlib + import uuid + + base = tmp_path / "spool" + with _catalog() as prefix: + config = _storage_config(fake_s3, prefix, reconcile_on_start=False) + sibling = _claim(base, config) + dead = Path(sibling.directory) + _stage_into(sibling.directory, STAGED_BY_THE_DEAD) + sibling.release() # its owner is gone + staged = sorted(dead.rglob("*.dmi-pack.ready")) + assert staged + lock = _claim(base, config) + refused = Path(lock.directory) / "fault" / "always-500" + refused.mkdir(parents=True) + body = b"never goes up" + own = refused / (f"{uuid.uuid4()}.1.1." + f"{hashlib.sha256(body).hexdigest()}.dmi-pack.ready") + own.write_bytes(body) + native = config._native_dict() + native.update( + spool_root=lock.directory, spool_max_bytes=1 << 30, + holder="own-first-test", poll_interval_ns=50_000_000, + max_backoff_ns=200_000_000, sweep_spool_on_start=True, + reconcile_on_start=False, spool_owner_lock="held_by_caller", + adopt_sibling_spools=True, uploader_max_attempts=1, + s3_max_attempts=1, **config._lease_native()) + service = _store().StorageService(native) + service.start() + try: + _wait_for(lambda: service.snapshot()["upload_failures"] >= 5) + snapshot = service.snapshot() + assert snapshot["adopted_packs"] == 0, snapshot + assert snapshot["adoption_owed"] is True, snapshot + assert sorted(dead.rglob("*.dmi-pack.ready")) == staged + assert _store().spool_owner(str(dead)) is None # not even taken + + own.unlink() + _wait_for(_adopted(service, 1)) + snapshot = service.snapshot() + assert snapshot["adopted_packs"] == len(staged), snapshot + assert not dead.exists() + finally: + service.stop() + lock.release_and_remove_if_empty() + + def test_a_latched_service_lets_go_of_the_sibling_it_was_adopting( fake_s3, tmp_path): """A service whose catalog another publisher keeps for 2 x TTL latches, From 5cc69fe9afc0253b763ba034b583cb9e36981efa Mon Sep 17 00:00:00 2001 From: Alan Liu Date: Tue, 29 Sep 2026 20:17:23 -0400 Subject: [PATCH 43/43] Say adoption lists a dead spool a pack a step, and how big its round is The uploader's note still had adoption listing a dead spool through Recover, and the flush docstring named a knob the Python config does not have: a round is at most four packs, one per upload worker. --- native/csrc/store/uploader.h | 3 ++- src/dmi/storage/native_capture.py | 2 +- 2 files changed, 3 insertions(+), 2 deletions(-) diff --git a/native/csrc/store/uploader.h b/native/csrc/store/uploader.h index 8666fb12f..843965e26 100644 --- a/native/csrc/store/uploader.h +++ b/native/csrc/store/uploader.h @@ -123,7 +123,8 @@ class SpoolUploader { // does after its listing: in their order, refs and failures positional. // A caller that uploads a listing in parts lists it once (the service's // own spool a chunk at a time, and adoption, which lists a dead spool - // once through Recover and uploads it a round at a time). + // once, a pack a step through Spool::BeginRecovery, and uploads it a + // round at a time). UploadBatchResult UploadStaged(std::vector pending); // Upload one staged entry with retry. Public for tests. *cancelled_out diff --git a/src/dmi/storage/native_capture.py b/src/dmi/storage/native_capture.py index 376e51a3d..d27bec6e7 100644 --- a/src/dmi/storage/native_capture.py +++ b/src/dmi/storage/native_capture.py @@ -872,7 +872,7 @@ def flush(self, timeout_s: float) -> None: Nor does a flush adopt, but it does wait, within ``timeout_s``, for the adoption step the loop is in when it is called: one of a dead spool's packs validated, which hashes it, or one round of their - uploads (at most ``uploader_max_workers`` packs) -- not a dead + uploads (at most four packs, one per upload worker) -- not a dead backlog's whole listing, which goes a pack a step. """ if not self._service.flush(float(timeout_s)):