Skip to content

bug(control): response_poller adaptive backoff is unreachable while responses keep arriving — one Control-Plane worker pins a core forever #376

Description

@EnRaiha

Version / build tested against

nodedb 0.5.0 @ 3b4a96da2 (local build/head-plus-fixes).

The affected function is byte-identical to origin/main @ 1ff35512b — verified against the running commit:

git show origin/main:nodedb/src/bootstrap/background_loops.rs   # spawn_response_poller
git show 3b4a96da2:nodedb/src/bootstrap/background_loops.rs     # spawn_response_poller
diff <(...)  <(...)   # empty

So this is not a stale-branch artifact: the code that spins is the code on main today.

Deployment mode

Origin — single node (local)

Engine(s) involved

Not engine-specific / unsure — this is the Control Plane response poller, not a storage engine.

Summary

response_poller (nodedb/src/bootstrap/background_loops.rs:383-424) never reaches its own backoff. Its idle backoff counter is reset to 0 by any routed response, so as long as the data-plane cores keep answering at a steady trickle (heartbeats, coordinated CRDT checkpoints, collection-GC sweeps, vector checkpoints) — which they do on an idle node with no clients — an idle_iters value above 1024 is never reached and the loop stays in its yield_now() fast path forever. The result is a permanently pinned Control-Plane worker: one tokio-rt-worker at ~100% of a core, ~1.39 cores across the process, sustained for the whole process lifetime (13 h+ observed), plus a cross-thread wake storm (futex + runtime waker eventfd) and ~23% of a 6-core host lost. The data-plane cores themselves sit at 0% — none of that CPU is doing work.

Steps to reproduce

No SQL and no clients are needed — the node reproduces it while completely idle.

  1. Run a single-node instance with the default 2 data-plane cores (RUST_LOG=nodedb=info is enough). Create no schema, connect no clients.
  2. Wait ~5–10 minutes, then look at the threads:
PID=$(pgrep -f 'nodedb --config')
top -H -b -n1 -p $PID | head -20
# one tokio-rt-worker sits at ~100.7% and never comes down; data-core-0/1 are at 0.0%
  1. Confirm the write rate into the runtime's waker eventfd (sample twice, 20 s apart):
ls -l /proc/$PID/fd | grep eventfd            # e.g. fd 4 in my run
grep eventfd-count /proc/$PID/fdinfo/4        # 1,634,967,162
sleep 20
grep eventfd-count /proc/$PID/fdinfo/4        # 1,635,596,931  -> +629,769 in 20.7 s ≈ 30,400/s

The count rises monotonically for hours: writes far outrun the one token consumed per edge, and the value is already > 1.6 billion.

  1. Confirm the shape of the spin (8 s on the pinned thread):
TID=$(ps -eLo pid,tid,pcpu,comm -p $PID | sort -k3 -rn | head -2 | tail -1 | awk '{print $2}')
strace -c -p $TID
% time     seconds  usecs/call     calls    errors syscall
------ ----------- ----------- --------- --------- ----------------
 56.45    0.179212           0    208622        28 futex
 42.83    0.135980           1    107602           write
  0.71    0.002247           1      1591           epoll_wait
  0.00    0.000012           0        16           sched_yield

epoll_wait(fd, [], 1024, 0) — always timeout 0, i.e. runtime turns with a task that is always ready; never a blocking wait.

  1. Attribute the writes to a loop registered through the shutdown loop registry:
strace -f -e trace=write -p $PID          # 2.5 s window
# 68,712 × write(<fd>, "\1\0\0\0\0\0\0\0", 8)      # eventfd notify value 1

strace -k on that write resolves the calling stack (symbol names only — the release binary ships without DWARF) into task-poll frames under nodedb::control::shutdown::spawn::spawn_loop::{{closure}}, i.e. one of the loops spawned via spawn_loop / spawn_loop_no_abort. Of the loops spawned from bootstrap/background_loops.rs, response_poller is the only one that contains yield_now() (:438, :443 in the local tree; :410, :415 on main).

Expected behavior

On an idle node the response poller routes nothing, goes through yield_now() 256 times, then settles into sleep(1ms) and finally sleep(10ms) — roughly 100 iterations/second, ~0% of a core. The Control Plane should not burn a core waiting for responses that are not arriving.

Actual behavior

One tokio-rt-worker runs at ~100% of a core for the whole process lifetime, with the rest of the runtime churning behind it:

per-thread CPU over a 10 s window (single node, 2 data cores, idle):
  tokio-rt-worker   100.7% of a core   (3.45 h CPU on that thread alone)
  tokio-rt-worker    12.2% of a core   (2.36 h)
  tokio-rt-worker    11.3% of a core   (2.85 h)
  tokio-rt-worker    11.2% of a core   (1.69 h)
  data-core-0         0.0%
  data-core-1         0.0%

whole process, 5 s window: 1.39 cores
host-wide: 227k context switches/s, ~23% of a 6-core box lost, process heap partly swapped out (VmSwap ≈ 1.5 GiB)

idle_iters is reset on every routed response, so the two backoff steps (1 ms after 256 idle iterations, 10 ms after 1024) are unreachable whenever responses arrive more often than ~1.8 s. With 2 cores emitting heartbeats (~1 s each) plus coordinated checkpoints every ~20–30 s, GC sweeps and vector checkpoints, that reset happens constantly — the loop just spins at yield_now speed (13–30k iterations/s measured).

Severity — facts

  • Acknowledged/committed data was lost, corrupted, or silently wrong
  • The server crashed, hung, or failed to start
  • A security or isolation boundary was crossed
  • Core functionality is broken with no acceptable workaround
  • A workaround exists (over-provision CPU / accept the burn)

Proposed severity

SEV-3 — Medium. No data loss, no crash, queries are served normally; stored data intact. Reporting it as Medium rather than Low only because the burn is unconditional and permanent (a 2-core node effectively loses ~half its CPU to it), and there is no in-product workaround. Maintainers to confirm at triage.

Reproducibility

Always — every attempt. Observed continuously since process start (13 h+ in the captured run); the pinned thread, the 30k/s eventfd growth and the 1.39-core process average are all reproducible on demand.

Last known-good version / commit (if a regression)

Unknown. The loop dates back to at least 25f6cea14 (2026-05-05, the file-splitting refactor that moved it into bootstrap/background_loops.rs) and only received shutdown-barrier touches since (1e59a59f7, a5becc434, 0e72933e1). Related closed issues: #20 (idle container at ~144% CPU) and #51 (async-runtime CPU/RAM hotspots) — this looks like the concrete instance of that class that survived those fixes.

Environment & logs

Linux x86_64, release build, single node, 2 data-plane cores
CPU: i5-8400 (6 cores), 32 GB RAM, LVM-thin storage
RUST_LOG=nodedb=info,nodedb_cluster=info,info
Uptime at capture: 13 h 08 m — the spin was present the whole time

Relevant server log cadence on that idle node (nothing anomalous, just the steady responses that keep the loop hot):

INFO nodedb::data::executor::handlers::control::checkpoint_crdt: CRDT checkpoint published   (≈22× / 5 min)
INFO nodedb::event::collection_gc::sweeper: collection-gc sweep complete                     (≈5× / 5 min)
INFO nodedb::data::executor::vector_checkpoint::write: vector checkpoint published           (≈4× / 5 min)
INFO nodedb::control::server::pgwire::listener: new/closed pgwire connection                 (≈4× / 5 min)

No ERROR/WARN spam accompanies the spin — it is silent.

Root cause (analysis)

nodedb/src/bootstrap/background_loops.rs (main line numbers):

399  let mut idle_iters: u32 = 0;
     loop {
         // ...
408      let routed = shared_poller.poll_and_route_responses();
409      if routed > 0 {
410          idle_iters = 0;                      // <- any response resets the backoff
411          tokio::task::yield_now().await;
412          continue;
         }
413      idle_iters = idle_iters.saturating_add(1);
414      if idle_iters <= 256 {
415          tokio::task::yield_now().await;
416      } else if idle_iters <= 1024 {
417          tokio::time::sleep(Duration::from_millis(1)).await;
418      } else {
419          tokio::time::sleep(Duration::from_millis(10)).await;
         }
     }

The fast path is correct for a burst, but the backoff state is coupled to success rather than to idleness: one routed response per ~1.8 s is enough to keep idle_iters pinned near zero forever. poll_and_route_responses() (control/state/methods.rs:455) also takes the dispatcher mutex and calls flush_wfq() per core on every call, so the spin is not free even when routed == 0.

Possible direction (not a patch — for maintainers)

  • Drive the backoff from consecutive empty polls (or from the observed response rate) so a steady trickle of responses cannot keep the loop in the hot branch permanently.
  • Or gate the poll on the cores' response-available eventfd instead of polling: an eventfd wait is exactly what the hot path is emulating, and the CoreChannel already holds a notifier per core.
  • Bound the hot branch to a fixed iteration budget per wake, so a mis-set rate degrades to latency rather than to a pinned core.

Before submitting

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

status:needs-triageAwaiting maintainer triage (severity + priority)type:bugA defect — broken, incorrect, or lost data

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions