Version / build tested against
origin/main @ 1ff3551
Deployment mode
Origin, single node (local)
Engine(s) involved
observability
Summary
Six of the per-database metric setters in DatabaseMetricsRegistry
(nodedb/src/control/metrics/database.rs) have no caller anywhere in the
tree, so the families they feed publish constants:
record_qps callers_outside=1 (pgwire sql_exec only)
set_mirror_lag_ms callers_outside=1 (mirror observer)
set_memory_bytes callers_outside=0
set_storage_bytes callers_outside=0
set_connections callers_outside=0
set_bridge_queue_depth callers_outside=0
set_wal_latency_p99 callers_outside=0
add_maintenance_cpu_secs callers_outside=0
render_prometheus is reached from the /metrics route, so
nodedb_database_memory_bytes, nodedb_database_storage_bytes,
nodedb_database_connections, nodedb_database_bridge_queue_depth,
nodedb_database_wal_commit_latency_p99_us, and
nodedb_database_maintenance_cpu_seconds_total are exported with values that
never move. There is no sampler task that fills them.
I found this while wiring loader-facing counters: a bulk writer wants to read
per-database queue depth and commit latency to size its pacing, and the
gauges it would read are always zero. qps is also partial: only the pgwire
statement path calls record_qps, so HTTP and native queries do not count.
Steps to verify
python3 - <<'PY'
import pathlib, re, subprocess
src = pathlib.Path('nodedb/src/control/metrics/database.rs').read_text()
for m in re.findall(r'pub fn (\w+)', src):
out = subprocess.run(['grep','-rn',f'.{m}(','nodedb/src','--include=*.rs'],
capture_output=True, text=True).stdout
sites = [l for l in out.splitlines() if 'control/metrics/database.rs' not in l]
print(f"{m:32} callers_outside={len(sites)}")
PY
Expected behavior
- A periodic sampler (order of ten seconds) fills each per-database gauge
from its owning subsystem: connections from the connection tracker,
memory and storage from the governor or catalog, bridge queue depth from
the bridge dispatcher, WAL commit latency from the group-commit path.
bridge_queue_depth becomes readable per database, so a bulk loader can
back off before the WFQ refuses (the refusal now reports a retryable
dispatch_capacity class).
- Where a source genuinely does not exist yet (maintenance CPU seconds is a
candidate), the metric is either removed or documented as not yet filled,
rather than exported as a constant.
Actual behavior
The setters exist and render; nothing calls them.
What actually happened? (check all that are true)
Proposed severity
SEV-4, observability. No correctness impact, but the containers that promise
backpressure visibility are wrong as published.
Reproducibility
Always; the audit above is deterministic.
Last known-good version / commit (if a regression)
Not a regression as far as I can find; the setters appear to have been added
without a sampler.
Environment & logs
Linux x86_64, single-node origin. Related: #371 added the retryable
dispatch_capacity class and its process-wide counter, and #372 wired the
graph write/delete counters; this issue covers the per-database family those
sit beside.
Before submitting
Version / build tested against
origin/main @ 1ff3551
Deployment mode
Origin, single node (local)
Engine(s) involved
observability
Summary
Six of the per-database metric setters in
DatabaseMetricsRegistry(
nodedb/src/control/metrics/database.rs) have no caller anywhere in thetree, so the families they feed publish constants:
render_prometheusis reached from the/metricsroute, sonodedb_database_memory_bytes,nodedb_database_storage_bytes,nodedb_database_connections,nodedb_database_bridge_queue_depth,nodedb_database_wal_commit_latency_p99_us, andnodedb_database_maintenance_cpu_seconds_totalare exported with values thatnever move. There is no sampler task that fills them.
I found this while wiring loader-facing counters: a bulk writer wants to read
per-database queue depth and commit latency to size its pacing, and the
gauges it would read are always zero.
qpsis also partial: only the pgwirestatement path calls
record_qps, so HTTP and native queries do not count.Steps to verify
Expected behavior
from its owning subsystem: connections from the connection tracker,
memory and storage from the governor or catalog, bridge queue depth from
the bridge dispatcher, WAL commit latency from the group-commit path.
bridge_queue_depthbecomes readable per database, so a bulk loader canback off before the WFQ refuses (the refusal now reports a retryable
dispatch_capacityclass).candidate), the metric is either removed or documented as not yet filled,
rather than exported as a constant.
Actual behavior
The setters exist and render; nothing calls them.
What actually happened? (check all that are true)
conservative)
Proposed severity
SEV-4, observability. No correctness impact, but the containers that promise
backpressure visibility are wrong as published.
Reproducibility
Always; the audit above is deterministic.
Last known-good version / commit (if a regression)
Not a regression as far as I can find; the setters appear to have been added
without a sampler.
Environment & logs
Linux x86_64, single-node origin. Related: #371 added the retryable
dispatch_capacityclass and its process-wide counter, and #372 wired thegraph write/delete counters; this issue covers the per-database family those
sit beside.
Before submitting