Conversation
Bulk loaders pay one GRAPH INSERT EDGE per edge, so a 66k edge import spends hours in round trips. The plan, staging, WAL, and Data Planes already carried EdgePutBatch / EdgeDeleteBatch; the statement layer above them was the missing piece. Add GRAPH INSERT EDGES / GRAPH DELETE EDGES taking a property-less VALUES list of (src, dst, label) triples, capped at 1000 edges per statement with the cap named in the error. Per-edge PROPERTIES stays on the single-edge form until BatchEdge can carry a property object. The handler resolves each edge (surrogates, write policy) and buckets by home vShard, so a single-home batch is one apply burst per home and the Calvin path sequences the whole batch as one tx class. build_static_tx_class now derives participant homes and lock identity from batch edge plans.
Three loader-facing gaps on one surface. SHOW SNAPSHOT returns the monotonic WAL pin (wal_next_lsn) with the node id and version, giving a reader that pages with a fresh snapshot per page a token to capture once and compare against a producer's commit marker. A saturated dispatch WFQ reported a generic dispatch error, which the classifier leaves unclassified, so a loader had no retry contract. It now reports a `dispatch_capacity:` detail that classifies as rate_exceeded in the same class vshard admission capacity uses, counted process-wide and rendered as nodedb_dispatch_capacity_exhausted_total. SystemMetrics gains graph_edges_written_total and graph_edges_deleted_total, incremented at the edge-write handler success points and rendered in /metrics and SHOW STATS, so pacing can read arrival versus apply rate instead of estimating from query counts.
Coverage for the batched edge statements: a batch delete under a FOR WRITE owner policy evaluates the policy per edge (the conforming edge is deleted, the denied one survives, and the statement reports an error), and a property-less batch insert under an owner policy is refused with nothing landing. That refusal matches the existing batch-edge decision in the RLS injection pass, pinning that a batch is never the write surface a policy does not reach.
The per-database bridge_queue_depth gauge had no writer: nothing filled it from the dispatch WFQ, so the exported line was a constant zero for every database. Add a ten-second background loop that reads each database's WFQ depth across all cores and publishes it, backed by a read-only Dispatcher accessor for the sum. The other per-database families stay unwritten on purpose, and the sampler documents why: the session registry that would feed connections has no production registration caller, and memory, storage, WAL commit latency, and maintenance CPU have no per-database source at all.
4509e64 to
a08af14
Compare
The connections gauge had no writer. Fill it from the admission registry's per-database live permit count, resolved by name the way SHOW DATABASE USAGE resolves it, and publish connections before the dispatcher lock is touched so a contended dispatch poller cannot stall the sample. An entry exists once a quota record applies, and a database with no entry stays unpublished: absence means unmeasured, never a fabricated zero. Removing a cap keeps the entry and its live count, so a measured zero still publishes.
|
Superseded by four focused pull requests, each rebased onto current Why this PR is superseded The branch is 68 commits behind
This PR reports a saturated dispatch queue by string-sniffing a Where each piece went
#373 was already closed while its work sat only in this branch, so the merge could not close it. Evidence carried over Each replacement PR was built from current
Thanks for the review time on this one. |
What
Three commits closing the loader-facing gaps from the ETL port review:
feat(graph):GRAPH INSERT EDGES/GRAPH DELETE EDGES(graph: bulk edge ingest has no batch form — every loader pays one statement per edge; add GRAPH INSERT/DELETE EDGES #369). The plan, staging, WAL, and Data Planes already carriedEdgePutBatch/EdgeDeleteBatch; the statement layer above them was missing. The batch form takes a property-lessVALUESlist of(src, dst, label)triples, capped at 1000 edges per statement with the cap named in the error. The handler resolves each edge (surrogates, write policy) and buckets by home vShard, so a single-home batch is one apply burst per home;build_static_tx_classnow derives participant homes and lock identity from batch edge plans for the Calvin path.feat(obsv):SHOW SNAPSHOTreturns the monotonic WAL pin (wal_next_lsn) with node id and version, so a reader that pages with a fresh snapshot per page has a token to capture once and compare against a producer's commit marker (pgwire: no way to pin a read snapshot across pages — multi-page readers can observe half-applied producer batches; add pg_snapshot_xmin / SHOW SNAPSHOT #370). A saturated dispatch WFQ reported a generic dispatch error that the classifier leaves unclassified; it now reports adispatch_capacity:detail classified asrate_exceeded, the same class vshard admission capacity uses, with a process-wide counter rendered asnodedb_dispatch_capacity_exhausted_total(cluster: a saturated dispatch queue has no documented retryable class — clients cannot back off correctly; add one and a counter #371).SystemMetricsgainsgraph_edges_written_totalandgraph_edges_deleted_total, incremented at the edge-write handler success points and rendered in/metricsandSHOW STATS(obsv: add loader-facing write-pressure counters (graph edge writes, dispatch capacity busy) #372).test(graph): wire coverage proving batch write paths honour RLS/GRANT (test(graph): batch write paths need RLS and GRANT probes #373). A batch delete under aFOR WRITEowner policy evaluates the policy per edge (conforming edge deleted, denied edge survives, statement reports an error); a property-less batch insert under an owner policy is refused with nothing landing, matching the existing batch-edge decision in the RLS injection pass.feat(obsv)(2nd): a ten-second sampler loop publishes per-database bridge queue depth from the dispatch WFQ, with a read-onlyDispatcher::db_queue_depthaccessor. The other per-database families stay unwritten on purpose and the sampler documents why (the session registry has no production registration caller; memory, storage, WAL commit latency, and maintenance CPU have no per-database source).Notes
BatchEdgecarries no property object, so per-edgePROPERTIESstays on the single-edge form. That is also why the RLS injection pass refuses batch writes when a policy applies, which the new tests pin.bridge_queue_depthgauge has no writer anywhere in the tree (six per-database setters have no callers). That is a separate finding and is not touched here.Evidence
nodedb-sql: 863 tests pass (5 new parser tests).error_classify: unit test for the retryable dispatch capacity class.graph_dsl_batch3/3,graph_batch_rls3/3,show_snapshot1/1; regressionsgraph_dsl44/44,graph_timeseries_rls_probe7/7,engine_surface_graph14/14,pgwire_show_dispatch19/19.cargo check -p nodedb,cargo fmt,clippy -p nodedb-sql --all-targetsclean.One machine-load flake was observed in this session on
graph_path_preserves_colon_containing_user_idunder a parallel run; it passes when run alone and in sequence, so it is noted rather than treated as a regression.What CI does not cover locally
run-cirequested for the full suite.Closes #369
Closes #370
Closes #371
Closes #373
Refs #372
Refs #375