Skip to content

Install coherent deployment plans with dataset-bound SDS identities - #749

Open
zzylol wants to merge 26 commits into
mainfrom
refactor/backend-plan-split
Open

zzylol wants to merge 26 commits into
mainfrom
refactor/backend-plan-split

Conversation

@zzylol

@zzylol zzylol commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Deployments need one coherent installed plan that binds queries and producers to stored state with known meaning. Routing fingerprints alone cannot distinguish equal metric expressions over different logical datasets, or represent two deployed outputs with the same semantics safely.

Before this PR

Backend configuration mixes plan concerns and does not establish the SDS identity contract from #737. A query binding cannot independently express the semantic computation and the authorized deployed output.

After this PR

Install a coherent PrecomputePlan, QueryPlan, transmission configuration and immutable catalog snapshot. Planner exports the persisted output's canonical typed dependency closure, including logical dataset identity; deployment binding connects it to a stored output.

Planner persisted-output semantics + logical dataset
    → SummaryDefinition / definition_id
    → deployed stored_output_id
    → installed writer and query bindings
    → validated stored records

Two KLL(latency) outputs named hot/rebuild may share a definition but are independently bound. Another dataset changes the definition; relocating the same dataset preserves it. Same-version restart restores valid state. A new plan version remains cold until fresh input is produced and never adopts previous-version payloads.

Catalog schema 6 and planning snapshot schema 3 establish this identity contract here. Typed Count/Rate support formerly in #771 is included because those semantics must remain distinct. Shared physical execution and its precompute execution design document continue in #774.

Obsolete installed-plan aliases, untyped projection decoders, the unused materialization wire adapter, and cross-version adoption metadata are removed. Recovery accepts only the current sidecar schema and propagates malformed or unsupported metadata errors. Producer partition rosters are explicit on the wire.

Dependency changes

  • asap_sketch_codec contains the envelope encoding/decoding helpers used by backend ingest, DDSketch/KLL accumulators and wire-format tests. It supports removing the asap-precompute-rs Collector runtime dependency; it adds no sketch algorithms. The workspace member and data-plane dependency are intentional. With the Collector dependency removed, its Sketchlib patch block is unused and is removed too.
  • Planner advances from cd7e9e0 to bccc837 because this implementation uses LogicalDatasetIdentity and SummarySemanticFragment::from_stored_output_in_dataset. Both APIs are absent at the old revision. All four Planner dependencies use the same immutable revision.

Validation and scope

The SDS implementation passed locally: 117 type-library tests, 431 control-plane library tests, 1,165 data-plane library tests, 8 HTTP API tests, 13 serving integration tests, the production-process restart/warm-up test, strict all-target Clippy and formatting.

Coverage includes dataset changes, endpoint relocation, independent hot/rebuild bindings, tampered definitions, writer/read consistency, same-version restart without re-ingestion, and new-version warm-up. Earlier downstream validation also passed Level 1, exhaustive synthetic Level 2 selection, 332 storage tests and 13 serving integration tests.

Discovery, calibration and workload replay now preserve version-3 dataset identity. Process fixtures use typed deployment configuration instead of removed flat aggregation documents. The 81 Python tests and seven focused sketch/component/monitor/forwarding process tests pass locally; the complete workspace and required CI are rerunning.

PR-specific evidence directories and generated archives are removed from the repository. Test summaries belong here; raw logs remain local or in CI.

Ad-hoc discovery, cross-version adoption and production-cost validation remain out of scope. Restricted native configuration helpers support explicit imported-state fixtures; production Planner compilation requires dataset identity. No manual deployment or human approval is claimed.

@zzylol zzylol changed the title refactor: split installed maintenance DAGs from query execution refactor: split backend plans and remove Collector dependency Sep 19, 2026
@zzylol zzylol changed the title refactor: split backend plans and remove Collector dependency refactor: split backend plans, unpin Planner, and remove Collector dependency Sep 21, 2026
@zzylol zzylol changed the title refactor: split backend plans, unpin Planner, and remove Collector dependency refactor: split backend plans and bind SDS state slots Sep 21, 2026
@zzylol
zzylol marked this pull request as ready for review September 21, 2026 16:13
@zzylol
zzylol changed the base branch from main to docs/physical-plan-design September 22, 2026 00:37
@zzylol zzylol changed the title refactor: split backend plans and bind SDS state slots refactor: split backend plans and bind stored summary outputs Sep 22, 2026
zzylol added a commit that referenced this pull request Sep 22, 2026
Integrate PR #749 reader/writer bindings and maintenance projections while preserving data-partition worker ownership and plan-derived configuration. Keep selected and maintenance DAG schemas distinct and adapt shared-sink execution to the projected graph.
@zzylol
zzylol force-pushed the docs/physical-plan-design branch from b0d77ce to 5f1eebf Compare September 28, 2026 13:47
@zzylol
zzylol force-pushed the refactor/backend-plan-split branch from 12896bd to 70a8d88 Compare September 28, 2026 16:14
@zzylol
zzylol changed the base branch from docs/physical-plan-design to main September 28, 2026 16:14
@zzylol zzylol changed the title refactor: split backend plans and bind stored summary outputs Install coherent deployment plans with dataset-bound SDS identities Sep 28, 2026
zzylol and others added 15 commits September 28, 2026 17:27
CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant