diff --git a/Cargo.lock b/Cargo.lock index a063dd78..7f0a3755 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -368,8 +368,10 @@ dependencies = [ "asap-aware-mapping", "asap-frontend-promql", "asap-frontend-sql", + "asap-physical-operators", "asap-types", "asap_sketchlib 0.3.0 (git+https://github.com/ProjectASAP/asap_sketchlib)", + "futures", "serde_json", "tokio", ] diff --git a/crates/asap-physical-operators/README.md b/crates/asap-physical-operators/README.md new file mode 100644 index 00000000..5111d316 --- /dev/null +++ b/crates/asap-physical-operators/README.md @@ -0,0 +1,130 @@ +# ASAP physical operators + +An independent Rust physical operator DAG runtime shared by ingestion time and +query time execution. The library requires neither backend engine, a server, +a storage implementation, Arrow nor DataFusion. DataFusion informed the design; +it is not the execution framework. + +`plan::PhysicalDag` binds typed operator inputs to node IDs. Each execution starts +one producer per reachable node, shares output batches among its consumers, and +bounds buffering. Dropping one consumer does not cancel other consumers. A +`RunContext` carries query or ingestion scope, cancellation and byte accounting. +Executions use the caller's worker and worker-local streams, with no internal +thread pool. Poll multiple root streams concurrently when they share inputs. + +`operators::Operator` implements native batch sources, scalar values, +projection, filtering, grouped exact aggregation, semi-join, grouped Sort and +Limit, vector-to-scalar conversion, Union, and summary construction/merge/readout. +Sort followed by Limit implements grouped ranking; no dedicated TopK physical +operator is needed. Summary construction updates state batch by batch. End of +input means the supplied query range or ingestion window is complete. + +```rust +use asap_physical_operators::{ + expressions::Expression, + operators::Operator, + values::Value, + plan::PhysicalDag, + runtime::{Limits, RunContext, Scope}, +}; +use asap_physical_operators::planner::pre_asap::DataType; +use futures::{executor::block_on, StreamExt}; + +let source = Operator::scalar(Value::Int64(7), DataType::Int64)?; +let negate = Operator::project(source.schema(), vec![ + ("value".into(), Expression::Negate(Box::new(Expression::Column(0)))), +])?; +let mut plan = PhysicalDag::default(); +plan.add(0, vec![], source)?; +plan.add(1, vec![0], negate)?; +let run = RunContext::new( + Scope::Query { evaluation_time_ms: 1000, revision: 1 }, + Limits::default(), +)?; +let mut output = plan.execute(&[1], run)?.remove(0); +let batch = block_on(output.next()).unwrap()?; +assert!(matches!(batch.rows()[0][0], Value::Int64(-7))); +# Ok::<(), asap_physical_operators::dag::Error>(()) +``` + +`physical_planner::compile` accepts a logical Post-ASAP DAG (`PostAsapDag`) and typed input contracts. +The resulting candidate is instantiated with deployment readers after selection. It rejects unsupported operations and +schema mismatches before starting a source. Implement `PhysicalOperator` for a +deployment source, including asynchronous I/O; computation operators remain in +the library. The public `planner` export identifies the exact Planner types used +by the crate. The physical compiler currently supports a subset of those types and +operations; it does not interpret an unknown node as external fallback. + +Plain values preserve Planner scalar/collection types and nullability. Numeric +arithmetic uses matching Int64 or Float64 inputs; integer overflow is an error. +Boolean predicates use three-valued logic. Native summary states currently cover +exact Sum/Count/Min/Max/Rate/Increase, KLL, DDSketch, HLL and Float64 weighted CMS and CountSketch with candidate heaps. Binding checks family, +parameters and readout compatibility; source batches also validate state payloads. +Existing accumulator algorithms are reused as kernels behind these operators. + +This crate is owned by ASAPPlanner. Its `planner-types` dependency is the local +IR crate, so a contract change and its execution tests belong in the same PR. +Deployments supply storage/ingestion sources and adapt output protocols. The +library has no ASAPQuery-backend dependency. Backend raw Scan remains a separate +deployment capability. + +See [the design](../../docs/design_docs/physical-planning-and-deployment.md). + +## Module boundaries + +- `plan`: immutable graph, operator interface, schemas and execution properties. +- `runtime`: per-run streams, shared producers, memory reservations and cancellation. +- `expressions`: scalar evaluation; typed builders and the Planner expression adapter. +- `operators`: projection, filter, joins, aggregate/window, sort, limit and summary implementations. +- `sources`: raw-source interface, Scan and the memory connector. +- `physical_planner`: native operator lowering, typed input contracts and checked instantiation. +- `summary_kernels`: in-memory summary state over `asap_sketchlib` and exact Planner state: merge, typed readout and update adapters. +- `readout`: readouts over merged exact summary states. +- `capability`: explicit kernel, native-batch and typed readout validation. + +The `dag`, `factory`, `traits` and `arithmetic` paths are re-exports. They contain no alternative +execution implementations. + +Sketch algorithms and their state encodings belong to `asap_sketchlib`. Wire decoding, delta +frames, edge sampling and storage statistics belong to deployments. Kernels hold one population's +state; operators own grouping. + +A source must declare `Boundedness::Bounded` to feed a blocking operator. +The default for a custom raw source is `Unknown`; query or ingestion scope alone +does not promise that its cursor ends. `PhysicalDag::properties` validates these +requirements before any source starts and returns boundedness and emission mode +for every reachable node. The memory connector declares finite input. Custom +physical sources expose the same facts through `PhysicalOperator::properties`. + +Blocking operators reserve estimated workspace and yield cooperatively during +row processing and sort merges. Cancellation releases reservations when the +stream is polled or dropped. Individual scalar evaluations, bounded sort chunks +and sketch kernel calls are synchronous; this is not preemptive execution. +There is no spill or partitioned parallel execution in this implementation. + +## Physical compilation and deployment inputs + +`physical_planner::compile` accepts a Planner `PostAsapDag`, typed +`InputContract`s and output roots. It returns a reusable `CompiledPhysicalDag` +containing selected native operators and no live readers. Compilation validates +schemas, input ordering, sharing and boundedness before deployment source access. + +A deployment calls `CompiledPhysicalDag::instantiate` with exactly the declared +inputs. This checks source schemas and execution properties and constructs the +runnable graph without repeating logical lowering. The graph executes through +the shared runtime with independent per-run state. Window coverage, revision and +maintenance-policy admission remain deployment/planning contracts; this compiler +does not discover storage or silently change a selected maintenance strategy. + +`physical_planner::compile_temporal_pane_candidate` lowers a selected continuous +KLL lifecycle and Sliding/Tumbling framework into maintenance and query DAGs. +`TemporalPaneMaintenance` supplies pane geometry and a resolved complete entity +identity contract. The compiler inserts population guards, scan predicates, +pane construction, ordered state slots, a shared merge and quantile readouts. +Pane outputs have distinct physical identities from the logical whole-window +summary, and the returned candidate retains the maintenance contract for binding. +Each run checks phase, pane timestamps and duplicate entity states. The initial +realization uses complete bounded snapshots; partial edges, exponential +histograms and cross-run delta accumulation are unsupported. Storage identities, +revision selection, completeness/readiness evidence and scheduling stay with +deployment. diff --git a/crates/asap-physical-operators/src/dag/mod.rs b/crates/asap-physical-operators/src/dag/mod.rs new file mode 100644 index 00000000..c7837672 --- /dev/null +++ b/crates/asap-physical-operators/src/dag/mod.rs @@ -0,0 +1,8 @@ +//! Compatibility imports. New code should use plan, runtime, operators, physical_planner and sources directly. +pub use crate::plan::{NodeId, PhysicalDag, PhysicalOperator}; +pub use crate::runtime::batch_execution; +pub use crate::runtime::{ + Input, Limits, OutputStream, Reservation, RunContext, Scope, SharedValue, +}; +pub use crate::Error; +pub use crate::{expressions, operators, physical_planner as planner, sources as scan, values}; diff --git a/crates/asap-physical-operators/src/lib.rs b/crates/asap-physical-operators/src/lib.rs index 0835159c..345ee762 100644 --- a/crates/asap-physical-operators/src/lib.rs +++ b/crates/asap-physical-operators/src/lib.rs @@ -1,5 +1,4 @@ -//! Shared physical operators: summary kernels, typed values, native operators -//! and the DAG runtime. +#![doc = include_str!("../README.md")] pub mod key_by_label_values; pub mod measurement; @@ -12,20 +11,23 @@ pub use measurement::Measurement; pub use statistic::Statistic; pub use traits::*; +pub use expressions::arithmetic; pub mod capability; pub use summary_kernels::factory; /// The exact Planner contract used by these kernels. pub use planner_types as planner; +pub mod dag; + +pub mod readout; + mod error; pub use error::Error; -pub mod values; - -pub use expressions::arithmetic; pub mod expressions; pub mod operators; +pub mod physical_planner; pub mod plan; -pub mod readout; pub mod runtime; pub mod sources; +pub mod values; diff --git a/crates/asap-physical-operators/src/operators/joins/mod.rs b/crates/asap-physical-operators/src/operators/joins/mod.rs index 608fd8d7..a4397033 100644 --- a/crates/asap-physical-operators/src/operators/joins/mod.rs +++ b/crates/asap-physical-operators/src/operators/joins/mod.rs @@ -32,6 +32,15 @@ impl Operator { } self } + pub(crate) fn certified_pruning_keys(&self) -> Option<&[(usize, usize)]> { + match &self.kind { + Kind::SemiJoin { + keys, + require_complete_right: true, + } => Some(keys), + _ => None, + } + } pub fn relational_join( left: Schema, right: Schema, diff --git a/crates/asap-physical-operators/src/operators/mod.rs b/crates/asap-physical-operators/src/operators/mod.rs index e4c6944f..0f20fa3d 100644 --- a/crates/asap-physical-operators/src/operators/mod.rs +++ b/crates/asap-physical-operators/src/operators/mod.rs @@ -135,6 +135,42 @@ pub struct Operator { output: Schema, } impl Operator { + pub(crate) fn row_preserving_input(&self) -> Option { + match self.kind { + Kind::Filter(_) | Kind::Sort { .. } | Kind::Limit { .. } | Kind::SemiJoin { .. } => { + Some(0) + } + _ => None, + } + } + + pub(crate) fn is_counter_readout(&self) -> bool { + matches!( + self.kind, + Kind::Readout { + query: ReadoutQuery::Exact(crate::summary_kernels::exact::ExactReadout { + statistic: crate::Statistic::Rate | crate::Statistic::Increase, + .. + }), + .. + } + ) + } + + pub(crate) fn with_counter_lookback(mut self, lookback: i64) -> Result { + if lookback <= 0 { + return Err(invalid("counter lookback must be positive")); + } + if let Kind::Readout { + query: ReadoutQuery::Exact(readout), + .. + } = &mut self.kind + { + readout.lookback_ms = Some(lookback); + } + Ok(self) + } + /// Resolve a counter readout's logical lookback to this run's evaluation range. pub(super) fn readout_range(&self, context: &RunContext) -> Result, Error> { let Kind::Readout { diff --git a/crates/asap-physical-operators/src/physical_planner/candidates.rs b/crates/asap-physical-operators/src/physical_planner/candidates.rs new file mode 100644 index 00000000..a8f476f3 --- /dev/null +++ b/crates/asap-physical-operators/src/physical_planner/candidates.rs @@ -0,0 +1,275 @@ +//! Compile maintenance-selected frontiers without deployment-specific graph rewrites. +use super::*; + +/// One computation realization; lifecycle/window/revision requirements accompany +/// it during optimization and deployment. Stored outputs have no storage identity. +/// Deserialization validates the producer/reader boundary. +#[derive(Clone, serde::Serialize, serde::Deserialize)] +#[serde(try_from = "UncheckedCandidate")] +pub struct PhysicalCandidate { + pub precompute: Option, + pub query: CompiledPhysicalDag, + pub materialized_outputs: BTreeMap, +} + +/// Compile an explicit materialization frontier selected by Planner maintenance +/// search. Operators upstream of that frontier run in precompute, including +/// readouts/reductions; query execution receives their typed output values. +/// Empty frontiers retain the full computation in the query DAG. +/// +/// Repeated windows must be instantiated with the same evaluation/population +/// contract used to build each output. This API never treats a result from a +/// different window or revision as interchangeable merely because types match. +pub fn compile_candidate( + dag: &PostAsapDag, + inputs: BTreeMap, + roots: &[NodeId], + frontier: &[NodeId], +) -> Result { + if frontier.is_empty() { + return Ok(PhysicalCandidate { + precompute: None, + query: compile(dag, inputs, roots)?, + materialized_outputs: BTreeMap::new(), + }); + } + let frontier_set: BTreeSet<_> = frontier.iter().copied().collect(); + if frontier_set.len() != frontier.len() || frontier.iter().any(|id| inputs.contains_key(id)) { + return Err(invalid("frontier must contain distinct computed outputs")); + } + let full = compile(dag, inputs.clone(), roots)?; + let precompute = compile(dag, inputs.clone(), frontier)?; + let mut materialized_outputs = BTreeMap::new(); + for &id in frontier { + // Also proves that the frontier is reachable from the requested roots. + full.output_contract(id)?; + let mut output = precompute.output_contract(id)?; + if output.properties.boundedness != Boundedness::Bounded { + return Err(invalid("materialized output requires bounded execution")); + } + // A stored reader may stream batches even when the producer blocked. + // Its timing is independent; the retained result still must be finite. + output.properties.emission = Emission::Unknown; + materialized_outputs.insert(id, output); + } + let mut query_inputs = inputs; + query_inputs.extend(materialized_outputs.clone()); + let query = compile(dag, query_inputs, roots)?; + let used: BTreeSet<_> = query.input_contracts().map(|(id, _)| id).collect(); + if !frontier.iter().all(|id| used.contains(id)) { + return Err(invalid( + "frontier contains an output shadowed by another boundary", + )); + } + Ok(PhysicalCandidate { + precompute: Some(precompute), + query, + materialized_outputs, + }) +} + +/// Enumerate bounded, reachable materialization frontiers above explicit inputs. +/// Each frontier is an antichain: storing an output and its ancestor together +/// would leave the ancestor unused by query execution. Lifecycle eligibility +/// and deployment feasibility are evaluated separately before cost selection. +/// Exceeding the search budget returns an error, never a partial inventory. +pub fn enumerate_frontiers( + dag: &PostAsapDag, + inputs: &BTreeMap, + roots: &[NodeId], + max_candidates: usize, +) -> Result>, Error> { + if max_candidates == 0 { + return Err(invalid( + "frontier search requires a positive candidate budget", + )); + } + let compiled = compile(dag, inputs.clone(), roots)?; + let mut ancestors = BTreeMap::>::new(); + let mut eligible = Vec::new(); + for node in &dag.nodes { + let id = u64::from(node.id.0); + if inputs.contains_key(&id) { + continue; + } + let Ok(contract) = compiled.output_contract(id) else { + continue; + }; + if contract.properties.boundedness != Boundedness::Bounded { + continue; + } + let mut seen = BTreeSet::new(); + let mut pending = vec![id]; + while let Some(current) = pending.pop() { + if !seen.insert(current) || inputs.contains_key(¤t) { + continue; + } + pending.extend( + dag.edges + .iter() + .filter(|edge| u64::from(edge.consumer.0) == current) + .map(|edge| u64::from(edge.producer.0)), + ); + } + ancestors.insert(id, seen); + eligible.push(id); + } + eligible.sort_unstable(); + let mut frontiers = vec![vec![]]; + for id in eligible { + let additions = frontiers + .iter() + .filter(|frontier| { + frontier.iter().all(|previous| { + !ancestors[&id].contains(previous) && !ancestors[previous].contains(&id) + }) + }) + .map(|frontier| { + let mut next = frontier.clone(); + next.push(id); + next + }) + .collect::>(); + if additions.len() > max_candidates.saturating_sub(frontiers.len()) { + return Err(invalid( + "materialization frontier search exceeds candidate budget", + )); + } + frontiers.extend(additions); + } + Ok(frontiers) +} + +/// Lower every maintenance candidate before feasibility/cost evaluation. Keep +/// individual failures visible; do not substitute another computation on error. +pub fn compile_candidates( + dag: &PostAsapDag, + inputs: BTreeMap, + roots: &[NodeId], + frontiers: &[Vec], +) -> Vec> { + frontiers + .iter() + .map(|frontier| compile_candidate(dag, inputs.clone(), roots, frontier)) + .collect() +} + +/// Complete workload cost supplied by scoped optimizer/deployment evidence. +/// The evaluator includes build/update work, retained state, shared producers +/// and recurrent reads over the same horizon; these are not per-query timings. +#[derive(Clone, Debug)] +pub struct CandidateCost { + pub workload_scope: String, + pub horizon_seconds: f64, + pub total_cost: f64, +} + +pub struct CandidateSelection { + pub candidate: T, + pub candidate_index: usize, + pub cost: CandidateCost, +} + +/// Select only compiled and deployment-feasible physical candidates. `None` +/// rejects an unbindable candidate before pricing. Comparable scoped costs are +/// required; deployment never rewrites the selected frontier after this step. +/// The payload is generic so deployments can retain binding/diagnostic metadata +/// alongside each compiled computation without duplicating winner selection. +pub fn select_candidate( + candidates: Vec>, + mut evaluate: impl FnMut(&T) -> Result, Error>, +) -> Result, Error> { + let mut scope: Option<(String, f64)> = None; + let mut selected: Option> = None; + for (candidate_index, candidate) in candidates.into_iter().enumerate() { + let Ok(candidate) = candidate else { continue }; + let Some(cost) = evaluate(&candidate)? else { + continue; + }; + if cost.workload_scope.is_empty() + || !cost.horizon_seconds.is_finite() + || cost.horizon_seconds <= 0. + || !cost.total_cost.is_finite() + || cost.total_cost < 0. + { + return Err(invalid( + "candidate cost lacks a valid workload scope/horizon", + )); + } + let current_scope = (cost.workload_scope.clone(), cost.horizon_seconds); + if scope.as_ref().is_some_and(|scope| scope != ¤t_scope) { + return Err(invalid( + "candidate costs describe different workloads or horizons", + )); + } + scope = Some(current_scope); + if selected + .as_ref() + .is_none_or(|selected| cost.total_cost < selected.cost.total_cost) + { + selected = Some(CandidateSelection { + candidate, + candidate_index, + cost, + }); + } + } + selected.ok_or_else(|| invalid("no feasible priced physical candidate")) +} + +#[derive(serde::Deserialize)] +#[serde(deny_unknown_fields)] +struct UncheckedCandidate { + precompute: Option, + query: CompiledPhysicalDag, + materialized_outputs: BTreeMap, +} +impl TryFrom for PhysicalCandidate { + type Error = Error; + fn try_from(candidate: UncheckedCandidate) -> Result { + let result = Self { + precompute: candidate.precompute, + query: candidate.query, + materialized_outputs: candidate.materialized_outputs, + }; + result.validate()?; + Ok(result) + } +} + +impl PhysicalCandidate { + /// Validate the physical handoff, including the producer/reader boundary. + pub fn validate(&self) -> Result<(), Error> { + self.query.validate()?; + let Some(precompute) = &self.precompute else { + return if self.materialized_outputs.is_empty() { + Ok(()) + } else { + Err(invalid("materialized outputs have no producer DAG")) + }; + }; + precompute.validate()?; + let outputs: BTreeSet<_> = self.materialized_outputs.keys().copied().collect(); + if outputs.is_empty() || outputs != precompute.roots().iter().copied().collect() { + return Err(invalid("physical frontier differs from precompute outputs")); + } + let readers: BTreeMap<_, _> = self.query.input_contracts().collect(); + for (&id, contract) in &self.materialized_outputs { + let produced = precompute.output_contract(id)?; + // Direct frontiers retain their node IDs. Temporal candidates can + // read several window instances through distinct input slots; + // their deployment bindings must validate those slots separately. + let reader = readers.get(&id); + if contract.schema != produced.schema + || reader.is_some_and(|reader| contract.schema != reader.schema) + || produced.properties.boundedness != Boundedness::Bounded + || contract.properties.boundedness != Boundedness::Bounded + || reader + .is_some_and(|reader| reader.properties.boundedness != Boundedness::Bounded) + { + return Err(invalid("physical frontier schema or boundedness mismatch")); + } + } + Ok(()) + } +} diff --git a/crates/asap-physical-operators/src/physical_planner/compiled.rs b/crates/asap-physical-operators/src/physical_planner/compiled.rs new file mode 100644 index 00000000..70d6a9ff --- /dev/null +++ b/crates/asap-physical-operators/src/physical_planner/compiled.rs @@ -0,0 +1,310 @@ +//! Reader-independent physical computation and checked deployment instantiation. +use super::*; + +/// A typed execution boundary, without storage identity or a live reader. +#[derive(Clone, Debug, serde::Serialize, serde::Deserialize)] +pub struct InputContract { + pub schema: Schema, + pub properties: PlanProperties, +} +impl InputContract { + pub fn bounded(schema: Schema) -> Self { + Self { + schema, + properties: PlanProperties { + boundedness: Boundedness::Bounded, + emission: Emission::Unknown, + }, + } + } + pub fn from_source(source: &dyn PhysicalOperator) -> Self { + Self { + schema: source.output_schema(), + properties: source.properties(&[]), + } + } +} +#[derive(Clone, serde::Serialize, serde::Deserialize)] +enum Node { + Input(InputContract), + Operator { + inputs: Vec, + operator: Operator, + }, +} + +/// Selected native operators and input slots. Rebinding never repeats lowering. +/// Serde is format-agnostic; deployments choose the encoding and its versioning. +/// Deserialization validates the graph before it is usable. +#[derive(Clone, serde::Serialize, serde::Deserialize)] +#[serde(try_from = "UncheckedDag")] +pub struct CompiledPhysicalDag { + nodes: BTreeMap, + roots: Vec, +} +#[derive(serde::Deserialize)] +#[serde(deny_unknown_fields)] +struct UncheckedDag { + nodes: BTreeMap, + roots: Vec, +} +impl TryFrom for CompiledPhysicalDag { + type Error = Error; + fn try_from(dag: UncheckedDag) -> Result { + let result = Self { + nodes: dag.nodes, + roots: dag.roots, + }; + result.validate()?; + Ok(result) + } +} + +impl CompiledPhysicalDag { + /// Link already-selected physical fragments without lowering operators again. + /// Fragment keys and source keys share a namespace; repeated dependency IDs + /// therefore remain one producer in the composed graph. + pub fn compose( + sources: BTreeMap, + fragments: BTreeMap, Self)>, + roots: Vec, + ) -> Result { + if sources.keys().any(|id| fragments.contains_key(id)) { + return Err(invalid("physical source and fragment IDs overlap")); + } + let mut contracts = sources.clone(); + for (&id, (_, fragment)) in &fragments { + fragment.validate()?; + let [root] = fragment.roots() else { + return Err(invalid("composed fragment requires one root")); + }; + if fragment.input_contracts().any(|(id, _)| id == *root) { + return Err(invalid("fragment root must be a computed output")); + } + contracts.insert(id, fragment.output_contract(*root)?); + } + let mut next = contracts + .keys() + .next_back() + .copied() + .unwrap_or(0) + .checked_add(1) + .ok_or_else(|| invalid("physical node ID overflow"))?; + let mut result = Self::new(roots); + for (id, contract) in sources { + result.add_input(id, contract)?; + } + for (id, (inputs, fragment)) in fragments { + if inputs.len() != fragment.input_contracts().count() { + return Err(invalid("physical fragment input arity mismatch")); + } + let mut mapping = BTreeMap::new(); + for ((local, expected), global) in fragment.input_contracts().zip(inputs) { + let actual = contracts + .get(&global) + .ok_or_else(|| invalid("missing physical fragment dependency"))?; + if expected.schema != actual.schema + || (expected.properties.boundedness == Boundedness::Bounded + && actual.properties.boundedness != Boundedness::Bounded) + { + return Err(invalid("physical fragment dependency contract mismatch")); + } + mapping.insert(local, global); + } + mapping.insert(fragment.roots[0], id); + for local in fragment.nodes.keys() { + if !mapping.contains_key(local) { + mapping.insert(*local, next); + next = next + .checked_add(1) + .ok_or_else(|| invalid("physical node ID overflow"))?; + } + } + for (local, node) in fragment.nodes { + if let Node::Operator { inputs, operator } = node { + result.add( + mapping[&local], + inputs.into_iter().map(|input| mapping[&input]).collect(), + operator, + )?; + } + } + } + result.validate()?; + Ok(result) + } + + /// Assemble already-lowered operators and typed external inputs. This is + /// useful for engines that compose multiple compiled computation fragments. + pub fn from_operators( + inputs: BTreeMap, + operators: BTreeMap, Operator)>, + roots: Vec, + ) -> Result { + let mut result = Self::new(roots); + for (id, contract) in inputs { + result.add_input(id, contract)?; + } + for (id, (inputs, operator)) in operators { + result.add(id, inputs, operator)?; + } + result.validate()?; + Ok(result) + } + pub(super) fn new(roots: Vec) -> Self { + Self { + nodes: BTreeMap::new(), + roots, + } + } + pub(super) fn add_input(&mut self, id: NodeId, contract: InputContract) -> Result<(), Error> { + self.insert(id, Node::Input(contract)) + } + pub(super) fn add( + &mut self, + id: NodeId, + inputs: Vec, + operator: Operator, + ) -> Result<(), Error> { + self.insert(id, Node::Operator { inputs, operator }) + } + fn insert(&mut self, id: NodeId, node: Node) -> Result<(), Error> { + if self.nodes.insert(id, node).is_some() { + return Err(invalid(format!("duplicate physical node {id}"))); + } + Ok(()) + } + /// Identify the external input whose rows survive unchanged at this output. + /// Protocol adapters can retain labels that are outside a closed physical schema. + pub fn row_source(&self, id: NodeId) -> Option { + match self.nodes.get(&id)? { + Node::Input(_) => Some(id), + Node::Operator { inputs, operator } => { + let index = operator.row_preserving_input()?; + self.row_source(*inputs.get(index)?) + } + } + } + + /// Selected operator name, for plan inspection without decoding its wire format. + /// Certified candidate pruning checks authoritative-key coverage inside this operator. + pub fn certified_pruning_keys(&self, id: NodeId) -> Option<&[(usize, usize)]> { + match self.nodes.get(&id)? { + Node::Operator { operator, .. } => operator.certified_pruning_keys(), + Node::Input(_) => None, + } + } + pub fn operator_name(&self, id: NodeId) -> Option<&str> { + match self.nodes.get(&id)? { + Node::Input(_) => Some("Input"), + Node::Operator { operator, .. } => Some(operator.name()), + } + } + + pub fn roots(&self) -> &[NodeId] { + &self.roots + } + pub fn input_contracts(&self) -> impl Iterator { + self.nodes.iter().filter_map(|(&id, node)| match node { + Node::Input(contract) => Some((id, contract)), + Node::Operator { .. } => None, + }) + } + /// Derive a reachable output contract without opening deployment readers. + pub fn output_contract(&self, id: NodeId) -> Result { + let sources = self + .input_contracts() + .map(|(id, contract)| (id, Box::new(contract.clone()) as Source<'_>)) + .collect(); + let graph = self.instantiate(sources)?; + let properties = graph.properties(&self.roots)?; + let properties = *properties + .get(&id) + .ok_or_else(|| invalid("output is not reachable"))?; + let schema = match self + .nodes + .get(&id) + .ok_or_else(|| invalid("missing output"))? + { + Node::Input(contract) => contract.schema.clone(), + Node::Operator { operator, .. } => operator.output_schema(), + }; + Ok(InputContract { schema, properties }) + } + /// Validate using contract-only sources. No deployment reader is available. + pub fn validate(&self) -> Result<(), Error> { + let sources = self + .input_contracts() + .map(|(id, c)| (id, Box::new(c.clone()) as Source<'_>)) + .collect(); + self.instantiate(sources).map(|_| ()) + } + /// Resolve exactly the declared inputs and validate before any source starts. + pub fn instantiate<'a>( + &self, + mut sources: BTreeMap>, + ) -> Result, Error> { + let mut graph = PhysicalDag::default(); + for (&id, node) in &self.nodes { + match node { + Node::Input(contract) => { + let source = sources + .remove(&id) + .ok_or_else(|| invalid(format!("missing physical input {id}")))?; + let actual = source.properties(&[]); + if !source.input_schemas().is_empty() + || source.output_schema() != contract.schema + || (contract.properties.boundedness != Boundedness::Unknown + && actual.boundedness != contract.properties.boundedness) + || (contract.properties.emission != Emission::Unknown + && actual.emission != contract.properties.emission) + { + return Err(invalid(format!( + "physical input {id} violates its compiled contract" + ))); + } + graph.add_boxed( + id, + vec![], + Box::new(CheckedSource { + source, + output: contract.schema.clone(), + }), + )?; + } + Node::Operator { inputs, operator } => { + graph.add(id, inputs.clone(), operator.clone())?; + } + } + } + if !sources.is_empty() { + return Err(invalid("unexpected physical input binding")); + } + graph.validate(&self.roots)?; + Ok(graph) + } +} +impl PhysicalOperator for InputContract { + fn name(&self) -> &str { + "UnresolvedInput" + } + fn input_schemas(&self) -> Vec { + vec![] + } + fn output_schema(&self) -> Schema { + self.schema.clone() + } + fn properties(&self, _: &[PlanProperties]) -> PlanProperties { + self.properties + } + fn output_bytes(&self, batch: &Batch) -> usize { + batch.bytes() + } + fn start<'a>( + &'a self, + _: Vec>, + _: crate::runtime::RunContext, + ) -> Result, Error> { + Err(invalid("physical input must be resolved before execution")) + } +} diff --git a/crates/asap-physical-operators/src/physical_planner/mod.rs b/crates/asap-physical-operators/src/physical_planner/mod.rs new file mode 100644 index 00000000..142a2c25 --- /dev/null +++ b/crates/asap-physical-operators/src/physical_planner/mod.rs @@ -0,0 +1,900 @@ +//! Compile logical computation to native operators with typed external inputs. +//! Compilation needs no readers; deployment resolves inputs after selection. +use crate::operators::ReadoutQuery; +use crate::summary_kernels::exact::ExactReadout; +use crate::{ + operators::{Expression, Operator, Reduction, SortKey}, + plan::{Boundedness, Emission, NodeId, PhysicalDag, PhysicalOperator, PlanProperties}, + values::{Batch, Schema}, + Error, +}; +use planner_types::{ + post_asap::{ + ExactOperation, PostAsapDag, PostAsapDagNode, PostAsapOperatorPayload as Payload, + SketchQuery, SummaryFamilyType, SummaryInputExpr, ValueOperation, + }, + pre_asap::{ + AggIntent, ColumnRef, CompareOpKind, GroupKeys, QueryExpr, Reduction as PlannerReduction, + }, +}; +use std::{ + collections::{BTreeMap, BTreeSet}, + sync::Arc, +}; +fn invalid(message: impl Into) -> Error { + Error::Invalid(message.into()) +} + +/// Source nodes cut the DAG at an installed storage/ingestion frontier. The +/// binding must have exactly the declared schema and no upstream dependencies. +/// A deployment must authorize these frontiers before calling this function. +pub type Source<'a> = Box + 'a>; + +pub mod precompute; +pub mod promql_rows; +pub mod promql_values; + +mod candidates; +pub use candidates::{ + compile_candidate, compile_candidates, enumerate_frontiers, select_candidate, CandidateCost, + CandidateSelection, PhysicalCandidate, +}; + +mod temporal_panes; +pub use temporal_panes::{ + compile_temporal_pane_candidate, TemporalEntityIdentity, TemporalPaneCandidate, + TemporalPaneMaintenance, +}; + +mod compiled; +pub use compiled::{CompiledPhysicalDag, InputContract}; + +/// Compile computation without opening or retaining deployment readers. +/// Input contracts identify explicit boundaries selected by maintenance planning. +pub fn compile( + dag: &PostAsapDag, + inputs: BTreeMap, + roots: &[NodeId], +) -> Result { + compile_internal(dag, inputs, roots) +} + +/// Convenience for callers that already resolved inputs. Lowering still uses +/// only their contracts, and instantiation checks those contracts again. +pub fn bind<'a>( + dag: &PostAsapDag, + sources: BTreeMap>, + roots: &[NodeId], +) -> Result, Error> { + let inputs = sources + .iter() + .map(|(&id, source)| (id, InputContract::from_source(source.as_ref()))) + .collect(); + compile(dag, inputs, roots)?.instantiate(sources) +} + +/// Resolve raw scan connectors before invoking the reader-independent compiler. +pub fn bind_with_data_sources<'a>( + dag: &PostAsapDag, + mut sources: BTreeMap>, + roots: &[NodeId], + data_sources: &crate::sources::DataSources, +) -> Result, Error> { + // Only resolve scans reachable below the selected input boundaries. + let mut pending = roots.to_vec(); + let mut seen = BTreeSet::new(); + while let Some(id) = pending.pop() { + if !seen.insert(id) || sources.contains_key(&id) { + continue; + } + let node = dag + .nodes + .iter() + .find(|n| u64::from(n.id.0) == id) + .ok_or_else(|| invalid(format!("missing node {id}")))?; + if let Payload::Fallback { + expression: expression @ QueryExpr::Scan { .. }, + } = &node.payload + { + sources.insert(id, Box::new(data_sources.bind(expression)?)); + } else { + pending.extend( + dag.edges + .iter() + .filter(|e| u64::from(e.consumer.0) == id) + .map(|e| u64::from(e.producer.0)), + ); + } + } + bind(dag, sources, roots) +} + +fn compile_internal( + dag: &PostAsapDag, + mut sources: BTreeMap, + roots: &[NodeId], +) -> Result { + preflight_depth(dag)?; + dag.validate().map_err(|e| invalid(e.to_string()))?; + let nodes = dag + .nodes + .iter() + .map(|node| (u64::from(node.id.0), node)) + .collect::>(); + let mut dependencies = BTreeMap::>::new(); + // Binary input order is semantic; serialized edge order is not. + let mut edges = dag.edges.iter().collect::>(); + edges.sort_by_key(|edge| { + ( + edge.consumer.0, + match edge.role { + planner_types::post_asap::EdgeRole::Left => 0, + planner_types::post_asap::EdgeRole::Input => 1, + planner_types::post_asap::EdgeRole::Right => 2, + }, + ) + }); + for edge in edges { + dependencies + .entry(u64::from(edge.consumer.0)) + .or_default() + .push(u64::from(edge.producer.0)); + } + if sources.keys().any(|id| !nodes.contains_key(id)) { + return Err(invalid("source binding names an unknown node")); + } + let mut ordered = Vec::new(); + let mut seen = BTreeSet::new(); + let mut pending = roots.iter().map(|&id| (id, false)).collect::>(); + while let Some((id, expanded)) = pending.pop() { + if expanded { + ordered.push(id); + continue; + } + if !seen.insert(id) { + continue; + } + if !nodes.contains_key(&id) { + return Err(invalid(format!("missing root {id}"))); + } + pending.push((id, true)); + if !sources.contains_key(&id) { + for &input in dependencies.get(&id).into_iter().flatten() { + pending.push((input, false)); + } + } + } + let mut graph = CompiledPhysicalDag::new(roots.to_vec()); + let mut auxiliary = u64::MAX; + for id in ordered { + let node = nodes[&id]; + let output = Arc::new(node.output_schema.clone()); + crate::values::validate_schema(&output)?; + if let Some(source) = sources.remove(&id) { + if source.schema != output { + return Err(invalid("frontier does not have the declared schema")); + } + graph.add_input(id, source)?; + } else { + let mut inputs = dependencies.get(&id).cloned().unwrap_or_default(); + let mut schemas = inputs + .iter() + .map(|id| Arc::new(nodes[id].output_schema.clone())) + .collect::>(); + if matches!(node.payload, Payload::SummaryMerge) && inputs.len() > 1 { + if schemas.iter().any(|s| s != &schemas[0]) { + return Err(invalid("summary merge inputs have different schemas")); + } + graph.add( + auxiliary, + inputs, + Operator::union(schemas[0].clone(), schemas.len())?, + )?; + inputs = vec![auxiliary]; + auxiliary -= 1; + schemas.truncate(1); + } + if let Payload::Value { + operation: ValueOperation::MaintainPopulation { population }, + } = &node.payload + { + use planner_types::post_asap::maintained_population::PopulationInput; + let PopulationInput::CurrentSeries(spec) = &population.input else { + return Err(invalid( + "native maintained population requires a current-series input", + )); + }; + let [input] = schemas.as_slice() else { + return Err(invalid("current-series population requires one input")); + }; + if spec.without { + return Err(invalid( + "dynamic without grouping requires label-set projection", + )); + } + let identity = named_column( + input, + &ColumnRef::Named(promql_rows::SERIES_IDENTITY_COLUMN.into()), + )?; + let coordinate = input + .time_index + .ok_or_else(|| invalid("current-series input lacks timestamp"))?; + let value = named_column(input, &ColumnRef::SampleValue)?; + let lookback = i64::try_from(spec.lookback_ms) + .map_err(|_| invalid("current-series lookback overflows"))?; + graph.add( + id, + inputs, + Operator::current_series(input.clone(), identity, coordinate, value, lookback)? + .with_output_schema(output)?, + )?; + continue; + } + if let Payload::Value { + operation: ValueOperation::ReadPopulation { readout }, + } = &node.payload + { + use planner_types::post_asap::maintained_population::{ + PopulationInput, PopulationReadout, + }; + let PopulationReadout::TopK { k } = readout else { + return Err(invalid( + "native population readout does not support this operation", + )); + }; + let [producer] = inputs.as_slice() else { + return Err(invalid("population readout requires one input")); + }; + let Payload::Value { + operation: ValueOperation::MaintainPopulation { population }, + } = &nodes[producer].payload + else { + return Err(invalid( + "population readout requires its declared population", + )); + }; + let PopulationInput::CurrentSeries(spec) = &population.input else { + return Err(invalid("current-series population required")); + }; + if spec.without { + return Err(invalid( + "dynamic without ranking requires label-set projection", + )); + } + let input = schemas[0].clone(); + let groups = spec + .grouping + .iter() + .map(|name| named_column(&input, &ColumnRef::Named(name.clone()))) + .collect::, _>>()?; + let value = named_column(&input, &ColumnRef::SampleValue)?; + graph.add( + auxiliary, + inputs, + Operator::sort( + input.clone(), + vec![SortKey { + column: value, + descending: true, + nulls_first: false, + }], + groups.clone(), + )?, + )?; + graph.add( + id, + vec![auxiliary], + Operator::limit(input, *k as u64, 0, groups)?.with_output_schema(output)?, + )?; + auxiliary -= 1; + continue; + } + // A closed row must include either all source labels or the explicit + // complete-label identity. Projected labels alone are insufficient. + if let Payload::SummaryAgg { + family, + input: update, + reduction: PlannerReduction::PerEntity, + grouping, + } = &node.payload + { + let [input_id] = inputs.as_slice() else { + return Err(invalid("per-entity summary requires one input")); + }; + let Payload::Fallback { + expression: QueryExpr::TimeRange { child, .. }, + } = &nodes[input_id].payload + else { + return Err(invalid( + "per-entity summary requires a resolved raw time range", + )); + }; + let QueryExpr::Scan { schema, .. } = child.as_ref() else { + return Err(invalid("per-entity summary requires a resolved source")); + }; + if !schema.closed || update.item.is_some() { + return Err(invalid( + "per-entity summary requires complete source identity", + )); + } + crate::capability::validate_summary_kernel(family, update, grouping) + .map_err(Error::Invalid)?; + let SummaryInputExpr::Column(value) = &update.weight else { + return Err(invalid( + "per-entity update requires a projected value column", + )); + }; + let input = schemas[0].clone(); + let value = named_column(&input, value)?; + let coordinate = input + .time_index + .ok_or_else(|| invalid("temporal input lacks time"))?; + let groups = (0..input.fields.len()) + .filter(|&column| column != value && column != coordinate) + .collect(); + let build = Operator::summary_build( + input, + family.clone(), + value, + Some(coordinate), + groups, + )?; + let compact = build.schema(); + graph.add(auxiliary, inputs, build)?; + graph.add( + id, + vec![auxiliary], + Operator::scope_timestamp(compact, output)?, + )?; + auxiliary -= 1; + continue; + } + let mut operator = compile_node(node, &schemas) + .map_err(|error| invalid(format!("node {id}: {error}")))?; + if operator.is_counter_readout() { + let mut pending = vec![id]; + let mut visited = BTreeSet::new(); + let mut ranges = BTreeSet::new(); + while let Some(ancestor) = pending.pop() { + if !visited.insert(ancestor) { + continue; + } + if let Payload::Fallback { + expression: QueryExpr::TimeRange { range, .. }, + } = &nodes[&ancestor].payload + { + ranges.insert( + i64::try_from(range.as_millis()) + .map_err(|_| invalid("counter lookback exceeds Int64"))?, + ); + continue; + } + pending.extend(dependencies.get(&ancestor).into_iter().flatten().copied()); + } + if ranges.len() > 1 { + return Err(invalid("counter readout has ambiguous logical windows")); + } + if let Some(lookback) = ranges.into_iter().next() { + operator = operator.with_counter_lookback(lookback)?; + } + } + graph.add(id, inputs, operator)?; + } + } + graph.validate()?; + Ok(graph) +} + +/// Bind a Planner node against the schemas supplied by its deployment edges. +/// This is the same checked path used by complete DAG binding. +pub fn compile_node(node: &PostAsapDagNode, inputs: &[Schema]) -> Result { + for schema in inputs { + crate::values::validate_schema(schema)?; + } + bind_operation(node, inputs)?.with_output_schema(Arc::new(node.output_schema.clone())) +} + +fn bind_operation(node: &PostAsapDagNode, inputs: &[Schema]) -> Result { + if let Payload::Binary { operator } = &node.payload { + let [left, right] = inputs else { + return Err(invalid("binary requires two inputs")); + }; + if node.output_state.timing == planner_types::post_asap::ExecutionTiming::IngestionTime { + let value = |schema: &Schema| -> Result { + let columns = schema + .fields + .iter() + .enumerate() + .filter(|(_, field)| { + field.dtype + == SummaryFamilyType::Plain(planner_types::pre_asap::DataType::Float64) + }) + .map(|(i, _)| i) + .collect::>(); + match columns.as_slice() { + [value] => Ok(*value), + _ => Err(invalid("aligned binary requires one value column")), + } + }; + let (l, r) = (value(left)?, value(right)?); + let keys = left + .fields + .iter() + .enumerate() + .filter(|(i, _)| *i != l) + .map(|(i, field)| { + right + .fields + .iter() + .position(|other| other.name == field.name && other.dtype == field.dtype) + .map(|j| (i, j)) + .ok_or_else(|| invalid("aligned input identities differ")) + }) + .collect::, _>>()?; + return Operator::aligned_binary( + left.clone(), + right.clone(), + keys, + (l, r), + operator.clone(), + ); + } + return Operator::vector_binary(left.clone(), right.clone(), operator.clone(), false); + } + if let Payload::RelationalJoin { + join_kind, + pred, + pruning, + } = &node.payload + { + use planner_types::{post_asap::CandidateCompleteness, pre_asap::JoinKind}; + if pruning.is_some() && *join_kind != JoinKind::Semi { + return Err(invalid("pruning certificate requires a semi-join")); + } + if matches!(pruning,Some(CandidateCompleteness::Certified { guarantee }) if guarantee.has_unknown() || guarantee.metric != planner_types::post_asap::ErrorMetric::TopKMembership) + { + return Err(invalid("invalid pruning certificate")); + } + let [left, right] = inputs else { + return Err(invalid("join requires two inputs")); + }; + if *join_kind == JoinKind::Semi { + if let Ok(keys) = equijoin_keys(pred, left, right) { + let operator = Operator::semi_join(left.clone(), right.clone(), keys)?; + return Ok( + if matches!(pruning, Some(CandidateCompleteness::Certified { .. })) { + operator.require_complete_right() + } else { + operator + }, + ); + } + } + if matches!(pruning, Some(CandidateCompleteness::Certified { .. })) { + return Err(invalid("certified pruning requires explicit equijoin keys")); + } + return Operator::relational_join( + left.clone(), + right.clone(), + join_kind.clone(), + pred, + Arc::new(node.output_schema.clone()), + ); + } + let [input] = inputs else { + return Err(invalid( + "native Planner binding currently requires a unary operation or an explicit source", + )); + }; + match &node.payload { + Payload::Value { operation, .. } => match operation { + ValueOperation::Project { cols, .. } => Operator::project( + input.clone(), + cols.iter() + .enumerate() + .map(|(i, col)| { + Ok(( + node.output_schema + .fields + .get(i) + .ok_or_else(|| invalid("projection width mismatch"))? + .name + .clone(), + match &col.expr { + QueryExpr::Column(index) => Expression::Column(*index), + expr => expression(expr, input)?, + }, + )) + }) + .collect::>()?, + ), + ValueOperation::Filter { pred } => { + Operator::filter(input.clone(), expression(&pred.0, input)?) + } + ValueOperation::Sort { keys, partition_by } => Operator::sort( + input.clone(), + keys.iter() + .map(|key| { + let QueryExpr::Column(column) = key.expr else { + return Err(invalid( + "sort expression must be projected before sorting", + )); + }; + Ok(SortKey { + column, + descending: !key.ascending, + nulls_first: key.nulls_first, + }) + }) + .collect::>()?, + groups(input, partition_by)?, + ), + ValueOperation::Limit { + n, + offset, + partition_by, + } => Operator::limit( + input.clone(), + *n as u64, + *offset as u64, + groups(input, partition_by)?, + ), + ValueOperation::Exact(ExactOperation::Aggregate { + reduction, + measures, + output_names, + having: None, + }) => { + if measures.len() != output_names.len() { + return Err(invalid("aggregate output names differ from measures")); + } + let PlannerReduction::Reduce(keys) = reduction else { + return Err(invalid( + "per-entity aggregate requires an explicit entity binding", + )); + }; + let measures = measures + .iter() + .zip(output_names) + .map(|(m, name)| { + let column = |col: Option| { + col.map(Ok) + .unwrap_or_else(|| named_column(input, &ColumnRef::SampleValue)) + }; + let m = match m { + AggIntent::Count { .. } => Reduction::Count, + AggIntent::Sum { col } => Reduction::Sum(column(*col)?), + AggIntent::Avg { col } => Reduction::Avg(column(*col)?), + AggIntent::Min { col } => Reduction::Min(column(*col)?), + AggIntent::Max { col } => Reduction::Max(column(*col)?), + _ => { + return Err(invalid( + "aggregate intent has no native implementation", + )) + } + }; + Ok((name.clone(), m)) + }) + .collect::>()?; + Operator::aggregate(input.clone(), groups(input, keys)?, measures) + } + ValueOperation::FinalizeExactAccumulator => { + let state = summary_column(input)?; + use crate::Statistic as S; + use planner_types::post_asap::ExactKind as E; + let statistic = match &input.fields[state].dtype { + SummaryFamilyType::ExactAggregate(kind, _) => match kind { + E::Sum => S::Sum, + E::Count => S::Count, + E::Min => S::Min, + E::Max => S::Max, + E::Rate => S::Rate, + E::Increase => S::Increase, + _ => return Err(invalid("exact family readout is unsupported")), + }, + _ => return Err(invalid("exact finalization requires exact state")), + }; + Operator::readout( + input.clone(), + state, + ReadoutQuery::Exact(ExactReadout { + statistic, + lookback_ms: None, + }), + ) + } + _ => Err(invalid("value operation has no native implementation")), + }, + Payload::SummaryAgg { + family, + input: update, + reduction, + grouping, + } => { + if let Some(item) = &update.item { + let PlannerReduction::Reduce(keys) = reduction else { + return Err(invalid("keyed summary requires explicit partitions")); + }; + let SummaryInputExpr::Column(weight) = &update.weight else { + return Err(invalid( + "keyed summary weight must be a finalized value column", + )); + }; + if matches!(family, SummaryFamilyType::Sketch(kind, _) if kind.algorithm() == &planner_types::post_asap::SketchAlgorithm::CmsWithHeap) + && !matches!( + update.weight_domain, + planner_types::post_asap::WeightDomain::NonNegative { .. } + ) + { + return Err(invalid("CMS requires a nonnegative weight contract")); + } + fn columns( + expr: &SummaryInputExpr, + input: &Schema, + result: &mut Vec, + ) -> Result<(), Error> { + match expr { + SummaryInputExpr::Column(column) => { + result.push(named_column(input, column)?) + } + SummaryInputExpr::Tuple(items) => { + for item in items { + columns(item, input, result)?; + } + } + _ => return Err(invalid("keyed summary needs explicit item columns")), + } + Ok(()) + } + let mut items = Vec::new(); + columns(item, input, &mut items)?; + return Operator::keyed_summary_build( + input.clone(), + family.clone(), + named_column(input, weight)?, + items, + groups(input, keys)?, + ); + } + crate::capability::validate_summary_kernel(family, update, grouping) + .map_err(Error::Invalid)?; + let SummaryInputExpr::Column(column) = &update.weight else { + return Err(invalid( + "summary update expression must be projected to a column", + )); + }; + let PlannerReduction::Reduce(keys) = reduction else { + return Err(invalid( + "summary construction requires explicit grouping columns", + )); + }; + Operator::summary_build( + input.clone(), + family.clone(), + named_column(input, column)?, + input.time_index, + groups(input, keys)?, + ) + } + Payload::SummaryMerge => { + let state = summary_column(input)?; + Operator::summary_merge( + input.clone(), + state, + (0..input.fields.len()) + .filter(|&i| i != state && Some(i) != input.time_index) + .collect(), + ) + } + Payload::SummaryEstimate { query } => { + if let SketchQuery::TopK { k } = query { + return Operator::keyed_readout( + input.clone(), + summary_column(input)?, + *k, + Arc::new(node.output_schema.clone()), + ); + } + Operator::readout( + input.clone(), + summary_column(input)?, + ReadoutQuery::Sketch(query.clone()), + ) + } + _ => Err(invalid( + "physical operation has no native binding; no fallback is installed", + )), + } +} +fn summary_column(input: &Schema) -> Result { + let columns = input + .fields + .iter() + .enumerate() + .filter(|(_, f)| !matches!(f.dtype, SummaryFamilyType::Plain(_))) + .map(|(i, _)| i) + .collect::>(); + match columns.as_slice() { + [column] => Ok(*column), + _ => Err(invalid("one summary state column required")), + } +} +fn named_column(input: &Schema, column: &ColumnRef) -> Result { + let name = match column { + // Executable SummarySchema retains column names, not table qualifiers. + // Frontend binding has resolved the qualifier; still reject ambiguous + // names here rather than guessing a join side. + ColumnRef::Named(name) | ColumnRef::Qualified { name, .. } => name.as_str(), + ColumnRef::SampleValue => "value", + _ => { + return Err(invalid( + "summary update requires an unambiguous bound column", + )) + } + }; + let matches = input + .fields + .iter() + .enumerate() + .filter(|(_, field)| field.name == name) + .map(|(i, _)| i) + .collect::>(); + match matches.as_slice() { + [column] => Ok(*column), + _ => Err(invalid("summary update column missing or ambiguous")), + } +} +fn groups(input: &Schema, groups: &GroupKeys) -> Result, Error> { + if groups.is_without() { + return Err(invalid("grouping without requires resolved label columns")); + } + if groups.keys().iter().any(|&i| i >= input.fields.len()) { + return Err(invalid("grouping column out of range")); + } + Ok(groups.keys().to_vec()) +} +fn expression(expr: &QueryExpr, input: &Schema) -> Result { + Ok(Expression::planner( + crate::expressions::CompiledExpression::compile(expr, input)?, + )) +} + +struct CheckedSource<'a> { + source: Source<'a>, + output: Schema, +} +impl PhysicalOperator for CheckedSource<'_> { + fn properties(&self, inputs: &[crate::plan::PlanProperties]) -> crate::plan::PlanProperties { + self.source.properties(inputs) + } + + fn name(&self) -> &str { + self.source.name() + } + fn input_schemas(&self) -> Vec { + vec![] + } + fn output_schema(&self) -> Schema { + self.output.clone() + } + fn output_bytes(&self, batch: &Batch) -> usize { + self.source.output_bytes(batch) + } + fn start<'a>( + &'a self, + inputs: Vec>, + context: crate::runtime::RunContext, + ) -> Result, Error> { + use futures::StreamExt; + Ok(self + .source + .start(inputs, context)? + .map(|batch| { + let batch = batch?; + if batch.schema() != &self.output { + return Err(invalid("source batch differs from its bound schema")); + } + Ok(batch) + }) + .boxed_local()) + } +} + +// Bound recursion before invoking the upstream recursive provenance validator. +fn preflight_depth(dag: &PostAsapDag) -> Result<(), Error> { + let mut remaining = dag + .nodes + .iter() + .map(|node| (node.id, 0usize)) + .collect::>(); + if remaining.len() != dag.nodes.len() { + return Err(invalid("duplicate Planner node")); + } + let mut consumers = BTreeMap::<_, Vec<_>>::new(); + for edge in &dag.edges { + if !remaining.contains_key(&edge.producer) { + return Err(invalid("missing Planner edge producer")); + } + *remaining + .get_mut(&edge.consumer) + .ok_or_else(|| invalid("missing Planner edge consumer"))? += 1; + consumers + .entry(edge.producer) + .or_default() + .push(edge.consumer); + } + let mut ready = remaining + .iter() + .filter(|(_, n)| **n == 0) + .map(|(id, _)| *id) + .collect::>(); + let mut depths = BTreeMap::new(); + let mut visited = 0; + while let Some(id) = ready.pop_front() { + visited += 1; + let depth = *depths.get(&id).unwrap_or(&1usize); + if depth > 128 { + return Err(invalid("DAG exceeds the supported execution depth of 128")); + } + for &consumer in consumers.get(&id).into_iter().flatten() { + let next = depths.entry(consumer).or_insert(1); + *next = (*next).max(depth + 1); + let count = remaining.get_mut(&consumer).expect("validated endpoint"); + *count -= 1; + if *count == 0 { + ready.push_back(consumer); + } + } + } + if visited != dag.nodes.len() { + return Err(invalid("Planner DAG contains a cycle")); + } + Ok(()) +} + +/// Join predicates address the concatenated left/right schema. +fn semi_join_keys( + expr: &QueryExpr, + left: usize, + right: usize, + keys: &mut Vec<(usize, usize)>, +) -> Result<(), Error> { + match expr { + QueryExpr::BoolAnd(parts) => { + for part in parts { + semi_join_keys(part, left, right, keys)?; + } + } + QueryExpr::Compare { + left: a, + op: CompareOpKind::Eq, + right: b, + } => { + let (QueryExpr::Column(a), QueryExpr::Column(b)) = (a.as_ref(), b.as_ref()) else { + return Err(invalid("semi-join requires column equality keys")); + }; + let (a, b) = if a < b { (*a, *b) } else { (*b, *a) }; + if a >= left || b < left || b >= left + right { + return Err(invalid("semi-join key must match left to right")); + } + keys.push((a, b - left)); + } + _ => return Err(invalid("unsupported semi-join predicate")), + } + Ok(()) +} + +/// Resolve equality keys against the Planner join's concatenated input schema. +/// Deployments may use these positions to bind their source columns. +pub fn equijoin_keys( + pred: &planner_types::pre_asap::Predicate, + left: &planner_types::post_asap::SummarySchema, + right: &planner_types::post_asap::SummarySchema, +) -> Result, Error> { + let mut keys = Vec::new(); + semi_join_keys(&pred.0, left.fields.len(), right.fields.len(), &mut keys)?; + if keys.is_empty() { + return Err(invalid("semi-join requires explicit matching keys")); + } + Ok(keys) +} diff --git a/crates/asap-physical-operators/src/physical_planner/precompute.rs b/crates/asap-physical-operators/src/physical_planner/precompute.rs new file mode 100644 index 00000000..063bd0b8 --- /dev/null +++ b/crates/asap-physical-operators/src/physical_planner/precompute.rs @@ -0,0 +1,384 @@ +//! Compile immutable summary-input computation with explicit population and pane identity. +use super::*; +use planner_types::{ + post_asap::{ExecutionTiming, GroupingStrategy, SummarySchema}, + pre_asap::DataType, +}; + +/// Physical rows carry the population and pane coordinate alongside the logical value. +/// These fields preserve identities which are implicit in a stored summary instance. +pub fn population_schema(family: SummaryFamilyType) -> Schema { + Arc::new(SummarySchema { + fields: vec![ + planner_types::post_asap::SummaryField { + name: "$population".into(), + dtype: SummaryFamilyType::Plain(DataType::Map { + key: Box::new(DataType::Utf8), + value: Box::new(DataType::Utf8), + value_nullable: false, + }), + nullable: false, + }, + planner_types::post_asap::SummaryField { + name: "$window_end".into(), + dtype: SummaryFamilyType::Plain(DataType::Timestamp), + nullable: false, + }, + planner_types::post_asap::SummaryField { + name: "value".into(), + dtype: family, + nullable: false, + }, + ], + time_index: Some(1), + }) +} + +/// Validate the adapter layout during installed-plan recovery without lowering operators. +pub fn source_schema(logical: &SummarySchema) -> Result { + let states = logical + .fields + .iter() + .filter(|f| !matches!(f.dtype, SummaryFamilyType::Plain(_))) + .collect::>(); + let [state] = states.as_slice() else { + return Err(invalid( + "stored population requires one typed summary state", + )); + }; + if logical.fields.iter().enumerate().any(|(i, field)| matches!(&field.dtype, SummaryFamilyType::Plain(dtype) + if field.nullable || !matches!(dtype, DataType::Utf8) && !(Some(i) == logical.time_index && *dtype == DataType::Timestamp))) { + return Err(invalid("stored population metadata cannot reconstruct extra value columns")); + } + if state.nullable { + return Err(invalid("stored population state cannot be null")); + } + Ok(population_schema(state.dtype.clone())) +} + +pub fn is_population_schema(schema: &Schema) -> bool { + schema + .fields + .get(2) + .is_some_and(|field| *schema == population_schema(field.dtype.clone())) +} + +/// Compile a complete selected precompute sub-DAG. Inputs are already-computed +/// state boundaries; the deployment supplies groups, panes and states, never operations. +pub fn compile( + dag: &PostAsapDag, + frontiers: &[NodeId], + roots: &[NodeId], +) -> Result { + preflight_depth(dag)?; + dag.validate().map_err(|e| invalid(e.to_string()))?; + let nodes = dag + .nodes + .iter() + .map(|n| (u64::from(n.id.0), n)) + .collect::>(); + let frontier = frontiers.iter().copied().collect::>(); + if frontier.len() != frontiers.len() || roots.iter().any(|r| frontier.contains(r)) { + return Err(invalid( + "precompute boundaries must be distinct from outputs", + )); + } + let mut dependencies = BTreeMap::>::new(); + let mut edges = dag.edges.iter().collect::>(); + edges.sort_by_key(|edge| { + ( + edge.consumer.0, + match edge.role { + planner_types::post_asap::EdgeRole::Left => 0, + planner_types::post_asap::EdgeRole::Input => 1, + planner_types::post_asap::EdgeRole::Right => 2, + }, + ) + }); + for edge in edges { + dependencies + .entry(u64::from(edge.consumer.0)) + .or_default() + .push(u64::from(edge.producer.0)); + } + let mut ordered = Vec::new(); + let mut seen = BTreeSet::new(); + let mut pending = roots.iter().map(|&id| (id, false)).collect::>(); + while let Some((id, expanded)) = pending.pop() { + if expanded { + ordered.push(id); + continue; + } + if !seen.insert(id) { + continue; + } + if !nodes.contains_key(&id) { + return Err(invalid("missing precompute node")); + } + pending.push((id, true)); + if !frontier.contains(&id) { + pending.extend( + dependencies + .get(&id) + .into_iter() + .flatten() + .map(|id| (*id, false)), + ); + } + } + let mut sources = BTreeMap::new(); + let mut fragments = BTreeMap::new(); + let mut outputs = BTreeMap::::new(); + for id in ordered { + let node = nodes[&id]; + if frontier.contains(&id) { + let schema = source_schema(&node.output_schema)?; + sources.insert(id, InputContract::bounded(schema.clone())); + outputs.insert(id, schema); + continue; + } + if node.output_state.timing != ExecutionTiming::IngestionTime { + return Err(invalid("precompute graph contains a query-time operation")); + } + let inputs = dependencies.get(&id).cloned().unwrap_or_default(); + let schemas = inputs + .iter() + .map(|id| { + outputs + .get(id) + .cloned() + .ok_or_else(|| invalid("missing precompute input")) + }) + .collect::, _>>()?; + let graph = fragment( + node, + &schemas, + &inputs.iter().map(|id| nodes[id]).collect::>(), + )?; + outputs.insert(id, graph.output_contract(graph.roots()[0])?.schema); + fragments.insert(id, (inputs, graph)); + } + CompiledPhysicalDag::compose(sources, fragments, roots.to_vec()) +} + +fn validate_value_output(node: &PostAsapDagNode) -> Result<(), Error> { + let schema = &node.output_schema; + let values = schema + .fields + .iter() + .enumerate() + .filter(|(i, _)| Some(*i) != schema.time_index) + .collect::>(); + if !matches!(values.as_slice(), [(_, field)] if !field.nullable && field.dtype == SummaryFamilyType::Plain(DataType::Float64)) + || schema.time_index.is_some_and(|i| { + schema.fields.get(i).is_none_or(|f| { + f.nullable || f.dtype != SummaryFamilyType::Plain(DataType::Timestamp) + }) + }) + { + return Err(invalid( + "precompute value schema requires Float64 and an optional declared timestamp", + )); + } + Ok(()) +} + +fn fragment( + node: &PostAsapDagNode, + schemas: &[Schema], + parents: &[&PostAsapDagNode], +) -> Result { + let sources = schemas + .iter() + .enumerate() + .map(|(id, schema)| (id as u64, InputContract::bounded(schema.clone()))) + .collect(); + let mut operators = BTreeMap::new(); + let mut next = schemas.len() as u64; + let mut add = |inputs: Vec, op: Operator| -> Result { + let id = next; + next += 1; + operators.insert(id, (inputs, op)); + Ok(id) + }; + let root = match &node.payload { + Payload::Binary { operator } => { + validate_value_output(node)?; + if node.output_schema.time_index.is_none() + || parents.iter().any(|p| p.output_schema.time_index.is_none()) + { + return Err(invalid( + "precompute binary requires declared window timestamps", + )); + } + let [left, right] = schemas else { + return Err(invalid("precompute binary requires two inputs")); + }; + add( + vec![0, 1], + Operator::aligned_binary( + left.clone(), + right.clone(), + vec![(0, 0), (1, 1)], + (2, 2), + operator.clone(), + )?, + )? + } + Payload::Value { + operation: ValueOperation::FinalizeExactAccumulator, + } => { + let [input] = schemas else { + return Err(invalid("finalize requires one state input")); + }; + validate_value_output(node)?; + let statistic = match &input.fields[2].dtype { + SummaryFamilyType::ExactAggregate(planner_types::post_asap::ExactKind::Sum, _) => { + crate::Statistic::Sum + } + SummaryFamilyType::ExactAggregate( + planner_types::post_asap::ExactKind::Count, + _, + ) => crate::Statistic::Count, + _ => { + return Err(invalid( + "precompute finalization requires explicit Sum or Count semantics", + )) + } + }; + let read = Operator::readout( + input.clone(), + 2, + ReadoutQuery::Exact(ExactReadout { + statistic, + lookback_ms: None, + }), + )?; + let output = read.schema(); + let read = add(vec![0], read)?; + let project = Operator::project( + output, + vec![ + ("$population".into(), Expression::Column(0)), + ("$window_end".into(), Expression::Column(1)), + ( + "value".into(), + Expression::FiniteFloat64(Box::new(Expression::ExactFloat64(2))), + ), + ], + )? + .with_output_schema(population_schema(SummaryFamilyType::Plain( + DataType::Float64, + )))?; + add(vec![read], project)? + } + Payload::SummaryAgg { + family, + input: update, + reduction, + grouping, + } => { + let [input] = schemas else { + return Err(invalid("summary update requires one input")); + }; + if update.item.is_some() + || !matches!(grouping, GroupingStrategy::PerSubpopulationInstance) + { + return Err(invalid( + "precompute keyed/shared update needs its dedicated physical candidate", + )); + } + crate::capability::validate_summary_kernel(family, update, grouping) + .map_err(Error::Invalid)?; + let labels = match reduction { + PlannerReduction::PerEntity => Expression::Column(0), + PlannerReduction::Reduce(keys) => Expression::LabelSet { + column: 0, + labels: keys + .keys() + .iter() + .map(|key| { + parents[0] + .output_schema + .fields + .get(*key) + .filter(|field| { + !field.nullable + && field.dtype == SummaryFamilyType::Plain(DataType::Utf8) + }) + .map(|f| f.name.clone()) + .ok_or_else(|| { + invalid("summary grouping must identify population labels") + }) + }) + .collect::, _>>()?, + without: keys.is_without(), + }, + }; + let weight = match &update.weight { + SummaryInputExpr::Constant(value) => Expression::Literal { + value: crate::values::Value::Float64(*value), + dtype: DataType::Float64, + }, + SummaryInputExpr::Column(ColumnRef::SampleValue) => Expression::Column(2), + SummaryInputExpr::Column(ColumnRef::Named(name)) + if parents[0].output_schema.fields.iter().any(|f| { + f.name == *name && f.dtype == SummaryFamilyType::Plain(DataType::Float64) + }) => + { + Expression::Column(2) + } + _ => { + return Err(invalid( + "summary weight does not resolve to the input value", + )) + } + }; + let project = Operator::project( + input.clone(), + vec![ + ("$population".into(), labels), + ("$window_end".into(), Expression::Column(1)), + ("value".into(), Expression::FiniteFloat64(Box::new(weight))), + ], + )? + .with_output_schema(population_schema(SummaryFamilyType::Plain( + DataType::Float64, + )))?; + let projected = project.schema(); + let project = add(vec![0], project)?; + let build = Operator::summary_build(projected, family.clone(), 2, Some(1), vec![0])?; + let built = build.schema(); + let build = add(vec![project], build)?; + add( + vec![build], + Operator::scope_timestamp(built, population_schema(family.clone()))?, + )? + } + Payload::SummaryMerge => { + let Some(input) = schemas.first() else { + return Err(invalid("summary merge requires inputs")); + }; + if schemas.iter().any(|s| s != input) { + return Err(invalid("summary merge inputs differ")); + } + let union = add( + (0..schemas.len() as u64).collect(), + Operator::union(input.clone(), schemas.len())?, + )?; + let merge = Operator::summary_merge(input.clone(), 2, vec![0])?; + let merged = merge.schema(); + let merge = add(vec![union], merge)?; + add( + vec![merge], + Operator::scope_timestamp(merged, input.clone())?, + )? + } + _ => { + return Err(invalid( + "precompute operation has no native population implementation", + )) + } + }; + CompiledPhysicalDag::from_operators(sources, operators, vec![root]) +} diff --git a/crates/asap-physical-operators/src/physical_planner/promql_rows.rs b/crates/asap-physical-operators/src/physical_planner/promql_rows.rs new file mode 100644 index 00000000..cb680096 --- /dev/null +++ b/crates/asap-physical-operators/src/physical_planner/promql_rows.rs @@ -0,0 +1,389 @@ +//! A bounded PromQL source row carries the entire label set, not just labels +//! mentioned by the query. The source adapter owns this lossless encoding. +use super::*; +use planner_types::pre_asap::{Column, DataType, Source as LogicalSource}; +use std::rc::Rc; + +/// Not a legal PromQL label name, so it cannot shadow a user label. +pub use planner_types::pre_asap::schema::PROMQL_SERIES_IDENTITY as SERIES_IDENTITY_COLUMN; + +/// Canonical, reversible identity. JSON object encoding preserves label names, +/// empty values and escaping; sorting makes ingestion order irrelevant. +pub fn encode_series_identity(labels: &BTreeMap) -> Result { + serde_json::to_string(labels).map_err(|error| invalid(error.to_string())) +} + +pub fn decode_series_identity(encoded: &str) -> Result, Error> { + let labels: BTreeMap = + serde_json::from_str(encoded).map_err(|error| invalid(error.to_string()))?; + if encode_series_identity(&labels)? != encoded { + return Err(invalid("series identity is not canonically encoded")); + } + Ok(labels) +} + +/// Resolve the row representation before candidate search. `closed` describes +/// physical columns here: the final column contains every dynamic source label. +/// It does not assert that the query's projected labels are the full label set. +/// +/// This realization supports explicit `by` grouping and per-series computation. +/// Operators that rewrite or implicitly match dynamic label sets require their +/// own realization; they must not accidentally treat the opaque identity as a +/// user label or silently discard it. +pub fn with_series_identity(root: &QueryExpr) -> Result { + let mut root = root.clone(); + fn visit(node: &mut QueryExpr) -> Result<(), Error> { + use planner_types::pre_asap::Reduction; + match node { + QueryExpr::Scan { + source: LogicalSource::TimeSeries { .. }, + schema, + .. + } => { + if schema + .columns + .iter() + .any(|column| column.name == SERIES_IDENTITY_COLUMN) + { + return Err(invalid( + "source already contains a physical series identity", + )); + } + if schema.closed { + return Err(invalid( + "dynamic series identity requires an open PromQL source", + )); + } + schema + .columns + .push(Column::new(SERIES_IDENTITY_COLUMN, DataType::Utf8, false)); + schema.closed = true; + Ok(()) + } + QueryExpr::TimeRange { child, .. } | QueryExpr::Limit { child, .. } => { + visit(Rc::make_mut(child)) + } + QueryExpr::Aggregate { + child, reduction, .. + } => { + if matches!(reduction, Reduction::Reduce(keys) if keys.is_without()) { + return Err(invalid( + "dynamic without grouping requires label-set projection", + )); + } + visit(Rc::make_mut(child)) + } + QueryExpr::Sort { + child, + partition_by, + .. + } => { + if partition_by.is_without() { + return Err(invalid( + "dynamic without ranking requires label-set projection", + )); + } + visit(Rc::make_mut(child)) + } + _ => Err(invalid( + "operator has no dynamic series-identity realization", + )), + } + } + visit(&mut root)?; + root.output_schema() + .map_err(|error| invalid(error.to_string()))?; + Ok(root) +} + +/// Construct source rows only from full identities. The named label columns +/// are projections of that same identity and cannot independently redefine it. +pub fn series_row( + schema: &Schema, + labels: &BTreeMap, + timestamp: i64, + value: f64, +) -> Result, Error> { + use crate::values::Value; + let identity = encode_series_identity(labels)?; + let mut found = false; + let row = schema + .fields + .iter() + .enumerate() + .map(|(index, field)| { + if field.name == SERIES_IDENTITY_COLUMN { + if field.dtype != SummaryFamilyType::Plain(DataType::Utf8) + || field.nullable + || found + { + return Err(invalid("invalid series identity column")); + } + found = true; + Ok(Value::Utf8(identity.clone().into())) + } else if Some(index) == schema.time_index { + Ok(Value::Timestamp(timestamp)) + } else if field.name == "value" + && field.dtype == SummaryFamilyType::Plain(DataType::Float64) + { + Ok(Value::Float64(value)) + } else if field.dtype == SummaryFamilyType::Plain(DataType::Utf8) { + Ok(labels.get(&field.name).map_or_else( + || Value::Utf8("".into()), + |value| Value::Utf8(value.clone().into()), + )) + } else { + Err(invalid("unsupported PromQL source column")) + } + }) + .collect::, _>>()?; + if !found { + return Err(invalid("source lacks its full series identity")); + } + Ok(row) +} + +/// Compile the selected TopK computation above an existing maintained-population +/// source. The boundary supplies the complete eligible vector, not a truncated +/// TopK result; ranking remains a native physical operator. +pub fn compile_current_series_readout( + selected: &Rc, +) -> Result { + use planner_types::post_asap::{ + compile_post_asap_dag, maintained_population::PopulationReadout, SummaryField, + }; + let mut dag = compile_post_asap_dag(selected).map_err(|error| invalid(error.to_string()))?; + // Typed snapshot candidates already carry full identity throughout the DAG. + // Cut at the population output, preserving all selected heap/readout nodes. + let populations = dag.nodes.iter().filter(|node| matches!(&node.payload, + Payload::Value { operation: ValueOperation::MaintainPopulation { population } } + if matches!(population.input, planner_types::post_asap::maintained_population::PopulationInput::CurrentSeries(_)) + )).collect::>(); + if let [population] = populations.as_slice() { + if population + .output_schema + .fields + .iter() + .any(|field| field.name == SERIES_IDENTITY_COLUMN) + { + return compile( + &dag, + BTreeMap::from([( + u64::from(population.id.0), + InputContract::bounded(Arc::new(population.output_schema.clone())), + )]), + &[u64::from(dag.root.0)], + ); + } + } + if dag.nodes.len() != 3 + || !dag.nodes.iter().any(|node| { + node.id == dag.root + && matches!( + node.payload, + Payload::Value { + operation: ValueOperation::ReadPopulation { + readout: PopulationReadout::TopK { .. } + } + } + ) + }) + { + return Err(invalid( + "expected one selected current-series TopK computation", + )); + } + let mut frontier = None; + for node in &mut dag.nodes { + match &mut node.payload { + Payload::Fallback { expression } => { + *expression = with_series_identity(expression)?; + } + Payload::Value { + operation: ValueOperation::MaintainPopulation { .. }, + } => { + frontier = Some(u64::from(node.id.0)); + } + Payload::Value { + operation: + ValueOperation::ReadPopulation { + readout: PopulationReadout::TopK { .. }, + }, + } => {} + _ => return Err(invalid("unsupported current-series readout dependency")), + } + if node + .output_schema + .fields + .iter() + .any(|field| field.name == SERIES_IDENTITY_COLUMN) + { + return Err(invalid( + "current-series input already has a physical identity column", + )); + } + node.output_schema.fields.push(SummaryField { + name: SERIES_IDENTITY_COLUMN.into(), + dtype: SummaryFamilyType::Plain(DataType::Utf8), + nullable: false, + }); + } + for edge in &mut dag.edges { + edge.intermediate_schema = dag + .nodes + .iter() + .find(|node| node.id == edge.producer) + .unwrap() + .output_schema + .clone(); + } + let frontier = frontier.ok_or_else(|| invalid("missing current-series population"))?; + let schema = Arc::new( + dag.nodes + .iter() + .find(|node| u64::from(node.id.0) == frontier) + .unwrap() + .output_schema + .clone(), + ); + compile( + &dag, + BTreeMap::from([(frontier, InputContract::bounded(schema))]), + &[u64::from(dag.root.0)], + ) +} + +/// Compile selected ranking or aggregation above an exact per-series Rate +/// readout. Deployments bind complete window readouts at this boundary; +/// the heap is rebuilt independently for each evaluation. This does not move +/// that frontier to ingestion time or authorize combining finalized rates. +pub fn compile_rate_ranking( + selected: &Rc, +) -> Result< + ( + Rc, + CompiledPhysicalDag, + ), + Error, +> { + use planner_types::post_asap::{ + compile_post_asap_dag_with_node_ids, ExactKind, SummaryExpr, SummaryNode, + }; + fn frontier(node: &Rc) -> Option> { + match &node.expr { + SummaryExpr::ValueOperation { + child, + operation: ValueOperation::FinalizeExactAccumulator, + timing: planner_types::post_asap::ExecutionTiming::QueryTime, + } if matches!(&child.expr, SummaryExpr::SummaryAgg { + family: SummaryFamilyType::ExactAggregate(ExactKind::Rate, _), + reduction: planner_types::pre_asap::Reduction::PerEntity, + child: raw, .. + } if matches!(&raw.expr, SummaryExpr::KeepPreAsap(expr) if matches!(expr.as_ref(), QueryExpr::TimeRange { .. }))) => + { + Some(Rc::clone(node)) + } + SummaryExpr::ValueOperation { child, .. } | SummaryExpr::SummaryAgg { child, .. } => { + frontier(child) + } + SummaryExpr::SummaryEstimate { summary_input, .. } => frontier(summary_input), + _ => None, + } + } + let source = frontier(selected) + .ok_or_else(|| invalid("ranking requires one exact per-series Rate frontier"))?; + if !source + .schema + .fields + .iter() + .any(|field| field.name == SERIES_IDENTITY_COLUMN) + { + return Err(invalid("Rate ranking requires complete series identity")); + } + let compiled = compile_post_asap_dag_with_node_ids(selected) + .map_err(|error| invalid(error.to_string()))?; + let id = u64::from( + compiled + .node_ids + .node_id(&source) + .ok_or_else(|| invalid("missing Rate frontier"))? + .0, + ); + let program = compile( + &compiled.dag, + BTreeMap::from([(id, InputContract::bounded(Arc::new(source.schema.clone())))]), + &[u64::from(compiled.dag.root.0)], + )?; + Ok((source, program)) +} + +/// The selected logical placement requires fresh aggregate state per closed window. +/// Compile both physical graphs before deployment chooses storage or scheduling. +/// The input is the complete collection of per-series exact counter states. +pub fn compile_fixed_window_rate_aggregation( + selected: &Rc, +) -> Result { + use planner_types::post_asap::{ + compile_post_asap_dag, ExactKind, ExecutionTiming, SketchAlgorithm, + }; + let dag = compile_post_asap_dag(selected).map_err(|e| invalid(e.to_string()))?; + let sources = dag + .nodes + .iter() + .filter(|n| { + matches!( + &n.payload, + Payload::SummaryAgg { + family: SummaryFamilyType::ExactAggregate(ExactKind::Rate, _), + reduction: planner_types::pre_asap::Reduction::PerEntity, + .. + } + ) + }) + .collect::>(); + let heaps = dag + .nodes + .iter() + .filter(|n| { + n.output_state.timing == ExecutionTiming::IngestionTime + && match &n.payload { + Payload::SummaryAgg { + family: SummaryFamilyType::Sketch(kind, _), + .. + } => matches!( + kind.algorithm(), + SketchAlgorithm::CmsWithHeap | SketchAlgorithm::CountSketchWithHeap + ), + Payload::SummaryAgg { + family: SummaryFamilyType::ExactAggregate(ExactKind::Sum, _), + .. + } => true, + _ => false, + } + }) + .collect::>(); + let ([source], [heap]) = (sources.as_slice(), heaps.as_slice()) else { + return Err(invalid( + "expected one selected fixed-window Rate aggregation", + )); + }; + if !source + .output_schema + .fields + .iter() + .any(|f| f.name == SERIES_IDENTITY_COLUMN) + { + return Err(invalid( + "fixed-window Rate aggregation requires complete series identity", + )); + } + compile_candidate( + &dag, + BTreeMap::from([( + u64::from(source.id.0), + InputContract::bounded(Arc::new(source.output_schema.clone())), + )]), + &[u64::from(dag.root.0)], + &[u64::from(heap.id.0)], + ) +} diff --git a/crates/asap-physical-operators/src/physical_planner/promql_values.rs b/crates/asap-physical-operators/src/physical_planner/promql_values.rs new file mode 100644 index 00000000..98032505 --- /dev/null +++ b/crates/asap-physical-operators/src/physical_planner/promql_values.rs @@ -0,0 +1,278 @@ +//! Physical scalar/vector contracts preserve complete label sets across native computation. +use super::*; + +pub fn scalar_schema() -> Schema { + crate::operators::vector_binary::value_schema(true) +} +pub fn vector_schema() -> Schema { + crate::operators::vector_binary::value_schema(false) +} + +pub fn matrix_schema() -> Schema { + crate::operators::vector_window::matrix_schema() +} + +pub fn compile_scalar(value: f64) -> Result { + let operator = Operator::scalar( + crate::values::Value::Float64(value), + planner_types::pre_asap::DataType::Float64, + )? + .with_output_schema(scalar_schema())?; + CompiledPhysicalDag::from_operators( + BTreeMap::new(), + BTreeMap::from([(0, (vec![], operator))]), + vec![0], + ) +} + +pub fn compile_temporal( + intent: &AggIntent, + preserve_metric_name: bool, +) -> Result { + let operator = Operator::range_window(intent.clone())?; + let mut operators = vec![operator]; + if !preserve_metric_name { + operators.push(Operator::project( + vector_schema(), + vec![ + ( + "labels".into(), + Expression::LabelSet { + column: 0, + labels: vec![], + without: true, + }, + ), + ("value".into(), Expression::Column(1)), + ], + )?); + } + unary(operators, matrix_schema()) +} + +pub fn compile_histogram_quantile() -> Result { + CompiledPhysicalDag::from_operators( + BTreeMap::from([ + (0, InputContract::bounded(scalar_schema())), + (1, InputContract::bounded(vector_schema())), + ]), + BTreeMap::from([(2, (vec![0, 1], Operator::histogram_quantile()))]), + vec![2], + ) +} + +/// Compile before deployment chooses readers. Input slots 0 and 1 retain operand order. +pub fn compile_binary( + operator: &planner_types::post_asap::BinaryOperator, + return_bool: bool, + left_scalar: bool, + right_scalar: bool, +) -> Result { + let left = crate::operators::vector_binary::value_schema(left_scalar); + let right = crate::operators::vector_binary::value_schema(right_scalar); + let op = Operator::vector_binary(left.clone(), right.clone(), operator.clone(), return_bool)?; + CompiledPhysicalDag::from_operators( + BTreeMap::from([ + (0, InputContract::bounded(left)), + (1, InputContract::bounded(right)), + ]), + BTreeMap::from([(2, (vec![0, 1], op))]), + vec![2], + ) +} + +fn unary(operators: Vec, input: Schema) -> Result { + let root = operators.len() as u64; + CompiledPhysicalDag::from_operators( + BTreeMap::from([(0, InputContract::bounded(input))]), + operators + .into_iter() + .enumerate() + .map(|(i, op)| ((i + 1) as u64, (vec![i as u64], op))) + .collect(), + vec![root], + ) +} + +fn grouped(grouping: &GroupKeys) -> Result { + let labels = grouping + .keys() + .iter() + .map(|key| match key { + ColumnRef::Named(label) => Ok(label.clone()), + _ => Err(invalid("vector grouping requires label names")), + }) + .collect::, _>>()?; + Operator::project( + vector_schema(), + vec![ + ("labels".into(), Expression::Column(0)), + ("value".into(), Expression::Column(1)), + ( + "group".into(), + Expression::LabelSet { + column: 0, + labels, + without: grouping.is_without(), + }, + ), + ], + ) +} + +fn vector_output(input: Schema, labels: usize, value: usize) -> Result { + let value = Expression::ExactFloat64(value); + Operator::project( + input, + vec![ + ("labels".into(), Expression::Column(labels)), + ("value".into(), value), + ], + ) +} + +pub fn compile_aggregate( + intent: &AggIntent, + grouping: &GroupKeys, +) -> Result { + let project = grouped(grouping)?; + let reduction = match intent { + AggIntent::Sum { .. } => Reduction::Sum(1), + AggIntent::Avg { .. } => Reduction::Avg(1), + AggIntent::Count { .. } => Reduction::Count, + AggIntent::Min { .. } => Reduction::Min(1), + AggIntent::Max { .. } => Reduction::Max(1), + _ => return Err(invalid("unsupported vector aggregate")), + }; + let aggregate = + Operator::aggregate(project.schema(), vec![2], vec![("value".into(), reduction)])?; + let output = vector_output(aggregate.schema(), 0, 1)?; + unary(vec![project, aggregate, output], vector_schema()) +} + +pub fn compile_sort( + descending: bool, + grouping: &GroupKeys, +) -> Result { + let project = grouped(grouping)?; + let sort = Operator::sort( + project.schema(), + vec![SortKey { + column: 1, + descending, + nulls_first: false, + }], + vec![2], + )?; + let output = vector_output(sort.schema(), 0, 1)?; + unary(vec![project, sort, output], vector_schema()) +} + +pub fn compile_limit( + n: u64, + offset: u64, + grouping: &GroupKeys, +) -> Result { + let project = grouped(grouping)?; + let limit = Operator::limit(project.schema(), n, offset, vec![2])?; + let output = vector_output(limit.schema(), 0, 1)?; + unary(vec![project, limit, output], vector_schema()) +} + +pub fn compile_negate(scalar: bool) -> Result { + let input = if scalar { + scalar_schema() + } else { + vector_schema() + }; + let mut columns = Vec::new(); + if !scalar { + columns.push(("labels".into(), Expression::Column(0))); + } + columns.push(( + if scalar { + "$promql_scalar".into() + } else { + "value".into() + }, + Expression::Negate(Box::new(Expression::Column(if scalar { 0 } else { 1 }))), + )); + unary(vec![Operator::project(input.clone(), columns)?], input) +} + +pub fn compile_vector_to_scalar() -> Result { + unary( + vec![Operator::vector_to_scalar(vector_schema(), 1)?.with_output_schema(scalar_schema())?], + vector_schema(), + ) +} + +/// A stored exact-state input retains the complete population identity. The +/// deployment supplies eligible panes; merging and finalization are computation. +pub fn exact_state_schema(family: SummaryFamilyType) -> Result { + if !matches!(family, SummaryFamilyType::ExactAggregate(..)) { + return Err(invalid("exact-state input requires an exact family")); + } + crate::values::validate_family(&family)?; + let mut schema = (*vector_schema()).clone(); + schema.fields[1].dtype = family; + Ok(Arc::new(schema)) +} + +/// Retain exact readout semantics before any deployment state is opened. +pub fn compile_exact_readout( + family: SummaryFamilyType, + lookback_ms: u64, + preserve_metric_name: bool, +) -> Result { + use planner_types::post_asap::ExactKind; + let statistic = match &family { + SummaryFamilyType::ExactAggregate(kind, _) => match kind { + ExactKind::Sum => crate::Statistic::Sum, + ExactKind::Count => crate::Statistic::Count, + ExactKind::Min => crate::Statistic::Min, + ExactKind::Max => crate::Statistic::Max, + ExactKind::Rate => crate::Statistic::Rate, + ExactKind::Increase => crate::Statistic::Increase, + ExactKind::IRate => return Err(invalid("instant-rate state readout is not supported")), + }, + _ => return Err(invalid("exact readout requires an exact family")), + }; + let input = exact_state_schema(family)?; + let merge = Operator::summary_merge(input.clone(), 1, vec![0])?; + let mut readout = Operator::readout( + merge.schema(), + 1, + ReadoutQuery::Exact(ExactReadout { + statistic, + lookback_ms: None, + }), + )?; + if matches!( + statistic, + crate::Statistic::Rate | crate::Statistic::Increase + ) { + readout = readout.with_counter_lookback( + i64::try_from(lookback_ms).map_err(|_| invalid("counter lookback exceeds Int64"))?, + )?; + } + let project = Operator::project( + readout.schema(), + vec![ + ( + "labels".into(), + if preserve_metric_name { + Expression::Column(0) + } else { + Expression::LabelSet { + column: 0, + labels: vec![], + without: true, + } + }, + ), + ("value".into(), Expression::ExactFloat64(1)), + ], + )?; + unary(vec![merge, readout, project], input) +} diff --git a/crates/asap-physical-operators/src/physical_planner/temporal_panes.rs b/crates/asap-physical-operators/src/physical_planner/temporal_panes.rs new file mode 100644 index 00000000..d535c777 --- /dev/null +++ b/crates/asap-physical-operators/src/physical_planner/temporal_panes.rs @@ -0,0 +1,331 @@ +//! Lower a selected temporal maintenance contract; deployment supplies readers. +use super::*; +use planner_types::post_asap::{ + EvaluationSchedule, OutputRepresentation, PaneLayout, SketchAlgorithm, + SummaryMaintenanceLifecycle, SummaryMaintenanceLifecycleGuarantee, SummaryMaintenanceMode, + SummaryWindowFramework, +}; + +/// Resolved source identity, supplied with physical capability evidence. +/// A schemaless PromQL projection cannot establish the complete label set. +#[derive(Clone, Debug)] +pub enum TemporalEntityIdentity { + /// The input resolver guarantees that the slot contains one entity. + SingleEntity, + /// All entity keys are represented by these columns; there are no hidden + /// labels distinguishing two rows with the same key. + Columns(Vec), +} + +/// Planner-selected lifecycle/window requirements for one temporal producer. +/// Pane geometry is semantic input, not a storage identity or scheduling policy. +#[derive(Clone, Debug)] +pub struct TemporalPaneMaintenance { + pub summary_node: NodeId, + pub lifecycle: SummaryMaintenanceLifecycleGuarantee, + pub framework: SummaryWindowFramework, + pub layout: PaneLayout, + pub entity_identity: TemporalEntityIdentity, +} + +/// Generated precompute and query computation. `pane_inputs` is ordered from +/// the oldest complete pane to the newest; each run checks actual timestamps. +#[derive(Clone)] +pub struct TemporalPaneCandidate { + pub physical: PhysicalCandidate, + pub maintenance: TemporalPaneMaintenance, + pub pane_inputs: Vec, + pub merged_state: NodeId, + pub window_width_ms: u64, +} + +/// Compile bounded pane construction and a shared pane merge for temporal KLL +/// quantile roots. The selected contract remains authoritative; unsupported +/// lifecycle/framework/operator shapes fail rather than being substituted. +/// This initial realization consumes complete pane populations and emits full +/// state snapshots. Cross-run delta accumulation belongs to other candidates. +pub fn compile_temporal_pane_candidate( + dag: &PostAsapDag, + inputs: BTreeMap, + roots: &[NodeId], + maintenance: &TemporalPaneMaintenance, +) -> Result { + dag.validate().map_err(|error| invalid(error.to_string()))?; + if maintenance.lifecycle.summary_maintenance_lifecycle + != SummaryMaintenanceLifecycle::ContinuouslyMaintained + || maintenance.lifecycle.summary_maintenance_mode != SummaryMaintenanceMode::Incremental + || maintenance.lifecycle.evaluation_schedule != EvaluationSchedule::PerUpdate + || maintenance.lifecycle.output_representation != OutputRepresentation::SummaryState + { + return Err(invalid( + "pane candidate requires continuous incremental summary maintenance", + )); + } + let build = dag + .nodes + .iter() + .find(|node| u64::from(node.id.0) == maintenance.summary_node) + .ok_or_else(|| invalid("unknown maintained producer"))?; + let Payload::SummaryAgg { + family, + input: update, + reduction: PlannerReduction::PerEntity, + grouping, + } = &build.payload + else { + return Err(invalid( + "pane candidate requires a temporal per-entity summary", + )); + }; + if !matches!(family, SummaryFamilyType::Sketch(kind, _) if kind.algorithm() == &SketchAlgorithm::Kll) + || update.item.is_some() + { + return Err(invalid("pane candidate supports unkeyed temporal KLL only")); + } + crate::capability::validate_summary_kernel(family, update, grouping).map_err(Error::Invalid)?; + let dependencies: Vec<_> = dag + .edges + .iter() + .filter(|edge| edge.consumer == build.id) + .map(|edge| edge.producer) + .collect(); + let [raw_id] = dependencies.as_slice() else { + return Err(invalid("temporal producer requires one raw input")); + }; + let raw = dag + .nodes + .iter() + .find(|node| node.id == *raw_id) + .ok_or_else(|| invalid("missing raw input"))?; + let Payload::Fallback { + expression: QueryExpr::TimeRange { range, child }, + } = &raw.payload + else { + return Err(invalid( + "temporal producer requires an explicit logical time range", + )); + }; + let QueryExpr::Scan { predicates, .. } = child.as_ref() else { + return Err(invalid("temporal pane source requires a raw scan")); + }; + let window_width_ms: u64 = range + .as_millis() + .try_into() + .map_err(|_| invalid("temporal window overflows"))?; + if window_width_ms == 0 + || window_width_ms > i64::MAX as u64 + || range.subsec_nanos() % 1_000_000 != 0 + { + return Err(invalid( + "temporal window requires positive integral milliseconds", + )); + } + let width = maintenance.layout.pane_width_ms; + if width == 0 || width > window_width_ms || !window_width_ms.is_multiple_of(width) { + return Err(invalid("temporal window must contain whole panes")); + } + match maintenance.framework { + SummaryWindowFramework::Sliding => {} + SummaryWindowFramework::Tumbling if width == window_width_ms => {} + _ => return Err(invalid("unsupported temporal window realization")), + } + let count = window_width_ms / width; + if count > 4096 { + return Err(invalid("temporal pane candidate exceeds input budget")); + } + let raw_id = u64::from(raw_id.0); + if inputs.len() != 1 { + return Err(invalid( + "pane candidate requires exactly its raw input contract", + )); + } + let contract = inputs + .get(&raw_id) + .ok_or_else(|| invalid("missing raw input contract"))?; + let raw_schema = Arc::new(raw.output_schema.clone()); + if contract.schema != raw_schema || contract.properties.boundedness != Boundedness::Bounded { + return Err(invalid("pane source requires its declared bounded schema")); + } + let coordinate = raw_schema + .time_index + .ok_or_else(|| invalid("temporal source requires a time index"))?; + let SummaryInputExpr::Column(value) = &update.weight else { + return Err(invalid("pane builder requires a value column")); + }; + let value = named_column(&raw_schema, value)?; + let groups: Vec<_> = (0..raw_schema.fields.len()) + .filter(|&index| index != coordinate && index != value) + .collect(); + match &maintenance.entity_identity { + TemporalEntityIdentity::SingleEntity if groups.is_empty() => {} + TemporalEntityIdentity::Columns(columns) + if !columns.is_empty() + && columns.len() == columns.iter().collect::>().len() + && columns.iter().copied().collect::>() + == groups.iter().copied().collect() => {} + _ => { + return Err(invalid( + "pane input requires its complete resolved entity identity", + )) + } + } + let mut next = dag + .nodes + .iter() + .map(|node| u64::from(node.id.0)) + .max() + .unwrap_or(0) + + 1; + let mut allocate = || { + let id = next; + next += 1; + id + }; + let mut operators = BTreeMap::new(); + let guard = allocate(); + operators.insert( + guard, + ( + vec![raw_id], + Operator::pane_input( + raw_schema.clone(), + coordinate, + maintenance.layout.clone(), + None, + )?, + ), + ); + let mut previous = guard; + for predicate in predicates { + let id = allocate(); + operators.insert( + id, + ( + vec![previous], + Operator::filter(raw_schema.clone(), expression(&predicate.0, &raw_schema)?)?, + ), + ); + previous = id; + } + let native = + Operator::summary_build(raw_schema, family.clone(), value, Some(coordinate), groups)?; + let compact_state = native.schema(); + let native_id = allocate(); + operators.insert(native_id, (vec![previous], native)); + let state_schema = Arc::new(build.output_schema.clone()); + let pane_output = allocate(); + operators.insert( + pane_output, + ( + vec![native_id], + Operator::scope_timestamp(compact_state, state_schema.clone())?, + ), + ); + let precompute = CompiledPhysicalDag::from_operators(inputs, operators, vec![pane_output])?; + let state_coordinate = state_schema + .time_index + .ok_or_else(|| invalid("pane state requires a time index"))?; + let state_column = summary_column(&state_schema)?; + let mut query_inputs = BTreeMap::new(); + let mut operators = BTreeMap::new(); + let mut pane_inputs = Vec::new(); + let mut guarded_inputs = Vec::new(); + for pane in 0..count { + let input = allocate(); + let guard = allocate(); + query_inputs.insert(input, InputContract::bounded(state_schema.clone())); + let offset = ((count - 1 - pane) * width) as i64; + operators.insert( + guard, + ( + vec![input], + Operator::pane_input( + state_schema.clone(), + state_coordinate, + maintenance.layout.clone(), + Some(offset), + )?, + ), + ); + pane_inputs.push(input); + guarded_inputs.push(guard); + } + let union = allocate(); + operators.insert( + union, + ( + guarded_inputs, + Operator::union(state_schema.clone(), count as usize)?, + ), + ); + let merge = Operator::summary_merge( + state_schema.clone(), + state_column, + (0..state_schema.fields.len()) + .filter(|&index| index != state_coordinate && index != state_column) + .collect(), + )?; + let merged_schema = merge.schema(); + let merged_state = allocate(); + operators.insert(merged_state, (vec![union], merge)); + if roots.is_empty() || roots.iter().copied().collect::>().len() != roots.len() { + return Err(invalid("temporal query requires distinct output roots")); + } + for &root in roots { + let node = dag + .nodes + .iter() + .find(|node| u64::from(node.id.0) == root) + .ok_or_else(|| invalid("unknown temporal output root"))?; + let Payload::SummaryEstimate { + query: SketchQuery::Quantile { q }, + } = node.payload + else { + return Err(invalid("temporal pane root must be a KLL quantile")); + }; + if !q.is_finite() || !(0. ..=1.).contains(&q) { + return Err(invalid("invalid temporal quantile")); + } + let dependencies: Vec<_> = dag + .edges + .iter() + .filter(|edge| edge.consumer == node.id) + .map(|edge| u64::from(edge.producer.0)) + .collect(); + if dependencies != [maintenance.summary_node] { + return Err(invalid( + "temporal readout must consume the maintained producer", + )); + } + let readout = Operator::readout( + merged_schema.clone(), + summary_column(&merged_schema)?, + ReadoutQuery::Sketch(SketchQuery::Quantile { q }), + )?; + let readout_schema = readout.schema(); + let readout_id = allocate(); + operators.insert(readout_id, (vec![merged_state], readout)); + operators.insert( + root, + ( + vec![readout_id], + Operator::scope_timestamp(readout_schema, Arc::new(node.output_schema.clone()))?, + ), + ); + } + let query = CompiledPhysicalDag::from_operators(query_inputs, operators, roots.to_vec())?; + let mut output = precompute.output_contract(pane_output)?; + // Persisted readers have independent timing from the blocking builder. + output.properties.emission = Emission::Unknown; + Ok(TemporalPaneCandidate { + physical: PhysicalCandidate { + precompute: Some(precompute), + query, + materialized_outputs: BTreeMap::from([(pane_output, output)]), + }, + maintenance: maintenance.clone(), + pane_inputs, + merged_state, + window_width_ms, + }) +} diff --git a/crates/asap-physical-operators/tests/current_series_heap.rs b/crates/asap-physical-operators/tests/current_series_heap.rs new file mode 100644 index 00000000..322bd6fe --- /dev/null +++ b/crates/asap-physical-operators/tests/current_series_heap.rs @@ -0,0 +1,401 @@ +//! Spatial heap weights come from a fresh instant vector, never sample history. +use asap_physical_operators::{ + operators::Operator, + physical_planner::{ + promql_rows::{decode_series_identity, series_row, SERIES_IDENTITY_COLUMN}, + CompiledPhysicalDag, InputContract, Source, + }, + runtime::{Limits, RunContext, Scope}, + values::{Batch, Value}, +}; +use futures::{executor::block_on, StreamExt}; +use planner_types::{post_asap::*, pre_asap::DataType}; +use std::{collections::BTreeMap, sync::Arc}; + +fn schema() -> Arc { + Arc::new(SummarySchema { + fields: [ + ("ts", DataType::Timestamp), + ("value", DataType::Float64), + ("job", DataType::Utf8), + (SERIES_IDENTITY_COLUMN, DataType::Utf8), + ] + .into_iter() + .map(|(name, dtype)| SummaryField { + name: name.into(), + dtype: SummaryFamilyType::Plain(dtype), + nullable: false, + }) + .collect(), + time_index: Some(0), + }) +} +fn run(program: &CompiledPhysicalDag, data: Batch, end: i64) -> Result, String> { + let recovered = serde_json::from_slice::( + &serde_json::to_vec(&program).map_err(|e| e.to_string())?, + ) + .map_err(|e| e.to_string())?; + let input_id = recovered.input_contracts().next().unwrap().0; + let graph = recovered + .instantiate(BTreeMap::from([( + input_id, + Box::new(Operator::source(data.schema().clone(), vec![data]).unwrap()) as Source<'_>, + )])) + .unwrap(); + let context = RunContext::new( + Scope::Query { + evaluation_time_ms: end, + revision: 0, + }, + Limits::default(), + ) + .unwrap(); + block_on(async { + let mut stream = graph.execute(recovered.roots(), context).unwrap().remove(0); + let mut batches = Vec::new(); + while let Some(batch) = stream.next().await { + batches.push((*batch.map_err(|e| e.to_string())?).clone()); + } + Ok(batches) + }) +} +fn input(samples: &[(&str, i64, f64)]) -> Batch { + let schema = schema(); + let rows = samples + .iter() + .map(|(instance, time, value)| { + series_row( + &schema, + &BTreeMap::from([ + ("job".into(), "api".into()), + ("hidden_instance".into(), (*instance).into()), + ]), + *time, + *value, + ) + .unwrap() + }) + .collect(); + Batch::try_new(schema, rows).unwrap() +} +fn snapshot_plan() -> CompiledPhysicalDag { + CompiledPhysicalDag::from_operators( + BTreeMap::from([(0, InputContract::bounded(schema()))]), + BTreeMap::from([( + 1, + ( + vec![0], + Operator::current_series(schema(), 3, 0, 1, 60_000).unwrap(), + ), + )]), + vec![1], + ) + .unwrap() +} + +// Replacement, expiry and stale markers act before sketch updates. Hidden labels +// survive even when every series has the same projected `job` value. +#[test] +fn latest_snapshot_replaces_decreases_expires_and_retains_full_identity() { + let plan = snapshot_plan(); + let batches = run( + &plan, + input(&[ + ("decrease", 10_000, 100.), + ("decrease", 50_000, 1.), + ("steady", 40_000, 20.), + ("expired", 0, 1_000.), + ("stale", 20_000, 500.), + ("stale", 55_000, f64::from_bits(0x7ff0_0000_0000_0002)), + ("future", 60_001, 2_000.), + ]), + 60_000, + ) + .unwrap(); + let values = batches + .iter() + .flat_map(|batch| batch.rows()) + .map(|row| { + let Value::Utf8(identity) = &row[3] else { + panic!() + }; + let Value::Float64(value) = row[1] else { + panic!() + }; + assert!(matches!(row[0], Value::Timestamp(60_000))); + ( + decode_series_identity(identity).unwrap()["hidden_instance"].clone(), + value, + ) + }) + .collect::>(); + assert_eq!( + values, + BTreeMap::from([("decrease".into(), 1.), ("steady".into(), 20.)]) + ); + assert!(run(&plan, input(&[("steady", 40_000, 20.)]), 100_000) + .unwrap() + .iter() + .all(|batch| batch.rows().is_empty())); + assert!(run( + &plan, + input(&[("conflict", 50_000, 1.), ("conflict", 50_000, 2.)]), + 60_000 + ) + .is_err()); +} + +#[test] +fn spatial_heap_ranks_latest_values_in_independent_runs() { + for algorithm in [ + SketchAlgorithm::CmsWithHeap, + SketchAlgorithm::CountSketchWithHeap, + ] { + let params = match algorithm { + SketchAlgorithm::CmsWithHeap => SketchParams::CmsWithHeap { + width: 2048, + depth: 5, + heap_size: 100, + }, + _ => SketchParams::CountSketchWithHeap { + width: 2048, + depth: 5, + heap_size: 100, + }, + }; + let family = + SummaryFamilyType::Sketch(SketchKind::new(algorithm, params), Default::default()); + let build = Operator::keyed_summary_build(schema(), family, 1, vec![3], vec![2]).unwrap(); + let output = Arc::new(SummarySchema { + fields: vec![ + schema().fields[2].clone(), + schema().fields[3].clone(), + schema().fields[1].clone(), + ], + time_index: None, + }); + let read = Operator::keyed_readout(build.schema(), 1, 1, output).unwrap(); + let plan = CompiledPhysicalDag::from_operators( + BTreeMap::from([(0, InputContract::bounded(schema()))]), + BTreeMap::from([ + ( + 1, + ( + vec![0], + Operator::current_series(schema(), 3, 0, 1, 60_000).unwrap(), + ), + ), + (2, (vec![1], build)), + (3, (vec![2], read)), + ]), + vec![3], + ) + .unwrap(); + for (samples, end, winner, score) in [ + ( + vec![("a", 10_000, 100.), ("a", 50_000, 1.), ("b", 50_000, 20.)], + 60_000, + "b", + 20., + ), + ( + vec![("a", 110_000, 3.), ("b", 50_000, 20.)], + 120_000, + "a", + 3., + ), + ] { + let batches = run(&plan, input(&samples), end).unwrap(); + let rows = batches + .iter() + .flat_map(|batch| batch.rows()) + .collect::>(); + assert_eq!(rows.len(), 1); + let Value::Utf8(encoded) = &rows[0][1] else { + panic!() + }; + assert_eq!( + decode_series_identity(encoded).unwrap()["hidden_instance"], + winner + ); + assert!(matches!(rows[0][2], Value::Float64(actual) if actual == score)); + } + } +} + +// Blocking membership selection shares the run's cancellation and byte budget. +#[test] +fn current_series_observes_resource_limits() { + use asap_physical_operators::Error; + let plan = snapshot_plan(); + for cancelled in [false, true] { + let data = input(&[("one", 50_000, 1.)]); + let graph = plan + .instantiate(BTreeMap::from([( + 0, + Box::new(Operator::source(data.schema().clone(), vec![data]).unwrap()) + as Source<'_>, + )])) + .unwrap(); + let context = RunContext::new( + Scope::Query { + evaluation_time_ms: 60_000, + revision: 0, + }, + Limits { + max_bytes: if cancelled { 1 << 20 } else { 1 }, + ..Limits::default() + }, + ) + .unwrap(); + if cancelled { + context.cancel(); + } + let result = match graph.execute(&[1], context.clone()) { + Err(error) => Err(error), + Ok(mut streams) => block_on(streams.remove(0).next()).unwrap().map(|_| ()), + }; + assert!(matches!( + (cancelled, result), + (true, Err(Error::Cancelled)) | (false, Err(Error::MemoryLimit)) + )); + assert_eq!(context.retained_bytes(), 0); + } +} + +#[test] +fn identity_encoding_is_lossless_and_rejects_noncanonical_inputs() { + use asap_physical_operators::physical_planner::promql_rows::encode_series_identity; + let labels = BTreeMap::from([ + ("a".into(), "quote\"slash\\".into()), + ("other".into(), "".into()), + ]); + assert_eq!( + decode_series_identity(&encode_series_identity(&labels).unwrap()).unwrap(), + labels + ); + for invalid in [ + "[]", + "{\"a\":1}", + "{\"a\":\"x\",\"a\":\"x\"}", + "{ \"a\":\"x\"}", + ] { + assert!(decode_series_identity(invalid).is_err(), "{invalid}"); + } +} + +// The actual Planner population candidate lowers to native operators; this +// test does not manually assemble the computation or its dependency edges. +#[test] +fn planner_current_series_candidate_compiles_with_dynamic_identity() { + use asap_physical_operators::physical_planner::{compile, promql_rows::with_series_identity}; + use planner_types::{types::AccuracyTarget, workload::*}; + use std::rc::Rc; + let workload = PlanningWorkload { + query_workload: QueryWorkload { + language: QueryLanguage::PromQL, + query_batch: Some(vec![BatchEntry { + query: Query("topk by(job)(1, m)".into()), + requirements: QueryRequirements { + accuracy: AccuracyRequirement::Explicit(AccuracyTarget::Exact), + ..Default::default() + }, + predictability: Predictability::Unknown, + invocations: 1, + execute_at: None, + time_selection: TimeSelection::default(), + }]), + repeating_queries: None, + }, + data_workload: Some(DataWorkload { + data_ingestion_interval: Evidence { + value: Some(DurationMs(60_000)), + ..Default::default() + }, + ..Default::default() + }), + }; + let original = asap_frontend_promql::lower_promql_workload(&workload, 0) + .unwrap() + .remove(0); + let open_root = Rc::new(original.clone()); + let open_selected = + asap_aware_mapping::maintained_population::MaintainedPopulationStrategy::new( + std::slice::from_ref(&open_root), + ) + .candidate(&open_root) + .unwrap(); + let snapshot_program = + asap_physical_operators::physical_planner::promql_rows::compile_current_series_readout( + &open_selected, + ) + .unwrap(); + let encoded = String::from_utf8(serde_json::to_vec(&snapshot_program).unwrap()).unwrap(); + assert!( + !encoded.contains("CurrentSeries"), + "maintained input must not be rebuilt" + ); + assert!(encoded.contains("Sort") && encoded.contains("Limit")); + assert_eq!(snapshot_program.input_contracts().count(), 1); + let root = Rc::new(with_series_identity(&original).unwrap()); + let selected = asap_aware_mapping::maintained_population::MaintainedPopulationStrategy::new( + std::slice::from_ref(&root), + ) + .candidate(&root) + .unwrap(); + let logical = compile_post_asap_dag(&selected).unwrap(); + let raw = logical + .nodes + .iter() + .find(|node| matches!(node.payload, PostAsapOperatorPayload::Fallback { .. })) + .unwrap(); + let raw_schema = Arc::new(raw.output_schema.clone()); + let physical = compile( + &logical, + BTreeMap::from([( + u64::from(raw.id.0), + InputContract::bounded(raw_schema.clone()), + )]), + &[u64::from(logical.root.0)], + ) + .unwrap(); + let bytes = String::from_utf8(serde_json::to_vec(&physical).unwrap()).unwrap(); + assert!(bytes.contains("CurrentSeries")); + assert!(bytes.contains("Sort")); + assert!(bytes.contains("Limit")); + let rows = [("a", 10_000, 100.), ("a", 50_000, 1.), ("b", 50_000, 20.)] + .into_iter() + .map(|(member, at, value)| { + series_row( + &raw_schema, + &BTreeMap::from([ + ("job".into(), "api".into()), + ("unreferenced".into(), member.into()), + ]), + at, + value, + ) + .unwrap() + }) + .collect(); + let batches = run(&physical, Batch::try_new(raw_schema, rows).unwrap(), 60_000).unwrap(); + let rows = batches + .iter() + .flat_map(|batch| batch.rows()) + .collect::>(); + assert_eq!(rows.len(), 1); + assert!(matches!(rows[0][1], Value::Float64(20.))); + let id = batches[0] + .schema() + .fields + .iter() + .position(|field| field.name == SERIES_IDENTITY_COLUMN) + .unwrap(); + let Value::Utf8(encoded) = &rows[0][id] else { + panic!() + }; + assert_eq!( + decode_series_identity(encoded).unwrap()["unreferenced"], + "b" + ); +} diff --git a/crates/asap-physical-operators/tests/physical_dag.rs b/crates/asap-physical-operators/tests/physical_dag.rs new file mode 100644 index 00000000..96e74e54 --- /dev/null +++ b/crates/asap-physical-operators/tests/physical_dag.rs @@ -0,0 +1,1443 @@ +//! Acceptance tests use the library directly, without either backend engine. +use asap_physical_operators::{ + dag::{ + operators::{Expression, Operator, Reduction, SortKey}, + values::{Batch, Schema, Value}, + Limits, PhysicalDag, RunContext, Scope, + }, + Statistic, +}; +use futures::{executor::block_on, StreamExt}; +use planner_types::{ + post_asap::{ExactKind, ExactParams, SummaryFamilyType, SummaryField, SummarySchema}, + pre_asap::DataType, +}; +use std::sync::Arc; +fn schema(fields: &[(&str, DataType, bool)]) -> Schema { + Arc::new(SummarySchema { + fields: fields + .iter() + .map(|(name, dtype, nullable)| SummaryField { + name: (*name).into(), + dtype: SummaryFamilyType::Plain(dtype.clone()), + nullable: *nullable, + }) + .collect(), + time_index: None, + }) +} +fn run(dag: &PhysicalDag<'_, Batch, Schema>, root: u64, scope: Scope) -> Vec> { + let context = RunContext::new( + scope, + Limits { + max_buffered_batches: 1, + ..Limits::default() + }, + ) + .unwrap(); + block_on(async { + let mut stream = dag.execute(&[root], context.clone()).unwrap().remove(0); + let mut rows = vec![]; + while let Some(batch) = stream.next().await { + rows.extend(batch.unwrap().rows().iter().cloned()); + } + assert_eq!(context.retained_bytes(), 0); + rows + }) +} +fn query() -> Scope { + Scope::Query { + evaluation_time_ms: 1000, + revision: 2, + } +} +fn floats(rows: &[Vec], column: usize) -> Vec { + rows.iter() + .map(|r| { + if let Value::Float64(v) = r[column] { + v + } else { + panic!("not Float64") + } + }) + .collect() +} + +// Sort followed by partitioned Limit implements ranking independently per group. +#[test] +fn grouped_sort_limit_across_batches() { + let schema = schema(&[ + ("group", DataType::Int64, false), + ("score", DataType::Float64, false), + ]); + let batches = [ + vec![(1, 1.), (2, 4.), (1, 9.)], + vec![(2, 8.), (1, 5.), (2, 2.)], + ] + .into_iter() + .map(|rows| { + Batch::try_new( + schema.clone(), + rows.into_iter() + .map(|(g, v)| vec![Value::Int64(g), Value::Float64(v)]) + .collect(), + ) + .unwrap() + }) + .collect(); + let mut dag = PhysicalDag::default(); + dag.add( + 0, + vec![], + Operator::source(schema.clone(), batches).unwrap(), + ) + .unwrap(); + dag.add( + 1, + vec![0], + Operator::sort( + schema.clone(), + vec![SortKey { + column: 1, + descending: true, + nulls_first: false, + }], + vec![0], + ) + .unwrap(), + ) + .unwrap(); + dag.add(2, vec![1], Operator::limit(schema, 1, 1, vec![0]).unwrap()) + .unwrap(); + assert_eq!(floats(&run(&dag, 2, query()), 1), vec![5., 4.]); +} + +// The same computation runs in either engine scope with fresh per-run state. +#[test] +fn summary_construction_merge_and_readout_at_both_phases() { + let schema = schema(&[("v", DataType::Float64, false)]); + let batches = (1..=20) + .map(|v| Batch::try_new(schema.clone(), vec![vec![Value::Float64(v as f64)]]).unwrap()) + .collect(); + let family = SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum); + let build = Operator::summary_build(schema.clone(), family, 0, None, vec![]).unwrap(); + let state = build.schema(); + let mut dag = PhysicalDag::default(); + dag.add(0, vec![], Operator::source(schema, batches).unwrap()) + .unwrap(); + dag.add(1, vec![0], build).unwrap(); + dag.add(2, vec![1, 1], Operator::union(state.clone(), 2).unwrap()) + .unwrap(); + dag.add( + 3, + vec![2], + Operator::summary_merge(state.clone(), 0, vec![]).unwrap(), + ) + .unwrap(); + dag.add( + 4, + vec![3], + Operator::readout( + state, + 0, + asap_physical_operators::operators::ReadoutQuery::Exact( + asap_physical_operators::summary_kernels::exact::ExactReadout { + statistic: Statistic::Sum, + lookback_ms: None, + }, + ), + ) + .unwrap(), + ) + .unwrap(); + for scope in [ + query(), + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 1000, + revision: 2, + }, + ] { + assert_eq!(floats(&run(&dag, 4, scope), 0), vec![420.]); + } +} + +// A semi-join can consume two branches of one producer with a one-batch buffer. +#[test] +fn diamond_semijoin_preserves_left_values_and_multiplicity() { + let schema = schema(&[("key", DataType::Int64, false)]); + let batches = [1, 2, 2, 3] + .into_iter() + .map(|v| Batch::try_new(schema.clone(), vec![vec![Value::Int64(v)]]).unwrap()) + .collect(); + let filter = Operator::filter( + schema.clone(), + Expression::Equal( + Box::new(Expression::Column(0)), + Box::new(Expression::Literal { + value: Value::Int64(2), + dtype: DataType::Int64, + }), + ), + ) + .unwrap(); + let mut dag = PhysicalDag::default(); + dag.add( + 0, + vec![], + Operator::source(schema.clone(), batches).unwrap(), + ) + .unwrap(); + dag.add(1, vec![0], filter).unwrap(); + dag.add( + 2, + vec![0, 1], + Operator::semi_join(schema.clone(), schema, vec![(0, 0)]).unwrap(), + ) + .unwrap(); + let rows = run(&dag, 2, query()); + assert_eq!(rows.len(), 2); + assert!(rows.iter().all(|r| matches!(r[0], Value::Int64(2)))); +} + +// Integer aggregation must not silently lose precision through Float64. +#[test] +fn exact_integer_and_empty_extrema() { + let schema = schema(&[("v", DataType::Int64, false)]); + let aggregate = Operator::aggregate( + schema.clone(), + vec![], + vec![("sum".into(), Reduction::Sum(0))], + ) + .unwrap(); + let mut dag = PhysicalDag::default(); + let value = 9_007_199_254_740_993; + dag.add( + 0, + vec![], + Operator::source( + schema.clone(), + vec![Batch::try_new( + schema.clone(), + vec![vec![Value::Int64(value)], vec![Value::Int64(2)]], + ) + .unwrap()], + ) + .unwrap(), + ) + .unwrap(); + dag.add(1, vec![0], aggregate).unwrap(); + assert!(matches!(run(&dag,1,query())[0][0],Value::Int64(v) if v==value+2)); + let mut empty = PhysicalDag::default(); + empty + .add(0, vec![], Operator::source(schema.clone(), vec![]).unwrap()) + .unwrap(); + empty + .add( + 1, + vec![0], + Operator::aggregate(schema, vec![], vec![("min".into(), Reduction::Min(0))]).unwrap(), + ) + .unwrap(); + assert!(matches!(run(&empty, 1, query())[0][0], Value::Null)); +} + +// Plain value operators are library implementations, including NaN comparison. +#[test] +fn scalar_negation_and_vector_conversion() { + let scalar = Operator::scalar(Value::Float64(7.), DataType::Float64).unwrap(); + let project = Operator::project( + scalar.schema(), + vec![( + "v".into(), + Expression::Negate(Box::new(Expression::Column(0))), + )], + ) + .unwrap(); + let convert = Operator::vector_to_scalar(project.schema(), 0).unwrap(); + let mut dag = PhysicalDag::default(); + dag.add(0, vec![], scalar).unwrap(); + dag.add(1, vec![0], project).unwrap(); + dag.add(2, vec![1], convert).unwrap(); + assert_eq!(floats(&run(&dag, 2, query()), 0), vec![-7.]); + let scalar = Operator::scalar(Value::Float64(f64::NAN), DataType::Float64).unwrap(); + let predicate = Expression::Equal( + Box::new(Expression::Column(0)), + Box::new(Expression::Column(0)), + ); + let filter = Operator::filter(scalar.schema(), predicate).unwrap(); + let mut dag = PhysicalDag::default(); + dag.add(0, vec![], scalar).unwrap(); + dag.add(1, vec![0], filter).unwrap(); + assert!(run(&dag, 1, query()).is_empty()); +} + +// Invalid operations fail at binding rather than becoming external fallbacks. +#[test] +fn binding_rejects_unsupported_operations() { + let schema = schema(&[("v", DataType::Float64, false)]); + assert!(Operator::summary_build( + schema.clone(), + SummaryFamilyType::ExactAggregate(ExactKind::Rate, ExactParams::Rate), + 0, + None, + vec![] + ) + .is_err()); + let sum = Operator::summary_build( + schema.clone(), + SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum), + 0, + None, + vec![], + ) + .unwrap(); + assert!(Operator::readout( + sum.schema(), + 0, + asap_physical_operators::operators::ReadoutQuery::Sketch( + planner_types::post_asap::SketchQuery::Quantile { q: 0.5 } + ) + ) + .is_err()); + assert!(Operator::filter(schema, Expression::Column(0)).is_err()); +} + +// KLL is one family example: precomputation changes input sources, not operators. +#[test] +fn kll_raw_partial_and_precomputed_are_native_dags() { + use planner_types::post_asap::{GroupingStrategy, SketchAlgorithm, SketchKind, SketchParams}; + let input = schema(&[("value", DataType::Float64, false)]); + let family = SummaryFamilyType::Sketch( + SketchKind::new(SketchAlgorithm::Kll, SketchParams::Kll { k: 512 }), + GroupingStrategy::PerSubpopulationInstance, + ); + let build = Operator::summary_build(input.clone(), family, 0, None, vec![]).unwrap(); + let state = build.schema(); + let build_range = |start: u32, end: u32| { + let mut dag = PhysicalDag::default(); + let batch = Batch::try_new( + input.clone(), + (start..end) + .map(|v| vec![Value::Float64(f64::from(v))]) + .collect(), + ) + .unwrap(); + dag.add( + 0, + vec![], + Operator::source(input.clone(), vec![batch]).unwrap(), + ) + .unwrap(); + dag.add(1, vec![0], build.clone()).unwrap(); + run( + &dag, + 1, + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 1000, + revision: 1, + }, + ) + }; + let prefix = build_range(0, 64); + let complete = build_range(0, 128); + let query_plan = |stored: Option>>, raw_start: Option| { + let mut dag = PhysicalDag::default(); + let mut states = vec![]; + if let Some(rows) = stored { + dag.add( + 0, + vec![], + Operator::source( + state.clone(), + vec![Batch::try_new(state.clone(), rows).unwrap()], + ) + .unwrap(), + ) + .unwrap(); + states.push(0); + } + if let Some(start) = raw_start { + dag.add( + 1, + vec![], + Operator::source( + input.clone(), + vec![Batch::try_new( + input.clone(), + (start..128) + .map(|v| vec![Value::Float64(f64::from(v))]) + .collect(), + ) + .unwrap()], + ) + .unwrap(), + ) + .unwrap(); + dag.add(2, vec![1], build.clone()).unwrap(); + states.push(2); + } + dag.add( + 3, + states.clone(), + Operator::union(state.clone(), states.len()).unwrap(), + ) + .unwrap(); + dag.add( + 4, + vec![3], + Operator::summary_merge(state.clone(), 0, vec![]).unwrap(), + ) + .unwrap(); + dag.add( + 5, + vec![4], + Operator::readout( + state.clone(), + 0, + asap_physical_operators::operators::ReadoutQuery::Sketch( + planner_types::post_asap::SketchQuery::Quantile { q: 0.5 }, + ), + ) + .unwrap(), + ) + .unwrap(); + floats(&run(&dag, 5, query()), 0)[0] + }; + let raw = query_plan(None, Some(0)); + let partial = query_plan(Some(prefix), Some(64)); + let full = query_plan(Some(complete), None); + assert_eq!(raw, partial); + assert_eq!(partial, full); + assert!((raw - 64.).abs() <= 1.); +} + +// Exact state must match its declared family; a mislabeled state is rejected. +#[test] +fn exact_state_and_family_validation() { + use asap_physical_operators::summary_kernels::exact::ExactAccumulator; + let family = SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum); + let mut acc = ExactAccumulator::new(family.clone(), false).unwrap(); + acc.update(None, 7., 0); + let schema = Arc::new(SummarySchema { + fields: vec![SummaryField { + name: "state".into(), + dtype: family.clone(), + nullable: false, + }], + time_index: None, + }); + let value = Value::Summary { + family: family.clone(), + state: Arc::new(acc), + }; + let mut dag = PhysicalDag::default(); + dag.add( + 0, + vec![], + Operator::source( + schema.clone(), + vec![Batch::try_new(schema.clone(), vec![vec![value]]).unwrap()], + ) + .unwrap(), + ) + .unwrap(); + dag.add( + 1, + vec![0], + Operator::readout( + schema.clone(), + 0, + asap_physical_operators::operators::ReadoutQuery::Exact( + asap_physical_operators::summary_kernels::exact::ExactReadout { + statistic: Statistic::Sum, + lookback_ms: None, + }, + ), + ) + .unwrap(), + ) + .unwrap(); + assert_eq!(floats(&run(&dag, 1, query()), 0), vec![7.]); + let wrong = ExactAccumulator::new( + SummaryFamilyType::ExactAggregate(ExactKind::Max, ExactParams::Max), + false, + ) + .unwrap(); + assert!(Batch::try_new( + schema, + vec![vec![Value::Summary { + family, + state: Arc::new(wrong) + }]] + ) + .is_err()); +} + +// Planner binding rejects unknown computation instead of accepting a fallback. +#[test] +fn bind_post_asap_before_execution() { + use asap_physical_operators::dag::planner::bind; + use planner_types::{ + post_asap::{ + EdgeRole, ExecutionDataState, GroupingEdgeCompatibility, PostAsapDag, PostAsapDagEdge, + PostAsapDagNode, PostAsapNodeId, PostAsapOperatorPayload, ValueOperation, + WindowEdgeCompatibility, + }, + pre_asap::{ArithmeticOpKind, ProjectItem, QueryExpr, ScalarValue}, + }; + use std::{collections::BTreeMap, rc::Rc}; + let schema = schema(&[("value", DataType::Float64, false)]); + let node = |id, payload| PostAsapDagNode { + id: PostAsapNodeId(id), + payload, + output_state: ExecutionDataState::QUERY_ROWS, + output_schema: (*schema).clone(), + guarantee: None, + }; + let mut dag = PostAsapDag { + nodes: vec![ + node( + 0, + PostAsapOperatorPayload::Fallback { + expression: QueryExpr::promql_scalar(1.), + }, + ), + node( + 1, + PostAsapOperatorPayload::Value { + operation: ValueOperation::Project { + cols: vec![ProjectItem { + alias: None, + expr: QueryExpr::Arithmetic { + op: ArithmeticOpKind::Add, + left: Rc::new(QueryExpr::Column(0)), + right: Rc::new(QueryExpr::Literal(ScalarValue::Float64(2.))), + }, + }], + qualifier: None, + }, + }, + ), + ], + edges: vec![PostAsapDagEdge { + producer: PostAsapNodeId(0), + consumer: PostAsapNodeId(1), + role: EdgeRole::Input, + intermediate_schema: (*schema).clone(), + data_state: ExecutionDataState::QUERY_ROWS, + grouping: GroupingEdgeCompatibility::NotApplicable, + window: WindowEdgeCompatibility::NotApplicable, + }], + root: PostAsapNodeId(1), + }; + let sources = || -> BTreeMap> { + BTreeMap::from([( + 0, + Box::new( + Operator::source( + schema.clone(), + vec![Batch::try_new(schema.clone(), vec![vec![Value::Float64(1.)]]).unwrap()], + ) + .unwrap(), + ) as asap_physical_operators::dag::planner::Source<'static>, + )]) + }; + let native = bind(&dag, sources(), &[1]).unwrap(); + assert_eq!(floats(&run(&native, 1, query()), 0), vec![3.]); + assert!(bind(&dag, BTreeMap::new(), &[1]).is_err()); + dag.nodes[1].payload = PostAsapOperatorPayload::Value { + operation: ValueOperation::Extension { + name: "unknown".into(), + }, + }; + assert!(bind(&dag, sources(), &[1]).is_err()); +} + +// A completed empty population has an exact zero count, with integer output. +#[test] +fn empty_exact_count_is_an_integer_state_readout() { + let input = schema(&[("value", DataType::Float64, false)]); + let build = Operator::summary_build( + input.clone(), + SummaryFamilyType::ExactAggregate(ExactKind::Count, ExactParams::Count), + 0, + None, + vec![], + ) + .unwrap(); + let read = Operator::readout( + build.schema(), + 0, + asap_physical_operators::operators::ReadoutQuery::Exact( + asap_physical_operators::summary_kernels::exact::ExactReadout { + statistic: Statistic::Count, + lookback_ms: None, + }, + ), + ) + .unwrap(); + let mut dag = PhysicalDag::default(); + dag.add(0, vec![], Operator::source(input, vec![]).unwrap()) + .unwrap(); + dag.add(1, vec![0], build).unwrap(); + dag.add(2, vec![1], read).unwrap(); + assert!(matches!(run(&dag, 2, query())[0][0], Value::Int64(0))); +} + +// A deployment source cannot pass a different row shape to bound expressions. +#[test] +fn source_batches_must_match_the_bound_schema() { + use asap_physical_operators::dag::{self, PhysicalOperator}; + use planner_types::{ + post_asap::{ + ExecutionDataState, PostAsapDag, PostAsapDagNode, PostAsapNodeId, + PostAsapOperatorPayload, + }, + pre_asap::QueryExpr, + }; + use std::{cell::Cell, collections::BTreeMap, rc::Rc}; + struct WrongSource { + schema: Schema, + starts: Rc>, + } + impl PhysicalOperator for WrongSource { + fn name(&self) -> &str { + "ExternalSource" + } + fn input_schemas(&self) -> Vec { + vec![] + } + fn output_schema(&self) -> Schema { + self.schema.clone() + } + fn output_bytes(&self, value: &Batch) -> usize { + value.bytes() + } + fn start<'a>( + &'a self, + _: Vec>, + _: RunContext, + ) -> Result, dag::Error> { + self.starts.set(self.starts.get() + 1); + Ok( + futures::stream::once(async { Batch::try_new(schema(&[]), vec![vec![]]) }) + .boxed_local(), + ) + } + } + let expected = schema(&[("value", DataType::Float64, false)]); + let starts = Rc::new(Cell::new(0)); + let plan = PostAsapDag { + nodes: vec![PostAsapDagNode { + id: PostAsapNodeId(0), + payload: PostAsapOperatorPayload::Fallback { + expression: QueryExpr::promql_scalar(1.), + }, + output_state: ExecutionDataState::QUERY_ROWS, + output_schema: (*expected).clone(), + guarantee: None, + }], + edges: vec![], + root: PostAsapNodeId(0), + }; + let source = Box::new(WrongSource { + schema: expected, + starts: starts.clone(), + }) as dag::planner::Source<'static>; + let native = dag::planner::bind(&plan, BTreeMap::from([(0, source)]), &[0]).unwrap(); + assert_eq!(starts.get(), 0); + let context = RunContext::new(query(), Limits::default()).unwrap(); + let mut output = native.execute(&[0], context).unwrap().remove(0); + assert!(matches!( + block_on(output.next()), + Some(Err(dag::Error::AtNode { node: 0, .. })) + )); + assert_eq!(starts.get(), 1); +} + +// Float extrema have the same NaN behavior as the exact summary kernels. +#[test] +fn extrema_preserve_numeric_values_in_the_presence_of_nan() { + let input = schema(&[("v", DataType::Float64, false)]); + let mut dag = PhysicalDag::default(); + dag.add( + 0, + vec![], + Operator::source( + input.clone(), + vec![Batch::try_new( + input.clone(), + vec![vec![Value::Float64(-f64::NAN)], vec![Value::Float64(5.)]], + ) + .unwrap()], + ) + .unwrap(), + ) + .unwrap(); + dag.add( + 1, + vec![0], + Operator::aggregate( + input, + vec![], + vec![ + ("min".into(), Reduction::Min(0)), + ("max".into(), Reduction::Max(0)), + ], + ) + .unwrap(), + ) + .unwrap(); + let rows = run(&dag, 1, query()); + assert_eq!(floats(&rows, 0), vec![5.]); + assert_eq!(floats(&rows, 1), vec![5.]); +} + +// Planner wire nodes, including grouping and edge roles, are executable at either phase. +#[test] +fn planner_semijoin_sort_limit_contract_at_both_phases() { + use asap_physical_operators::dag::planner::{bind, Source}; + use planner_types::{ + post_asap::*, + pre_asap::{CompareOpKind, GroupKeys, JoinKind, Predicate, QueryExpr, SortKey}, + }; + use std::{collections::BTreeMap, rc::Rc}; + let rows_schema = schema(&[ + ("group", DataType::Utf8, false), + ("key", DataType::Utf8, false), + ("score", DataType::Float64, false), + ]); + let keys_schema = schema(&[("key", DataType::Utf8, false)]); + let node = |id, payload, schema: &Schema| PostAsapDagNode { + id: PostAsapNodeId(id), + payload, + output_schema: (**schema).clone(), + output_state: ExecutionDataState::QUERY_ROWS, + guarantee: None, + }; + let edge = |producer, consumer, role, schema: &Schema| PostAsapDagEdge { + producer: PostAsapNodeId(producer), + consumer: PostAsapNodeId(consumer), + role, + intermediate_schema: (**schema).clone(), + data_state: ExecutionDataState::QUERY_ROWS, + grouping: GroupingEdgeCompatibility::NotApplicable, + window: WindowEdgeCompatibility::NotApplicable, + }; + let groups = GroupKeys::by(vec![0]); + let dag = PostAsapDag { + nodes: vec![ + node( + 0, + PostAsapOperatorPayload::Fallback { + expression: QueryExpr::promql_scalar(0.), + }, + &rows_schema, + ), + node( + 1, + PostAsapOperatorPayload::Fallback { + expression: QueryExpr::promql_scalar(0.), + }, + &keys_schema, + ), + node( + 2, + PostAsapOperatorPayload::RelationalJoin { + join_kind: JoinKind::Semi, + pruning: None, + pred: Predicate(Rc::new(QueryExpr::Compare { + left: Rc::new(QueryExpr::Column(1)), + op: CompareOpKind::Eq, + right: Rc::new(QueryExpr::Column(3)), + })), + }, + &rows_schema, + ), + node( + 3, + PostAsapOperatorPayload::Value { + operation: ValueOperation::Sort { + keys: vec![SortKey { + expr: QueryExpr::Column(2), + ascending: false, + nulls_first: false, + }], + partition_by: groups.clone(), + }, + }, + &rows_schema, + ), + node( + 4, + PostAsapOperatorPayload::Value { + operation: ValueOperation::Limit { + n: 1, + offset: 0, + partition_by: groups, + }, + }, + &rows_schema, + ), + ], + // Deliberately put Right before Left: list order must not swap inputs. + edges: vec![ + edge(1, 2, EdgeRole::Right, &keys_schema), + edge(0, 2, EdgeRole::Left, &rows_schema), + edge(2, 3, EdgeRole::Input, &rows_schema), + edge(3, 4, EdgeRole::Input, &rows_schema), + ], + root: PostAsapNodeId(4), + }; + let text = |v: &str| Value::Utf8(v.into()); + for (phase, scope) in [ + (ExecutionTiming::QueryTime, query()), + ( + ExecutionTiming::IngestionTime, + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 1000, + revision: 2, + }, + ), + ] { + let dag = dag + .with_execution_phases(&dag.nodes.iter().map(|node| (node.id, phase)).collect()) + .unwrap(); + let sources: BTreeMap> = BTreeMap::from([ + ( + 0, + Box::new( + Operator::source( + rows_schema.clone(), + vec![Batch::try_new( + rows_schema.clone(), + vec![ + vec![text("a"), text("x"), Value::Float64(8.)], + vec![text("a"), text("y"), Value::Float64(9.)], + vec![text("b"), text("x"), Value::Float64(2.)], + vec![text("b"), text("z"), Value::Float64(99.)], + ], + ) + .unwrap()], + ) + .unwrap(), + ) as Source<'static>, + ), + ( + 1, + Box::new( + Operator::source( + keys_schema.clone(), + vec![Batch::try_new( + keys_schema.clone(), + vec![vec![text("x")], vec![text("y")]], + ) + .unwrap()], + ) + .unwrap(), + ) as Source<'static>, + ), + ]); + let native = bind(&dag, sources, &[4]).unwrap(); + let mut scores = floats(&run(&native, 4, scope), 2); + scores.sort_by(f64::total_cmp); + assert_eq!(scores, vec![2., 9.]); + } +} + +// Planner scalar signatures, collection access and null predicates share native execution. +#[test] +fn planner_expressions_preserve_collection_and_nullable_types() { + use asap_physical_operators::dag::expressions::CompiledExpression; + use planner_types::pre_asap::{CompareOpKind, QueryExpr, ScalarValue}; + use std::rc::Rc; + let input_schema = schema(&[( + "items", + DataType::Map { + key: Box::new(DataType::Utf8), + value: Box::new(DataType::Int64), + value_nullable: false, + }, + false, + )]); + let access = QueryExpr::FunctionCall { + name: "asap_element_access".into(), + args: vec![ + QueryExpr::Column(0), + QueryExpr::Literal(ScalarValue::Utf8("count".into())), + ], + }; + let project = Operator::project( + input_schema.clone(), + vec![( + "count".into(), + Expression::planner(CompiledExpression::compile(&access, &input_schema).unwrap()), + )], + ) + .unwrap(); + let mut dag = PhysicalDag::default(); + dag.add( + 0, + vec![], + Operator::source( + input_schema.clone(), + vec![Batch::try_new( + input_schema.clone(), + vec![ + vec![Value::Map( + vec![(Value::Utf8("count".into()), Value::Int64(7))].into(), + )], + vec![Value::Map(Arc::from([]))], + ], + ) + .unwrap()], + ) + .unwrap(), + ) + .unwrap(); + let projected = project.schema(); + dag.add(1, vec![0], project).unwrap(); + let predicate = QueryExpr::Compare { + left: Rc::new(QueryExpr::Column(0)), + op: CompareOpKind::Ge, + right: Rc::new(QueryExpr::Literal(ScalarValue::Int64(1))), + }; + dag.add( + 2, + vec![1], + Operator::filter( + projected.clone(), + Expression::planner(CompiledExpression::compile(&predicate, &projected).unwrap()), + ) + .unwrap(), + ) + .unwrap(); + let rows = run(&dag, 2, query()); + assert!(matches!(rows.as_slice(),[row] if matches!(row.as_slice(),[Value::Int64(7)]))); + let unknown = QueryExpr::FunctionCall { + name: "unregistered_function".into(), + args: vec![QueryExpr::Column(0)], + }; + assert!(CompiledExpression::compile(&unknown, &input_schema).is_err()); +} + +// Outer, semi and anti joins share Planner predicates and preserve SQL null behavior. +#[test] +fn native_relational_join_kinds_preserve_unmatched_rows() { + use planner_types::pre_asap::{CompareOpKind, JoinKind, Predicate, QueryExpr}; + use std::rc::Rc; + let input = schema(&[("key", DataType::Int64, true)]); + let predicate = Predicate(Rc::new(QueryExpr::Compare { + left: Rc::new(QueryExpr::Column(0)), + op: CompareOpKind::Eq, + right: Rc::new(QueryExpr::Column(1)), + })); + for (kind, count) in [ + (JoinKind::Inner, 1), + (JoinKind::Left, 3), + (JoinKind::Right, 3), + (JoinKind::Full, 5), + (JoinKind::Semi, 1), + (JoinKind::Anti, 2), + (JoinKind::Cross, 9), + ] { + let output = if matches!(kind, JoinKind::Semi | JoinKind::Anti) { + input.clone() + } else { + schema(&[ + ("left", DataType::Int64, true), + ("right", DataType::Int64, true), + ]) + }; + let mut dag = PhysicalDag::default(); + for (id, rows) in [ + ( + 0, + vec![ + vec![Value::Int64(1)], + vec![Value::Int64(2)], + vec![Value::Null], + ], + ), + ( + 1, + vec![ + vec![Value::Int64(2)], + vec![Value::Int64(3)], + vec![Value::Null], + ], + ), + ] { + dag.add( + id, + vec![], + Operator::source( + input.clone(), + vec![Batch::try_new(input.clone(), rows).unwrap()], + ) + .unwrap(), + ) + .unwrap(); + } + dag.add( + 2, + vec![0, 1], + Operator::relational_join( + input.clone(), + input.clone(), + kind.clone(), + &predicate, + output, + ) + .unwrap(), + ) + .unwrap(); + assert_eq!(run(&dag, 2, query()).len(), count, "{kind:?}"); + } +} + +// Per-series fractional rates feed either weighted frequency family per job, in either scope. +#[test] +fn weighted_rate_topk_preserves_partitions_fractional_scores_and_evaluation_scope() { + for count_sketch in [false, true] { + assert_weighted_rate_topk(count_sketch); + } +} +fn assert_weighted_rate_topk(count_sketch: bool) { + use planner_types::post_asap::{SketchAlgorithm, SketchKind, SketchParams}; + let raw = schema(&[ + ("service", DataType::Utf8, false), + ("job", DataType::Utf8, false), + ("instance", DataType::Int64, false), + ("t", DataType::Timestamp, false), + ("value", DataType::Float64, false), + ]); + let mut rows = Vec::new(); + // Multiple instances of auth accumulate. Batch has a very different scale. + for (service, job, instance, rate) in [ + ("auth", "api", 1, 0.125), + ("auth", "api", 2, 0.25), + ("checkout", "api", 1, 0.3125), + ("search", "api", 1, 0.0625), + ("ingest", "batch", 1, 100.0), + ("export", "batch", 1, 80.0), + ("cleanup", "batch", 1, 20.0), + ] { + for (t, value) in [(0, 0.0), (30_000, rate * 30.0), (60_000, rate * 60.0)] { + rows.push(vec![ + Value::Utf8(service.into()), + Value::Utf8(job.into()), + Value::Int64(instance), + Value::Timestamp(t), + Value::Float64(value), + ]); + } + } + let rates = Operator::window( + raw.clone(), + planner_types::pre_asap::AggIntent::Rate, + 3, + 4, + vec![0, 1, 2], + Some((0, 60_000)), + ) + .unwrap(); + let family = SummaryFamilyType::Sketch( + SketchKind::new( + if count_sketch { + SketchAlgorithm::CountSketchWithHeap + } else { + SketchAlgorithm::CmsWithHeap + }, + if count_sketch { + SketchParams::CountSketchWithHeap { + width: 4096, + depth: 5, + heap_size: 8, + } + } else { + SketchParams::CmsWithHeap { + width: 4096, + depth: 5, + heap_size: 8, + } + }, + ), + Default::default(), + ); + let build = Operator::keyed_summary_build(rates.schema(), family, 3, vec![0], vec![1]).unwrap(); + let output = schema(&[ + ("job", DataType::Utf8, false), + ("service", DataType::Utf8, false), + ("score", DataType::Float64, false), + ]); + let readout = Operator::keyed_readout(build.schema(), 1, 8, output.clone()).unwrap(); + let mut dag = PhysicalDag::default(); + dag.add( + 0, + vec![], + Operator::source(raw.clone(), vec![Batch::try_new(raw, rows).unwrap()]).unwrap(), + ) + .unwrap(); + dag.add(1, vec![0], rates).unwrap(); + dag.add(2, vec![1], build).unwrap(); + dag.add(3, vec![2], readout).unwrap(); + dag.add( + 4, + vec![3], + Operator::sort( + output.clone(), + vec![SortKey { + column: 2, + descending: true, + nulls_first: false, + }], + vec![0], + ) + .unwrap(), + ) + .unwrap(); + dag.add(5, vec![4], Operator::limit(output, 2, 0, vec![0]).unwrap()) + .unwrap(); + for scope in [ + query(), + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 60_000, + revision: 2, + }, + query(), + ] { + let result = run(&dag, 5, scope); + assert_eq!(result.len(), 4); + assert_eq!(floats(&result, 2), vec![0.375, 0.3125, 100.0, 80.0]); + let services = result + .iter() + .map(|row| match &row[1] { + Value::Utf8(v) => v.as_ref(), + _ => panic!("service"), + }) + .collect::>(); + assert_eq!(services, vec!["auth", "checkout", "ingest", "export"]); + } +} + +// The grouped temporal reducer's sample schema must survive physical Sort/Limit binding. +#[test] +fn grouped_temporal_schema_compiles_and_executes_topk() { + use asap_physical_operators::physical_planner::{ + compile_node, CompiledPhysicalDag, InputContract, Source, + }; + use planner_types::post_asap::{ + ExecutionDataState, PostAsapDagNode, PostAsapNodeId, PostAsapOperatorPayload, + ValueOperation, + }; + use planner_types::pre_asap::{ + aggregate_output_schema, AggIntent, Column, GroupKeys, QueryExpr, Reduction as IrReduction, + Schema as IrSchema, + }; + let grouped = IrSchema::new(vec![ + Column::new("job", DataType::Utf8, false), + Column::new("sum", DataType::Float64, false), + ]); + let output = aggregate_output_schema( + &grouped, + &IrReduction::PerEntity, + &[AggIntent::Avg { col: None }], + &[], + ) + .unwrap(); + let input = schema( + &output + .columns + .iter() + .map(|c| (c.name.as_str(), c.dtype.clone(), c.nullable)) + .collect::>(), + ); + let node = |id, operation| PostAsapDagNode { + id: PostAsapNodeId(id), + payload: PostAsapOperatorPayload::Value { operation }, + output_state: ExecutionDataState::QUERY_ROWS, + output_schema: (*input).clone(), + guarantee: None, + }; + let sort = compile_node( + &node( + 1, + ValueOperation::Sort { + keys: vec![planner_types::pre_asap::SortKey { + expr: QueryExpr::Column(1), + ascending: false, + nulls_first: false, + }], + partition_by: GroupKeys::none(), + }, + ), + std::slice::from_ref(&input), + ) + .unwrap(); + let limit = compile_node( + &node( + 2, + ValueOperation::Limit { + n: 1, + offset: 0, + partition_by: GroupKeys::none(), + }, + ), + std::slice::from_ref(&input), + ) + .unwrap(); + let compiled = CompiledPhysicalDag::from_operators( + [(0, InputContract::bounded(input.clone()))].into(), + [(1, (vec![0], sort)), (2, (vec![1], limit))].into(), + vec![2], + ) + .unwrap(); + let recovered = + serde_json::from_slice::(&serde_json::to_vec(&compiled).unwrap()) + .unwrap(); + assert_eq!(recovered.row_source(2), Some(0)); + assert_eq!(recovered.operator_name(2), Some("Limit")); + let expected = vec![Value::Utf8("api".into()), Value::Float64(9.)]; + let batch = Batch::try_new( + input.clone(), + vec![ + vec![Value::Utf8("worker".into()), Value::Float64(2.)], + expected.clone(), + ], + ) + .unwrap(); + let source = Box::new(Operator::source(input, vec![batch]).unwrap()) as Source<'_>; + let physical = recovered.instantiate([(0, source)].into()).unwrap(); + let mut stream = physical + .execute(&[2], RunContext::new(query(), Limits::default()).unwrap()) + .unwrap() + .remove(0); + let rows = block_on(async { + let mut rows = vec![]; + while let Some(batch) = stream.next().await { + rows.extend_from_slice(batch.unwrap().rows()); + } + rows + }); + assert_eq!(rows.len(), 1); + assert!(matches!(&rows[0][0], Value::Utf8(label) if label.as_ref() == "api")); + assert!(matches!(rows[0][1], Value::Float64(9.))); +} + +// A certified candidate set must have authoritative values for every key, including after recovery. +#[test] +fn certified_pruning_rejects_missing_authoritative_values_after_recovery() { + use asap_physical_operators::physical_planner::{ + compile_node, CompiledPhysicalDag, InputContract, Source, + }; + use planner_types::{ + post_asap::*, + pre_asap::{CompareOpKind, JoinKind, Predicate, QueryExpr}, + }; + use std::{collections::BTreeMap, rc::Rc}; + let schema = schema(&[("key", DataType::Utf8, false)]); + for certified in [false, true] { + let node = PostAsapDagNode { + id: PostAsapNodeId(2), + output_schema: (*schema).clone(), + output_state: ExecutionDataState::QUERY_ROWS, + guarantee: None, + payload: PostAsapOperatorPayload::RelationalJoin { + join_kind: JoinKind::Semi, + pred: Predicate(Rc::new(QueryExpr::Compare { + left: Rc::new(QueryExpr::Column(0)), + op: CompareOpKind::Eq, + right: Rc::new(QueryExpr::Column(1)), + })), + pruning: certified.then_some(CandidateCompleteness::Certified { + guarantee: ResultGuarantee { + metric: ErrorMetric::TopKMembership, + bound: BoundExpr::Zero, + failure_probability: ProbabilityExpr::Constant { value: 0.01 }, + provenance: vec![], + }, + }), + }, + }; + let graph = CompiledPhysicalDag::from_operators( + [ + (0, InputContract::bounded(schema.clone())), + (1, InputContract::bounded(schema.clone())), + ] + .into(), + [( + 2, + ( + vec![0, 1], + compile_node(&node, &[schema.clone(), schema.clone()]).unwrap(), + ), + )] + .into(), + vec![2], + ) + .unwrap(); + let graph = + serde_json::from_slice::(&serde_json::to_vec(&graph).unwrap()) + .unwrap(); + assert_eq!( + graph.certified_pruning_keys(2), + certified.then_some(&[(0, 0)][..]) + ); + for complete in [false, true] { + let sources = [ + vec!["a"], + if complete { + vec!["a"] + } else { + vec!["a", "missing"] + }, + ] + .into_iter() + .enumerate() + .map(|(i, keys)| { + let batch = Batch::try_new( + schema.clone(), + keys.into_iter() + .map(|k| vec![Value::Utf8(k.into())]) + .collect(), + ) + .unwrap(); + ( + i as u64, + Box::new(Operator::source(schema.clone(), vec![batch]).unwrap()) as Source<'_>, + ) + }) + .collect::>(); + let bound = graph.instantiate(sources).unwrap(); + let result = block_on(async { + let mut stream = bound + .execute( + graph.roots(), + RunContext::new(query(), Limits::default()).unwrap(), + ) + .unwrap() + .remove(0); + let mut rows = vec![]; + while let Some(batch) = stream.next().await { + rows.extend(batch?.rows().iter().cloned()); + } + Ok::<_, asap_physical_operators::Error>(rows) + }); + if certified && !complete { + assert!(result + .unwrap_err() + .to_string() + .contains("no authoritative value")); + } else { + let rows = result.unwrap(); + assert_eq!(rows.len(), 1); + assert!(matches!(&rows[0][0], Value::Utf8(key) if key.as_ref() == "a")); + } + } + } +} + +// Precompute arithmetic must match population/window identities, never zip arrival order. +#[test] +fn compiled_ingestion_binary_preserves_alignment_and_rejects_missing_updates() { + use asap_physical_operators::physical_planner::{ + compile_node, CompiledPhysicalDag, InputContract, Source, + }; + use planner_types::{ + post_asap::*, + pre_asap::{ArithmeticOpKind, BinaryOpKind}, + }; + use std::collections::BTreeMap; + let input = schema(&[ + ("population", DataType::Utf8, false), + ("time", DataType::Timestamp, false), + ("value", DataType::Float64, false), + ]); + let node = PostAsapDagNode { + id: PostAsapNodeId(2), + output_schema: (*input).clone(), + output_state: ExecutionDataState::INGESTION_ROWS, + guarantee: None, + payload: PostAsapOperatorPayload::Binary { + operator: BinaryOperator { + kind: BinaryOpKind::Arithmetic(ArithmeticOpKind::Sub), + vector_match: None, + checked_relative_division: false, + checked_finite_division: false, + }, + }, + }; + let program = CompiledPhysicalDag::from_operators( + [ + (0, InputContract::bounded(input.clone())), + (1, InputContract::bounded(input.clone())), + ] + .into(), + [( + 2, + ( + vec![0, 1], + compile_node(&node, &[input.clone(), input.clone()]).unwrap(), + ), + )] + .into(), + vec![2], + ) + .unwrap(); + let program = + serde_json::from_slice::(&serde_json::to_vec(&program).unwrap()) + .unwrap(); + for (right, expected) in [ + (vec![("b", 2, 3.), ("a", 1, 2.)], Some(vec![8., 17.])), + (vec![("b", 2, 3.)], None), + (vec![("a", 1, 2.), ("a", 1, 2.)], None), + (vec![("a", 2, 2.), ("b", 1, 3.)], None), + (vec![("a", 1, f64::NAN), ("b", 2, 3.)], None), + ] { + let sources = [vec![("a", 1, 10.), ("b", 2, 20.)], right] + .into_iter() + .enumerate() + .map(|(i, rows)| { + let rows = rows + .into_iter() + .map(|(group, time, value)| { + vec![ + Value::Utf8(group.into()), + Value::Timestamp(time), + Value::Float64(value), + ] + }) + .collect(); + let batch = Batch::try_new(input.clone(), rows).unwrap(); + ( + i as u64, + Box::new(Operator::source(input.clone(), vec![batch]).unwrap()) as Source<'_>, + ) + }) + .collect::>(); + let graph = program.instantiate(sources).unwrap(); + let result = block_on(async { + let mut stream = graph + .execute( + program.roots(), + RunContext::new(query(), Limits::default()).unwrap(), + ) + .unwrap() + .remove(0); + let mut rows = Vec::new(); + while let Some(batch) = stream.next().await { + rows.extend(batch?.rows().iter().cloned()); + } + Ok::<_, asap_physical_operators::Error>(rows) + }); + match expected { + Some(values) => assert_eq!(floats(&result.unwrap(), 2), values), + None => assert!(result.is_err()), + } + } +} diff --git a/crates/asap-physical-operators/tests/physical_plan_recovery.rs b/crates/asap-physical-operators/tests/physical_plan_recovery.rs new file mode 100644 index 00000000..5a920a8d --- /dev/null +++ b/crates/asap-physical-operators/tests/physical_plan_recovery.rs @@ -0,0 +1,105 @@ +//! Deserialized physical plans recover selected operators without logical lowering. +//! Deployments choose the encoding; JSON is used here only as a test format. +use asap_physical_operators::{ + operators::{Operator, SortKey}, + physical_planner::{CompiledPhysicalDag, InputContract}, +}; +use planner_types::{ + post_asap::{SummaryFamilyType, SummaryField, SummarySchema}, + pre_asap::DataType, +}; +use std::{collections::BTreeMap, sync::Arc}; + +fn sorted() -> CompiledPhysicalDag { + let schema = Arc::new(SummarySchema { + fields: vec![SummaryField { + name: "value".into(), + dtype: SummaryFamilyType::Plain(DataType::Float64), + nullable: false, + }], + time_index: None, + }); + CompiledPhysicalDag::from_operators( + BTreeMap::from([(0, InputContract::bounded(schema.clone()))]), + BTreeMap::from([( + 1, + ( + vec![0], + Operator::sort( + schema, + vec![SortKey { + column: 0, + descending: true, + nulls_first: false, + }], + vec![], + ) + .unwrap(), + ), + )]), + vec![1], + ) + .unwrap() +} + +#[test] +fn recovery_retains_selected_operator_and_rejects_invalid_contracts() { + let bytes = serde_json::to_vec(&sorted()).unwrap(); + let recovered = serde_json::from_slice::(&bytes).unwrap(); + assert_eq!(serde_json::to_vec(&recovered).unwrap(), bytes); + for mutation in ["column", "edge", "output"] { + let mut wire: serde_json::Value = serde_json::from_slice(&bytes).unwrap(); + match mutation { + "column" => { + wire["nodes"]["1"]["Operator"]["operator"]["kind"]["Sort"]["keys"][0]["column"] = + 7.into() + } + "edge" => wire["nodes"]["1"]["Operator"]["inputs"][0] = 999.into(), + "output" => { + wire["nodes"]["1"]["Operator"]["operator"]["output"]["fields"][0]["dtype"] = + serde_json::json!({"Plain":"utf8"}) + } + _ => unreachable!(), + } + assert!( + serde_json::from_slice::(&serde_json::to_vec(&wire).unwrap()) + .is_err(), + "accepted {mutation}" + ); + } +} + +#[test] +fn candidate_recovery_preserves_materialization_boundary() { + use asap_physical_operators::physical_planner::PhysicalCandidate; + let precompute = sorted(); + let output = InputContract::bounded(precompute.output_contract(1).unwrap().schema); + let query = CompiledPhysicalDag::from_operators( + BTreeMap::from([(1, output.clone())]), + BTreeMap::from([( + 2, + ( + vec![1], + Operator::limit(output.schema.clone(), 3, 0, vec![]).unwrap(), + ), + )]), + vec![2], + ) + .unwrap(); + let candidate = PhysicalCandidate { + precompute: Some(precompute), + query, + materialized_outputs: BTreeMap::from([(1, output)]), + }; + let bytes = serde_json::to_vec(&candidate).unwrap(); + let restored = serde_json::from_slice::(&bytes).unwrap(); + assert_eq!(restored.precompute.as_ref().unwrap().roots(), &[1]); + assert_eq!(restored.query.roots(), &[2]); + assert_eq!(serde_json::to_vec(&restored).unwrap(), bytes); + let mut wire: serde_json::Value = serde_json::from_slice(&bytes).unwrap(); + wire["materialized_outputs"]["1"]["schema"]["fields"][0]["dtype"] = + serde_json::json!({"Plain":"utf8"}); + assert!( + serde_json::from_slice::(&serde_json::to_vec(&wire).unwrap()).is_err() + ); +} diff --git a/crates/asap-physical-operators/tests/physical_semantics.rs b/crates/asap-physical-operators/tests/physical_semantics.rs new file mode 100644 index 00000000..5770d7d6 --- /dev/null +++ b/crates/asap-physical-operators/tests/physical_semantics.rs @@ -0,0 +1,695 @@ +//! Contract tests inspired by DataFusion's limit, sort and join test matrices. +//! Expectations follow ASAP's IR (notably row-count and IEEE NaN equality). +//! Reference: apache/datafusion e2ca7f3, physical-plan/src/{limit.rs,sorts/sort.rs}. +use asap_physical_operators::{ + expressions::CompiledExpression, + operators::{Expression, Operator, Reduction, SortKey}, + plan::PhysicalDag, + runtime::{Limits, RunContext, Scope}, + values::{Batch, Schema, Value}, +}; +use futures::{executor::block_on, StreamExt}; +use planner_types::{ + post_asap::{SummaryFamilyType, SummaryField, SummarySchema}, + pre_asap::{CompareOpKind, DataType, JoinKind, Predicate, QueryExpr}, +}; +use std::{rc::Rc, sync::Arc}; + +fn schema(fields: &[(&str, DataType, bool)]) -> Schema { + Arc::new(SummarySchema { + fields: fields + .iter() + .map(|(name, dtype, nullable)| SummaryField { + name: (*name).into(), + dtype: SummaryFamilyType::Plain(dtype.clone()), + nullable: *nullable, + }) + .collect(), + time_index: None, + }) +} +fn context() -> RunContext { + RunContext::new( + Scope::Query { + evaluation_time_ms: 0, + revision: 1, + }, + Limits { + max_buffered_batches: 1, + ..Limits::default() + }, + ) + .unwrap() +} +fn collect(dag: &PhysicalDag<'_, Batch, Schema>, root: u64) -> Vec> { + let run = context(); + let rows = block_on(async { + let mut stream = dag.execute(&[root], run.clone()).unwrap().remove(0); + let mut rows = vec![]; + while let Some(batch) = stream.next().await { + rows.extend_from_slice(batch.unwrap().rows()); + } + rows + }); + assert_eq!(run.retained_bytes(), 0); + rows +} +fn unary(input: Schema, batches: Vec>>, op: Operator) -> Vec> { + let mut dag = PhysicalDag::default(); + let batches = batches + .into_iter() + .map(|rows| Batch::try_new(input.clone(), rows).unwrap()) + .collect(); + dag.add(0, vec![], Operator::source(input, batches).unwrap()) + .unwrap(); + dag.add(1, vec![0], op).unwrap(); + collect(&dag, 1) +} +fn keys(rows: &[Vec]) -> Vec>> { + rows.iter() + .map(|r| r.iter().map(|v| v.key().unwrap()).collect()) + .collect() +} +fn eq_predicate() -> Predicate { + Predicate(Rc::new(QueryExpr::Compare { + left: Rc::new(QueryExpr::Column(0)), + op: CompareOpKind::Eq, + right: Rc::new(QueryExpr::Column(1)), + })) +} +fn join(left: Vec, right: Vec, kind: JoinKind, keyed: bool) -> Vec> { + let input = schema(&[("key", DataType::Float64, true)]); + let output = if matches!(kind, JoinKind::Semi | JoinKind::Anti) { + input.clone() + } else { + schema(&[ + ("left", DataType::Float64, true), + ("right", DataType::Float64, true), + ]) + }; + let op = if keyed { + Operator::semi_join(input.clone(), input.clone(), vec![(0, 0)]).unwrap() + } else { + Operator::relational_join(input.clone(), input.clone(), kind, &eq_predicate(), output) + .unwrap() + }; + let mut dag = PhysicalDag::default(); + for (id, values) in [(0, left), (1, right)] { + let batches = values + .into_iter() + .map(|v| Batch::try_new(input.clone(), vec![vec![v]]).unwrap()) + .collect(); + dag.add( + id, + vec![], + Operator::source(input.clone(), batches).unwrap(), + ) + .unwrap(); + } + dag.add(2, vec![0, 1], op).unwrap(); + collect(&dag, 2) +} + +// OFFSET/FETCH must be invariant to empty batches and input batch boundaries. +#[test] +fn limit_offset_fetch_matrix() { + let input = schema(&[("v", DataType::Int64, false)]); + for chunk in [1, 2, 5, 12] { + let values = (0..9).map(|n| vec![Value::Int64(n)]).collect::>(); + let mut batches = vec![vec![]]; + for rows in values.chunks(chunk) { + batches.push(rows.to_vec()); + batches.push(vec![]); + } + for offset in [0, 1, 8, 9, 10, u64::MAX] { + for n in [0, 1, 3, 12, u64::MAX] { + let rows = unary( + input.clone(), + batches.clone(), + Operator::limit(input.clone(), n, offset, vec![]).unwrap(), + ); + let expected = values + .iter() + .skip(offset.min(9) as usize) + .take(n.min(9) as usize) + .cloned() + .collect::>(); + assert_eq!( + keys(&rows), + keys(&expected), + "chunk={chunk}, offset={offset}, n={n}" + ); + } + } + } +} + +// Zero-column batches still have rows: LIMIT must not infer cardinality from columns. +#[test] +fn limit_preserves_zero_column_row_count() { + let input = schema(&[]); + let rows = unary( + input.clone(), + vec![vec![vec![]; 5], vec![vec![]; 5]], + Operator::limit(input, 4, 3, vec![]).unwrap(), + ); + assert_eq!(rows.len(), 4); +} + +// NULL placement is independent of sort direction; ties retain original row order. +#[test] +fn sort_direction_null_placement_and_ties() { + let input = schema(&[("v", DataType::Int64, true), ("id", DataType::Int64, false)]); + let values = [Some(2), None, Some(1), Some(2), None]; + let rows = values + .iter() + .enumerate() + .map(|(i, v)| { + vec![ + v.map(Value::Int64).unwrap_or(Value::Null), + Value::Int64(i as i64), + ] + }) + .collect::>(); + for (descending, nulls_first, expected) in [ + (false, false, vec![2, 0, 3, 1, 4]), + (false, true, vec![1, 4, 2, 0, 3]), + (true, false, vec![0, 3, 2, 1, 4]), + (true, true, vec![1, 4, 0, 3, 2]), + ] { + let op = Operator::sort( + input.clone(), + vec![SortKey { + column: 0, + descending, + nulls_first, + }], + vec![], + ) + .unwrap(); + let result = unary( + input.clone(), + vec![rows[..2].to_vec(), vec![], rows[2..].to_vec()], + op, + ); + let ids = result + .iter() + .map(|r| match r[1] { + Value::Int64(n) => n, + _ => unreachable!(), + }) + .collect::>(); + assert_eq!(ids, expected); + } +} + +// Outer joins preserve unmatched NULLs, while semi/anti joins preserve left multiplicity. +#[test] +fn joins_nulls_duplicates_and_empty_sides() { + for (kind, expected_len) in [ + (JoinKind::Inner, 4), + (JoinKind::Left, 6), + (JoinKind::Right, 6), + (JoinKind::Full, 8), + (JoinKind::Semi, 2), + (JoinKind::Anti, 2), + ] { + let left = vec![ + Value::Float64(1.), + Value::Float64(1.), + Value::Float64(2.), + Value::Null, + ]; + let right = vec![ + Value::Float64(1.), + Value::Float64(1.), + Value::Float64(3.), + Value::Null, + ]; + let result = join(left, right, kind.clone(), false); + assert_eq!(result.len(), expected_len, "{kind:?}"); + } + for (kind, expected_len) in [ + (JoinKind::Inner, 0), + (JoinKind::Left, 1), + (JoinKind::Right, 0), + (JoinKind::Full, 1), + (JoinKind::Semi, 0), + (JoinKind::Anti, 1), + ] { + assert_eq!( + join(vec![Value::Float64(7.)], vec![], kind.clone(), false).len(), + expected_len, + "{kind:?}" + ); + } + let result = join(vec![Value::Float64(7.)], vec![], JoinKind::Left, false); + assert!(matches!( + result[0].as_slice(), + [Value::Float64(7.), Value::Null] + )); +} + +// Changing the semi-join algorithm must not turn IEEE NaN != NaN into a match. +#[test] +fn keyed_semijoin_obeys_ieee_equality_for_nan_and_zero() { + let left = vec![ + Value::Float64(f64::NAN), + Value::Float64(-0.), + Value::Float64(0.), + Value::Null, + ]; + let right = vec![Value::Float64(f64::NAN), Value::Float64(0.), Value::Null]; + let keyed = join(left, right, JoinKind::Semi, true); + let expected = vec![vec![Value::Float64(-0.)], vec![Value::Float64(0.)]]; + assert_eq!(keys(&keyed), keys(&expected)); +} + +// Group equality intentionally differs from predicate equality: NULL and NaNs group together. +#[test] +fn grouping_canonicalizes_null_nan_and_signed_zero() { + let input = schema(&[("v", DataType::Float64, true)]); + let op = Operator::aggregate( + input.clone(), + vec![0], + vec![("count".into(), Reduction::Count)], + ) + .unwrap(); + let values = vec![ + Value::Null, + Value::Null, + Value::Float64(0.), + Value::Float64(-0.), + Value::Float64(f64::NAN), + Value::Float64(f64::from_bits(0x7ff8000000000001)), + ]; + let result = unary( + input, + values.into_iter().map(|v| vec![vec![v]]).collect(), + op, + ); + assert_eq!(result.len(), 3); + assert!(result.iter().all(|r| matches!(r[1], Value::Int64(2)))); +} + +// Global empty input yields one aggregate row; grouped empty input yields none. +#[test] +fn aggregate_empty_and_all_null_follow_asap_contract() { + let input = schema(&[("v", DataType::Int64, true)]); + for batches in [ + vec![], + vec![vec![]], + vec![vec![vec![Value::Null], vec![Value::Null]]], + ] { + let n = batches.iter().map(Vec::len).sum::(); + let op = Operator::aggregate( + input.clone(), + vec![], + vec![ + ("count".into(), Reduction::Count), + ("min".into(), Reduction::Min(0)), + ("max".into(), Reduction::Max(0)), + ], + ) + .unwrap(); + let result = unary(input.clone(), batches, op); + assert_eq!(result.len(), 1); + assert!(matches!(result[0][0], Value::Int64(v) if v == n as i64)); + assert!(matches!(result[0][1], Value::Null)); + assert!(matches!(result[0][2], Value::Null)); + } + let op = Operator::aggregate( + input.clone(), + vec![0], + vec![("count".into(), Reduction::Count)], + ) + .unwrap(); + assert!(unary(input, vec![], op).is_empty()); +} + +// A precompiled expression with a different input contract must fail during binding. +#[test] +fn projection_rejects_expression_bound_to_another_schema() { + let original = schema(&[("a", DataType::Int64, false), ("b", DataType::Int64, false)]); + let current = schema(&[("a", DataType::Int64, false)]); + let expr = CompiledExpression::compile(&QueryExpr::Column(1), &original).unwrap(); + assert!(Operator::project(current, vec![("b".into(), Expression::planner(expr))]).is_err()); +} + +// A valid Planner MIN/MAX schema must bind even for a non-null input column. +#[test] +fn global_extrema_bind_with_planner_derived_schema() { + use asap_physical_operators::physical_planner::compile_node; + use planner_types::{ + post_asap::*, + pre_asap::{AggIntent, Column, GroupKeys, Reduction as PlanReduction}, + }; + let input = schema(&[("v", DataType::Int64, false)]); + for measure in [ + AggIntent::Min { col: Some(0) }, + AggIntent::Max { col: Some(0) }, + ] { + let planner_input = + planner_types::pre_asap::Schema::new(vec![Column::new("v", DataType::Int64, false)]); + let derived = planner_types::pre_asap::query_expr::aggregate_output_schema( + &planner_input, + &PlanReduction::Reduce(GroupKeys::by(vec![])), + std::slice::from_ref(&measure), + &[], + ) + .unwrap(); + let result = derived.columns[0].clone(); + let output = schema(&[(&result.name, result.dtype, result.nullable)]); + let node = PostAsapDagNode { + id: PostAsapNodeId(1), + payload: PostAsapOperatorPayload::Value { + operation: ValueOperation::Exact(ExactOperation::Aggregate { + reduction: PlanReduction::Reduce(GroupKeys::by(vec![])), + measures: vec![measure], + output_names: vec![result.name], + having: None, + }), + }, + output_state: ExecutionDataState::QUERY_ROWS, + output_schema: (*output).clone(), + guarantee: None, + }; + let operator = compile_node(&node, std::slice::from_ref(&input)) + .expect("global extremum should bind to its Planner schema"); + assert!(operator.schema().fields[0].nullable); + let empty = unary(input.clone(), vec![], operator.clone()); + assert!(matches!(empty[0][0], Value::Null)); + let nonempty = unary(input.clone(), vec![vec![vec![Value::Int64(7)]]], operator); + assert!(matches!(nonempty[0][0], Value::Int64(7))); + } +} + +// NaN is a valid numeric input, not a schema error; all six comparisons obey IEEE rules. +#[test] +fn planner_comparisons_handle_nan_without_execution_errors() { + let input = schema(&[ + ("a", DataType::Float64, false), + ("b", DataType::Float64, false), + ]); + for op in [ + CompareOpKind::Eq, + CompareOpKind::Ne, + CompareOpKind::Lt, + CompareOpKind::Le, + CompareOpKind::Gt, + CompareOpKind::Ge, + ] { + let expression = QueryExpr::Compare { + left: Rc::new(QueryExpr::Column(0)), + op: op.clone(), + right: Rc::new(QueryExpr::Column(1)), + }; + let compiled = CompiledExpression::compile(&expression, &input).unwrap(); + for row in [ + [Value::Float64(f64::NAN), Value::Float64(1.)], + [Value::Float64(1.), Value::Float64(f64::NAN)], + [Value::Float64(f64::NAN), Value::Float64(f64::NAN)], + ] { + let actual = compiled.evaluate(&row).unwrap(); + assert!(matches!(actual,Value::Bool(value) if value == (op == CompareOpKind::Ne))); + } + } +} + +// A bounded LIMIT branch must unsubscribe so another branch can drain the producer. +#[test] +fn limit_branch_finishes_without_blocking_shared_sibling() { + let input = schema(&[("v", DataType::Int64, false)]); + let mut dag = PhysicalDag::default(); + let batches = (0..100) + .map(|v| Batch::try_new(input.clone(), vec![vec![Value::Int64(v)]]).unwrap()) + .collect(); + dag.add(0, vec![], Operator::source(input.clone(), batches).unwrap()) + .unwrap(); + dag.add( + 1, + vec![0], + Operator::limit(input.clone(), 1, 0, vec![]).unwrap(), + ) + .unwrap(); + dag.add(2, vec![0, 1], Operator::union(input, 2).unwrap()) + .unwrap(); + // Bound polls as well as rows so a backpressure regression cannot hang the suite. + use futures::{task::noop_waker_ref, Stream}; + use std::{ + pin::Pin, + task::{Context, Poll}, + }; + let run = context(); + let mut stream = dag.execute(&[2], run.clone()).unwrap().remove(0); + let mut cx = Context::from_waker(noop_waker_ref()); + let mut count = 0; + for _ in 0..2000 { + match Pin::new(&mut stream).poll_next(&mut cx) { + Poll::Ready(Some(batch)) => count += batch.unwrap().rows().len(), + Poll::Ready(None) => { + assert_eq!(count, 101); + drop(stream); + assert_eq!(run.retained_bytes(), 0); + return; + } + Poll::Pending => {} + } + } + panic!("shared LIMIT/Union failed to make progress"); +} + +// Mixed numeric comparisons must not round Int64 values through f64 before comparing. +#[test] +fn mixed_numeric_comparisons_preserve_large_integer_precision() { + let input = schema(&[ + ("a", DataType::Int64, false), + ("b", DataType::Float64, false), + ]); + let expr = QueryExpr::Compare { + left: Rc::new(QueryExpr::Column(0)), + op: CompareOpKind::Gt, + right: Rc::new(QueryExpr::Column(1)), + }; + let compiled = CompiledExpression::compile(&expr, &input).unwrap(); + for (a, b, expected) in [ + (9_007_199_254_740_993, 9_007_199_254_740_992.0, true), + (i64::MAX, 9_223_372_036_854_775_808.0, false), + (i64::MIN, f64::NEG_INFINITY, true), + ] { + assert!( + matches!(compiled.evaluate(&[Value::Int64(a),Value::Float64(b)]).unwrap(), Value::Bool(v) if v == expected) + ); + } +} + +// Both expression paths must implement all nine combinations of three-valued booleans. +#[test] +fn boolean_truth_tables_agree_between_expression_paths() { + let input = schema(&[("a", DataType::Bool, true), ("b", DataType::Bool, true)]); + for and in [true, false] { + for a in [None, Some(false), Some(true)] { + for b in [None, Some(false), Some(true)] { + let parts = vec![QueryExpr::Column(0), QueryExpr::Column(1)]; + let planner = if and { + QueryExpr::BoolAnd(parts) + } else { + QueryExpr::BoolOr(parts) + }; + let native = if and { + Expression::And( + Box::new(Expression::Column(0)), + Box::new(Expression::Column(1)), + ) + } else { + Expression::Or( + Box::new(Expression::Column(0)), + Box::new(Expression::Column(1)), + ) + }; + let expected = match (a, b, and) { + (Some(false), _, true) | (_, Some(false), true) => Some(false), + (Some(true), _, false) | (_, Some(true), false) => Some(true), + (None, _, _) | (_, None, _) => None, + (Some(a), Some(b), true) => Some(a && b), + (Some(a), Some(b), false) => Some(a || b), + } + .map(Value::Bool) + .unwrap_or(Value::Null); + let row = vec![ + a.map(Value::Bool).unwrap_or(Value::Null), + b.map(Value::Bool).unwrap_or(Value::Null), + ]; + let compiled = CompiledExpression::compile(&planner, &input).unwrap(); + assert_eq!( + compiled.evaluate(&row).unwrap().key().unwrap(), + expected.key().unwrap() + ); + let op = Operator::project(input.clone(), vec![("result".into(), native)]).unwrap(); + let result = unary(input.clone(), vec![vec![row]], op); + assert_eq!(result[0][0].key().unwrap(), expected.key().unwrap()); + } + } + } +} + +// Partial/final execution must agree with one build for an uncompacted KLL population. +#[test] +fn kll_partial_merge_and_multiple_readouts_preserve_population() { + use planner_types::post_asap::{SketchAlgorithm, SketchKind, SketchParams}; + let input = schema(&[("v", DataType::Float64, false)]); + let family = SummaryFamilyType::Sketch( + SketchKind::new(SketchAlgorithm::Kll, SketchParams::Kll { k: 512 }), + Default::default(), + ); + let mut dag = PhysicalDag::default(); + for (id, range) in [(0, 0..64), (1, 64..128), (2, 0..128)] { + let rows = range.map(|n| vec![Value::Float64(n as f64)]).collect(); + dag.add( + id, + vec![], + Operator::source( + input.clone(), + vec![Batch::try_new(input.clone(), rows).unwrap()], + ) + .unwrap(), + ) + .unwrap(); + dag.add( + id + 3, + vec![id], + Operator::summary_build(input.clone(), family.clone(), 0, None, vec![]).unwrap(), + ) + .unwrap(); + } + let state = Operator::summary_build(input, family, 0, None, vec![]) + .unwrap() + .schema(); + dag.add(6, vec![3, 4], Operator::union(state.clone(), 2).unwrap()) + .unwrap(); + dag.add( + 7, + vec![6], + Operator::summary_merge(state.clone(), 0, vec![]).unwrap(), + ) + .unwrap(); + let mut roots = vec![]; + for (i, q) in [0.0, 0.5, 1.0].into_iter().enumerate() { + for (j, build) in [5, 7].into_iter().enumerate() { + let id = 8 + (i * 2 + j) as u64; + dag.add( + id, + vec![build], + Operator::readout( + state.clone(), + 0, + asap_physical_operators::operators::ReadoutQuery::Sketch( + planner_types::post_asap::SketchQuery::Quantile { q }, + ), + ) + .unwrap(), + ) + .unwrap(); + roots.push(id); + } + } + for _ in 0..2 { + let run = context(); + let outputs = block_on(futures::future::join_all( + dag.execute(&roots, run.clone()) + .unwrap() + .into_iter() + .map(|s| s.collect::>()), + )); + for (pair, expected) in outputs.chunks(2).zip([0., 64., 127.]) { + let value = |batches: &[Result< + asap_physical_operators::runtime::SharedValue, + asap_physical_operators::Error, + >]| { + assert_eq!(batches.len(), 1); + match batches[0].as_ref().unwrap().rows()[0][0] { + Value::Float64(v) => v, + _ => panic!("quantile must be Float64"), + } + }; + assert_eq!(value(&pair[0]), value(&pair[1])); + assert!((value(&pair[0]) - expected).abs() <= 1.); + } + drop(outputs); + assert_eq!(run.retained_bytes(), 0); + } +} + +// Retained zero-column rows still own Vec headers and must consume the output budget. +#[test] +fn zero_column_output_obeys_memory_limit() { + use asap_physical_operators::Error; + let input = schema(&[]); + let batch = Batch::try_new(input.clone(), vec![vec![]; 200]).unwrap(); + let mut dag = PhysicalDag::default(); + dag.add(0, vec![], Operator::source(input, vec![batch]).unwrap()) + .unwrap(); + let run = RunContext::new( + Scope::Query { + evaluation_time_ms: 0, + revision: 0, + }, + Limits { + max_bytes: 1024, + ..Limits::default() + }, + ) + .unwrap(); + let mut stream = dag.execute(&[0], run.clone()).unwrap().remove(0); + assert!(matches!( + block_on(stream.next()), + Some(Err(Error::MemoryLimit)) + )); + drop(stream); + assert_eq!(run.retained_bytes(), 0); +} + +// Empty exact-state finalization must preserve ordinary global MIN/MAX null semantics. +#[test] +fn empty_exact_summary_extrema_agree_with_ordinary_aggregation() { + use asap_physical_operators::Statistic; + use planner_types::post_asap::{ExactKind, ExactParams}; + let input = schema(&[("v", DataType::Float64, false)]); + for (kind, params, statistic) in [ + (ExactKind::Min, ExactParams::Min, Statistic::Min), + (ExactKind::Max, ExactParams::Max, Statistic::Max), + ] { + let build = Operator::summary_build( + input.clone(), + SummaryFamilyType::ExactAggregate(kind, params), + 0, + None, + vec![], + ) + .unwrap(); + let state = build.schema(); + let mut dag = PhysicalDag::default(); + dag.add(0, vec![], Operator::source(input.clone(), vec![]).unwrap()) + .unwrap(); + dag.add(1, vec![0], build).unwrap(); + dag.add( + 2, + vec![1], + Operator::readout( + state, + 0, + asap_physical_operators::operators::ReadoutQuery::Exact( + asap_physical_operators::summary_kernels::exact::ExactReadout { + statistic, + lookback_ms: None, + }, + ), + ) + .unwrap(), + ) + .unwrap(); + let rows = collect(&dag, 2); + assert_eq!(rows.len(), 1); + assert!(matches!(rows[0][0], Value::Null)); + } +} diff --git a/crates/asap-physical-operators/tests/precompute_candidates.rs b/crates/asap-physical-operators/tests/precompute_candidates.rs new file mode 100644 index 00000000..c00b388f --- /dev/null +++ b/crates/asap-physical-operators/tests/precompute_candidates.rs @@ -0,0 +1,548 @@ +//! Materialized frontiers are compiled by Planner, never rewritten by deployment. +use asap_aware_mapping::{cost_model::DefaultCostModel, search_workload}; +use asap_physical_operators::{ + factory::create_planner_accumulator, + operators::Operator, + physical_planner::{ + compile_candidates, select_candidate, CandidateCost, CompiledPhysicalDag, InputContract, + Source, + }, + runtime::{Limits, RunContext, Scope}, + values::{Batch, Value}, +}; +use futures::{executor::block_on, StreamExt}; +use planner_types::{post_asap::*, pre_asap::DataType, types::AccuracyTarget, workload::*}; +use std::{collections::BTreeMap, rc::Rc, sync::Arc}; + +fn grouped_rate_space() -> asap_aware_mapping::PlanSpace<&'static str> { + let workload = PlanningWorkload { + query_workload: QueryWorkload { + language: QueryLanguage::PromQL, + query_batch: Some(vec![BatchEntry { + query: Query("sum by(job)(rate(m[1m]))".into()), + requirements: QueryRequirements { + accuracy: AccuracyRequirement::Explicit(AccuracyTarget::Exact), + ..Default::default() + }, + predictability: Predictability::Unknown, + invocations: 1, + execute_at: None, + time_selection: TimeSelection::default(), + }]), + repeating_queries: None, + }, + data_workload: Some(DataWorkload { + data_ingestion_interval: Evidence { + value: Some(DurationMs(1000)), + ..Default::default() + }, + ..Default::default() + }), + }; + let root = Rc::new( + asap_frontend_promql::lower_promql_workload(&workload, 0) + .unwrap() + .remove(0), + ); + let root = Rc::new( + asap_physical_operators::physical_planner::promql_rows::with_series_identity(&root) + .unwrap(), + ); + search_workload(vec![("grouped-rate", root)]) +} + +fn grouped_rate() -> PostAsapDag { + let space = grouped_rate_space(); + let selected = space + .global_selection(&DefaultCostModel) + .assemble_selected_dag(&space.roots[0].1) + .unwrap() + .unwrap(); + compile_post_asap_dag(&selected).unwrap() +} +fn run(plan: &CompiledPhysicalDag, inputs: BTreeMap, scope: Scope) -> Vec { + let sources = inputs + .into_iter() + .map(|(id, batch)| { + let source = Operator::source(batch.schema().clone(), vec![batch]).unwrap(); + (id, Box::new(source) as Source<'_>) + }) + .collect(); + let dag = plan.instantiate(sources).unwrap(); + block_on(async { + let context = RunContext::new(scope, Limits::default()).unwrap(); + let mut output = dag.execute(plan.roots(), context).unwrap().remove(0); + let mut batches = vec![]; + while let Some(batch) = output.next().await { + batches.push((*batch.unwrap()).clone()); + } + batches + }) +} + +/// Rate readouts and grouped Sum can run together during bounded precompute; +/// storing per-series rates instead leaves the same Sum in the query DAG. +#[test] +fn grouped_rate_can_be_materialized_before_or_after_grouped_sum() { + let dag = grouped_rate(); + let state = dag + .nodes + .iter() + .find(|node| { + matches!( + node.payload, + PostAsapOperatorPayload::SummaryAgg { + family: SummaryFamilyType::ExactAggregate(ExactKind::Rate, _), + .. + } + ) + }) + .unwrap(); + let readout = dag + .nodes + .iter() + .find(|node| { + matches!( + node.payload, + PostAsapOperatorPayload::Value { + operation: ValueOperation::FinalizeExactAccumulator + } + ) + }) + .unwrap(); + let input_schema = Arc::new(state.output_schema.clone()); + let (family, update, grouping) = match &state.payload { + PostAsapOperatorPayload::SummaryAgg { + family, + input, + grouping, + .. + } => (family, input, grouping), + _ => unreachable!(), + }; + let range_ms = Some((-58_000, 2_000)); + let mut expected_rate_sum = 0.; + let rows = [[100., 0., 100.], [100., 200., 0.]] + .into_iter() + .enumerate() + .map(|(index, values)| { + let mut accumulator = create_planner_accumulator(family, update, grouping).unwrap(); + for (i, value) in values.into_iter().enumerate() { + accumulator.update_single(value, i as i64 * 1000); + } + let state = accumulator.into_accumulator(); + expected_rate_sum += state + .as_any() + .downcast_ref::() + .unwrap() + .readout(asap_physical_operators::Statistic::Rate, range_ms, None) + .unwrap() + .unwrap(); + let summary = Value::Summary { + family: family.clone(), + state: Arc::from(state), + }; + input_schema + .fields + .iter() + .map(|field| match &field.dtype { + SummaryFamilyType::ExactAggregate(..) => summary.clone(), + SummaryFamilyType::Plain(DataType::Timestamp) => Value::Timestamp(2000), + SummaryFamilyType::Plain(DataType::Utf8) => { + Value::Utf8(if field.name == "job" { + "api".into() + } else { + format!("series-{index}").into() + }) + } + _ => panic!("unexpected input field {field:?}"), + }) + .collect() + }) + .collect(); + let batch = Batch::try_new(input_schema.clone(), rows).unwrap(); + let root = u64::from(dag.root.0); + let state_id = u64::from(state.id.0); + let rate_id = u64::from(readout.id.0); + let frontiers = asap_physical_operators::physical_planner::enumerate_frontiers( + &dag, + &BTreeMap::from([(state_id, InputContract::bounded(input_schema.clone()))]), + &[root], + 128, + ) + .unwrap(); + assert!(frontiers.contains(&vec![])); + assert!(frontiers.contains(&vec![rate_id])); + assert!(frontiers.contains(&vec![root])); + assert!(!frontiers.contains(&vec![root, rate_id])); + assert!( + asap_physical_operators::physical_planner::enumerate_frontiers( + &dag, + &BTreeMap::from([(state_id, InputContract::bounded(input_schema.clone()))]), + &[root], + 1, + ) + .is_err() + ); + let candidates = compile_candidates( + &dag, + BTreeMap::from([(state_id, InputContract::bounded(input_schema))]), + &[root], + &[vec![], vec![rate_id], vec![root]], + ); + // Scoped cost fixtures select either precompute boundary. No readers are + // opened during candidate construction or selection. + for prefer_grouped in [false, true] { + let inventory = compile_candidates( + &dag, + BTreeMap::from([( + state_id, + InputContract::bounded(Arc::new(state.output_schema.clone())), + )]), + &[root], + &[vec![999], vec![rate_id], vec![root]], + ); + assert!(inventory[0].is_err()); + let mut evaluated = 0; + let selected = select_candidate(inventory, |candidate| { + evaluated += 1; + let grouped = candidate.materialized_outputs.contains_key(&root); + Ok(Some(CandidateCost { + workload_scope: "reset-counter-workload".into(), + horizon_seconds: 300., + total_cost: if grouped == prefer_grouped { 1. } else { 100. }, + })) + }) + .unwrap(); + assert_eq!( + selected.candidate.materialized_outputs.contains_key(&root), + prefer_grouped + ); + assert_eq!(selected.cost.total_cost, 1.); + assert_eq!(evaluated, 2, "uncompilable candidates must never be priced"); + let candidate = selected.candidate; + let precompute = candidate.precompute.as_ref().unwrap(); + let stored = run( + precompute, + BTreeMap::from([(state_id, batch.clone())]), + Scope::Ingestion { + window_start_ms: -58_000, + window_end_ms: 2000, + revision: 1, + }, + ); + let output = run( + &candidate.query, + BTreeMap::from([(precompute.roots()[0], stored[0].clone())]), + Scope::Query { + evaluation_time_ms: 2000, + revision: 1, + }, + ); + assert!( + matches!(output[0].rows()[0][1], Value::Float64(value) if value == expected_rate_sum) + ); + } + let contracts = BTreeMap::from([( + state_id, + InputContract::bounded(Arc::new(state.output_schema.clone())), + )]); + for frontier in [vec![rate_id, rate_id], vec![root, rate_id], vec![999]] { + assert!( + asap_physical_operators::physical_planner::compile_candidate( + &dag, + contracts.clone(), + &[root], + &frontier + ) + .is_err() + ); + } + let inventory = compile_candidates( + &dag, + contracts.clone(), + &[root], + &[vec![rate_id], vec![root]], + ); + let selected = select_candidate(inventory, |candidate| { + if candidate.materialized_outputs.contains_key(&root) { + return Ok(None); + } + Ok(Some(CandidateCost { + workload_scope: "same-workload".into(), + horizon_seconds: 300., + total_cost: 100., + })) + }) + .unwrap(); + assert!(selected + .candidate + .materialized_outputs + .contains_key(&rate_id)); + let inventory = compile_candidates(&dag, contracts, &[root], &[vec![rate_id], vec![root]]); + assert!( + select_candidate(inventory, |candidate| Ok(Some(CandidateCost { + workload_scope: "same-workload".into(), + horizon_seconds: if candidate.materialized_outputs.contains_key(&root) { + 60. + } else { + 300. + }, + total_cost: 1., + }))) + .is_err() + ); + let query_scope = Scope::Query { + evaluation_time_ms: 2000, + revision: 1, + }; + let maintenance_scope = Scope::Ingestion { + window_start_ms: -58_000, + window_end_ms: 2000, + revision: 1, + }; + let mut results = vec![]; + for candidate in candidates { + let candidate = candidate.unwrap(); + let inputs = if let Some(precompute) = &candidate.precompute { + let source = Operator::source(batch.schema().clone(), vec![batch.clone()]).unwrap(); + let invalid = precompute + .instantiate(BTreeMap::from([(state_id, Box::new(source) as Source<'_>)])) + .unwrap(); + let context = RunContext::new( + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 2000, + revision: 1, + }, + Limits::default(), + ) + .unwrap(); + assert!(invalid.execute(precompute.roots(), context).is_err()); + let stored = run( + precompute, + BTreeMap::from([(state_id, batch.clone())]), + maintenance_scope.clone(), + ); + assert_eq!(stored.len(), 1); + let boundary = precompute.roots()[0]; + assert_eq!( + candidate.materialized_outputs[&boundary].schema, + *stored[0].schema() + ); + BTreeMap::from([(boundary, stored[0].clone())]) + } else { + BTreeMap::from([(state_id, batch.clone())]) + }; + let output = run(&candidate.query, inputs, query_scope.clone()); + assert_eq!(output.len(), 1); + assert_eq!(output[0].rows().len(), 1); + assert!(matches!(&output[0].rows()[0][0], Value::Utf8(job) if job.as_ref() == "api")); + assert!( + matches!(output[0].rows()[0][1], Value::Float64(value) if value == expected_rate_sum) + ); + results.push( + output[0].rows()[0] + .iter() + .map(|value| value.key().unwrap()) + .collect::>(), + ); + } + assert_eq!(results[0], results[1]); + assert_eq!(results[1], results[2]); + let mut wrong_order = create_planner_accumulator(family, update, grouping).unwrap(); + for (i, value) in [200., 200., 100.].into_iter().enumerate() { + wrong_order.update_single(value, i as i64 * 1000); + } + let rate_of_sum = wrong_order + .into_accumulator() + .as_any() + .downcast_ref::() + .unwrap() + .readout(asap_physical_operators::Statistic::Rate, range_ms, None) + .unwrap() + .unwrap(); + assert_ne!( + expected_rate_sum, rate_of_sum, + "counter resets prohibit moving Sum before Rate" + ); +} + +/// Enumerated frontiers include both grouped-result and per-series readout +/// persistence; an explicit Rate-state input retains its original semantics. +#[test] +fn bounded_inventory_exposes_grouped_rate_physical_frontiers() { + use asap_physical_operators::physical_planner::enumerate_frontiers; + let dag = grouped_rate(); + let state = dag + .nodes + .iter() + .find(|node| { + matches!( + &node.payload, + PostAsapOperatorPayload::SummaryAgg { + family: SummaryFamilyType::ExactAggregate(ExactKind::Rate, _), + .. + } + ) + }) + .unwrap(); + let inputs = BTreeMap::from([( + u64::from(state.id.0), + InputContract::bounded(Arc::new(state.output_schema.clone())), + )]); + let roots = [u64::from(dag.root.0)]; + let frontiers = enumerate_frontiers(&dag, &inputs, &roots, 4096).unwrap(); + let candidates = compile_candidates(&dag, inputs.clone(), &roots, &frontiers) + .into_iter() + .collect::, _>>() + .unwrap(); + assert!(candidates.iter().any(|c| c.precompute.is_none())); + assert!(candidates + .iter() + .any(|c| c.materialized_outputs.contains_key(&roots[0]))); + assert!(candidates + .iter() + .any(|c| !c.materialized_outputs.is_empty() + && !c.materialized_outputs.contains_key(&roots[0]))); + assert!(enumerate_frontiers(&dag, &inputs, &roots, 1).is_err()); +} + +#[test] +fn enumerated_grouped_rate_candidates_execute_numeric_query_outputs() { + let inventory = grouped_rate_space().enumerate_candidate_dags(4096).unwrap(); + let mut executed = 0; + for forest in inventory.candidates { + let root = &forest[0].1; + let dag = compile_post_asap_dag(root).unwrap(); + let Some(state) = dag.nodes.iter().find(|node| { + matches!( + node.payload, + PostAsapOperatorPayload::SummaryAgg { + family: SummaryFamilyType::ExactAggregate(ExactKind::Rate, _), + .. + } + ) + }) else { + continue; + }; + let boundary = dag + .nodes + .iter() + .find(|node| { + matches!( + node.payload, + PostAsapOperatorPayload::SummaryAgg { + family: SummaryFamilyType::ExactAggregate(ExactKind::Sum, _), + .. + } + ) + }) + .map(|node| u64::from(node.id.0)) + .unwrap_or(u64::from(dag.root.0)); + let physical_candidates = compile_candidates( + &dag, + BTreeMap::from([( + u64::from(state.id.0), + InputContract::bounded(Arc::new(state.output_schema.clone())), + )]), + &[u64::from(dag.root.0)], + &[vec![], vec![boundary]], + ); + let (family, input, grouping) = match &state.payload { + PostAsapOperatorPayload::SummaryAgg { + family, + input, + grouping, + .. + } => (family, input, grouping), + _ => unreachable!(), + }; + let schema = Arc::new(state.output_schema.clone()); + let rows = ["a", "b"] + .into_iter() + .map(|instance| { + let mut accumulator = create_planner_accumulator(family, input, grouping).unwrap(); + for (timestamp, value) in [(1_000, 1.), (31_000, 31.), (59_000, 59.)] { + accumulator.update_single(value, timestamp); + } + let summary = Value::Summary { + family: family.clone(), + state: Arc::from(accumulator.into_accumulator()), + }; + schema + .fields + .iter() + .map(|field| match &field.dtype { + SummaryFamilyType::ExactAggregate(..) => summary.clone(), + SummaryFamilyType::Plain(DataType::Timestamp) => Value::Timestamp(60_000), + SummaryFamilyType::Plain(DataType::Utf8) + if field.name == "$promql_series_identity" => + { + Value::Utf8( + serde_json::to_string(&BTreeMap::from([ + ("job", "api"), + ("instance", instance), + ])) + .unwrap() + .into(), + ) + } + SummaryFamilyType::Plain(DataType::Utf8) => Value::Utf8("api".into()), + _ => panic!("unexpected input field {field:?}"), + }) + .collect() + }) + .collect(); + let batch = Batch::try_new(schema, rows).unwrap(); + for physical in physical_candidates { + let physical = physical.unwrap(); + let inputs = if let Some(precompute) = &physical.precompute { + let source_id = precompute.input_contracts().next().unwrap().0; + let stored = run( + precompute, + BTreeMap::from([(source_id, batch.clone())]), + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 60_000, + revision: 1, + }, + ); + assert_eq!(stored.len(), 1); + BTreeMap::from([(precompute.roots()[0], stored[0].clone())]) + } else { + BTreeMap::from([( + physical.query.input_contracts().next().unwrap().0, + batch.clone(), + )]) + }; + let output = run( + &physical.query, + inputs, + Scope::Query { + evaluation_time_ms: 60_000, + revision: 1, + }, + ); + assert_eq!(output.len(), 1); + assert_eq!(output[0].rows().len(), 1); + assert!(output[0] + .schema() + .fields + .iter() + .all(|field| matches!(field.dtype, SummaryFamilyType::Plain(_)))); + assert!( + output[0].rows()[0] + .iter() + .any(|value| matches!(value, Value::Float64(x) if (*x - 2.).abs() < 1e-12)), + "{:?}", + output[0].rows() + ); + executed += 1; + } + } + assert!( + executed >= 2, + "must execute both stored and query-time grouped Rate candidates: {executed}" + ); +} diff --git a/crates/asap-physical-operators/tests/precompute_population.rs b/crates/asap-physical-operators/tests/precompute_population.rs new file mode 100644 index 00000000..3bb3af6c --- /dev/null +++ b/crates/asap-physical-operators/tests/precompute_population.rs @@ -0,0 +1,425 @@ +//! Persisted precompute graphs preserve group/window identity and execute state-to-state computation. +use asap_physical_operators::{ + factory::create_planner_accumulator, + operators::Operator, + physical_planner::{precompute, CompiledPhysicalDag, Source}, + runtime::{Limits, RunContext, Scope}, + values::{Batch, Value}, + Statistic, +}; +use futures::{executor::block_on, StreamExt}; +use planner_types::{ + post_asap::*, + pre_asap::{ArithmeticOpKind, BinaryOpKind, ColumnRef, DataType, GroupKeys, Reduction}, +}; +use std::{collections::BTreeMap, sync::Arc}; + +#[test] +fn finalized_shared_panes_rebuild_one_global_summary_after_recovery() { + let family = SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum); + let schema = |dtype| SummarySchema { + fields: vec![SummaryField { + name: "value".into(), + dtype, + nullable: false, + }], + time_index: None, + }; + let state_schema = schema(family.clone()); + let mut value_schema = schema(SummaryFamilyType::Plain(DataType::Float64)); + value_schema.fields.push(SummaryField { + name: "time".into(), + dtype: SummaryFamilyType::Plain(DataType::Timestamp), + nullable: false, + }); + value_schema.time_index = Some(1); + for (weight, expected) in [ + (SummaryInputExpr::Column(ColumnRef::SampleValue), 60.), + (SummaryInputExpr::Constant(1.), 4.), + ] { + let nodes = vec![ + PostAsapDagNode { + id: PostAsapNodeId(0), + payload: PostAsapOperatorPayload::SummaryMerge, + output_state: ExecutionDataState::INGESTION_SUMMARY, + output_schema: state_schema.clone(), + guarantee: None, + }, + PostAsapDagNode { + id: PostAsapNodeId(1), + payload: PostAsapOperatorPayload::Value { + operation: ValueOperation::FinalizeExactAccumulator, + }, + output_state: ExecutionDataState::INGESTION_ROWS, + output_schema: value_schema.clone(), + guarantee: None, + }, + PostAsapDagNode { + id: PostAsapNodeId(2), + payload: PostAsapOperatorPayload::Binary { + operator: BinaryOperator { + kind: BinaryOpKind::Arithmetic(ArithmeticOpKind::Add), + vector_match: None, + checked_relative_division: false, + checked_finite_division: false, + }, + }, + output_state: ExecutionDataState::INGESTION_ROWS, + output_schema: value_schema.clone(), + guarantee: None, + }, + PostAsapDagNode { + id: PostAsapNodeId(3), + payload: PostAsapOperatorPayload::SummaryAgg { + family: family.clone(), + input: SummaryUpdate { + weight, + ..SummaryUpdate::column(ColumnRef::SampleValue) + }, + reduction: Reduction::Reduce(GroupKeys::by(vec![])), + grouping: GroupingStrategy::PerSubpopulationInstance, + }, + output_state: ExecutionDataState::INGESTION_SUMMARY, + output_schema: state_schema.clone(), + guarantee: None, + }, + ]; + let edges = [ + (0, 1, EdgeRole::Input), + (1, 2, EdgeRole::Left), + (1, 2, EdgeRole::Right), + (2, 3, EdgeRole::Input), + ] + .into_iter() + .map(|(producer, consumer, role)| PostAsapDagEdge { + producer: PostAsapNodeId(producer), + consumer: PostAsapNodeId(consumer), + role, + intermediate_schema: nodes[producer as usize].output_schema.clone(), + data_state: nodes[producer as usize].output_state, + grouping: GroupingEdgeCompatibility::NotApplicable, + window: WindowEdgeCompatibility::NotApplicable, + }) + .collect(); + let dag = PostAsapDag { + nodes, + edges, + root: PostAsapNodeId(3), + }; + let mut invalid_grouping = dag.clone(); + let PostAsapOperatorPayload::SummaryAgg { reduction, .. } = + &mut invalid_grouping.nodes[3].payload + else { + unreachable!() + }; + *reduction = Reduction::Reduce(GroupKeys::by(vec![0])); + assert!( + precompute::compile(&invalid_grouping, &[0], &[3]).is_err(), + "numeric values cannot be reinterpreted as population labels" + ); + let program = precompute::compile(&dag, &[0], &[3]).unwrap(); + let program = + serde_json::from_slice::(&serde_json::to_vec(&program).unwrap()) + .unwrap(); + assert_eq!(program.input_contracts().count(), 1); + for revision in [1, 2] { + let rows = [ + ("a", 1000, 2.), + ("a", 2000, 4.), + ("b", 1000, 8.), + ("b", 2000, 16.), + ] + .into_iter() + .map(|(group, time, value)| { + let mut state = create_planner_accumulator( + &family, + &SummaryUpdate::column(ColumnRef::SampleValue), + &GroupingStrategy::PerSubpopulationInstance, + ) + .unwrap(); + state.update_single(value, time); + vec![ + Value::Map( + vec![(Value::Utf8("instance".into()), Value::Utf8(group.into()))].into(), + ), + Value::Timestamp(time), + Value::Summary { + family: family.clone(), + state: Arc::from(state.into_accumulator()), + }, + ] + }) + .collect(); + let input = + Batch::try_new(precompute::population_schema(family.clone()), rows).unwrap(); + let sources = BTreeMap::from([( + 0, + Box::new(Operator::source(input.schema().clone(), vec![input]).unwrap()) + as Source<'_>, + )]); + let graph = program.instantiate(sources).unwrap(); + let context = RunContext::new( + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 2000, + revision, + }, + Limits::default(), + ) + .unwrap(); + let output = block_on(async { + let mut stream = graph.execute(program.roots(), context).unwrap().remove(0); + let batch = stream.next().await.unwrap().unwrap(); + assert!(stream.next().await.is_none()); + batch + }); + assert_eq!(output.rows().len(), 1); + assert!(matches!(&output.rows()[0][0], Value::Map(labels) if labels.is_empty())); + assert!(matches!(&output.rows()[0][1], Value::Timestamp(2000))); + let Value::Summary { state, .. } = &output.rows()[0][2] else { + panic!("state expected") + }; + assert_eq!( + state + .as_any() + .downcast_ref::() + .unwrap() + .readout(Statistic::Sum, None, None) + .unwrap() + .unwrap(), + expected + ); + } + } +} + +fn logical_schema(family: SummaryFamilyType) -> SummarySchema { + SummarySchema { + fields: vec![SummaryField { + name: "value".into(), + dtype: family, + nullable: false, + }], + time_index: None, + } +} +fn state_graph( + family: SummaryFamilyType, + target: Option, + merge: bool, +) -> CompiledPhysicalDag { + let mut nodes = vec![PostAsapDagNode { + id: PostAsapNodeId(0), + payload: PostAsapOperatorPayload::SummaryMerge, + output_state: ExecutionDataState::INGESTION_SUMMARY, + output_schema: logical_schema(family.clone()), + guarantee: None, + }]; + if merge { + nodes.push(PostAsapDagNode { + id: PostAsapNodeId(1), + payload: PostAsapOperatorPayload::SummaryMerge, + ..nodes[0].clone() + }); + } + let read_id = nodes.len() as u32; + nodes.push(PostAsapDagNode { + id: PostAsapNodeId(read_id), + payload: PostAsapOperatorPayload::Value { + operation: ValueOperation::FinalizeExactAccumulator, + }, + output_state: ExecutionDataState::INGESTION_ROWS, + output_schema: logical_schema(SummaryFamilyType::Plain(DataType::Float64)), + guarantee: None, + }); + if let Some(target) = target { + nodes.push(PostAsapDagNode { + id: PostAsapNodeId(nodes.len() as u32), + payload: PostAsapOperatorPayload::SummaryAgg { + family: target.clone(), + input: SummaryUpdate::column(ColumnRef::SampleValue), + reduction: Reduction::by(vec![]), + grouping: GroupingStrategy::default(), + }, + output_state: ExecutionDataState::INGESTION_SUMMARY, + output_schema: logical_schema(target), + guarantee: None, + }); + } + let edges = (1..nodes.len()) + .map(|i| PostAsapDagEdge { + producer: nodes[i - 1].id, + consumer: nodes[i].id, + role: EdgeRole::Input, + intermediate_schema: nodes[i - 1].output_schema.clone(), + data_state: nodes[i - 1].output_state, + grouping: GroupingEdgeCompatibility::NotApplicable, + window: WindowEdgeCompatibility::NotApplicable, + }) + .collect(); + let root = nodes.last().unwrap().id; + precompute::compile( + &PostAsapDag { nodes, edges, root }, + &[0], + &[u64::from(root.0)], + ) + .unwrap() +} +fn native_run( + program: &CompiledPhysicalDag, + family: SummaryFamilyType, + states: Vec>, + context: RunContext, +) -> Result>, asap_physical_operators::Error> { + let program = + serde_json::from_slice::(&serde_json::to_vec(&program).unwrap()) + .unwrap(); + let rows = states + .into_iter() + .enumerate() + .map(|(i, state)| { + vec![ + Value::Map(vec![].into()), + Value::Timestamp((i as i64 + 1) * 1000), + Value::Summary { + family: family.clone(), + state, + }, + ] + }) + .collect(); + let input = Batch::try_new(precompute::population_schema(family), rows)?; + let graph = program.instantiate(BTreeMap::from([( + 0, + Box::new(Operator::source(input.schema().clone(), vec![input])?) as Source<'_>, + )]))?; + block_on(async { + let mut rows = Vec::new(); + let mut stream = graph.execute(program.roots(), context)?.remove(0); + while let Some(batch) = stream.next().await { + rows.extend(batch?.rows().iter().cloned()); + } + Ok(rows) + }) +} +fn ingestion_context(limits: Limits) -> RunContext { + RunContext::new( + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 2000, + revision: 1, + }, + limits, + ) + .unwrap() +} +fn sum_state(value: f64) -> Arc { + let mut state = asap_physical_operators::summary_kernels::exact::ExactAccumulator::new( + planner_types::post_asap::SummaryFamilyType::ExactAggregate( + planner_types::post_asap::ExactKind::Sum, + planner_types::post_asap::ExactParams::Sum, + ), + false, + ) + .unwrap(); + state.update(None, value, 0); + Arc::new(state) +} + +// Only an explicit merge may collapse distinct pane updates before finalization. +#[test] +fn explicit_merge_changes_pane_cardinality() { + let family = SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum); + for (merge, expected) in [(false, vec![2., 7.]), (true, vec![9.])] { + let program = state_graph(family.clone(), None, merge); + let rows = native_run( + &program, + family.clone(), + vec![sum_state(2.), sum_state(7.)], + ingestion_context(Limits::default()), + ) + .unwrap(); + let values = rows + .iter() + .map(|row| match row[2] { + Value::Float64(v) => v, + _ => panic!("numeric readout expected"), + }) + .collect::>(); + assert_eq!(values, expected); + assert!(matches!(rows.last().unwrap()[1], Value::Timestamp(2000))); + } +} + +// Typed updates reject invalid domains before publishing any target state. +#[test] +fn precompute_rejects_nonfinite_and_nonpositive_dds_updates() { + let source = SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum); + let target = SummaryFamilyType::Sketch( + SketchKind::new( + SketchAlgorithm::DDSketch, + SketchParams::DDSketch { alpha: 0.01 }, + ), + GroupingStrategy::default(), + ); + let program = state_graph(source.clone(), Some(target), false); + assert!(native_run( + &program, + source.clone(), + vec![sum_state(20.)], + ingestion_context(Limits::default()) + ) + .is_ok()); + for value in [-20., 0., f64::MAX, f64::NAN, f64::INFINITY] { + assert!(native_run( + &program, + source.clone(), + vec![sum_state(value)], + ingestion_context(Limits::default()) + ) + .is_err()); + } +} + +// An exact count must not silently lose units when exposed through Float64 rows. +#[test] +fn precompute_count_conversion_checks_precision() { + use asap_physical_operators::summary_kernels::exact::ExactAccumulator; + let family = SummaryFamilyType::ExactAggregate(ExactKind::Count, ExactParams::Count); + let program = state_graph(family.clone(), None, false); + for (count, valid) in [(3u64, true), ((1u64 << 53) + 1, false)] { + let mut state = + serde_json::to_value(ExactAccumulator::new(family.clone(), false).unwrap()).unwrap(); + state["scalar"]["Count"] = count.into(); + let state: ExactAccumulator = serde_json::from_value(state).unwrap(); + let result = native_run( + &program, + family.clone(), + vec![Arc::new(state)], + ingestion_context(Limits::default()), + ); + assert_eq!(result.is_ok(), valid); + } +} + +// Graph execution retains terminal cancellation and shared workspace limits. +#[test] +fn precompute_graph_enforces_cancellation_and_budget() { + let family = SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum); + let program = state_graph(family.clone(), None, true); + let context = ingestion_context(Limits::default()); + context.cancel(); + let error = native_run(&program, family.clone(), vec![sum_state(1.)], context).unwrap_err(); + assert!(format!("{error:?}").contains("Cancelled")); + let error = native_run( + &program, + family, + vec![sum_state(1.)], + ingestion_context(Limits { + max_bytes: 1, + ..Limits::default() + }), + ) + .unwrap_err(); + assert!(format!("{error:?}").contains("MemoryLimit")); +} diff --git a/crates/asap-physical-operators/tests/promql_binary.rs b/crates/asap-physical-operators/tests/promql_binary.rs new file mode 100644 index 00000000..24844b5e --- /dev/null +++ b/crates/asap-physical-operators/tests/promql_binary.rs @@ -0,0 +1,270 @@ +//! Binary computation must be fully compiled before deployment binds values. +use asap_physical_operators::{ + operators::Operator, + physical_planner::{compile_node, CompiledPhysicalDag, InputContract, Source}, + runtime::{Limits, RunContext, Scope}, + values::{Batch, Schema, Value}, +}; +use futures::{executor::block_on, StreamExt}; +use planner_types::{ + post_asap::{ + BinaryOperator, ExecutionDataState, PostAsapDagNode, PostAsapNodeId, + PostAsapOperatorPayload, SummaryFamilyType, SummaryField, SummarySchema, + }, + pre_asap::{ArithmeticOpKind, BinaryOpKind, DataType}, +}; +use std::{collections::BTreeMap, sync::Arc}; + +fn schema() -> Schema { + Arc::new(SummarySchema { + fields: vec![ + SummaryField { + name: "labels".into(), + dtype: SummaryFamilyType::Plain(DataType::Map { + key: Box::new(DataType::Utf8), + value: Box::new(DataType::Utf8), + value_nullable: false, + }), + nullable: false, + }, + SummaryField { + name: "value".into(), + dtype: SummaryFamilyType::Plain(DataType::Float64), + nullable: false, + }, + ], + time_index: None, + }) +} +fn row(name: &str, job: &str, value: f64) -> Vec { + vec![ + Value::Map( + vec![ + (Value::Utf8("__name__".into()), Value::Utf8(name.into())), + (Value::Utf8("job".into()), Value::Utf8(job.into())), + ] + .into(), + ), + Value::Float64(value), + ] +} +fn program() -> CompiledPhysicalDag { + let schema = schema(); + let node = PostAsapDagNode { + id: PostAsapNodeId(2), + payload: PostAsapOperatorPayload::Binary { + operator: BinaryOperator { + kind: BinaryOpKind::Arithmetic(ArithmeticOpKind::Div), + vector_match: None, + checked_relative_division: true, + checked_finite_division: false, + }, + }, + output_state: ExecutionDataState::QUERY_ROWS, + output_schema: (*schema).clone(), + guarantee: None, + }; + let operator = compile_node(&node, &[schema.clone(), schema.clone()]).unwrap(); + let graph = CompiledPhysicalDag::from_operators( + BTreeMap::from([ + (0, InputContract::bounded(schema.clone())), + (1, InputContract::bounded(schema)), + ]), + BTreeMap::from([(2, (vec![0, 1], operator))]), + vec![2], + ) + .unwrap(); + serde_json::from_slice::(&serde_json::to_vec(&graph).unwrap()).unwrap() +} +fn evaluate( + left: Vec>, + right: Vec>, +) -> Result>, asap_physical_operators::Error> { + let graph = program(); + let sources = [left, right] + .into_iter() + .enumerate() + .map(|(id, rows)| { + let batch = Batch::try_new(schema(), rows).unwrap(); + ( + id as u64, + Box::new(Operator::source(schema(), vec![batch]).unwrap()) as Source<'_>, + ) + }) + .collect(); + let bound = graph.instantiate(sources)?; + let ctx = RunContext::new( + Scope::Query { + evaluation_time_ms: 1, + revision: 0, + }, + Limits::default(), + )?; + block_on(async { + let mut stream = bound.execute(&[2], ctx)?.remove(0); + let mut rows = Vec::new(); + while let Some(batch) = stream.next().await { + rows.extend(batch?.rows().iter().cloned()); + } + Ok(rows) + }) +} + +#[test] +fn compiled_binary_matches_series_and_preserves_checked_division() { + let rows = evaluate( + vec![row("left", "api", 6.), row("left", "unmatched", 8.)], + vec![row("right", "api", 2.)], + ) + .unwrap(); + assert_eq!( + serde_json::to_value(&rows).unwrap(), + serde_json::to_value(vec![vec![ + Value::Map(vec![(Value::Utf8("job".into()), Value::Utf8("api".into()))].into()), + Value::Float64(3.) + ]]) + .unwrap() + ); + assert!(evaluate(vec![row("a", "api", 1.)], vec![row("b", "api", 0.)]).is_err()); +} + +#[test] +fn duplicate_matching_identity_is_rejected() { + assert!(evaluate( + vec![row("a", "api", 1.)], + vec![row("b", "api", 2.), row("c", "api", 3.)] + ) + .is_err()); +} + +// Scalar broadcasting and comparison filtering keep the vector operand's value. +#[test] +fn scalar_broadcast_and_bool_comparison_are_distinct() { + use asap_physical_operators::physical_planner::promql_values; + use planner_types::pre_asap::CompareOpKind; + for return_bool in [false, true] { + let graph = promql_values::compile_binary( + &BinaryOperator { + kind: BinaryOpKind::Compare(CompareOpKind::Lt), + vector_match: None, + checked_relative_division: false, + checked_finite_division: false, + }, + return_bool, + true, + false, + ) + .unwrap(); + let graph = + serde_json::from_slice::(&serde_json::to_vec(&graph).unwrap()) + .unwrap(); + let scalar = promql_values::scalar_schema(); + let vector = promql_values::vector_schema(); + let sources = BTreeMap::from([ + ( + 0, + Box::new( + Operator::source( + scalar.clone(), + vec![Batch::try_new(scalar, vec![vec![Value::Float64(2.)]]).unwrap()], + ) + .unwrap(), + ) as Source<'_>, + ), + ( + 1, + Box::new( + Operator::source( + vector.clone(), + vec![Batch::try_new( + vector, + vec![row("requests", "api", 4.), row("requests", "worker", 1.)], + ) + .unwrap()], + ) + .unwrap(), + ) as Source<'_>, + ), + ]); + let bound = graph.instantiate(sources).unwrap(); + let context = RunContext::new( + Scope::Query { + evaluation_time_ms: 1, + revision: 0, + }, + Limits::default(), + ) + .unwrap(); + let result = block_on(async { + bound + .execute(&[2], context) + .unwrap() + .remove(0) + .next() + .await + .unwrap() + .unwrap() + }); + assert_eq!(result.rows().len(), if return_bool { 2 } else { 1 }); + assert!( + matches!(result.rows()[0].last(), Some(Value::Float64(v)) if *v == if return_bool { 1. } else { 4. }) + ); + let Value::Map(labels) = &result.rows()[0][0] else { + panic!("missing labels") + }; + assert_eq!( + labels + .iter() + .any(|(key, _)| matches!(key, Value::Utf8(s) if s.as_ref() == "__name__")), + !return_bool + ); + } +} + +// Terminal request controls retain their native error classification. +#[test] +fn binary_obeys_memory_and_cancellation() { + for cancel in [false, true] { + let graph = program(); + let sources = (0..2) + .map(|id| { + ( + id, + Box::new( + Operator::source( + schema(), + vec![Batch::try_new(schema(), vec![row("x", "api", 1.)]).unwrap()], + ) + .unwrap(), + ) as Source<'_>, + ) + }) + .collect(); + let bound = graph.instantiate(sources).unwrap(); + let context = RunContext::new( + Scope::Query { + evaluation_time_ms: 1, + revision: 0, + }, + Limits { + max_bytes: if cancel { 10000 } else { 1 }, + ..Limits::default() + }, + ) + .unwrap(); + if cancel { + context.cancel(); + } + let result = block_on(async { + match bound.execute(&[2], context) { + Err(error) => Err(error), + Ok(mut streams) => streams.remove(0).next().await.unwrap().map(|_| ()), + } + }); + assert!(matches!( + (cancel, result), + (true, Err(asap_physical_operators::Error::Cancelled)) + | (false, Err(asap_physical_operators::Error::MemoryLimit)) + )); + } +} diff --git a/crates/asap-physical-operators/tests/promql_values.rs b/crates/asap-physical-operators/tests/promql_values.rs new file mode 100644 index 00000000..886b7d00 --- /dev/null +++ b/crates/asap-physical-operators/tests/promql_values.rs @@ -0,0 +1,492 @@ +//! Compile, persist and rebind dynamic-label computation without deployment lowering. +use asap_physical_operators::{ + operators::Operator, + physical_planner::{promql_values::*, CompiledPhysicalDag, Source}, + runtime::{Limits, RunContext, Scope}, + values::{Batch, Value}, +}; +use futures::{executor::block_on, StreamExt}; +use planner_types::pre_asap::{AggIntent, ColumnRef, GroupKeys}; +use std::collections::BTreeMap; + +fn row(labels: &[(&str, &str)], value: f64) -> Vec { + vec![ + Value::Map( + labels + .iter() + .map(|(k, v)| (Value::Utf8((*k).into()), Value::Utf8((*v).into()))) + .collect::>() + .into(), + ), + Value::Float64(value), + ] +} +fn run(graph: CompiledPhysicalDag, rows: Vec>) -> Vec> { + run_inputs(graph, vec![Batch::try_new(vector_schema(), rows).unwrap()]).unwrap() +} +fn run_inputs( + graph: CompiledPhysicalDag, + batches: Vec, +) -> Result>, asap_physical_operators::Error> { + let graph = serde_json::from_slice::(&serde_json::to_vec(&graph).unwrap()) + .unwrap(); + let sources = batches + .into_iter() + .enumerate() + .map(|(id, batch)| { + ( + id as u64, + Box::new(Operator::source(batch.schema().clone(), vec![batch]).unwrap()) + as Source<'_>, + ) + }) + .collect::>(); + let bound = graph.instantiate(sources).unwrap(); + let context = RunContext::new( + Scope::Query { + evaluation_time_ms: 0, + revision: 1, + }, + Limits::default(), + ) + .unwrap(); + block_on(async { + let mut stream = bound.execute(graph.roots(), context).unwrap().remove(0); + let mut rows = Vec::new(); + while let Some(batch) = stream.next().await { + rows.extend(batch?.rows().iter().cloned()); + } + Ok(rows) + }) +} +fn equal_rows(actual: Vec>, expected: Vec>) { + let mut actual = actual + .into_iter() + .map(|r| serde_json::to_string(&r).unwrap()) + .collect::>(); + let mut expected = expected + .into_iter() + .map(|r| serde_json::to_string(&r).unwrap()) + .collect::>(); + actual.sort(); + expected.sort(); + assert_eq!(actual, expected); +} + +#[test] +fn grouping_preserves_unenumerated_labels_and_empty_label_semantics() { + let rows = vec![ + row(&[("__name__", "m"), ("instance", "a"), ("job", "api")], 1.), + row(&[("__name__", "m"), ("instance", "b"), ("job", "api")], 2.), + row(&[("instance", "c"), ("job", "")], 4.), + row(&[("instance", "d")], 8.), + ]; + equal_rows( + run( + compile_aggregate( + &AggIntent::Sum { col: None }, + &GroupKeys::without(vec![ColumnRef::Named("instance".into())]), + ) + .unwrap(), + rows.clone(), + ), + vec![row(&[("job", "api")], 3.), row(&[], 12.)], + ); + equal_rows( + run( + compile_aggregate( + &AggIntent::Count { + accuracy: planner_types::types::AccuracyTarget::Exact, + }, + &GroupKeys::by(vec![ColumnRef::Named("job".into())]), + ) + .unwrap(), + rows, + ), + vec![row(&[("job", "api")], 2.), row(&[], 2.)], + ); +} + +#[test] +fn ranking_and_grouped_limit_preserve_full_selected_series() { + let grouping = GroupKeys::by(vec![ColumnRef::Named("job".into())]); + let rows = vec![ + row(&[("instance", "a"), ("job", "api")], 1.), + row(&[("instance", "b"), ("job", "api")], 3.), + row(&[("instance", "c"), ("job", "worker")], 2.), + ]; + let sorted = run(compile_sort(true, &grouping).unwrap(), rows); + let selected = run(compile_limit(1, 0, &grouping).unwrap(), sorted); + equal_rows( + selected, + vec![ + row(&[("instance", "b"), ("job", "api")], 3.), + row(&[("instance", "c"), ("job", "worker")], 2.), + ], + ); + assert!(run(compile_limit(0, 0, &grouping).unwrap(), vec![row(&[], 1.)]).is_empty()); +} + +#[test] +fn empty_vector_aggregation_stays_empty() { + assert!(run( + compile_aggregate(&AggIntent::Sum { col: None }, &GroupKeys::default()).unwrap(), + vec![] + ) + .is_empty()); + let scalar = run(compile_vector_to_scalar().unwrap(), vec![]); + assert!(matches!(scalar[0][0],Value::Float64(v) if v.is_nan())); +} + +// One persisted temporal graph accepts different request windows and detects resets. +#[test] +fn temporal_graph_uses_bound_window_without_recompilation() { + let graph = compile_temporal(&AggIntent::Rate, false).unwrap(); + for start in [0, 60_000] { + let labels = row(&[("__name__", "counter"), ("job", "api")], 0.)[0].clone(); + let samples = [(0, 5.), (30_000, 1.), (60_000, 7.)]; + let rows = samples + .into_iter() + .map(|(time, value)| { + vec![ + labels.clone(), + Value::Timestamp(start + time), + Value::Float64(value), + Value::Timestamp(start), + Value::Timestamp(start + 60_000), + ] + }) + .collect(); + let output = run_inputs( + graph.clone(), + vec![Batch::try_new(matrix_schema(), rows).unwrap()], + ) + .unwrap(); + equal_rows(output, vec![row(&[("job", "api")], 7. / 60.)]); + } + let labels = row(&[("job", "api")], 0.)[0].clone(); + let rows = vec![ + vec![ + labels.clone(), + Value::Timestamp(0), + Value::Float64(1.), + Value::Timestamp(0), + Value::Timestamp(1000), + ], + vec![ + labels, + Value::Timestamp(1000), + Value::Float64(2.), + Value::Timestamp(0), + Value::Timestamp(2000), + ], + ]; + assert!(run_inputs(graph, vec![Batch::try_new(matrix_schema(), rows).unwrap()]).is_err()); +} + +// The quantile is an ordinary scalar input, and bucket labels are native computation. +#[test] +fn histogram_quantile_keeps_each_label_group() { + let graph = compile_histogram_quantile().unwrap(); + let buckets = vec![ + row(&[("job", "api"), ("le", "1")], 2.), + row(&[("job", "api"), ("le", "2")], 4.), + row(&[("job", "api"), ("le", "+Inf")], 4.), + ]; + let output = run_inputs( + graph, + vec![ + Batch::try_new(scalar_schema(), vec![vec![Value::Float64(0.75)]]).unwrap(), + Batch::try_new(vector_schema(), buckets).unwrap(), + ], + ) + .unwrap(); + equal_rows(output, vec![row(&[("job", "api")], 1.5)]); +} + +// Linking an ensemble preserves its shared producer and every selected operator. +#[test] +fn composed_ensemble_shares_a_producer_across_roots() { + use asap_physical_operators::{ + physical_planner::InputContract, + plan::{PhysicalOperator, PlanProperties}, + runtime::{Input, OutputStream}, + values::Schema, + }; + use planner_types::{ + post_asap::BinaryOperator, + pre_asap::{ArithmeticOpKind, BinaryOpKind}, + }; + struct Counted { + source: Operator, + starts: std::rc::Rc>, + } + impl PhysicalOperator for Counted { + fn name(&self) -> &str { + "CountedInput" + } + fn input_schemas(&self) -> Vec { + vec![] + } + fn output_schema(&self) -> Schema { + self.source.schema() + } + fn output_bytes(&self, batch: &Batch) -> usize { + batch.bytes() + } + fn properties(&self, inputs: &[PlanProperties]) -> PlanProperties { + self.source.properties(inputs) + } + fn start<'a>( + &'a self, + inputs: Vec>, + context: RunContext, + ) -> Result, asap_physical_operators::Error> { + self.starts.set(self.starts.get() + 1); + self.source.start(inputs, context) + } + } + let aggregate = compile_aggregate( + &AggIntent::Sum { col: None }, + &GroupKeys::by(vec![ColumnRef::Named("job".into())]), + ) + .unwrap(); + let binary = compile_binary( + &BinaryOperator { + kind: BinaryOpKind::Arithmetic(ArithmeticOpKind::Add), + vector_match: None, + checked_finite_division: false, + checked_relative_division: false, + }, + false, + false, + false, + ) + .unwrap(); + let graph = CompiledPhysicalDag::compose( + BTreeMap::from([(0, InputContract::bounded(vector_schema()))]), + BTreeMap::from([ + (10, (vec![0], aggregate)), + (20, (vec![10, 10], binary)), + (30, (vec![10], compile_negate(false).unwrap())), + ]), + vec![20, 30], + ) + .unwrap(); + let graph = serde_json::from_slice::(&serde_json::to_vec(&graph).unwrap()) + .unwrap(); + assert_eq!(graph.input_contracts().count(), 1); + let starts = std::rc::Rc::new(std::cell::Cell::new(0)); + for _ in 0..2 { + let input = Batch::try_new(vector_schema(), vec![row(&[("job", "api")], 3.)]).unwrap(); + let source = Counted { + source: Operator::source(vector_schema(), vec![input]).unwrap(), + starts: starts.clone(), + }; + let bound = graph + .instantiate(BTreeMap::from([(0, Box::new(source) as Source<'_>)])) + .unwrap(); + let context = RunContext::new( + Scope::Query { + evaluation_time_ms: 0, + revision: 0, + }, + Limits { + max_buffered_batches: 1, + ..Limits::default() + }, + ) + .unwrap(); + let results = + block_on(futures::future::join_all( + bound + .execute(graph.roots(), context) + .unwrap() + .into_iter() + .map(|mut stream| async move { + stream.next().await.unwrap().unwrap().rows().to_vec() + }), + )); + equal_rows(results[0].clone(), vec![row(&[("job", "api")], 6.)]); + equal_rows(results[1].clone(), vec![row(&[("job", "api")], -3.)]); + } + assert_eq!(starts.get(), 2); +} + +#[test] +fn compiled_constant_needs_no_deployment_source() { + let graph = compile_scalar(3.).unwrap(); + assert_eq!(graph.input_contracts().count(), 0); + let result = run_inputs(graph, vec![]).unwrap(); + assert!(matches!(result[0][0], Value::Float64(3.))); +} + +// Scalar broadcasting cannot silently create duplicate result identities when +// arithmetic or bool comparisons remove the metric name. +#[test] +fn scalar_broadcast_rejects_colliding_result_labels_after_recovery() { + use planner_types::{ + post_asap::BinaryOperator, + pre_asap::{ArithmeticOpKind, BinaryOpKind, CompareOpKind}, + }; + for left_scalar in [false, true] { + for names in [["a", "a"], ["a", "b"]] { + for (kind, return_bool) in [ + (BinaryOpKind::Arithmetic(ArithmeticOpKind::Add), false), + (BinaryOpKind::Compare(CompareOpKind::Gt), true), + ] { + let graph = compile_binary( + &BinaryOperator { + kind, + vector_match: None, + checked_relative_division: false, + checked_finite_division: false, + }, + return_bool, + left_scalar, + !left_scalar, + ) + .unwrap(); + let vector = Batch::try_new( + vector_schema(), + vec![ + row(&[("__name__", names[0]), ("job", "api")], 2.), + row(&[("__name__", names[1]), ("job", "api")], 3.), + ], + ) + .unwrap(); + let scalar = + Batch::try_new(scalar_schema(), vec![vec![Value::Float64(1.)]]).unwrap(); + let result = run_inputs( + graph, + if left_scalar { + vec![scalar, vector] + } else { + vec![vector, scalar] + }, + ); + assert!(result.is_err(), "duplicate output label sets were accepted"); + } + } + } + let graph = compile_binary( + &BinaryOperator { + kind: BinaryOpKind::Compare(CompareOpKind::Gt), + vector_match: None, + checked_relative_division: false, + checked_finite_division: false, + }, + false, + false, + true, + ) + .unwrap(); + let rows = vec![ + row(&[("__name__", "a"), ("job", "api")], 2.), + row(&[("__name__", "b"), ("job", "api")], 3.), + ]; + equal_rows( + run_inputs( + graph, + vec![ + Batch::try_new(vector_schema(), rows.clone()).unwrap(), + Batch::try_new(scalar_schema(), vec![vec![Value::Float64(1.)]]).unwrap(), + ], + ) + .unwrap(), + rows, + ); +} + +// Persisted exact readout graphs, rather than the storage adapter, merge panes, +// finalize each population, and preserve the requested metric-name semantics. +#[test] +fn exact_state_readouts_recover_and_finalize_panes() { + use asap_physical_operators::factory::create_planner_accumulator; + use planner_types::post_asap::*; + use std::sync::Arc; + for (kind, params, expected) in [ + (ExactKind::Sum, ExactParams::Sum, 12.), + (ExactKind::Count, ExactParams::Count, 4.), + (ExactKind::Min, ExactParams::Min, 1.), + (ExactKind::Max, ExactParams::Max, 5.), + ] { + let family = SummaryFamilyType::ExactAggregate(kind, params); + for preserve in [false, true] { + let rows = [[1., 2.], [4., 5.]] + .into_iter() + .map(|samples| { + let mut state = create_planner_accumulator( + &family, + &SummaryUpdate::column(ColumnRef::SampleValue), + &GroupingStrategy::PerSubpopulationInstance, + ) + .unwrap(); + for sample in samples { + state.update_single(sample, 0); + } + let labels = row(&[("__name__", "m"), ("instance", "a")], 0.).remove(0); + vec![ + labels, + Value::Summary { + family: family.clone(), + state: Arc::from(state.into_accumulator()), + }, + ] + }) + .collect(); + let output = run_inputs( + compile_exact_readout(family.clone(), 60_000, preserve).unwrap(), + vec![Batch::try_new(exact_state_schema(family.clone()).unwrap(), rows).unwrap()], + ) + .unwrap(); + let labels = if preserve { + vec![("__name__", "m"), ("instance", "a")] + } else { + vec![("instance", "a")] + }; + equal_rows(output, vec![row(&labels, expected)]); + } + } +} + +#[test] +fn recovered_exact_counter_uses_window_and_omits_insufficient_samples() { + use asap_physical_operators::factory::create_planner_accumulator; + use planner_types::post_asap::*; + use std::sync::Arc; + for (kind, params, expected) in [ + (ExactKind::Rate, ExactParams::Rate, 1.), + (ExactKind::Increase, ExactParams::Increase, 60.), + ] { + let family = SummaryFamilyType::ExactAggregate(kind, params); + let rows = [1, 2] + .into_iter() + .map(|count| { + let mut state = create_planner_accumulator( + &family, + &SummaryUpdate::column(ColumnRef::SampleValue), + &GroupingStrategy::PerSubpopulationInstance, + ) + .unwrap(); + state.update_single(100., -50_000); + if count == 2 { + state.update_single(140., -10_000); + } + vec![ + row(&[("instance", if count == 1 { "one" } else { "two" })], 0.).remove(0), + Value::Summary { + family: family.clone(), + state: Arc::from(state.into_accumulator()), + }, + ] + }) + .collect(); + let output = run_inputs( + compile_exact_readout(family.clone(), 60_000, false).unwrap(), + vec![Batch::try_new(exact_state_schema(family).unwrap(), rows).unwrap()], + ) + .unwrap(); + equal_rows(output, vec![row(&[("instance", "two")], expected)]); + } +} diff --git a/crates/asap-physical-operators/tests/raw_scan.rs b/crates/asap-physical-operators/tests/raw_scan.rs new file mode 100644 index 00000000..78e4c8c0 --- /dev/null +++ b/crates/asap-physical-operators/tests/raw_scan.rs @@ -0,0 +1,386 @@ +//! Scan acceptance uses the public connector contract and Planner physical DAGs. +use asap_physical_operators::dag::{ + planner::bind_with_data_sources, + scan::{DataSources, MemorySource, RawSource}, + values::{Batch, Schema, Value}, + Error, Limits, OutputStream, RunContext, Scope, +}; +use futures::{executor::block_on, stream, StreamExt}; +use planner_types::{ + post_asap::*, + pre_asap::{Column, DataType, GroupKeys, Predicate, QueryExpr, Source}, +}; +use std::{ + collections::BTreeMap, + rc::Rc, + sync::{ + atomic::{AtomicUsize, Ordering}, + Arc, + }, +}; + +fn fixture() -> (QueryExpr, Schema, Vec) { + let schema = + planner_types::pre_asap::Schema::new(vec![Column::new("value", DataType::Int64, true)]); + let output = Arc::new(SummarySchema { + fields: vec![SummaryField { + name: "value".into(), + dtype: SummaryFamilyType::Plain(DataType::Int64), + nullable: true, + }], + time_index: None, + }); + let scan = QueryExpr::Scan { + source: Source::Table { + table_ref: "numbers".into(), + }, + predicates: vec![Predicate(Rc::new(QueryExpr::IsNotNull(Rc::new( + QueryExpr::Column(0), + ))))], + schema, + }; + let batches = vec![ + Batch::try_new( + output.clone(), + vec![vec![Value::Int64(3)], vec![Value::Null]], + ) + .unwrap(), + Batch::try_new( + output.clone(), + vec![vec![Value::Int64(9)], vec![Value::Int64(2)]], + ) + .unwrap(), + ]; + (scan, output, batches) +} +fn plan(scan: QueryExpr, schema: &Schema, state: ExecutionDataState) -> PostAsapDag { + let node = |id, payload| PostAsapDagNode { + id: PostAsapNodeId(id), + payload, + output_state: state, + output_schema: (**schema).clone(), + guarantee: None, + }; + let edge = |producer, consumer| PostAsapDagEdge { + producer: PostAsapNodeId(producer), + consumer: PostAsapNodeId(consumer), + role: EdgeRole::Input, + intermediate_schema: (**schema).clone(), + data_state: state, + grouping: GroupingEdgeCompatibility::NotApplicable, + window: WindowEdgeCompatibility::NotApplicable, + }; + PostAsapDag { + nodes: vec![ + node(0, PostAsapOperatorPayload::Fallback { expression: scan }), + node( + 1, + PostAsapOperatorPayload::Value { + operation: ValueOperation::Sort { + keys: vec![planner_types::pre_asap::SortKey { + expr: QueryExpr::Column(0), + ascending: false, + nulls_first: false, + }], + partition_by: GroupKeys::by(vec![]), + }, + }, + ), + node( + 2, + PostAsapOperatorPayload::Value { + operation: ValueOperation::Limit { + n: 2, + offset: 0, + partition_by: GroupKeys::by(vec![]), + }, + }, + ), + ], + edges: vec![edge(0, 1), edge(1, 2)], + root: PostAsapNodeId(2), + } +} +fn registry(source: Arc) -> DataSources { + let mut r = DataSources::default(); + r.register( + Source::Table { + table_ref: "numbers".into(), + }, + source, + ) + .unwrap(); + r +} +fn context() -> RunContext { + RunContext::new( + Scope::Query { + evaluation_time_ms: 0, + revision: 1, + }, + Limits::default(), + ) + .unwrap() +} + +// Raw-only execution filters nulls and ranks across batches at either phase. +#[test] +fn raw_scan_to_sort_limit_at_both_phases() { + let (scan, schema, batches) = fixture(); + let sources = registry(Arc::new( + MemorySource::new(schema.clone(), batches).unwrap(), + )); + for state in [ + ExecutionDataState::QUERY_ROWS, + ExecutionDataState::INGESTION_ROWS, + ] { + let dag = plan(scan.clone(), &schema, state); + let bound = bind_with_data_sources(&dag, BTreeMap::new(), &[2], &sources).unwrap(); + let ctx = if state == ExecutionDataState::QUERY_ROWS { + context() + } else { + RunContext::new( + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 1, + revision: 1, + }, + Limits::default(), + ) + .unwrap() + }; + let rows = block_on(async { + let mut output = bound.execute(&[2], ctx.clone()).unwrap().remove(0); + let mut rows = vec![]; + while let Some(batch) = output.next().await { + rows.extend(batch.unwrap().rows().iter().cloned()); + } + rows + }); + assert!( + matches!(rows.as_slice(), [a,b] if matches!(a.as_slice(), [Value::Int64(9)]) && matches!(b.as_slice(), [Value::Int64(3)])) + ); + assert_eq!(ctx.retained_bytes(), 0); + } +} +struct CountingSource { + schema: Schema, + opened: Arc, + fail: bool, +} +impl RawSource for CountingSource { + fn boundedness(&self) -> asap_physical_operators::plan::Boundedness { + asap_physical_operators::plan::Boundedness::Bounded + } + fn schema(&self) -> Schema { + self.schema.clone() + } + fn scan(&self, _: RunContext) -> Result, Error> { + self.opened.fetch_add(1, Ordering::SeqCst); + if self.fail { + return Err(Error::Operator("reader failed".into())); + } + Ok(stream::iter(vec![Batch::try_new( + self.schema.clone(), + vec![vec![Value::Int64(7)]], + )]) + .boxed_local()) + } +} +// Binding and cancellation do not perform I/O; fan-out opens one cursor per run. +#[test] +fn lazy_open_shared_producer_and_cancellation() { + let (scan, schema, _) = fixture(); + let opened = Arc::new(AtomicUsize::new(0)); + let sources = registry(Arc::new(CountingSource { + schema: schema.clone(), + opened: opened.clone(), + fail: false, + })); + let plan = plan(scan, &schema, ExecutionDataState::QUERY_ROWS); + let bound = bind_with_data_sources(&plan, BTreeMap::new(), &[0, 2], &sources).unwrap(); + let ctx = context(); + let streams = bound.execute(&[0, 2], ctx.clone()).unwrap(); + assert_eq!(opened.load(Ordering::SeqCst), 0); + ctx.cancel(); + drop(streams); + assert_eq!(opened.load(Ordering::SeqCst), 0); + for _ in 0..2 { + block_on(async { + let streams = bound.execute(&[0, 2], context()).unwrap(); + let all = + futures::future::join_all(streams.into_iter().map(|s| s.collect::>())).await; + assert!(all.iter().flatten().all(Result::is_ok)); + }); + } + assert_eq!(opened.load(Ordering::SeqCst), 2); +} +// Unavailable sources and unsupported predicates fail before opening any cursor. +#[test] +fn binding_errors_and_reader_errors_are_not_empty_results() { + let (mut scan, schema, _) = fixture(); + assert!(DataSources::default().bind(&scan).is_err()); + let opened = Arc::new(AtomicUsize::new(0)); + let sources = registry(Arc::new(CountingSource { + schema: schema.clone(), + opened: opened.clone(), + fail: true, + })); + if let QueryExpr::Scan { predicates, .. } = &mut scan { + predicates.push(Predicate(Rc::new(QueryExpr::Column(0)))); + } + assert!(sources.bind(&scan).is_err()); + assert_eq!(opened.load(Ordering::SeqCst), 0); + let (scan, _, _) = fixture(); + let plan = plan(scan, &schema, ExecutionDataState::QUERY_ROWS); + let bound = bind_with_data_sources(&plan, BTreeMap::new(), &[2], &sources).unwrap(); + block_on(async { + let mut stream = bound.execute(&[2], context()).unwrap().remove(0); + assert!(stream.next().await.unwrap().is_err()); + }); +} + +// Schema drift cannot enter the DAG, and connector batches obey execution limits. +#[test] +fn schema_drift_and_memory_limits_fail_the_scan() { + struct Drift { + expected: Schema, + batch: Batch, + } + impl RawSource for Drift { + fn schema(&self) -> Schema { + self.expected.clone() + } + fn scan(&self, _: RunContext) -> Result, Error> { + Ok(stream::once(async { Ok(self.batch.clone()) }).boxed_local()) + } + } + let (scan, schema, batches) = fixture(); + let mut different = (*schema).clone(); + different.fields[0].name = "wrong".into(); + let bad = Batch::try_new(Arc::new(different), vec![vec![Value::Int64(1)]]).unwrap(); + let sources = registry(Arc::new(Drift { + expected: schema.clone(), + batch: bad, + })); + let plan = plan(scan, &schema, ExecutionDataState::QUERY_ROWS); + let graph = bind_with_data_sources(&plan, BTreeMap::new(), &[0], &sources).unwrap(); + block_on(async { + let mut s = graph.execute(&[0], context()).unwrap().remove(0); + assert!(s.next().await.unwrap().is_err()); + }); + let sources = registry(Arc::new(MemorySource::new(schema, batches).unwrap())); + let graph = bind_with_data_sources(&plan, BTreeMap::new(), &[0], &sources).unwrap(); + let ctx = RunContext::new( + Scope::Query { + evaluation_time_ms: 0, + revision: 1, + }, + Limits { + max_bytes: 1, + max_buffered_batches: 1, + }, + ) + .unwrap(); + block_on(async { + let mut s = graph.execute(&[0], ctx.clone()).unwrap().remove(0); + assert!(s.next().await.unwrap().is_err()); + }); + assert_eq!(ctx.retained_bytes(), 0); +} + +// An empty table is a valid empty scan; nullable comparisons retain only TRUE. +#[test] +fn empty_sources_and_three_valued_predicates() { + use planner_types::pre_asap::{CompareOpKind, ScalarValue}; + let (mut scan, schema, batches) = fixture(); + if let QueryExpr::Scan { + predicates, source, .. + } = &mut scan + { + *source = Source::TimeSeries { + metric: "samples".into(), + }; + *predicates = vec![Predicate(Rc::new(QueryExpr::Compare { + left: Rc::new(QueryExpr::Column(0)), + op: CompareOpKind::Gt, + right: Rc::new(QueryExpr::Literal(ScalarValue::Int64(2))), + }))]; + } + for (batches, expected) in [(vec![], 0), (batches, 2)] { + let mut sources = DataSources::default(); + sources + .register( + Source::TimeSeries { + metric: "samples".into(), + }, + Arc::new(MemorySource::new(schema.clone(), batches).unwrap()), + ) + .unwrap(); + let plan = plan(scan.clone(), &schema, ExecutionDataState::QUERY_ROWS); + let graph = bind_with_data_sources(&plan, BTreeMap::new(), &[0], &sources).unwrap(); + block_on(async { + let mut s = graph.execute(&[0], context()).unwrap().remove(0); + let mut count = 0; + while let Some(b) = s.next().await { + count += b.unwrap().rows().len(); + } + assert_eq!(count, expected); + }); + } +} + +// A physical candidate can be compiled once without readers and rebound per run. +#[test] +fn compile_without_readers_and_rebind_inputs() { + use asap_physical_operators::{ + operators::Operator, + physical_planner::{compile, InputContract, Source}, + }; + let (scan, schema, batches) = fixture(); + let dag = plan(scan, &schema, ExecutionDataState::QUERY_ROWS); + let compiled = compile( + &dag, + BTreeMap::from([(0, InputContract::bounded(schema.clone()))]), + &[2], + ) + .unwrap(); + assert_eq!(compiled.input_contracts().count(), 1); + for _ in 0..2 { + let sources = BTreeMap::from([( + 0, + Box::new(Operator::source(schema.clone(), batches.clone()).unwrap()) as Source<'_>, + )]); + let graph = compiled.instantiate(sources).unwrap(); + let mut outputs = graph.execute(compiled.roots(), context()).unwrap(); + let result = block_on(outputs.remove(0).collect::>()); + assert!(result.iter().all(Result::is_ok)); + assert_eq!( + result + .iter() + .map(|b| b.as_ref().unwrap().rows().len()) + .sum::(), + 2 + ); + } + assert!(compiled.instantiate(BTreeMap::new()).is_err()); +} + +// Input boundedness must be proved during compilation, before readers exist. +#[test] +fn compilation_rejects_unknown_boundedness_for_sort() { + use asap_physical_operators::{ + physical_planner::{compile, InputContract}, + plan::{Boundedness, Emission, PlanProperties}, + }; + let (scan, schema, _) = fixture(); + let dag = plan(scan, &schema, ExecutionDataState::QUERY_ROWS); + let input = InputContract { + schema, + properties: PlanProperties { + boundedness: Boundedness::Unknown, + emission: Emission::Unknown, + }, + }; + assert!(compile(&dag, BTreeMap::from([(0, input)]), &[2]).is_err()); +} diff --git a/crates/asap-physical-operators/tests/summary_projection.rs b/crates/asap-physical-operators/tests/summary_projection.rs new file mode 100644 index 00000000..a61a596c --- /dev/null +++ b/crates/asap-physical-operators/tests/summary_projection.rs @@ -0,0 +1,161 @@ +//! Opaque state travels through a retained physical projection without scalar decoding. +use asap_physical_operators::{ + expressions::Expression, + factory::create_planner_accumulator, + operators::Operator, + physical_planner::{compile, CompiledPhysicalDag, InputContract, Source}, + runtime::{Limits, RunContext, Scope}, + values::{Batch, Value}, +}; +use futures::{executor::block_on, StreamExt}; +use planner_types::{ + post_asap::*, + pre_asap::{ColumnRef, DataType, ProjectItem, QueryExpr}, +}; +use std::{collections::BTreeMap, sync::Arc}; + +// A Post-ASAP projection may reorder/rename summary columns; recovery must retain +// the family and pass through the same immutable state, without decoding the payload. +#[test] +fn post_asap_summary_projection_survives_recovery() { + let family = SummaryFamilyType::ExactAggregate(ExactKind::Sum, ExactParams::Sum); + let schema = Arc::new(SummarySchema { + fields: vec![ + SummaryField { + name: "state".into(), + dtype: family.clone(), + nullable: false, + }, + SummaryField { + name: "service".into(), + dtype: SummaryFamilyType::Plain(DataType::Utf8), + nullable: false, + }, + ], + time_index: None, + }); + let output = SummarySchema { + fields: vec![ + schema.fields[1].clone(), + SummaryField { + name: "renamed".into(), + ..schema.fields[0].clone() + }, + ], + time_index: None, + }; + let dag = PostAsapDag { + nodes: vec![ + PostAsapDagNode { + id: PostAsapNodeId(0), + payload: PostAsapOperatorPayload::SummaryMerge, + output_schema: (*schema).clone(), + output_state: ExecutionDataState::INGESTION_SUMMARY, + guarantee: None, + }, + PostAsapDagNode { + id: PostAsapNodeId(1), + payload: PostAsapOperatorPayload::Value { + operation: ValueOperation::Project { + cols: vec![1, 0] + .into_iter() + .map(|index| ProjectItem { + alias: None, + expr: QueryExpr::Column(index), + }) + .collect(), + qualifier: None, + }, + }, + output_schema: output.clone(), + output_state: ExecutionDataState::INGESTION_SUMMARY, + guarantee: None, + }, + ], + edges: vec![PostAsapDagEdge { + producer: PostAsapNodeId(0), + consumer: PostAsapNodeId(1), + role: EdgeRole::Input, + intermediate_schema: (*schema).clone(), + data_state: ExecutionDataState::INGESTION_SUMMARY, + grouping: GroupingEdgeCompatibility::NotApplicable, + window: WindowEdgeCompatibility::NotApplicable, + }], + root: PostAsapNodeId(1), + }; + let program = compile( + &dag, + BTreeMap::from([(0, InputContract::bounded(schema.clone()))]), + &[1], + ) + .unwrap(); + let encoded = serde_json::to_vec(&program).unwrap(); + let program = serde_json::from_slice::(&encoded).unwrap(); + let mut forged: serde_json::Value = serde_json::from_slice(&encoded).unwrap(); + forged["nodes"]["1"]["Operator"]["operator"]["output"]["fields"][1]["dtype"] = + serde_json::json!({"Plain": "float64"}); + assert!( + serde_json::from_slice::(&serde_json::to_vec(&forged).unwrap()) + .is_err() + ); + assert!(Operator::project( + schema.clone(), + vec![("invalid".into(), Expression::Column(2))] + ) + .is_err()); + assert!(Operator::project( + schema.clone(), + vec![( + "invalid".into(), + Expression::Negate(Box::new(Expression::Column(0))) + )] + ) + .is_err()); + assert_eq!(*program.output_contract(1).unwrap().schema, output); + let mut accumulator = create_planner_accumulator( + &family, + &SummaryUpdate::column(ColumnRef::SampleValue), + &GroupingStrategy::PerSubpopulationInstance, + ) + .unwrap(); + accumulator.update_single(7., 1); + let state = Arc::from(accumulator.into_accumulator()); + let batch = Batch::try_new( + schema.clone(), + vec![vec![ + Value::Summary { + family, + state: Arc::clone(&state), + }, + Value::Utf8("api".into()), + ]], + ) + .unwrap(); + let graph = program + .instantiate(BTreeMap::from([( + 0, + Box::new(Operator::source(schema, vec![batch]).unwrap()) as Source<'_>, + )])) + .unwrap(); + let context = RunContext::new( + Scope::Query { + evaluation_time_ms: 1, + revision: 1, + }, + Limits::default(), + ) + .unwrap(); + block_on(async { + let mut output = graph.execute(&[1], context).unwrap().remove(0); + let batch = output.next().await.unwrap().unwrap(); + assert!(matches!(&batch.rows()[0][0], Value::Utf8(label) if label.as_ref() == "api")); + let Value::Summary { + state: projected, .. + } = &batch.rows()[0][1] + else { + panic!("missing summary") + }; + assert!(Arc::ptr_eq(&state, projected)); + assert!(output.next().await.is_none()); + }); +} diff --git a/crates/asap-physical-operators/tests/weighted_topk_binding.rs b/crates/asap-physical-operators/tests/weighted_topk_binding.rs new file mode 100644 index 00000000..664ae799 --- /dev/null +++ b/crates/asap-physical-operators/tests/weighted_topk_binding.rs @@ -0,0 +1,1018 @@ +//! Planner output binds directly to the shared runtime at a declared rate-value frontier. +use asap_aware_mapping::{ + accuracy::{ + AccuracyEvidenceProvider, DefaultAccuracyModel, EqualSplitAllocator, PropagationStats, + }, + cost_model::DefaultCostModel, + Replacement, ReplacementStrategy, SketchAlgorithmStrategy, TargetSubDAG, +}; +use asap_physical_operators::dag::{ + operators::Operator, + planner::{compile, InputContract, Source}, + values::{Batch, Value}, + Limits, RunContext, Scope, +}; +use futures::{executor::block_on, StreamExt}; +use planner_types::{ + post_asap::*, + pre_asap::{DataType, QueryExpr}, + types::AccuracyTarget, +}; +use std::{collections::BTreeMap, rc::Rc, sync::Arc}; +struct Evidence; +impl AccuracyEvidenceProvider for Evidence { + fn topk_max_distinct_items(&self, _: &QueryExpr) -> Option { + Some(1000) + } + fn propagation_stats( + &self, + op: &CompositionOperator, + _: &SummaryFamilyType, + _: Option<&SketchQuery>, + ) -> PropagationStats { + if matches!(op, CompositionOperator::TopKSelection) { + PropagationStats { + topk_selected_lower_bound: Some(101.), + topk_excluded_upper_bound: Some(100.), + topk_interval_failure_probability: Some(0.001), + ..Default::default() + } + } else { + Default::default() + } + } +} +// The evidence here exercises binding; it is not inferred from the sample data. +#[test] +fn planner_weighted_topk_binds_at_either_deployment_phase() { + assert_weighted_binding(&Evidence, SketchAlgorithm::CmsWithHeap); + assert_weighted_binding(&Evidence, SketchAlgorithm::CountSketchWithHeap); +} + +// Binding validates representation, while deployment owns evidence acceptance. +#[test] +fn physical_binding_does_not_impose_an_accuracy_acceptance_policy() { + assert_weighted_binding( + &asap_aware_mapping::accuracy::NoAccuracyEvidence, + SketchAlgorithm::CmsWithHeap, + ); + assert_weighted_binding( + &asap_aware_mapping::accuracy::NoAccuracyEvidence, + SketchAlgorithm::CountSketchWithHeap, + ); +} + +fn assert_weighted_binding(evidence: &dyn AccuracyEvidenceProvider, algorithm: SketchAlgorithm) { + let root = Rc::new( + lower_promql( + "topk by(job)(2, sum by(service, job)(rate(m[1m])))", + AccuracyTarget::Epsilon(0.1), + ) + .unwrap(), + ); + let strategy = SketchAlgorithmStrategy::new_with_planning_inputs_and_evidence( + &DefaultCostModel, + &DefaultAccuracyModel, + &EqualSplitAllocator, + evidence, + ); + let plan = strategy + .replacements(&TargetSubDAG::new(&root)) + .into_iter() + .find_map(|candidate| match candidate.replacement { + Replacement::Summary(node) + if candidate.rationale.contains(&format!("{algorithm:?}")) => + { + Some(node) + } + _ => None, + }) + .unwrap(); + let dag = compile_post_asap_dag(&plan).unwrap(); + let build=dag.nodes.iter().find(|node|matches!(&node.payload,PostAsapOperatorPayload::SummaryAgg{family:SummaryFamilyType::Sketch(kind,_),..}if kind.algorithm()==&algorithm)).unwrap(); + let rate_id = dag + .edges + .iter() + .find(|edge| edge.consumer == build.id) + .unwrap() + .producer; + let rates = Arc::new( + dag.nodes + .iter() + .find(|node| node.id == rate_id) + .unwrap() + .output_schema + .clone(), + ); + let rows = [ + ("auth", "api", 0.125), + ("auth", "api", 0.25), + ("checkout", "api", 0.3125), + ("search", "api", 0.0625), + ("ingest", "batch", 100.), + ("export", "batch", 80.), + ("cleanup", "batch", 20.), + ] + .into_iter() + .map(|(service, job, value)| { + rates + .fields + .iter() + .map(|field| match field.name.as_str() { + "service" => Value::Utf8(service.into()), + "job" => Value::Utf8(job.into()), + "value" => Value::Float64(value), + _ => match field.dtype { + SummaryFamilyType::Plain(DataType::Timestamp) => Value::Timestamp(60_000), + _ => panic!("unexpected rate column {field:?}"), + }, + }) + .collect() + }) + .collect(); + let batch = Batch::try_new(rates.clone(), rows).unwrap(); + for (phase, scope) in [ + ( + ExecutionTiming::IngestionTime, + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 60_000, + revision: 1, + }, + ), + ( + ExecutionTiming::QueryTime, + Scope::Query { + evaluation_time_ms: 60_000, + revision: 1, + }, + ), + ] { + let placed = dag + .with_execution_phases(&dag.nodes.iter().map(|node| (node.id, phase)).collect()) + .unwrap(); + let source = Box::new(Operator::source(rates.clone(), vec![batch.clone()]).unwrap()) + as Source<'static>; + let compiled = compile( + &placed, + BTreeMap::from([(rate_id.0 as u64, InputContract::bounded(rates.clone()))]), + &[dag.root.0 as u64], + ) + .unwrap(); + let graph = compiled + .instantiate(BTreeMap::from([(rate_id.0 as u64, source)])) + .unwrap(); + let context = RunContext::new(scope, Limits::default()).unwrap(); + let output = block_on(async { + let mut output = Vec::new(); + let mut stream = graph + .execute(&[dag.root.0 as u64], context) + .unwrap() + .remove(0); + while let Some(batch) = stream.next().await { + output.extend(batch.unwrap().rows().iter().cloned()); + } + output + }); + assert_eq!(output.len(), 4); + let mut scores = output + .iter() + .map(|row| { + row.iter() + .find_map(|v| { + if let Value::Float64(v) = v { + Some(*v) + } else { + None + } + }) + .unwrap() + }) + .collect::>(); + scores.sort_by(f64::total_cmp); + assert_eq!(scores, vec![0.3125, 0.375, 80., 100.]); + } +} + +use asap_frontend_promql::lower_promql_workload; +use planner_types::workload::{ + AccuracyRequirement, BatchEntry, DataWorkload, DurationMs, Evidence as WorkloadEvidence, + PlanningWorkload, Predictability, Query, QueryLanguage, QueryRequirements, QueryWorkload, + TimeSelection, +}; +pub fn lower_promql( + query: &str, + accuracy: AccuracyTarget, +) -> Result { + let workload = PlanningWorkload { + query_workload: QueryWorkload { + language: QueryLanguage::PromQL, + query_batch: Some(vec![BatchEntry { + query: Query(query.into()), + requirements: QueryRequirements { + accuracy: AccuracyRequirement::Explicit(accuracy), + ..Default::default() + }, + predictability: Predictability::Unknown, + invocations: 1, + execute_at: None, + time_selection: TimeSelection::default(), + }]), + repeating_queries: None, + }, + data_workload: Some(DataWorkload { + data_ingestion_interval: WorkloadEvidence { + value: Some(DurationMs(1_000)), + ..Default::default() + }, + ..Default::default() + }), + }; + let mut lowered = lower_promql_workload(&workload, 0)?; + Ok(lowered.remove(0)) +} + +// The old untyped heap updater must not silently round a Planner rate update. +#[test] +fn rate_updates_cannot_enter_integer_heap_factory() { + let family = SummaryFamilyType::Sketch( + SketchKind::new( + SketchAlgorithm::CmsWithHeap, + SketchParams::CmsWithHeap { + width: 272, + depth: 5, + heap_size: 100, + }, + ), + Default::default(), + ); + let input = SummaryUpdate { + item: Some(SummaryInputExpr::Column( + planner_types::pre_asap::ColumnRef::Named("service".into()), + )), + weight: SummaryInputExpr::Column(planner_types::pre_asap::ColumnRef::SampleValue), + weight_domain: WeightDomain::NonNegative { + proof: NonNegativeWeightProof::ResetAwareCounterDerivative, + }, + }; + assert!( + asap_physical_operators::factory::create_planner_accumulator( + &family, + &input, + &Default::default() + ) + .is_err() + ); +} + +/// A catalog-resolved per-series rate can feed a heap sketch directly, without +/// requiring an otherwise unnecessary grouped Sum between Rate and TopK. +#[test] +fn direct_rate_topk_exposes_heap_candidates_with_complete_series_identity() { + check_direct_rate_topk(false); +} + +// Unreferenced labels still distinguish series throughout Rate and heap readout. +#[test] +fn direct_rate_topk_preserves_dynamic_unreferenced_labels() { + check_direct_rate_topk(true); +} + +fn check_direct_rate_topk(dynamic: bool) { + use asap_physical_operators::physical_planner::promql_rows::{ + decode_series_identity, series_row, with_series_identity, SERIES_IDENTITY_COLUMN, + }; + let mut logical = + lower_promql("topk by(job)(2, rate(m[1m]))", AccuracyTarget::Epsilon(0.1)).unwrap(); + fn resolve_catalog(node: &mut QueryExpr) { + match node { + QueryExpr::Aggregate { child, .. } | QueryExpr::TimeRange { child, .. } => { + resolve_catalog(Rc::make_mut(child)) + } + QueryExpr::Scan { schema, .. } => { + schema.closed = true; + schema + .columns + .push(planner_types::pre_asap::schema::Column::new( + "service", + DataType::Utf8, + false, + )); + } + _ => panic!("unexpected input shape: {node:?}"), + } + } + if dynamic { + logical = with_series_identity(&logical).unwrap(); + } else { + resolve_catalog(&mut logical); + } + let root = Rc::new(logical); + let strategy = SketchAlgorithmStrategy::new_with_planning_inputs_and_evidence( + &DefaultCostModel, + &DefaultAccuracyModel, + &EqualSplitAllocator, + &Evidence, + ); + let candidates = strategy.replacements(&TargetSubDAG::new(&root)); + for algorithm in [ + SketchAlgorithm::CmsWithHeap, + SketchAlgorithm::CountSketchWithHeap, + ] { + let candidate = candidates + .iter() + .find_map(|candidate| match &candidate.replacement { + Replacement::Summary(node) + if candidate.rationale.contains(&format!("{algorithm:?}")) => + { + Some(node) + } + _ => None, + }) + .unwrap_or_else(|| panic!("missing {algorithm:?} over direct Rate")); + if dynamic { + let (source, ranked) = + asap_physical_operators::physical_planner::promql_rows::compile_rate_ranking( + candidate, + ) + .unwrap(); + assert!(matches!( + source.expr, + SummaryExpr::ValueOperation { + operation: ValueOperation::FinalizeExactAccumulator, + .. + } + )); + assert_eq!(ranked.input_contracts().count(), 1); + let encoded = String::from_utf8(serde_json::to_vec(&ranked).unwrap()).unwrap(); + assert!(encoded.contains("KeyedSummaryBuild")); + assert!(encoded.contains("KeyedReadout")); + assert!( + !encoded.contains("\"Rate\""), + "Rate must be supplied by its exact stored-state readout" + ); + } + let dag = compile_post_asap_dag(candidate).unwrap(); + assert!(dag.nodes.iter().any(|node| matches!(&node.payload, + PostAsapOperatorPayload::SummaryAgg { family: SummaryFamilyType::Sketch(kind, _), .. } if kind.algorithm() == &algorithm))); + let build = dag.nodes.iter().find(|node| matches!(&node.payload, + PostAsapOperatorPayload::SummaryAgg { family: SummaryFamilyType::Sketch(kind, _), .. } if kind.algorithm() == &algorithm)).unwrap(); + let input_id = dag + .edges + .iter() + .find(|edge| edge.consumer == build.id) + .unwrap() + .producer; + let schema = Arc::new( + dag.nodes + .iter() + .find(|node| node.id == input_id) + .unwrap() + .output_schema + .clone(), + ); + let raw = dag + .nodes + .iter() + .find(|node| { + matches!( + &node.payload, + PostAsapOperatorPayload::Fallback { + expression: QueryExpr::TimeRange { .. } + } + ) + }) + .unwrap_or_else(|| panic!("no raw counter source: {dag:?}")); + let raw_schema = Arc::new(raw.output_schema.clone()); + let raw_compiled = compile( + &dag, + BTreeMap::from([( + u64::from(raw.id.0), + InputContract::bounded(raw_schema.clone()), + )]), + &[u64::from(dag.root.0)], + ) + .unwrap(); + let bytes = serde_json::to_vec(&raw_compiled).unwrap(); + let raw_compiled = serde_json::from_slice::< + asap_physical_operators::physical_planner::CompiledPhysicalDag, + >(&bytes) + .unwrap(); + // Each evaluation receives a complete raw window. A reset, a stopped + // series and an expired leader must not retain last run's heap weights. + for (end, series, expected) in [ + ( + 60_000, + vec![ + ("auth", vec![10., 30., 50.]), + ("checkout", vec![10., 50., 90.]), + ("search", vec![10., 70., 130.]), + ], + vec![11. / 6., 8. / 3.], + ), + ( + 120_000, + vec![ + ("auth", vec![100., 10., 50.]), + ("checkout", vec![100., 100., 100.]), + ], + vec![0., 1.25], + ), + ] { + let mut raw_rows = Vec::new(); + for (service, samples) in series { + for (offset, value) in [10_000, 30_000, 50_000].into_iter().zip(samples) { + if dynamic { + raw_rows.push( + series_row( + &raw_schema, + &BTreeMap::from([ + ("job".into(), "api".into()), + ("service".into(), service.into()), + ("unreferenced".into(), format!("{service}-extra")), + ]), + end - 60_000 + offset, + value, + ) + .unwrap(), + ); + continue; + } + raw_rows.push( + raw_schema + .fields + .iter() + .map(|field| match field.name.as_str() { + "service" => Value::Utf8(service.into()), + "job" => Value::Utf8("api".into()), + "value" => Value::Float64(value), + "ts" => Value::Timestamp(end - 60_000 + offset), + _ => panic!("unexpected raw field"), + }) + .collect(), + ); + } + } + let raw_batch = Batch::try_new(raw_schema.clone(), raw_rows).unwrap(); + for scope in [ + Scope::Ingestion { + window_start_ms: end - 60_000, + window_end_ms: end, + revision: 1, + }, + Scope::Query { + evaluation_time_ms: end, + revision: 1, + }, + ] { + let source = Box::new( + Operator::source(raw_schema.clone(), vec![raw_batch.clone()]).unwrap(), + ) as Source<'static>; + let graph = raw_compiled + .instantiate(BTreeMap::from([(u64::from(raw.id.0), source)])) + .unwrap(); + let context = RunContext::new(scope, Limits::default()).unwrap(); + let mut raw_scores = block_on(async { + let mut scores = Vec::new(); + let mut stream = graph + .execute(&[u64::from(dag.root.0)], context) + .unwrap() + .remove(0); + while let Some(batch) = stream.next().await { + let batch = batch.unwrap(); + for row in batch.rows() { + if dynamic { + let column = batch + .schema() + .fields + .iter() + .position(|field| field.name == SERIES_IDENTITY_COLUMN) + .unwrap(); + let Value::Utf8(encoded) = &row[column] else { + panic!("identity lost"); + }; + let labels = decode_series_identity(encoded).unwrap(); + assert_eq!(labels["job"], "api"); + assert_eq!( + labels["unreferenced"], + format!("{}-extra", labels["service"]) + ); + } + assert!(row.iter().any( + |value| matches!(value, Value::Timestamp(time) if *time == end) + )); + scores.extend(row.iter().filter_map(|value| match value { + Value::Float64(value) => Some(*value), + _ => None, + })); + } + } + scores + }); + raw_scores.sort_by(f64::total_cmp); + assert_eq!(raw_scores.len(), expected.len()); + for (actual, expected) in raw_scores.iter().zip(&expected) { + assert!( + (actual - expected).abs() < 1e-12, + "raw counter semantics must precede heap ranking: {raw_scores:?}" + ); + } + } + } + let compiled = compile( + &dag, + BTreeMap::from([( + u64::from(input_id.0), + InputContract::bounded(schema.clone()), + )]), + &[u64::from(dag.root.0)], + ) + .unwrap(); + for (time, values, expected) in [ + ( + 60_000, + vec![("auth", 3.), ("checkout", 2.), ("search", 1.)], + vec![2., 3.], + ), + ( + 61_000, + vec![("auth", 0.), ("checkout", 2.), ("search", 4.)], + vec![2., 4.], + ), + (62_000, vec![("auth", 0.), ("checkout", 2.)], vec![0., 2.]), + ] { + let rows = values + .into_iter() + .map(|(service, value)| { + if dynamic { + return series_row( + &schema, + &BTreeMap::from([ + ("job".into(), "api".into()), + ("service".into(), service.into()), + ]), + time, + value, + ) + .unwrap(); + } + schema + .fields + .iter() + .map(|field| match field.name.as_str() { + "service" => Value::Utf8(service.into()), + "job" => Value::Utf8("api".into()), + "value" => Value::Float64(value), + "ts" => Value::Timestamp(time), + _ => panic!("unexpected rate field {field:?}"), + }) + .collect() + }) + .collect(); + let batch = Batch::try_new(schema.clone(), rows).unwrap(); + for scope in [ + Scope::Query { + evaluation_time_ms: time, + revision: 1, + }, + Scope::Ingestion { + window_start_ms: time - 60_000, + window_end_ms: time, + revision: 1, + }, + ] { + let source = + Box::new(Operator::source(schema.clone(), vec![batch.clone()]).unwrap()) + as Source<'static>; + let graph = compiled + .instantiate(BTreeMap::from([(u64::from(input_id.0), source)])) + .unwrap(); + let context = RunContext::new(scope, Limits::default()).unwrap(); + let mut scores = block_on(async { + let mut scores = vec![]; + let mut stream = graph + .execute(&[u64::from(dag.root.0)], context) + .unwrap() + .remove(0); + while let Some(batch) = stream.next().await { + let batch = batch.unwrap(); + for row in batch.rows() { + assert!(row.iter().any( + |value| matches!(value, Value::Timestamp(actual) if *actual == time) + )); + scores.push( + row.iter() + .find_map(|value| { + if let Value::Float64(value) = value { + Some(*value) + } else { + None + } + }) + .unwrap(), + ); + } + } + scores + }); + scores.sort_by(f64::total_cmp); + assert_eq!( + scores, expected, + "heap snapshots must not accumulate across evaluations" + ); + } + } + } +} + +// Spatial ranking consumes one eligible instant vector. Signed values require +// CountSketch; a raw metric does not establish the non-negative CMS contract. +#[test] +fn spatial_topk_exposes_signed_heap_candidate_over_complete_snapshot() { + use asap_physical_operators::physical_planner::promql_rows::{ + decode_series_identity, series_row, with_series_identity, SERIES_IDENTITY_COLUMN, + }; + let logical = lower_promql("topk by(job)(1, m)", AccuracyTarget::Epsilon(0.1)).unwrap(); + let root = Rc::new(with_series_identity(&logical).unwrap()); + let strategy = SketchAlgorithmStrategy::new_with_planning_inputs_and_evidence( + &DefaultCostModel, + &DefaultAccuracyModel, + &EqualSplitAllocator, + &Evidence, + ); + let candidates = strategy + .current_series_topk_candidates(&root, &AccuracyTarget::Epsilon(0.1)) + .candidates; + assert!(!candidates + .iter() + .any(|c| c.rationale.contains("CmsWithHeap"))); + let selected = candidates + .iter() + .find_map(|candidate| match &candidate.replacement { + Replacement::Summary(node) if candidate.rationale.contains("CountSketchWithHeap") => { + Some(node) + } + _ => None, + }) + .expect("signed spatial TopK must expose CountSketch with heap"); + let dag = compile_post_asap_dag(selected).unwrap(); + let raw = dag + .nodes + .iter() + .find(|node| { + matches!( + &node.payload, + PostAsapOperatorPayload::Fallback { + expression: QueryExpr::TimeRange { .. } + } + ) + }) + .unwrap(); + let schema = Arc::new(raw.output_schema.clone()); + let program = compile( + &dag, + BTreeMap::from([(u64::from(raw.id.0), InputContract::bounded(schema.clone()))]), + &[u64::from(dag.root.0)], + ) + .unwrap(); + let snapshot_program = + asap_physical_operators::physical_planner::promql_rows::compile_current_series_readout( + selected, + ) + .unwrap(); + let encoded: serde_json::Value = + serde_json::from_slice(&serde_json::to_vec(&snapshot_program).unwrap()).unwrap(); + assert!(!encoded.to_string().contains("CurrentSeries")); + assert!(encoded.to_string().contains("KeyedSummaryBuild")); + assert!(encoded.to_string().contains("KeyedReadout")); + for (values, expected, score) in [ + ([100., 20.], "a", 100.), + ([1., 20.], "b", 20.), + ([-10., -2.], "b", -2.), + ] { + let rows = ["a", "b"] + .into_iter() + .zip(values) + .map(|(instance, value)| { + series_row( + &schema, + &BTreeMap::from([ + ("job".into(), "api".into()), + ("unreferenced".into(), instance.into()), + ]), + 60_000, + value, + ) + .unwrap() + }) + .collect(); + let batch = Batch::try_new(schema.clone(), rows).unwrap(); + let graph = program + .instantiate(BTreeMap::from([( + u64::from(raw.id.0), + Box::new(Operator::source(schema.clone(), vec![batch]).unwrap()) as Source<'_>, + )])) + .unwrap(); + block_on(async { + let context = RunContext::new( + Scope::Query { + evaluation_time_ms: 60_000, + revision: 0, + }, + Limits::default(), + ) + .unwrap(); + let mut stream = graph.execute(program.roots(), context).unwrap().remove(0); + let mut result = Vec::new(); + while let Some(batch) = stream.next().await { + let batch = batch.unwrap(); + let identity = batch + .schema() + .fields + .iter() + .position(|f| f.name == SERIES_IDENTITY_COLUMN) + .unwrap(); + let value = batch + .schema() + .fields + .iter() + .position(|f| f.name == "value") + .unwrap(); + for row in batch.rows() { + let Value::Utf8(labels) = &row[identity] else { + panic!() + }; + let Value::Float64(v) = row[value] else { + panic!() + }; + result.push(( + decode_series_identity(labels).unwrap()["unreferenced"].clone(), + v, + )); + } + } + assert_eq!(result, vec![(expected.into(), score)]); + }); + } +} + +// Placement changes execution ownership only. Every fixed-window candidate +// contains Rate finalization before a fresh heap, with query readout downstream. +#[test] +fn planner_exposes_fixed_window_rate_heap_precompute_candidates() { + use asap_physical_operators::physical_planner::{ + compile_candidate, promql_rows::with_series_identity, + }; + let root = Rc::new( + with_series_identity( + &lower_promql("topk by(job)(2, rate(m[1m]))", AccuracyTarget::Epsilon(0.1)).unwrap(), + ) + .unwrap(), + ); + let strategy = SketchAlgorithmStrategy::new_with_planning_inputs_and_evidence( + &DefaultCostModel, + &DefaultAccuracyModel, + &EqualSplitAllocator, + &Evidence, + ); + let candidates = strategy.fixed_window_rate_candidates(&root).candidates; + assert_eq!(candidates.len(), 2); + for candidate in candidates { + let Replacement::Summary(root) = candidate.replacement else { + panic!() + }; + let dag = compile_post_asap_dag(&root).unwrap(); + let state = dag + .nodes + .iter() + .find(|node| { + matches!( + &node.payload, + PostAsapOperatorPayload::SummaryAgg { + family: SummaryFamilyType::ExactAggregate(ExactKind::Rate, _), + .. + } + ) + }) + .unwrap(); + let heap = dag + .nodes + .iter() + .find(|node| { + matches!( + &node.payload, + PostAsapOperatorPayload::SummaryAgg { + family: SummaryFamilyType::Sketch(..), + .. + } + ) + }) + .unwrap(); + assert_eq!(heap.output_state.timing, ExecutionTiming::IngestionTime); + let physical = compile_candidate( + &dag, + BTreeMap::from([( + u64::from(state.id.0), + InputContract::bounded(Arc::new(state.output_schema.clone())), + )]), + &[u64::from(dag.root.0)], + &[u64::from(heap.id.0)], + ) + .unwrap(); + let exported = asap_physical_operators::physical_planner::promql_rows::compile_fixed_window_rate_aggregation(&root).unwrap(); + assert_eq!( + serde_json::to_vec(&exported).unwrap(), + serde_json::to_vec(&physical).unwrap() + ); + assert!( + asap_physical_operators::physical_planner::promql_rows::compile_rate_ranking(&root) + .is_err(), + "query binding must not move the selected precompute frontier" + ); + // Execute the selected split across a state serialization boundary. + // Each run builds fresh weights from that window's counters. + let execute = |plan: &asap_physical_operators::physical_planner::CompiledPhysicalDag, + input: Batch, + scope: Scope| { + let id = plan.input_contracts().next().unwrap().0; + let source = Box::new(Operator::source(input.schema().clone(), vec![input]).unwrap()) + as Source<'static>; + let graph = plan.instantiate(BTreeMap::from([(id, source)])).unwrap(); + block_on(async { + let mut stream = graph + .execute( + plan.roots(), + RunContext::new(scope, Limits::default()).unwrap(), + ) + .unwrap() + .remove(0); + let mut batches = Vec::new(); + while let Some(batch) = stream.next().await { + batches.push((*batch.unwrap()).clone()); + } + assert_eq!(batches.len(), 1); + batches.remove(0) + }) + }; + let (family, input, grouping) = match &state.payload { + PostAsapOperatorPayload::SummaryAgg { + family, + input, + grouping, + .. + } => (family, input, grouping), + _ => unreachable!(), + }; + for (end, samples, leader) in [ + ( + 60_000, + [[0., 100., 200.], [0., 10., 20.], [0., 1., 2.]], + "a", + ), + ( + 120_000, + [[200., 200., 200.], [100., 0., 300.], [2., 3., 4.]], + "b", + ), + ] { + let schema = Arc::new(state.output_schema.clone()); + let rows = samples + .into_iter() + .zip(["a", "b", "c"]) + .map(|(samples, label)| { + let mut accumulator = + asap_physical_operators::factory::create_planner_accumulator( + family, input, grouping, + ) + .unwrap(); + for (offset, value) in [10_000, 30_000, 50_000].into_iter().zip(samples) { + accumulator.update_single(value, end - 60_000 + offset); + } + let summary = Value::Summary { + family: family.clone(), + state: Arc::from(accumulator.into_accumulator()), + }; + schema + .fields + .iter() + .map(|field| match &field.dtype { + SummaryFamilyType::ExactAggregate(..) => summary.clone(), + SummaryFamilyType::Plain(DataType::Timestamp) => Value::Timestamp(end), + SummaryFamilyType::Plain(DataType::Utf8) + if field.name == "$promql_series_identity" => + { + Value::Utf8( + serde_json::to_string(&BTreeMap::from([ + ("job", "api"), + ("instance", label), + ])) + .unwrap() + .into(), + ) + } + SummaryFamilyType::Plain(DataType::Utf8) => Value::Utf8("api".into()), + _ => panic!("unexpected state field {field:?}"), + }) + .collect() + }) + .collect(); + let batch = Batch::try_new(schema, rows).unwrap(); + let precompute = physical.precompute.as_ref().unwrap(); + let heap = execute( + precompute, + batch, + Scope::Ingestion { + window_start_ms: end - 60_000, + window_end_ms: end, + revision: 1, + }, + ); + let result = execute( + &physical.query, + heap, + Scope::Query { + evaluation_time_ms: end, + revision: 1, + }, + ); + let identity = result + .schema() + .fields + .iter() + .position(|f| f.name == "$promql_series_identity") + .unwrap(); + let Value::Utf8(encoded) = &result.rows()[0][identity] else { + panic!() + }; + let labels: BTreeMap = serde_json::from_str(encoded).unwrap(); + assert_eq!(labels["instance"], leader); + assert_eq!(result.rows().len(), 2); + } + let precompute = + String::from_utf8(serde_json::to_vec(&physical.precompute.unwrap()).unwrap()).unwrap(); + assert!(precompute.contains("KeyedSummaryBuild")); + assert!(precompute.contains("Rate")); + assert!( + !String::from_utf8(serde_json::to_vec(&physical.query).unwrap()) + .unwrap() + .contains("KeyedSummaryBuild") + ); + } +} + +// Grouped Rate has a legal stored Sum candidate as well as query-time reduction. +#[test] +fn grouped_rate_exposes_precomputed_sum_with_query_readout() { + let root = Rc::new( + asap_physical_operators::physical_planner::promql_rows::with_series_identity( + &lower_promql("sum by(job)(rate(m[1m]))", AccuracyTarget::Exact).unwrap(), + ) + .unwrap(), + ); + let strategy = SketchAlgorithmStrategy::new_with_planning_inputs_and_evidence( + &DefaultCostModel, + &DefaultAccuracyModel, + &EqualSplitAllocator, + &Evidence, + ); + let direct = strategy.query_time_rate_aggregation_candidates(&root); + assert!( + direct.candidates.iter().any(|candidate| { + let Replacement::Summary(root) = &candidate.replacement else { + return false; + }; + let Ok((_, program)) = + asap_physical_operators::physical_planner::promql_rows::compile_rate_ranking(root) + else { + return false; + }; + let output = program.output_contract(program.roots()[0]).unwrap(); + output + .schema + .fields + .iter() + .all(|field| matches!(field.dtype, SummaryFamilyType::Plain(_))) + }), + "query-time grouped Rate must finalize Sum inside the physical graph" + ); + let candidates = strategy.fixed_window_rate_candidates(&root).candidates; + assert!( + !candidates.is_empty(), + "Planner must expose Rate -> grouped Sum at ingestion" + ); + for candidate in candidates { + let Replacement::Summary(root) = candidate.replacement else { + panic!() + }; + let physical = asap_physical_operators::physical_planner::promql_rows::compile_fixed_window_rate_aggregation(&root).unwrap(); + let precompute = + String::from_utf8(serde_json::to_vec(&physical.precompute.unwrap()).unwrap()).unwrap(); + assert!( + precompute.contains("SummaryBuild") + && precompute.contains("Rate") + && precompute.contains("Sum") + ); + let query = String::from_utf8(serde_json::to_vec(&physical.query).unwrap()).unwrap(); + assert!(query.contains("Readout") && !query.contains("SummaryBuild")); + } +} diff --git a/crates/integration-tests/Cargo.toml b/crates/integration-tests/Cargo.toml index 7b0c0d3e..5a5de990 100644 --- a/crates/integration-tests/Cargo.toml +++ b/crates/integration-tests/Cargo.toml @@ -13,3 +13,6 @@ asap-aware-mapping = { path = "../asap-aware-mapping" } asap_sketchlib = { workspace = true } serde_json = "1" tokio = { version = "1", features = ["rt", "macros", "rt-multi-thread"] } + +asap-physical-operators = { path = "../asap-physical-operators" } +futures = "0.3" diff --git a/crates/integration-tests/tests/kll_pane_execution.rs b/crates/integration-tests/tests/kll_pane_execution.rs new file mode 100644 index 00000000..e818796f --- /dev/null +++ b/crates/integration-tests/tests/kll_pane_execution.rs @@ -0,0 +1,286 @@ +//! Maintenance -> stored pane state -> independently bound query execution. +mod physical_common; +use asap_physical_operators::{ + operators::{Operator, ReadoutQuery}, + physical_planner::{CompiledPhysicalDag, InputContract, Source}, + plan::{PhysicalDag, PhysicalOperator, PlanProperties}, + runtime::{Input, Limits, OutputStream, RunContext, Scope}, + summary_kernels::datasketches_kll::DatasketchesKLLAccumulator, + values::{Batch, Schema, Value}, + AggregateCore, Error, +}; +use asap_types::{ + post_asap::{ + SketchAlgorithm, SketchKind, SketchParams, SketchQuery, SummaryFamilyType, SummaryField, + SummarySchema, + }, + pre_asap::DataType, +}; +use futures::{executor::block_on, StreamExt}; +use std::{ + collections::BTreeMap, + sync::{ + atomic::{AtomicUsize, Ordering}, + Arc, + }, +}; + +fn family(k: u32) -> SummaryFamilyType { + SummaryFamilyType::Sketch( + SketchKind::new(SketchAlgorithm::Kll, SketchParams::Kll { k }), + Default::default(), + ) +} +fn raw_schema() -> Schema { + Arc::new(SummarySchema { + fields: vec![SummaryField { + name: "value".into(), + dtype: SummaryFamilyType::Plain(DataType::Float64), + nullable: false, + }], + time_index: None, + }) +} +fn query_scope() -> Scope { + Scope::Query { + evaluation_time_ms: 300_000, + revision: 1, + } +} +fn fixture() -> (CompiledPhysicalDag, Operator, Schema) { + let raw = raw_schema(); + let build = Operator::summary_build(raw.clone(), family(200), 0, None, vec![]).unwrap(); + let state = build.schema(); + let maintenance = CompiledPhysicalDag::from_operators( + BTreeMap::from([(0, InputContract::bounded(raw))]), + BTreeMap::from([(1, (vec![0], build))]), + vec![1], + ) + .unwrap(); + let merge = Operator::summary_merge(state.clone(), 0, vec![]).unwrap(); + (maintenance, merge, state) +} +fn pane_state(maintenance: &CompiledPhysicalDag, pane: i64) -> Arc { + // Twenty samples in each (start,end] one-minute pane; k=200 avoids + // compaction so quantiles and sample counts have deterministic oracles. + let raw = raw_schema(); + let rows = (0..20) + .map(|i| vec![Value::Float64((pane * 20 + i) as f64)]) + .collect(); + let state = physical_common::execute( + maintenance, + BTreeMap::from([(0, Batch::try_new(raw, rows).unwrap())]), + Scope::Ingestion { + window_start_ms: pane * 60_000, + window_end_ms: (pane + 1) * 60_000, + revision: 1, + }, + ); + let Value::Summary { state, .. } = &state[0][0].rows()[0][0] else { + panic!("missing KLL") + }; + state.clone() +} +fn restore(schema: Schema, states: &[Arc]) -> Batch { + Batch::try_new( + schema, + states + .iter() + .map(|state| { + vec![Value::Summary { + family: family(200), + state: state.clone(), + }] + }) + .collect(), + ) + .unwrap() +} +fn readout(schema: Schema, q: f64) -> Operator { + Operator::readout(schema, 0, ReadoutQuery::Sketch(SketchQuery::Quantile { q })).unwrap() +} +struct CountStarts { + operator: Operator, + starts: Arc, +} +impl PhysicalOperator for CountStarts { + fn name(&self) -> &str { + self.operator.name() + } + fn properties(&self, inputs: &[PlanProperties]) -> PlanProperties { + self.operator.properties(inputs) + } + fn requires_bounded_input(&self) -> bool { + self.operator.requires_bounded_input() + } + fn input_schemas(&self) -> Vec { + self.operator.input_schemas() + } + fn output_schema(&self) -> Schema { + self.operator.output_schema() + } + fn output_bytes(&self, batch: &Batch) -> usize { + self.operator.output_bytes(batch) + } + fn start<'a>( + &'a self, + inputs: Vec>, + context: RunContext, + ) -> Result, Error> { + self.starts.fetch_add(1, Ordering::SeqCst); + self.operator.start(inputs, context) + } +} + +/// Actual codec bytes survive destruction of maintenance state; one shared +/// native merge supplies p50, p99 and the population-count oracle per run. +#[test] +fn five_panes_roundtrip_and_shared_merge_runs_once() { + let (maintenance, merge, schema) = fixture(); + let panes: Vec<_> = (0..6).map(|pane| pane_state(&maintenance, pane)).collect(); + drop(maintenance); + let compiled = CompiledPhysicalDag::from_operators( + (0..5) + .map(|id| (id, InputContract::bounded(schema.clone()))) + .collect(), + BTreeMap::from([ + ( + 5, + ( + vec![0, 1, 2, 3, 4], + Operator::union(schema.clone(), 5).unwrap(), + ), + ), + (6, (vec![5], merge.clone())), + (7, (vec![6], readout(schema.clone(), 0.5))), + (8, (vec![6], readout(schema.clone(), 0.99))), + ]), + vec![6, 7, 8], + ) + .unwrap(); + let layout = asap_types::post_asap::PaneLayout { + pane_width_ms: 60_000, + pane_origin_ms: Some(0), + }; + assert!( + asap_types::post_asap::validate_pane_coverage( + &layout, + Some(330_000), + &asap_types::post_asap::WindowEdgeCoverage::PaneAligned + ) + .is_err(), + "moving window edges require residual computation" + ); + for offset in [0, 1] { + let restored = restore(schema.clone(), &panes[offset..offset + 5]); + let evaluation_time_ms = (5 + offset as i64) * 60_000; + asap_types::post_asap::validate_pane_coverage( + &layout, + Some(evaluation_time_ms), + &asap_types::post_asap::WindowEdgeCoverage::PaneAligned, + ) + .unwrap(); + let inputs: BTreeMap<_, _> = (0..5) + .map(|id| { + ( + id as u64, + restore(schema.clone(), &panes[offset + id..offset + id + 1]), + ) + }) + .collect(); + // A five-pane deployment cannot bind only four state slots. + let incomplete: BTreeMap<_, _> = inputs + .iter() + .take(4) + .map(|(&id, batch)| { + ( + id, + Box::new(Operator::source(schema.clone(), vec![batch.clone()]).unwrap()) + as Source<'_>, + ) + }) + .collect(); + assert!(compiled.instantiate(incomplete).is_err()); + let result = physical_common::execute( + &compiled, + inputs, + Scope::Query { + evaluation_time_ms, + revision: 1, + }, + ); + let Value::Summary { state, .. } = &result[0][0].rows()[0][0] else { + panic!("missing merged state") + }; + let kll = state + .as_any() + .downcast_ref::() + .unwrap(); + assert_eq!(kll.inner.count(), 100); + let value = |index: usize| match result[index][0].rows()[0][0] { + Value::Float64(value) => value, + _ => panic!("missing quantile"), + }; + assert!((value(1) - (50 + offset * 20) as f64).abs() <= 1.); + assert!((value(2) - (99 + offset * 20) as f64).abs() <= 1.); + let starts = Arc::new(AtomicUsize::new(0)); + let mut dag = PhysicalDag::default(); + dag.add( + 0, + vec![], + Operator::source(schema.clone(), vec![restored]).unwrap(), + ) + .unwrap(); + dag.add( + 1, + vec![0], + CountStarts { + operator: merge.clone(), + starts: starts.clone(), + }, + ) + .unwrap(); + dag.add(2, vec![1], readout(schema.clone(), 0.5)).unwrap(); + dag.add(3, vec![1], readout(schema.clone(), 0.99)).unwrap(); + let outputs = block_on(futures::future::join_all( + dag.execute( + &[2, 3], + RunContext::new(query_scope(), Limits::default()).unwrap(), + ) + .unwrap() + .into_iter() + .map(|stream| stream.collect::>()), + )); + assert!(outputs + .iter() + .all(|output| output.len() == 1 && output[0].is_ok())); + assert_eq!(starts.load(Ordering::SeqCst), 1); + } +} + +/// Relabelled KLL parameters and missing bindings fail explicitly. +#[test] +fn panes_reject_parameters_schema_and_missing_binding() { + let (_maintenance, merge, schema) = fixture(); + let wrong = Value::Summary { + family: family(200), + state: Arc::new(DatasketchesKLLAccumulator::new(128)), + }; + assert!(Batch::try_new(schema.clone(), vec![vec![wrong]]).is_err()); + let compiled = CompiledPhysicalDag::from_operators( + BTreeMap::from([(0, InputContract::bounded(schema))]), + BTreeMap::from([(1, (vec![0], merge))]), + vec![1], + ) + .unwrap(); + assert!(compiled.instantiate(BTreeMap::new()).is_err()); + let raw = raw_schema(); + let source = Operator::source( + raw.clone(), + vec![Batch::try_new(raw, vec![vec![Value::Float64(1.)]]).unwrap()], + ) + .unwrap(); + assert!(compiled + .instantiate(BTreeMap::from([(0, Box::new(source) as Source<'_>)])) + .is_err()); +} diff --git a/crates/integration-tests/tests/physical_common/mod.rs b/crates/integration-tests/tests/physical_common/mod.rs new file mode 100644 index 00000000..aca4d4ef --- /dev/null +++ b/crates/integration-tests/tests/physical_common/mod.rs @@ -0,0 +1,42 @@ +use asap_physical_operators::{ + operators::Operator, + physical_planner::{CompiledPhysicalDag, Source}, + runtime::{Limits, RunContext, Scope}, + values::Batch, +}; +use futures::{executor::block_on, StreamExt}; +use std::collections::BTreeMap; + +pub fn execute( + plan: &CompiledPhysicalDag, + inputs: BTreeMap, + scope: Scope, +) -> Vec> { + let sources = inputs + .into_iter() + .map(|(id, batch)| { + ( + id, + Box::new(Operator::source(batch.schema().clone(), vec![batch]).unwrap()) + as Source<'_>, + ) + }) + .collect(); + let dag = plan.instantiate(sources).unwrap(); + block_on(async { + let streams = dag + .execute( + plan.roots(), + RunContext::new(scope, Limits::default()).unwrap(), + ) + .unwrap(); + futures::future::join_all(streams.into_iter().map(|mut stream| async move { + let mut batches = Vec::new(); + while let Some(batch) = stream.next().await { + batches.push((*batch.unwrap()).clone()); + } + batches + })) + .await + }) +} diff --git a/crates/integration-tests/tests/sql_to_physical.rs b/crates/integration-tests/tests/sql_to_physical.rs new file mode 100644 index 00000000..be96107e --- /dev/null +++ b/crates/integration-tests/tests/sql_to_physical.rs @@ -0,0 +1,165 @@ +//! SQL frontend, candidate selection, physical compilation and fresh-run execution. +use asap_aware_mapping::{search_workload, DefaultCostModel}; +use asap_frontend_sql::{lower_sql, SqlCatalog}; +use asap_physical_operators::{ + physical_planner::{compile, InputContract, Source}, + runtime::{Limits, RunContext, Scope}, + sources::{DataSources, MemorySource}, + values::{Batch, Value}, +}; +use asap_types::{ + post_asap::{compile_post_asap_dag, PostAsapOperatorPayload, SummaryFamilyType}, + pre_asap::{Column, DataType, QueryExpr, Schema}, + types::AccuracyTarget, +}; +use futures::StreamExt; +use std::{collections::BTreeMap, rc::Rc, sync::Arc}; + +/// SQL filtering and grouped aggregation survive logical/physical lowering; +/// rebinding the compiled DAG runs against new data rather than cached results. +#[tokio::test] +async fn sql_filter_grouped_sum_executes_and_rebinds() { + let catalog = SqlCatalog::new().with_table( + "metrics", + Schema::new(vec![ + Column::new("service", DataType::Utf8, false), + Column::new("value", DataType::Float64, true), + ]), + ); + for query in [ + "SELECT service, SUM(value) AS total FROM metrics WHERE value > 1 GROUP BY service", + "SELECT service, SUM(value) AS total FROM metrics GROUP BY service", + ] { + let logical = Rc::new( + lower_sql(query, &catalog, AccuracyTarget::Exact) + .await + .unwrap(), + ); + let space = search_workload(vec![("sql", logical)]); + let selected = space + .global_selection(&DefaultCostModel) + .assemble_selected_dag(&space.roots[0].1) + .unwrap() + .unwrap(); + let dag = compile_post_asap_dag(&selected).unwrap(); + let scan = dag + .nodes + .iter() + .find(|node| { + matches!( + &node.payload, + PostAsapOperatorPayload::Fallback { + expression: QueryExpr::Scan { .. } + } + ) + }) + .expect("raw SQL scan"); + let schema = Arc::new(scan.output_schema.clone()); + assert!(schema + .fields + .iter() + .all(|field| matches!(field.dtype, SummaryFamilyType::Plain(_)))); + let plan = compile( + &dag, + BTreeMap::from([(u64::from(scan.id.0), InputContract::bounded(schema.clone()))]), + &[u64::from(dag.root.0)], + ) + .unwrap(); + for multiplier in [1., 2.] { + let rows = [ + ("api", Some(2.)), + ("api", Some(3.)), + ("api", None), + ("batch", Some(4.)), + ("batch", Some(1.)), + ] + .into_iter() + .map(|(service, value)| { + schema + .fields + .iter() + .map(|field| match field.name.as_str() { + "service" => Value::Utf8(service.into()), + "value" => { + value.map_or(Value::Null, |value| Value::Float64(value * multiplier)) + } + _ => panic!("unexpected field {field:?}"), + }) + .collect() + }) + .collect(); + let PostAsapOperatorPayload::Fallback { expression } = &scan.payload else { + unreachable!() + }; + let QueryExpr::Scan { source, .. } = expression else { + unreachable!() + }; + let mut sources = DataSources::default(); + sources + .register( + source.clone(), + Arc::new( + MemorySource::new( + schema.clone(), + vec![Batch::try_new(schema.clone(), rows).unwrap()], + ) + .unwrap(), + ), + ) + .unwrap(); + let bound = plan + .instantiate(BTreeMap::from([( + u64::from(scan.id.0), + Box::new(sources.bind(expression).unwrap()) as Source<'_>, + )])) + .unwrap(); + let mut stream = bound + .execute( + plan.roots(), + RunContext::new( + Scope::Query { + evaluation_time_ms: 300_000, + revision: 1, + }, + Limits::default(), + ) + .unwrap(), + ) + .unwrap() + .remove(0); + let mut batches = Vec::new(); + while let Some(batch) = stream.next().await { + batches.push(batch.unwrap()); + } + let mut actual: Vec<_> = batches + .iter() + .flat_map(|batch| batch.rows()) + .map(|row| { + let Value::Utf8(service) = &row[0] else { + panic!("missing service") + }; + let Value::Float64(value) = row[1] else { + panic!("missing sum") + }; + (service.to_string(), value) + }) + .collect(); + actual.sort_by(|a, b| a.0.cmp(&b.0)); + assert_eq!( + actual, + vec![ + ("api".into(), 5. * multiplier), + ( + "batch".into(), + 4. * multiplier + + if query.contains("WHERE") && multiplier == 1. { + 0. + } else { + multiplier + } + ) + ] + ); + } + } +} diff --git a/crates/integration-tests/tests/summary_maintenance_lifecycle_e2e.rs b/crates/integration-tests/tests/summary_maintenance_lifecycle_e2e.rs index 186dd04b..41e5f1b3 100644 --- a/crates/integration-tests/tests/summary_maintenance_lifecycle_e2e.rs +++ b/crates/integration-tests/tests/summary_maintenance_lifecycle_e2e.rs @@ -119,52 +119,7 @@ fn dashboard_workload() -> PlanningWorkload { #[test] fn promql_dashboard_materializes_continuous_summary_with_explained_rejections() { let workload = dashboard_workload(); - workload.validate().unwrap(); - - let lowered = lower_promql_workload(&workload, 0) - .expect("valid PromQL workload") - .into_iter() - .next() - .expect("one normalized workload entry"); - let root = Rc::new(lowered); - let strategies = asap_aware_mapping::default_strategies_with(&FullyCostedRuntime); - let space = search_workload_with(vec![("dashboard", Rc::clone(&root))], &strategies); - let target = Rc::clone(&space.roots[0].1); - let capabilities = SummaryMaintenanceLifecycleCapabilities { - supports_ephemeral: true, - supports_prepared: false, - supports_shared: false, - supports_continuously_maintained: true, - }; - - let selection = global_selection_with_summary_maintenance_lifecycles( - &space, - WorkloadDemand { - workload: &workload.query_workload, - data_workload: workload.data_workload.as_ref(), - entry_indices: &[1], - }, - NOW_MS, - Some(Horizon(100.0)), - capabilities, - &FullyCostedRuntime, - ) - .unwrap(); - let plan = assemble_selected_dag_with_summary_maintenance_lifecycles( - &selection, - &target, - WorkloadDemand::new_with_data( - &workload.query_workload, - workload.data_workload.as_ref().unwrap(), - &[1], - ), - NOW_MS, - Some(Horizon(100.0)), - capabilities, - &FullyCostedRuntime, - ) - .unwrap() - .expect("selected summary plan"); + let plan = selected_plan(&workload); assert!(!plan.selected_raw_recompute); assert_eq!(plan.expected_reads, Some(100.0)); @@ -231,3 +186,686 @@ fn promql_dashboard_materializes_continuous_summary_with_explained_rejections() "continuously_maintained" ); } + +fn selected_plan( + workload: &PlanningWorkload, +) -> asap_aware_mapping::SummaryMaintenanceLifecyclePlan { + selected_plan_with_model(workload, &FullyCostedRuntime) +} + +fn selected_plan_with_model( + workload: &PlanningWorkload, + model: &dyn CostModel, +) -> asap_aware_mapping::SummaryMaintenanceLifecyclePlan { + selected_plan_with_horizon(workload, model, Horizon(100.)) +} + +fn selected_plan_with_horizon( + workload: &PlanningWorkload, + model: &dyn CostModel, + horizon: Horizon, +) -> asap_aware_mapping::SummaryMaintenanceLifecyclePlan { + workload.validate().unwrap(); + + let lowered = lower_promql_workload(workload, 0) + .expect("valid PromQL workload") + .into_iter() + .next() + .expect("one normalized workload entry"); + let root = Rc::new(lowered); + let strategies = asap_aware_mapping::default_strategies_with(model); + let space = search_workload_with(vec![("dashboard", Rc::clone(&root))], &strategies); + let target = Rc::clone(&space.roots[0].1); + let capabilities = SummaryMaintenanceLifecycleCapabilities { + supports_ephemeral: true, + supports_prepared: false, + supports_shared: false, + supports_continuously_maintained: true, + }; + + let selection = global_selection_with_summary_maintenance_lifecycles( + &space, + WorkloadDemand { + workload: &workload.query_workload, + data_workload: workload.data_workload.as_ref(), + entry_indices: &[1], + }, + NOW_MS, + Some(horizon), + capabilities, + model, + ) + .unwrap(); + assemble_selected_dag_with_summary_maintenance_lifecycles( + &selection, + &target, + WorkloadDemand::new_with_data( + &workload.query_workload, + workload.data_workload.as_ref().unwrap(), + &[1], + ), + NOW_MS, + Some(horizon), + capabilities, + model, + ) + .unwrap() + .expect("selected summary plan") +} + +mod physical_common; + +/// A selected continuous lifecycle supplies a materialization boundary; its +/// maintenance and query DAGs execute the selected KLL computation in fresh runs. +#[test] +fn continuous_lifecycle_compiles_and_executes_spatial_kll() { + use asap_physical_operators::{ + physical_planner::{compile_candidate, InputContract}, + runtime::Scope, + values::{Batch, Value}, + }; + use asap_types::{ + post_asap::{compile_post_asap_dag, PostAsapOperatorPayload, SummaryFamilyType}, + pre_asap::DataType, + }; + use std::{collections::BTreeMap, sync::Arc}; + let mut workload = dashboard_workload(); + workload.query_workload.query_batch.as_mut().unwrap()[0].query = + Query("quantile(0.99, latency)".into()); + workload.query_workload.repeating_queries.as_mut().unwrap()[0].query = + Query("quantile(0.99, latency)".into()); + let selected = selected_plan(&workload); + assert_eq!( + selected.deployments[0] + .summary_maintenance_lifecycle_guarantee + .as_ref() + .unwrap() + .summary_maintenance_lifecycle, + SummaryMaintenanceLifecycle::ContinuouslyMaintained + ); + let dag = compile_post_asap_dag(&selected.root).unwrap(); + let build = dag + .nodes + .iter() + .find(|node| matches!(node.payload, PostAsapOperatorPayload::SummaryAgg { .. })) + .unwrap(); + let input = dag + .edges + .iter() + .find(|edge| edge.consumer == build.id) + .unwrap() + .producer; + let raw = dag.nodes.iter().find(|node| node.id == input).unwrap(); + let schema = Arc::new(raw.output_schema.clone()); + let candidate = compile_candidate( + &dag, + BTreeMap::from([(u64::from(input.0), InputContract::bounded(schema.clone()))]), + &[u64::from(dag.root.0)], + &[u64::from(build.id.0)], + ) + .unwrap(); + + // A continuous input without a finite pane boundary cannot implement this + // blocking builder. Retain lifecycle ownership in the candidate payload; + // only the legal bounded request candidate reaches workload pricing. + let mut unbounded = InputContract::bounded(schema.clone()); + unbounded.properties.boundedness = asap_physical_operators::plan::Boundedness::Unbounded; + let rejected = compile_candidate( + &dag, + BTreeMap::from([(u64::from(input.0), unbounded)]), + &[u64::from(dag.root.0)], + &[u64::from(build.id.0)], + ); + assert!(rejected.is_err()); + let request = compile_candidate( + &dag, + BTreeMap::from([(u64::from(input.0), InputContract::bounded(schema.clone()))]), + &[u64::from(dag.root.0)], + &[], + ) + .unwrap(); + let mut priced = 0; + let feedback = asap_physical_operators::physical_planner::select_candidate( + vec![ + rejected.map(|candidate| { + ( + SummaryMaintenanceLifecycle::ContinuouslyMaintained, + candidate, + ) + }), + Ok((SummaryMaintenanceLifecycle::Ephemeral, request)), + ], + |_| { + priced += 1; + Ok(Some( + asap_physical_operators::physical_planner::CandidateCost { + workload_scope: "dashboard".into(), + horizon_seconds: 100., + total_cost: 1000., + }, + )) + }, + ) + .unwrap(); + assert_eq!(priced, 1); + assert_eq!(feedback.candidate.0, SummaryMaintenanceLifecycle::Ephemeral); + for revision in [1, 2] { + let rows = (1..=100) + .map(|value| { + schema + .fields + .iter() + .map(|field| match field.dtype { + SummaryFamilyType::Plain(DataType::Float64) => { + Value::Float64(f64::from(value)) + } + SummaryFamilyType::Plain(DataType::Timestamp) => Value::Timestamp(300_000), + _ => panic!("unexpected field {field:?}"), + }) + .collect() + }) + .collect(); + let raw_batch = Batch::try_new(schema.clone(), rows).unwrap(); + let direct = physical_common::execute( + &feedback.candidate.1.query, + BTreeMap::from([(u64::from(input.0), raw_batch.clone())]), + Scope::Query { + evaluation_time_ms: 300_000, + revision, + }, + ); + let state = physical_common::execute( + candidate.precompute.as_ref().unwrap(), + BTreeMap::from([(u64::from(input.0), raw_batch)]), + Scope::Ingestion { + window_start_ms: 0, + window_end_ms: 300_000, + revision, + }, + ); + let result = physical_common::execute( + &candidate.query, + BTreeMap::from([(u64::from(build.id.0), state[0][0].clone())]), + Scope::Query { + evaluation_time_ms: 300_000, + revision, + }, + ); + let values: Vec<_> = result[0] + .iter() + .flat_map(|batch| batch.rows()) + .flat_map(|row| row.iter()) + .filter_map(|value| { + if let Value::Float64(value) = value { + Some(*value) + } else { + None + } + }) + .collect(); + let direct_values: Vec<_> = direct[0] + .iter() + .flat_map(|batch| batch.rows()) + .flat_map(|row| row.iter()) + .filter_map(|value| { + if let Value::Float64(value) = value { + Some(*value) + } else { + None + } + }) + .collect(); + assert_eq!( + values, direct_values, + "maintenance and request candidates preserve the same population" + ); + assert_eq!(values.len(), 1); + assert!( + (98. ..=100.).contains(&values[0]), + "p99 rank must reflect the supplied population" + ); + } +} + +struct SlidingPaneModel; +impl CostModel for SlidingPaneModel { + fn raw_query_recompute_total_cost( + &self, + target: &asap_types::pre_asap::QueryExpr, + reads: f64, + ) -> Option { + let _ = (target, reads); + Some(Cost(100_000.0)) + } + fn rank_candidates( + &self, + intent: &AggIntent, + candidates: &[asap_types::post_asap::SketchAlgorithm], + ) -> Vec { + FullyCostedRuntime.rank_candidates(intent, candidates) + } + fn summary_maintenance_lifecycle_cost_inputs( + &self, + summary: &SummaryNode, + ) -> SummaryMaintenanceLifecycleCostInputs { + let mut costs = FullyCostedRuntime.summary_maintenance_lifecycle_cost_inputs(summary); + // Controlled workload evidence makes repeated raw construction more + // expensive than retaining and updating the same temporal population. + costs.build_cost = Some(Cost(1000.)); + costs + } + fn summary_maintenance_capabilities( + &self, + summary: &SummaryNode, + ) -> SummaryMaintenanceCapabilities { + FullyCostedRuntime.summary_maintenance_capabilities(summary) + } + fn complete_summary_candidate_estimate( + &self, + _root: &SummaryNode, + _target: Option<&asap_types::pre_asap::QueryExpr>, + deployments: &[asap_aware_mapping::cost_model::CostedSummaryDeployment<'_>], + _horizon: Option, + _reads: Option, + _accuracy: &[AccuracyTarget], + ) -> Option { + Some(asap_aware_mapping::CompleteSummaryCandidateEstimate { + cost: Cost( + deployments + .iter() + .map(|deployment| deployment.selected_cost.0) + .sum(), + ), + physical_plan_id: Some("bounded-sliding-pane-evidence".into()), + window_frameworks: deployments + .iter() + .map(|deployment| { + (deployment.guarantee.summary_maintenance_lifecycle + == SummaryMaintenanceLifecycle::ContinuouslyMaintained) + .then_some(asap_types::post_asap::SummaryWindowFramework::Sliding) + }) + .collect(), + window_accuracy_guarantee: None, + }) + } +} + +/// Workload and optimizer-selected lifecycle generate both physical DAGs. +/// No computational operators or graph edges are constructed by this fixture. +#[test] +fn selected_temporal_lifecycle_compiles_panes_and_executes() { + use asap_physical_operators::{ + operators::Operator, + physical_planner::{ + compile_temporal_pane_candidate, InputContract, Source, TemporalEntityIdentity, + TemporalPaneMaintenance, + }, + runtime::{Limits, RunContext, Scope}, + summary_kernels::datasketches_kll::DatasketchesKLLAccumulator, + values::{Batch, Value}, + }; + use asap_types::{ + post_asap::{ + compile_post_asap_dag, plan_pane_phase, PostAsapOperatorPayload, SummaryFamilyType, + SummaryWindowFramework, + }, + pre_asap::DataType, + workload::TimestampMs, + }; + use futures::{executor::block_on, StreamExt}; + use std::{collections::BTreeMap, sync::Arc}; + for quantile in [0.5, 0.99] { + let mut workload = dashboard_workload(); + let query = Query(format!( + "quantile_over_time({quantile}, latency{{job=\"api\"}}[5m])" + )); + workload.query_workload.query_batch.as_mut().unwrap()[0].query = query.clone(); + workload.query_workload.repeating_queries.as_mut().unwrap()[0].query = query; + + workload + .data_workload + .as_mut() + .unwrap() + .data_ingestion_interval + .value = Some(DurationMs(60_000)); + workload.query_workload.repeating_queries.as_mut().unwrap()[0].demand = + RepeatedDemand::FixedIntervalAt { + interval: RepetitionInterval(60_000), + evaluation_phase: TimestampMs(300_000), + }; + let plan = selected_plan_with_horizon(&workload, &SlidingPaneModel, Horizon(1000.)); + assert!(!plan.selected_raw_recompute); + assert_eq!(plan.deployments.len(), 1); + let deployment = &plan.deployments[0]; + assert_eq!( + deployment.selected_window_framework, + Some(SummaryWindowFramework::Sliding) + ); + let dag = compile_post_asap_dag(&plan.root).unwrap(); + let build = dag + .nodes + .iter() + .find(|node| matches!(node.payload, PostAsapOperatorPayload::SummaryAgg { .. })) + .unwrap(); + assert_eq!(build.id, deployment.post_asap_node_id); + let raw = dag + .nodes + .iter() + .find(|node| matches!(node.payload, PostAsapOperatorPayload::Fallback { .. })) + .unwrap(); + let schema = Arc::new(raw.output_schema.clone()); + let width = workload + .data_workload + .as_ref() + .unwrap() + .data_ingestion_interval + .value + .unwrap() + .0; + let layout = plan_pane_phase( + &workload.query_workload.repeating_queries.as_ref().unwrap()[0].demand, + width, + ) + .unwrap(); + let maintenance = TemporalPaneMaintenance { + summary_node: u64::from(build.id.0), + lifecycle: deployment + .summary_maintenance_lifecycle_guarantee + .clone() + .unwrap(), + framework: deployment.selected_window_framework.clone().unwrap(), + layout, + // The memory source has exactly the declared label columns; a + // schemaless deployment must resolve all entity keys first. + entity_identity: TemporalEntityIdentity::Columns( + schema + .fields + .iter() + .enumerate() + .filter(|(_, field)| field.name == "job") + .map(|(index, _)| index) + .collect(), + ), + }; + let candidate = compile_temporal_pane_candidate( + &dag, + BTreeMap::from([(u64::from(raw.id.0), InputContract::bounded(schema.clone()))]), + &[u64::from(dag.root.0)], + &maintenance, + ) + .unwrap(); + assert_eq!(candidate.window_width_ms, 300_000); + assert_eq!(candidate.pane_inputs.len(), 5); + assert_ne!( + candidate.physical.precompute.as_ref().unwrap().roots()[0], + maintenance.summary_node, + "a one-minute pane is not the logical five-minute summary output" + ); + let source_opens = Arc::new(std::sync::atomic::AtomicUsize::new(0)); + let mut stored = Vec::new(); + for pane in 0..6 { + let rows = (0..20) + .flat_map(|sample| { + ["api", "batch"] + .into_iter() + .map(move |entity| (sample, entity)) + }) + .map(|(sample, entity)| { + schema + .fields + .iter() + .map(|field| match field.dtype { + SummaryFamilyType::Plain(DataType::Timestamp) => { + Value::Timestamp(pane * 60_000 + (sample + 1) * 3000) + } + SummaryFamilyType::Plain(DataType::Float64) => Value::Float64( + (pane * 20 + sample) as f64 + + if entity == "batch" { 100_000. } else { 0. }, + ), + SummaryFamilyType::Plain(DataType::Utf8) => Value::Utf8(entity.into()), + _ => panic!("unexpected raw field {field:?}"), + }) + .collect() + }) + .collect(); + let result = physical_common::execute( + candidate.physical.precompute.as_ref().unwrap(), + BTreeMap::from([( + u64::from(raw.id.0), + Batch::try_new(schema.clone(), rows).unwrap(), + )]), + Scope::Ingestion { + window_start_ms: pane * 60_000, + window_end_ms: (pane + 1) * 60_000, + revision: 1, + }, + ); + let batch = &result[0][0]; + stored.push(batch.clone()); + } + for offset in [0, 1] { + let inputs: BTreeMap<_, _> = candidate + .pane_inputs + .iter() + .enumerate() + .map(|(index, &id)| (id, stored[offset + index].clone())) + .collect(); + let scope = Scope::Query { + evaluation_time_ms: (5 + offset as i64) * 60_000, + revision: 1, + }; + let result = + physical_common::execute(&candidate.physical.query, inputs.clone(), scope.clone()); + let rows = result[0][0].rows(); + assert_eq!(rows.len(), 1); + assert!(rows[0] + .iter() + .any(|value| matches!(value, Value::Utf8(label) if label.as_ref() == "api"))); + let Value::Float64(value) = rows[0][1] else { + panic!("missing p99") + }; + assert!( + (value - ((if quantile == 0.5 { 50 } else { 99 }) + offset * 20) as f64).abs() + <= 1. + ); + assert!( + matches!(rows[0][0], Value::Timestamp(timestamp) if timestamp == (5 + offset as i64) * 60_000) + ); + let sources = |inputs: BTreeMap| -> BTreeMap> { + inputs + .into_iter() + .map(|(id, batch)| { + ( + id, + Box::new(TemporalCountingSource { + operator: Operator::source(batch.schema().clone(), vec![batch]) + .unwrap(), + opens: source_opens.clone(), + }) as Source<'static>, + ) + }) + .collect() + }; + let bound = candidate + .physical + .query + .instantiate(sources(inputs.clone())) + .unwrap(); + let population = block_on( + bound + .execute( + &[candidate.merged_state], + RunContext::new(scope.clone(), Limits::default()).unwrap(), + ) + .unwrap() + .remove(0) + .collect::>(), + ); + let Value::Summary { state, .. } = population[0].as_ref().unwrap().rows()[0] + .iter() + .find(|value| matches!(value, Value::Summary { .. })) + .unwrap() + else { + panic!("missing merged state") + }; + assert_eq!( + state + .as_any() + .downcast_ref::() + .unwrap() + .inner + .count(), + 100 + ); + let mut missing = inputs.clone(); + missing.remove(&candidate.pane_inputs[0]); + assert!(candidate + .physical + .query + .instantiate(sources(missing)) + .is_err()); + let mut duplicate = inputs.clone(); + duplicate.insert(candidate.pane_inputs[1], stored[offset].clone()); + let bad = candidate + .physical + .query + .instantiate(sources(duplicate)) + .unwrap(); + let errors = block_on( + bad.execute( + candidate.physical.query.roots(), + RunContext::new(scope.clone(), Limits::default()).unwrap(), + ) + .unwrap() + .remove(0) + .collect::>(), + ); + assert!( + errors.iter().any(Result::is_err), + "duplicate pane must not be merged twice" + ); + let mut duplicate_entity = inputs.clone(); + let pane = &stored[offset]; + duplicate_entity.insert( + candidate.pane_inputs[0], + Batch::try_new( + pane.schema().clone(), + vec![pane.rows()[0].clone(), pane.rows()[0].clone()], + ) + .unwrap(), + ); + let bad = candidate + .physical + .query + .instantiate(sources(duplicate_entity)) + .unwrap(); + let errors = block_on( + bad.execute( + candidate.physical.query.roots(), + RunContext::new(scope.clone(), Limits::default()).unwrap(), + ) + .unwrap() + .remove(0) + .collect::>(), + ); + assert!( + errors.iter().any(Result::is_err), + "duplicate snapshots within a pane must fail" + ); + let bound = candidate + .physical + .query + .instantiate(sources(inputs)) + .unwrap(); + source_opens.store(0, std::sync::atomic::Ordering::SeqCst); + assert!(bound + .execute( + candidate.physical.query.roots(), + RunContext::new( + Scope::Query { + evaluation_time_ms: 330_000, + revision: 1 + }, + Limits::default() + ) + .unwrap() + ) + .is_err()); + assert_eq!( + source_opens.load(std::sync::atomic::Ordering::SeqCst), + 0, + "invalid phase must fail before opening readers" + ); + } + // A selected framework cannot be silently replaced by physical planning. + let mut wrong_framework = maintenance.clone(); + wrong_framework.framework = SummaryWindowFramework::ExponentialHistogram; + let mut wrong_identity = maintenance.clone(); + wrong_identity.entity_identity = TemporalEntityIdentity::SingleEntity; + let mut unknown_phase = maintenance.clone(); + unknown_phase.layout.pane_origin_ms = None; + let mut partial_panes = maintenance.clone(); + partial_panes.layout.pane_width_ms = 90_000; + let mut wrong_lifecycle = maintenance.clone(); + wrong_lifecycle.lifecycle.summary_maintenance_lifecycle = + SummaryMaintenanceLifecycle::Ephemeral; + for unsupported in [ + wrong_framework, + wrong_identity, + unknown_phase, + partial_panes, + wrong_lifecycle, + ] { + assert!(compile_temporal_pane_candidate( + &dag, + BTreeMap::from([(u64::from(raw.id.0), InputContract::bounded(schema.clone()))]), + &[u64::from(dag.root.0)], + &unsupported + ) + .is_err()); + } + } +} + +struct TemporalCountingSource { + operator: asap_physical_operators::operators::Operator, + opens: std::sync::Arc, +} +impl + asap_physical_operators::plan::PhysicalOperator< + asap_physical_operators::values::Batch, + asap_physical_operators::values::Schema, + > for TemporalCountingSource +{ + fn name(&self) -> &str { + "TemporalCountingSource" + } + fn properties( + &self, + inputs: &[asap_physical_operators::plan::PlanProperties], + ) -> asap_physical_operators::plan::PlanProperties { + self.operator.properties(inputs) + } + fn input_schemas(&self) -> Vec { + self.operator.input_schemas() + } + fn output_schema(&self) -> asap_physical_operators::values::Schema { + self.operator.output_schema() + } + fn output_bytes(&self, batch: &asap_physical_operators::values::Batch) -> usize { + batch.bytes() + } + fn start<'a>( + &'a self, + inputs: Vec< + asap_physical_operators::runtime::Input<'a, asap_physical_operators::values::Batch>, + >, + context: asap_physical_operators::runtime::RunContext, + ) -> Result< + asap_physical_operators::runtime::OutputStream<'a, asap_physical_operators::values::Batch>, + asap_physical_operators::Error, + > { + self.opens.fetch_add(1, std::sync::atomic::Ordering::SeqCst); + self.operator.start(inputs, context) + } +} diff --git a/docs/design_docs/physical-planning-and-deployment.md b/docs/design_docs/physical-planning-and-deployment.md new file mode 100644 index 00000000..72313a9d --- /dev/null +++ b/docs/design_docs/physical-planning-and-deployment.md @@ -0,0 +1,512 @@ +# Physical Planning, Summary Maintenance, and Deployment + +## 1. Architecture + +A Post-ASAP computation is progressively realized through four layers: + +```mermaid +flowchart LR + L["Logical Post-ASAP DAG
What computation?"] + M["Summary Maintenance Lifecycle
How is state maintained?"] + P["Physical DAG(s)
How is it executed?"] + D["Deployment Plan / DAG
How is it instantiated?"] + + L -->|"Summary Maintenance
Candidate Generation"| M + M -->|"Physical Plan
Compiler"| P + P -->|"Deployment Plan
Compiler"| D +``` + +| Layer | Defines | +| --- | --- | +| **Logical Post-ASAP DAG** | Computation semantics | +| **Summary Maintenance Lifecycle** | Build, retention, reuse, and window strategy | +| **Physical DAG(s)** | Supported physical candidates, executable operators and typed input boundaries | +| **Deployment Plan / DAG** | Selected candidate, concrete data/state bindings and operational lifecycle | + +ASAPPlanner owns the first three layers and the shared physical operator +implementation library. Deployment systems such as ASAPQuery and asap-fusion +own deployment compilation and operation. The lifecycle is a planning contract +associated with the logical DAG, not a separate computation IR. + +The Logical Post-ASAP DAG is preceded by the Pre-ASAP DAG (`QueryExpr`), the +language-independent query semantics before summary selection. Both are +logical. Planning builds Post-ASAP `SummaryNode` trees; `compile_post_asap_dag` +exports the selected tree as a `PostAsapDag`, which is the Physical Plan +Compiler's input. Its per-node execution phase (ingestion or query time) is an +initial placement: compilation places ingestion-time nodes in the precompute DAG, +while frontier enumeration proposes alternative materialization splits. Which +layer owns placement is an open design question, deferred to a later change. + +### Candidate generation and deployment selection + +Planner exposes the supported, semantically legal **physical plan candidates**. +It does not discard a computation family or materialization placement merely +because a deployment-independent cost estimate prefers another candidate. +Logical candidates are an internal search stage, not the deployment handoff. + +```text +Query semantics + accuracy and lifecycle requirements + ↓ Planner +Supported Physical DAG candidates + typed inputs/outputs + requirements + ↓ backend +Binding feasibility + runtime statistics + resource limits + ERP + ↓ backend deployment compiler +Selected PrecomputePlan + QueryPlan + StoredOutputReferences +``` + +Planner owns operators, dependencies, sharing, and each candidate's +materialization frontier. The backend rejects candidates it cannot realize and +prices feasible candidates over a comparable workload and time horizon. It binds +the selected candidate; it does not lower the logical computation again, exchange +operators, or move an operator across the selected frontier. A missing quote is +not a zero-cost implementation. ERP evidence cannot authorize an illegal rewrite. + +The candidate inventory must identify its supported search scope and budget. +If a configured exhaustive enumeration exceeds its budget, planning fails +explicitly instead of selecting from an undisclosed partial inventory. Reports +separate unsupported compilation, deployment infeasibility, missing evidence, +and a feasible candidate that loses on cost. Absence is not a cost comparison. + +For `sum by(job)(rate(m[1m]))`, Rate remains per series before grouped Sum. +When lifecycle requirements permit it, a candidate may finalize Rate and Sum +within a bounded precompute run and persist the grouped value. Another may leave +those operators in the query DAG. Storing a value requires its exact evaluation +window, revision, readiness and serving cadence to match the query contract. + +For instant-vector TopK, CMS/CountSketch with a candidate heap requires explicit +series identity and a supported latest-value input protocol. Appending historical +sample values does not preserve instant-vector semantics. Replacement, rank +decrease, expiry, grouping and the required approximation guarantee must be +validated before admitting that physical candidate. + +This is the target ownership contract. A backend path that still reconstructs +operators from logical candidates has not completed this integration. + +### Input semantics and summary semantics + +`source`, `filter`, `grouping` and `window` describe input-data semantics: +where records originate, which records qualify, how they are grouped and which +time interval applies. They are not a complete description of arbitrary summary +computation. In particular, the same four fields can summarize different value +expressions or produce different states. + +| Concern | Required semantic information | +| --- | --- | +| Input computation | Source identities and schemas, filters, joins/transforms and their order, or a reference to the canonical input sub-DAG | +| Values and grouping | Value expressions, item identities and weights where applicable, group keys and types, and operation-defined null/duplicate handling | +| Time | Time column and interpretation, interval bounds, evaluation alignment, and distinction between query range and maintained panes | +| Summary computation | Exact operation or sketch family, algorithm and parameters, and supported build/merge behavior | +| Output | State versus finalized value, output schema/type, and readout parameters when part of the output computation | + +For example, KLL over `latency_seconds` and KLL over `log(latency_seconds)` differ +even with identical source, filter, grouping and window. Likewise, weighted +frequency state needs both item and weight expressions. More complex inputs +must retain their computation DAG; four descriptive fields cannot replace it. + +The canonical selected computation is authoritative. These categories describe +what must be preserved, not a new flat IR or a second expression language. +Operator-defined behavior should be referenced through its canonical contract, +not independently configured in deployment metadata. Unsupported or unresolved +semantics cannot be treated as compatible. + +Logical planning defines the semantics; physical compilation realizes them as +operators and typed boundaries. Deployment binds concrete readers and state +records that satisfy those requirements. A stored summary definition records or +references the relevant semantics for compatibility checks. Matching a definition +alone does not establish actual window coverage, revision compatibility or +readiness; those require runtime checks. Physical location, encoding, scheduling +and retention are separate execution/deployment contracts. + +Persisted semantic identity, its wire format and any tenant or dataset binding +belong to the deployment. Planner provides the typed `PostAsapDag` that a +deployment canonicalizes; it does not define a stored-definition format. + +### Running example + +Suppose p50 and p99 are requested over the same latency samples in a five-minute window, +and one Planner candidate uses KLL with `k=200`. Assume query windows align with one-minute +pane boundaries and that the selected parameters satisfy the required guarantees. +Operator names below are illustrative; the example defines the design, not a +claim that the entire deployment integration is implemented. + +The data source identifies where samples come from. Filters, grouping and the +window determine which samples enter each summary. Here `pane_duration: 1m` +means each stored pane covers one minute; the query range is five minutes. +Neither duration specifies how often maintenance runs or how long state is kept. + +The example evolves through the architecture as follows: + +```text +1. Logical Post-ASAP DAG + +raw latency + ↓ +KLLBuild(k=200) + ↓ +KLLMerge + ┌─┴─────┐ + ↓ ↓ + p50 p99 + + │ + │ Summary Maintenance Candidate Generation + ▼ + +2. Summary Maintenance Lifecycle + +KLLBuild(k=200) + strategy = continuously maintain + window = 1-minute panes + reuse = p50 + p99 + query = merge panes covering requested aligned 5 minutes + + │ + │ Physical Plan Compiler + ▼ + +3. Physical DAGs + +Precompute DAG: +RawInput + ↓ +NativeKllBuild(k=200) + ↓ +KllStateOutput + +Query DAG: +InputSlot[5 panes] + ↓ +NativeKllMerge(k=200) + ┌─┴────────┐ + ↓ ↓ + NativeP50 NativeP99 + + │ + │ Deployment Plan Compiler + ▼ + +4. Deployment Plan / DAG + +Precompute: +OTLP latency source + ↓ +run KLL build over each complete 1-minute input pane + ↓ +store as latency-kll-1m/ + +Query: +resolve five latency-kll-1m states + ↓ +execute query Physical DAG + ↓ +return p50 / p99 +``` + +Each stage adds a different class of decision while preserving the preceding +contracts. Here, continuous maintenance means recurring production of pane state; +the bounded build DAG does not itself implement an unbounded streaming window. + +## 2. Logical Post-ASAP DAG → Summary Maintenance Lifecycle + +The **Logical Post-ASAP DAG** defines computation semantics: + +```text +Scan(latency) + ↓ +KLLBuild(k=200) + ↓ +KLLMerge + ┌─┴────────────┐ + ↓ ↓ +Quantile(.5) Quantile(.99) +``` + +It establishes that KLL with `k=200` is used and that the merge is shared by the +two readouts. It does not determine when KLL states are built or retained. + +**Summary Maintenance Candidate Generation** enumerates legal lifecycle choices +using workload demand, window/freshness requirements and supported physical +implementations. Backend selection uses runtime feasibility and cost after +physical compilation. The following example follows one candidate. + +For the running example, assume it selects: + +```text +producer: KLLBuild(k=200) + +strategy: + continuously maintain + +window realization: + 1-minute panes + +query requirement: + combine panes covering the requested aligned 5-minute range + +reuse: + one merged state serves p50 and p99 +``` + +This produces the **Summary Maintenance Lifecycle**. + +The lifecycle specifies how the selected logical summary should be maintained, +but not its concrete operator implementation or storage location. + +Physical feasibility may feed back into selection. For example, if the required +pane-based maintenance cannot be implemented, this lifecycle candidate cannot be +selected. One-minute panes alone also cannot cover an arbitrarily phased query +window; that requires supported boundary handling or a different candidate. + +## 3. Summary Maintenance Lifecycle → Physical DAG + +The **Physical Plan Compiler** consumes both computation semantics and maintenance +requirements: + +```text +Logical Post-ASAP DAG (PostAsapDag) ++ Summary Maintenance Lifecycle ++ physical capabilities + ↓ +Physical Plan Compiler + ↓ +Physical DAG(s) +``` + +For the running example, the lifecycle creates two execution boundaries. + +These two halves are named as `PhysicalCandidate` names them, `precompute` +and `query`. *Maintenance* stays the lifecycle's word (section 2): it covers +how state is built, retained, reused and scheduled. A precompute DAG is the +physical object that a maintenance lifecycle compiles to, so reusing +*maintenance* for it collapses two layers that the crates keep apart: +`asap-aware-mapping::summary_maintenance_*` owns the lifecycle, and +`asap-physical-operators::physical_planner` owns the DAGs. + +### Precompute Physical DAG + +```text +RawInputSlot( + window = 1m, + bounded = true +) + ↓ +NativeKllBuild(k=200) + ↓ +KllStateOutput(k=200) +``` + +This DAG implements construction of each maintained one-minute pane. Its input +contract requires all input samples matching the source, filters and group within that pane; the deployment supplies that +bounded input from its source integration. + +### Query Physical DAG + +```text +InputSlot( + k = 200, + coverage = requested aligned 5m +) + ↓ +NativeKllMerge(k=200) + ┌─┴──────────────────┐ + ↓ ↓ +NativeQuantile(.50) NativeQuantile(.99) +``` + +The Physical Plan Compiler chooses `NativeKllBuild`, `NativeKllMerge`, and the +physical quantile implementations, validates state compatibility, and preserves +the shared merge. It also resolves expressions, schemas, ordered dependencies +and execution properties. + +The resulting Physical DAGs know that compatible KLL states are required, but +do not know where those states are stored. + +For example: + +```text +InputSlot +``` + +is physical, while: + +```text +s3://.../latency-kll/12:01 +``` + +is deployment-specific. Placement and scheduling also remain outside the Physical +DAG. If the required behavior cannot be realized, physical compilation fails. + +### Physical candidates include precompute computation + +Materialization frontiers are Planner decisions. A candidate records both the +precompute Physical DAG and the query Physical DAG, with typed outputs connecting +them. The deployment compiler binds those outputs; it does not move operators. + +For `sum by(job)(rate(m[1m]))`, legal physical candidates can include: + +```text +Candidate A: + precompute: compatible per-series counter states → per-series Rate + materialized output: per-series rate values for window/evaluation/revision + query: stored per-series rate values → grouped Sum + +Candidate B: + precompute: compatible per-series counter states → per-series Rate → grouped Sum + materialized output: grouped values for window/evaluation/revision + query: stored grouped values → result +``` + +Both preserve reset-aware Rate before Sum. Summing raw counters before Rate is +not equivalent. The counter-state build may be another precompute DAG; typed +state inputs do not imply that a deployment can construct or bind those states. + +The shared library exposes `physical_planner::compile_candidates(...)` to lower +explicit frontier candidates to `PhysicalCandidate { precompute, query, +materialized_outputs }`. `select_candidate(...)` accepts deployment feasibility +and scoped complete-workload costs and chooses the lowest-cost feasible +candidate. Costs must describe the same workload and planning horizon; missing +feasibility is rejected before pricing. The optimizer supplies candidate +frontiers and cost evidence, including updates, retention, recurrence and sharing. +`enumerate_frontiers` constructs bounded, reachable antichain frontiers above explicit input boundaries, including query-only and fully precomputed results. It fails explicitly when the candidate budget is exceeded. Maintenance selection must still reject frontiers that violate window, freshness, or reuse requirements; deployment feasibility is checked before pricing. + +Physical compilation opens no readers. Bounded precompute outputs become typed +query inputs. Their source, filters, grouping, build window, evaluation time, readiness and +revision contracts must accompany the selected lifecycle and be checked during +deployment binding. Type compatibility alone does not establish reuse legality. + +The Planner integration test executes both candidates through the shared runtime +and reverses the selected frontier with two controlled cost fixtures. It also +rejects shadowed/duplicate boundaries and incomparable planning horizons. This +establishes Planner capability; it does not establish that ASAPQuery currently +supports persisting every scalar/result-output frontier. + +## 4. Physical DAG → Deployment Plan / DAG + +The **Deployment Plan Compiler** binds the Physical DAGs to the concrete deployment: + +```text +Physical DAGs ++ Summary Maintenance Lifecycle ++ deployment catalog/state ++ sources/materializations ++ operational policy + ↓ +Deployment Plan Compiler + ↓ +Deployment Plan / DAG +``` + +For the precompute DAG, it may produce: + +```text +Source: + RawInputSlot + → complete bounded panes from the OTLP latency source + +Schedule: + each 1-minute pane, once its completion requirements are met + +Execution: + RawInput → NativeKllBuild(k=200) + +Output: + KllStateOutput + → latency-kll-1m/ +``` + +For a query over `(12:00, 12:05]`, its input-binding rule resolves: + +```text +InputSlot[5 panes] + ├── latency-kll-1m/(12:00,12:01] + ├── latency-kll-1m/(12:01,12:02] + ├── latency-kll-1m/(12:02,12:03] + ├── latency-kll-1m/(12:03,12:04] + └── latency-kll-1m/(12:04,12:05] + ↓ + Query Physical DAG + ↓ + p50, p99 +``` + +The Deployment Plan Compiler establishes bindings and checks that their contracts +satisfy the physical inputs and selected lifecycle, including KLL parameters, +source, filters, grouping, window coverage and revision scope. The deployment engine +resolves request-specific states and checks their actual coverage, revisions and +readiness at execution time. A compiled plan cannot establish future readiness. + +The compiler does not replace `NativeKllMerge`, choose another sketch, or decide +to maintain different windows. Such changes require replanning. A Deployment +Plan / DAG is an operational instantiation, not another computation IR. + +## 5. Responsibility Boundary + +The complete example makes the ownership boundary explicit: + +| Stage | KLL example decision | +| --- | --- | +| **Logical Post-ASAP DAG** | Use `KLL(k=200)` with shared merge for p50/p99 | +| **Summary Maintenance Candidate Generation** | Maintain 1-minute panes and reuse them for aligned five-minute queries | +| **Summary Maintenance Lifecycle** | Record pane/window/freshness/reuse requirements | +| **Physical Plan Compiler** | Lower to native KLL build, merge, and readout operators | +| **Physical DAG** | Define precompute and query DAGs with typed input/output boundaries | +| **Deployment Plan Compiler** | Bind raw input and KLL state slots to concrete sources/materializations | +| **Deployment Plan / DAG** | Specify maintenance schedules, stored-pane resolution and query execution | + +```text +Logical: + "Use KLL for p50/p99." + +Lifecycle: + "Maintain reusable 1-minute KLL panes." + +Physical: + "Execute NativeKllBuild and + NativeKllMerge → {p50, p99}." + +Deployment: + "Read OTLP here, store panes here, + and bind these five panes for this aligned query." +``` + +The deployment engine executes the bound Physical DAGs through ASAPPlanner's +shared physical operator implementation library, `asap-physical-operators`, and +its DAG runtime. The merge executes once per run for both consumers. Execution +does not introduce additional planning decisions. + +Each maintained pane contributes its input samples once. A replacement snapshot +replaces that pane's state; query merging must not count both the old and new +snapshots as separate inputs. + +## 6. Executable acceptance coverage + +The tests cover optimizer-selected lifecycle execution and automatic temporal +pane compilation, alongside independent operator/runtime fixtures: + +| Test | Contract exercised | +| --- | --- | +| `summary_maintenance_lifecycle_e2e::selected_temporal_lifecycle_compiles_panes_and_executes` | PromQL p50/p99 workloads → selected continuous lifecycle and Sliding framework → automatically generated precompute/query DAGs → real codec round-trip → adjacent aligned windows; checks filters, entity identity, sample counts, missing/duplicate panes and phase rejection before opening readers | +| `summary_maintenance_lifecycle_e2e::continuous_lifecycle_compiles_and_executes_spatial_kll` | PromQL workload → selected continuous lifecycle → logical DAG → compiled precompute/query candidate → results in independent revisions; an unbounded candidate fails before pricing, and a bounded request candidate summarizes the same input samples | +| `kll_pane_execution::five_panes_roundtrip_and_shared_merge_runs_once` | Explicit one-minute precompute DAGs → real MessagePack state bytes → five required query inputs → shared native merge → p50/p99; counts every sample once, checks adjacent aligned windows and instruments one merge start per run | +| `kll_pane_execution::restored_panes_reject_corruption_parameters_schema_and_missing_binding` | Corrupt bytes, parameter relabelling, incompatible schemas and absent bindings fail explicitly | +| `precompute_candidates::grouped_rate_can_be_materialized_before_or_after_grouped_sum` | Cost changes select different legal precompute frontiers; both selected candidates execute with the same reset-sensitive result; uncompilable candidates are not priced | +| `sql_to_physical::sql_filter_grouped_sum_executes_and_rebinds` | SQL text → candidate search → physical compilation → shared Scan predicates and grouped summary execution; NULL samples are ignored and fresh bindings produce new results | + +`physical_planner::compile_temporal_pane_candidate` consumes the logical DAG, +selected lifecycle/framework and a generic pane/entity input contract. It +generates pane construction, scan predicates, ordered state slots, a shared +merge, quantile readouts and run-scoped timestamps. A physical pane output has +its own identity: one minute of state cannot masquerade as the logical +five-minute summary. The returned candidate retains the maintenance contract. + +This initial realization supports bounded, complete KLL panes with known phase +and resolved entity identity, for Sliding windows or a single Tumbling window. +Source capability evidence must declare all entity keys or isolate one entity; +usage-derived PromQL columns alone cannot establish that identity. Partial edge +panes, exponential histograms and cross-run delta accumulation require further +physical candidates and are rejected by this entry point. + +Physical execution checks pane timestamps and duplicate entity states. Concrete +stored identity, revisions, readiness and complete coverage of required input samples remain +deployment responsibilities. Real storage and HTTP execution belong to +deployment-repository E2E tests. diff --git a/docs/develop_docs/native-promql-inputs.md b/docs/develop_docs/native-promql-inputs.md new file mode 100644 index 00000000..aa1a52b6 --- /dev/null +++ b/docs/develop_docs/native-promql-inputs.md @@ -0,0 +1,55 @@ +# Native PromQL source rows + +Audience: source-adapter and physical-executor developers. + +A PromQL query only names some labels. Those columns cannot establish series +identity for Rate or TopK: two series with the same `job` may have different +unreferenced instance labels. + +`physical_planner::promql_rows::with_series_identity` resolves supported unary +PromQL computations to a bounded row representation before candidate search. +It appends `$promql_series_identity`, a non-null UTF-8 column containing the +canonical JSON encoding of the full label map. The name cannot collide with a +legal PromQL label. The resulting schema is closed over physical columns; the +label map remains dynamic and is not restricted to labels named in the query. +This realization rejects unsupported label rewriting, implicit vector matching, +and `without` operations rather than dropping hidden labels. + +Source adapters construct batches with `series_row`. Named label columns are +projections of the same complete identity; absent named labels project to empty +strings. `decode_series_identity` restores all labels on result conversion and +rejects noncanonical encodings. A query adapter must still apply the selected +operator's metric-name/result-label rules. Source selection, complete window +coverage and revision admission remain deployment responsibilities. + +Planner's maintained-population candidate recognizes this explicit identity +representation. Its TopK readout compiles automatically to `CurrentSeries`, +`Sort`, and `Limit`; deployment supplies the raw boundary or an already maintained +population boundary. Compilation does not open either source. + +The native `CurrentSeries` operator selects the latest sample per complete +identity in `(evaluation_time - lookback, evaluation_time]`. It removes stale +markers after selecting the latest sample, so an older value cannot reappear. +It rejects conflicting values at one series timestamp and emits the evaluation +timestamp. Each run builds a new snapshot; decreased values and expired series +cannot retain earlier heap weights. It reserves workspace and observes the +run's cancellation and byte budget. Precompute scopes must match the declared +lookback before any input is polled. + +CMS/CountSketch heap operators can consume this snapshot. CMS still requires +nonnegative weights; legal approximate TopK admission still requires the +Planner's accuracy/membership evidence. Executing a heap does not establish +that its result satisfies a query's accuracy requirements. + +Tests cover open-label Rate → CMS/CountSketch heaps, hidden-label round trips, +reset and zero-rate cases, snapshot replacement/decrease/expiry/staleness, +serialized physical recovery, and resource rejection. These are shared-library +tests, not proof of Backend candidate selection or durable deployment execution. + +Spatial heap candidates use the same complete series identity. Planner's +`current_series_topk_candidates` explores a CountSketch-with-heap realization +of canonical Sort/Limit under an explicit accuracy target. The physical graph +selects the latest eligible samples before building a fresh heap. A maintained +population boundary can supply that snapshot directly. Arbitrary signed metric +values do not authorize CMS; counter Rate's non-negative proof is separate. +These candidates still require membership/score evidence for deployment admission.