Skip to content

feat(planner): recognize normalized SQL entropy in nats - #564

Draft
zzylol wants to merge 1 commit into
stack/509-18-exact-cardinalityfrom
stack/509-19-sql-frequency-entropy
Draft

zzylol wants to merge 1 commit into
stack/509-18-exact-cardinalityfrom
stack/509-19-sql-frequency-entropy

Conversation

@zzylol

@zzylol zzylol commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Problem: the #509 Example 2 entropy query (Q2) has no Entropy intent, so Pass 1 cannot offer summary candidates for it

#509 §Pass 1: Local candidate generation lists Entropy(x) with three local candidates: "Exact entropy, a specialized entropy summary, UnivMon". #509 §Example 2 ("One summary for several computations") then writes Q2 in SQL and says:

The SQL frontend does not yet recognize the Q2 and Q3 forms as Entropy and L2. TODO: add this recognition to the per-language frontends.

Q3 (L2) was recognized earlier in this stack. Q2 was not. Q2 is:

SELECT -SUM(p * LN(p))
FROM (
  SELECT COUNT(*) * 1.0 / SUM(COUNT(*)) OVER () AS p
  FROM flows
  WHERE ts >= now() - INTERVAL '1 minute'
  GROUP BY src_ip
);

Before this PR, the lowered graph is only relational. Nothing in it says "entropy":

Project(-col0)
└── Aggregate(SUM(p * LN(p)))                    no group keys
    └── Project(p = COUNT * 1.0 / total)
        └── SQLWindowFunc(SUM(COUNT) OVER ())    appends `total`
            └── Aggregate(COUNT(*) GROUP BY src_ip)
                └── Scan(flows) …

So Pass 1 has no FrequencyEntropy intent to attach the entropy-summary or UnivMon candidates to, and the Pass 2 summary-capability rule cannot put Q2 on a shared UnivMon with Q1 and Q3.

There are also two semantic traps in a naive rewrite:

  1. Units. AggIntent::FrequencyEntropy is "Shannon entropy in bits". The SQL uses LN, so the answer is in nats. A rewrite that returns the intent directly is off by a factor of ln 2.
  2. Empty input. SQL SUM over zero groups returns NULL. The L2 rewrite added earlier in this stack restored this with CASE WHEN l2 = 0.0 THEN NULL. That is fine while the statistic is exact. It is wrong once the statistic is approximate: an estimate may be 0.0 for nonempty input, and then the query would return NULL instead of a number. Entropy has the same problem worse, because exact entropy is 0 for any single-identity population.

Scope covered here. Recognition of the Q2 natural-log form as FrequencyEntropy (the "TODO" above, for entropy), with SQL units, NULL, and negative-zero behavior kept. The same exact population guard is applied to the existing L2 rewrite.

Left out. Native execution of the original Q2 graph (the SQL window and LN) is the next PR, #565. UnivMon accuracy for entropy is not certified. Pass 2 sharing partitions (Example 2's 37 candidates) are not enumerated.

Proposed method

The rule is a Pass 1 logical rewrite. It runs inside the existing SemanticEquivalentRewriteStrategy in asap-aware-mapping, on the common IR after the SQL frontend. It adds one alternative and keeps the original exact graph as another candidate.

Matching, top-down (all in frequency_rewrite.rs):

  1. Root is a Project with one item. After following projections (expand), the item is Negative(SQL) of column 0 (casts to Float64 are ignored).
  2. Below it is an ungrouped Aggregate with exactly one measure Sum { col }, no measure filters, no HAVING.
  3. Following projections from that column gives a product x * LN(x) (either order, SQL semantics). probability_term checks that the LN argument is structurally equal to the other factor.
  4. x is unit_count / Column(2) with SQL division and Float64 type. unit_count_term accepts Column(1) or Column(1) * 1.0 (either order). Any other scale, such as * 2.0, is refused.
  5. Column(2) comes from an SQLWindowFunc with func: Sum, args == [Column(1)], empty partition_by, empty order_by, and a frame of UNBOUNDED PRECEDING to UNBOUNDED FOLLOWING.
  6. The window's child is a grouped unit count (grouped_unit_count, shared with L2): one group key, one Count measure, no measure filter, no HAVING, and the key is a non-nullable Bool, Int64 or Utf8 field.

If every check passes, the rewrite builds (sql_frequency_result):

Project(alias = original name,
        CASE WHEN population_count = 0 THEN CAST(NULL AS Float64)
             ELSE -(0.0 - frequency_entropy * LN_2) END)
└── Join(Cross, true)
    ├── Aggregate(FrequencyEntropy { col: key, accuracy })   → frequency_entropy
    └── Aggregate(Count { accuracy: Exact })                 → population_count
        (both read the grouped count's input, so WHERE filters are kept)
  • * LN_2 converts bits to nats.
  • -(0.0 - x) equals x, except that x = 0 gives -0.0. SQL's -SUM(1 * LN(1)) is also -0.0, so a single-identity population keeps the same bit pattern.
  • population_count is always exact. The approximate statistic never decides whether the result is NULL.
  • accuracy is copied from the original COUNT(*) measure, which carries the query's target.
  • The rewrite is returned only if its output schema equals the original root schema.

The L2 rewrite now uses the same grouped_unit_count and sql_frequency_result. Its output changes from CASE WHEN l2 = 0.0 … to the exact-count guard. Its value expression is Column(0) (no unit change).

Key code interfaces

All new functions are crate-private. The public entry point is unchanged: SemanticEquivalentRewriteStrategy (in crates/asap-aware-mapping/src/rewrite.rs) now also calls the entropy rule in matches and replacements. In replacements the entropy rule is tried first.

crates/asap-aware-mapping/src/frequency_rewrite.rs:

pub(super) fn frequency_entropy_rewrite(root: &Rc<OperatorNode>) -> Option<Rc<OperatorNode>>;

// Shared by L2 and entropy.
fn grouped_unit_count(
    node: &Rc<OperatorNode>,
) -> Option<(usize, AccuracyTarget, Rc<OperatorNode>)>;

fn sql_frequency_result(
    root: &Rc<OperatorNode>,
    input: Rc<OperatorNode>,
    measure: AggIntent,
    name: &str,
    value: ScalarExpr,
) -> Option<Rc<OperatorNode>>;

fn probability_term(product: &ScalarExpr) -> Option<&ScalarExpr>;
fn unit_count_term(expr: &ScalarExpr) -> bool;

The replacement it emits (rewrite.rs):

ReplacementSubDAG {
    strategy: "SemanticEquivalentRewriteStrategy",
    replacement: Replacement::SubDAG(rewritten),
    provenance: ReplacementProvenance::LogicalRewrite,
    rationale: "recognize SQL natural-log entropy with explicit bits-to-nats conversion and exact empty-population guard".into(),
}

substitute (projection lineage) now also passes through ScalarExpr::Literal and ScalarExpr::Negative.

Usage (from crates/frontend-sql/tests/frequency_entropy.rs):

let root = lower_sql("SELECT -SUM(p * LN(p)) AS entropy FROM (SELECT COUNT(*) * 1.0 / SUM(COUNT(*)) OVER () AS p FROM flows GROUP BY src_ip) f", &catalog(false), AccuracyTarget::Exact).await?;
let replacements = SemanticEquivalentRewriteStrategy.replacements(&TargetSubDAG::new(&root));

Fields

frequency_entropy_rewrite

Item Type Meaning
root &Rc<OperatorNode> Candidate target root. Must be the outer Project of the Q2 shape.
return Option<Rc<OperatorNode>> The guarded entropy graph, or None if any check in "Proposed method" fails. None means "not this rule", not an error.

grouped_unit_count

Item Type Meaning
node &Rc<OperatorNode> Must be Aggregate { reduction: Reduce(keys), measures: [Count], having: None }.
return .0 usize Index of the single group key in the aggregate's input. Becomes col of the frequency intent.
return .1 AccuracyTarget The Count measure's accuracy. Becomes the statistic's accuracy.
return .2 Rc<OperatorNode> The aggregate's input (rows before grouping). Both new aggregates read it.

Refused when: keys.is_without(), any measure filter, key field nullable, or key type not Bool/Int64/Utf8. A nullable key is refused because GROUP BY makes a real NULL group that the frequency intent would skip.

sql_frequency_result

Item Type Meaning
root &Rc<OperatorNode> Original root. Supplies the output alias and qualifier, and the schema the result must equal.
input Rc<OperatorNode> Rows to count (from grouped_unit_count).
measure AggIntent FrequencyEntropy { col, accuracy } or FrequencyL2 { col, accuracy }.
name &str Output name of the statistic column: "frequency_entropy" or "frequency_l2".
value ScalarExpr The ELSE branch over the joined row (Column(0) = statistic, Column(1) = population_count). Entropy passes the nats/negative-zero expression; L2 passes Column(0).
return Option<Rc<OperatorNode>> None if node construction fails or the schema differs from root.schema.

probability_term / unit_count_term

Item Meaning
probability_term(product) For a * LN(b) or LN(b) * a with SQL semantics and a == b, returns a. Function name match is case-insensitive.
unit_count_term(expr) True for Column(1) or Column(1) * 1.0 / 1.0 * Column(1), ignoring Float64 casts. Column 1 is the COUNT output of the grouped aggregate.

AggIntent::FrequencyEntropy (existing, not changed): col: Option<C> is the identity column (a column index here); accuracy: AccuracyTarget is the target; the value is in bits.

Examples

End-to-end (crates/integration-tests/tests/sql_frequency_entropy.rs). Table flows(src_ip Utf8, keep Bool). Query:

SELECT -SUM(p*LN(p)) AS entropy
FROM (SELECT COUNT(*)*1.0/SUM(COUNT(*)) OVER () AS p FROM flows WHERE keep GROUP BY src_ip) f

Each case also has one ("discard", false) row that the WHERE keep must drop. The rewritten graph is executed natively:

Kept src_ip values Expected result
none NULL (population_count = 0)
a, a -0.0 (bit pattern checked)
a, a, b, b ln 2
a, a, a, b -0.75·ln 0.75 − 0.25·ln 0.25

Accepted vs refused shapes (crates/frontend-sql/tests/frequency_entropy.rs, declines_non_equivalent_entropy_shapes):

Shape Result Reason
COUNT(*)*1.0 / SUM(COUNT(*)) OVER (), LN, non-null src_ip rewrite Q2 form
COUNT(*)*2.0 / … refused probabilities scaled
-SUM(p*LOG2(p)) refused different log base
… OVER (PARTITION BY src_ip) refused partial total
… OVER (ORDER BY src_ip) refused running total
… GROUP BY src_ip HAVING COUNT(*) > 1 refused population changed
nullable src_ip refused NULL group

Other tests:

  • recognizes_sql_entropy_in_nats: rewritten schema equals the original; original has no entropy intent.
  • entropy_search_preserves_relational_alternative: search_workload_with_targets with default_strategies() returns candidates both with and without FrequencyEntropy.
  • entropy_accuracy_and_population_guard_are_separate: with EpsilonDelta { epsilon: 0.05, delta: 0.01 } (Q2's target), the entropy aggregate gets that target and the count aggregate gets Exact.
  • frequency_l2.rs, frequency_empty_input_guard_uses_an_exact_count: L2 with EpsilonDelta { 0.01, 0.01 } now has an exact Count on the right side of the join.

Out of scope

  • Native execution of the original Q2 graph (SUM(…) OVER (), LN): feat(runtime): execute exact SQL entropy fallback #565.
  • Other log bases, scaled probabilities, partitioned/ordered/finite windows, filtered measures, HAVING, nullable or other key types.
  • UnivMon accuracy for entropy; Pass 2 sharing partitions and resizing to the strictest consumer.
  • docs/develop_docs/planner-layering-status.md is updated to record what this step does and does not cover.

Stack and validation

Stacked on #563 (stack/509-18-exact-cardinality). Head: stack/509-19-sql-frequency-entropy. Next: #565.

Validation (from the current PR body):

  • The entropy recognition and exact L2 population-guard regressions fail before their changes.
  • Frontend recognition, refusal, accuracy and inventory tests pass.
  • Native wire execution verifies filtering, empty input, nats, negative zero and unequal frequencies.
  • Existing L2/distinct acceptance passes.
  • Formatting and affected all-target Clippy with warnings denied pass.

🤖 Generated with Claude Code

@zzylol

zzylol commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Parked as draft: PR priorities changed (see #528). Order is now (A) finish #511 operator sharing, (B) the #572 crate/module reorganization, (C) #509 end-to-end stages. This PR sits on the old stack/528-legacy-physical-base chain, and Phase B moves the files it touches. Its content will be re-scoped onto the new layout in Phase C.

🤖 Generated with Claude Code

@zzylol

zzylol commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Ported onto the current stack in #598

zzylol added a commit that referenced this pull request Oct 4, 2026
Port of the parked #562 and #564 for #509 Example 2. The SQL frontend
names `SQRT(SUM(c*c))` over Float64 grouped unit counts as FrequencyL2,
and `-SUM(p*LN(p))` with `p = COUNT(*)*1.0 / SUM(COUNT(*)) OVER ()` as
FrequencyEntropy converted to nats. An exact population count guards
SQL's empty-input NULL. Integer products, nullable keys, filters, HAVING,
other log bases and partial windows are refused.

The rules are the old Pass 1 SemanticEquivalentRewriteStrategy rules, run
by lower_sql at the query root so the Stage 1 pipeline sees the intents;
the exact candidate is the native exact reducer. The executor gains SQL
sqrt (from #562).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant