Skip to content

fix(stats): retry only retriable stats failures, with jittered backoff - #1387

Open
lucaspimentel wants to merge 2 commits into
mainfrom
lpimentel/stats-flusher-hardening
Open

lucaspimentel wants to merge 2 commits into
mainfrom
lpimentel/stats-flusher-hardening

Conversation

@lucaspimentel

@lucaspimentel lucaspimentel commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Overview

Hardens the stats flusher's retry behavior in bottlecap/src/traces/stats_flusher.rs:

  • Any 2xx counts as delivered. Previously only 202 was accepted, so a 200 or 204 from a proxy was treated as failure and the same bytes were re-sent even though the intake had already processed them. This removes one real stats double-count path.
  • Permanent 4xx responses stop the retry loop immediately. Statuses other than 408, 425, 429 and 5xx are no longer retried, instead of re-sending the same bytes FLUSH_RETRY_COUNT times.
  • Full-jitter exponential backoff between retries. Base 50 ms, upper bound doubles per retry (worst case 150 ms per send round, ~300 ms including the redrive round). 429 is retriable; Retry-After is ignored for pacing (far smaller than any realistic value) but logged at debug.

Internally, send_stats_payload now returns a typed SendOutcome (Success / Retriable / Permanent) instead of anyhow errors, with status classification in a pure classify_status function mirroring the Go agent's isRetriableStatus. The retry loop moved into a free send_with_retry function. The 512-byte response-body preview is preserved in failure details.

Scope: this does NOT fix the ambiguous-accept stats double-count

A local experiment (the statsv2 double-count rig) reproduced a stats double-count end to end: the intake processed a payload, the response was lost, and the flusher re-sent the same bytes; the intake counted both, because it does not deduplicate retried POST /api/v0.2/stats requests.

This PR does not fix that. Transport errors and per-attempt timeouts are exactly the failures where the intake may already have the payload, so they must stay retriable (matching the Go agent). The real fix is intake-side deduplication of retried stats requests, or an idempotency key the intake can dedupe on. Both are out of scope here. The rig's ambiguous-accept arm still double-counts with this change applied (verified: 3/3 rig tests pass, cherry-picked onto the experiment branch).

Testing

  • 11 new unit tests in stats_flusher.rs:
    • classify_status: 2xx → success; 408/425/429/500/502/503/599 → retriable; 400/401/403/404/413 → permanent.
    • Backoff: upper bound doubles per retry (50 ms, 100 ms, 200 ms); sampled delays stay within [0, upper].
    • Retry loop against httpmock: 202 → 1 hit, delivered; 200 → 1 hit, delivered (regression for the removed double-count path); 400 → 1 hit, not delivered (permanent, no retry); 503 and 429 → 3 hits, not delivered; connection-refused → not delivered.
  • cargo fmt --all -- --check, cargo clippy --workspace --all-targets (default, fips, default,test-mode), cargo nextest run --workspace: 701/701 passed.
  • Regression check: cherry-picked onto the statsv2-double-count-rig branch; all 3 double_count_integration_test arms pass (baseline: 1 payload, safe rejection: 1 payload, ambiguous accept: still 2, as expected).

Why not libdatadog's send_with_retry?

The status classification mirrors the Go tracer's isRetriableStatus (dd-trace-go/internal/llmobs/transport/transport.go:462). It deliberately diverges from libdatadog's send_with_retry (the agent-side Rust mechanism), which would be a behavior regression for this endpoint:

Behavior libdatadog send_with_retry This PR (stats flusher)
4xx classification All 4xx and 5xx retried, no permanent/retriable split (send_with_retry/mod.rs:170) Only 408/425/429/5xx retried; other 4xx stop after one attempt
Success Any non-4xx/5xx, so 3xx counts as success (send_with_retry/mod.rs:178) Only 2xx; 3xx is permanent
Jitter Deterministic exponential + optional additive jitter (retry_strategy.rs:74) Full jitter over the exponential bound
Default attempts 5 retries (6 attempts), 100 ms base 3 attempts (FLUSH_RETRY_COUNT), 50 ms base
Retry-After Not read Ignored for pacing, logged at debug

The permanent-4xx split is the notable divergence: libdatadog would burn all its retries on a 401; this flusher stops after one. The Go trace-agent itself does not do in-flush status classification; it retries failed payloads through a bounded retry buffer, which is conceptually closer to this flusher's existing redrive path (unchanged by this PR).

Scope: this does NOT fix the ambiguous-accept stats double-count
reproduced by the statsv2 double-count rig. In that experiment the
intake processed a payload, the response was lost, and the flusher
re-sent the same bytes; the intake counted both, because it does not
deduplicate retried POST /api/v0.2/stats requests. Transport errors
and per-attempt timeouts are exactly the failures where the intake may
already have the payload, so they must stay retriable (matching the Go
agent's isRetriableStatus). The real fix is intake-side deduplication
of retried stats requests, or an idempotency key the intake can dedupe
on. Both are out of scope here.

What this change does:

- Removes one real double-count path: any 2xx other than 202 (for
  example a 200 or 204 returned by a proxy) was treated as failure and
  retried even though the intake accepted it. Any 2xx now counts as
  delivered.
- Stops wasted retries on permanent 4xx responses (400, 401, 403, 404,
  413, ...): the retry loop stops immediately instead of re-sending
  the same bytes FLUSH_RETRY_COUNT times.
- Spreads out retries under 429/5xx storms with full-jitter backoff
  over an exponential base (base 50 ms, upper bound doubles per retry;
  worst case 150 ms per send round). Retry-After on 429 is ignored for
  pacing (far smaller than any realistic value) but logged at debug.

What changed internally: send_stats_payload now returns a typed
SendOutcome (Success / Retriable / Permanent) instead of anyhow
errors, with status classification in a pure classify_status function
mirroring dd-trace-go. The retry loop moved into a free
send_with_retry function. The 512-byte response-body preview is kept
in failure details.

What did not change:

- Redrive semantics: after any failed round, including a permanent
  failure, send still returns Some(stats) for the one extra flush
  round in flushing/service.rs. Open question (noted in code): should
  permanent statuses skip redrive?
- trace_flusher.rs and the shared FLUSH_RETRY_COUNT (= 3).
- Empty-input, missing-API-key, and serialization-error handling.
@datadog-datadog-prod-us1-2

datadog-datadog-prod-us1-2 Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Pipelines

✨ Unblock PR with BitsAI

❌ Errors

Your PR has failed checks. Please review the issues below and take necessary action before merging.

🚦 2 Pipeline jobs failed

Bottlecap (Rust) | Format (ubuntu-22.04) — 🔧 Needs a code fix, caused by this PR

View more details · View in GitHub Actions

DataDog/datadog-lambda-extension | cargo fmt — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

Useful? React with 👍 / 👎

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 223d1ea | Docs | View more details | Give us feedback!

@lucaspimentel
lucaspimentel marked this pull request as ready for review September 25, 2026 18:06
Copilot AI lite review requested due to automatic review settings September 25, 2026 18:06
@lucaspimentel
lucaspimentel requested review from a team as code owners September 25, 2026 18:06

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bfa507c69f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread bottlecap/src/traces/stats_flusher.rs Outdated
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-25T18:08:28.402326Z bfa507c Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved retry redrive, response-data logging, and Retry-After handling issues remain.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity · 1 Medium severity

Open (2)
What changed in this PR

This PR hardens stats flushing with HTTP outcome classification and jittered retry backoff.

Changes:

  • Accepts all 2xx responses.
  • Retries only transient failures.
  • Adds jittered exponential backoff and focused tests.
File Summary Findings
bottlecap/​src/​traces/​stats_flusher.rs Implements status classification, retry backoff, and tests. Permanent failures are still redriven (moderate, 3 votes); response details may leak into logs (critical, 4 votes); Retry-After is ignored (moderate, 1 vote).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread bottlecap/src/traces/stats_flusher.rs
Comment thread bottlecap/src/traces/stats_flusher.rs Outdated
let retry = u32::try_from(attempt).unwrap_or(u32::MAX);
let delay = backoff.delay(&mut rand::thread_rng(), retry);
debug!("STATS | Retrying stats flush in {} ms", delay.as_millis());
tokio::time::sleep(delay).await;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

StatsFlusher::flush holds the aggregator mutex for the entire retry loop in send() (stats_flusher.rs:150-157), and this PR adds up to ~150ms of tokio::time::sleep per retriable failure inside that same call path (stats_flusher.rs:288-292) — pre-PR, retries were back-to-back with no sleep, so this is new lock-hold time. This causes lock contention with the background task that drains tracer-submitted stats into the aggregator (trace_agent.rs:230-233), which can add latency to the tracer's stats requests during an intake outage. Suggest scoping the guard to just the get_batch call (drop it before send() runs) in stats_flusher.rs:150-157 so retry sleeps happen outside the critical section:

loop {
    let stats = {
        let mut guard = self.aggregator.lock().await;
        guard.get_batch(force_flush).await   // lock held only for this line
    };                                          // guard dropped here, lock released
    if stats.is_empty() {
        break;
    }
    if let Some(failed) = self.send(stats).await {   // now runs with no lock held
        all_failed.extend(failed);
    }
}

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants