Skip to content

[metrics] Add per-error TabletServer request metrics - #4213

Merged
swuferhong merged 1 commit into
apache:mainfrom
fxbing:feature/20260902-request-error-dimension
Sep 14, 2026
Merged

swuferhong merged 1 commit into
apache:mainfrom
fxbing:feature/20260902-request-error-dimension

Conversation

@fxbing

@fxbing fxbing commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Purpose

Linked issue: close #4212

Expose failed TabletServer RPC requests by error type while preserving the existing aggregate metric.

Brief change log

  • Add lazy fluss_tabletserver_request_error_errorsPerSecond metrics with request and error labels.
  • Record one event for each failed RPC response and document the metric.

Tests

  • MetricGroupTest, RequestsMetricsTest, NettyServerHandlerTest
  • spotless:check and compilation of fluss-common and fluss-rpc
  • Live /metrics smoke test was not run.

API and Format

No public API or storage-format changes.

Documentation

Updated the metrics reference.

Generative AI disclosure

  • Yes — OpenAI Codex and Anthropic Claude

- Preserve the existing aggregate request metric and API.\n- Lazily register error-scoped meters through MetricGroup.addGroup.\n- Cover the sendError wiring and document the new metric scope.

@swuferhong swuferhong left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, approving. One follow-up suggestion, not a blocker.

The metric currently only fires from sendError(), i.e. when the whole request fails. But produceLog, fetchLog, lookup and putKv are batch APIs that return a successful envelope with per-bucket error_code fields inside, and that's where almost all steady-state errors live — a NOT_LEADER_OR_FOLLOWER during leader migration never reaches sendError(). So on the hottest paths the new series will usually be absent while buckets are in fact failing, which reads as "no errors" to whoever is on call.

Kafka counts these. Its docs for ErrorsPerSec say "if a response contains multiple errors, all are counted", and ProduceResponse.errorCounts() walks data.responses() → partitionResponses() to aggregate every errorCode.

We can cover this more cheaply than Kafka, because protogen already knows where the errors are: ProtobufMessage.hasErrorFields() detects the error_code/error_message pair and makes those messages implement ErrorMessage, so PbProduceLogRespForBucket, PbFetchLogRespForBucketetc. are already tagged. Since the generator also knows the nested-message structure statically, it could emit acollectErrorCounts(Map<Errors, Integer>)that self-reports its own code and recurses into nested message fields.NettyServerHandlerwould call it on the success path too, mapping throughErrors.forCode`. No hand-written per-response code, and no way for a new response type to silently drift out of sync — which is better than Kafka's manual overrides.

@swuferhong
swuferhong merged commit 14f9aae into apache:main Sep 14, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[metrics] Add per-error TabletServer request metrics

2 participants