Repository navigation
feat(observability): OpenTelemetry export surface for gateway operators #2507
Description
Activity
- addedtopic:observabilityLogging, metrics, and observability workLogging, metrics, and observability workarea:gatewayGateway server and control-plane workGateway server and control-plane workarea:clusterRelated to running OpenShell on k3s/dockerRelated to running OpenShell on k3s/docker
on Jul 27, 2026 - added a parent issue
on Jul 27, 2026 Additional sub-issue: compute driver and gateway interceptor spans
Adding an eighth proposed sub-issue, surfaced while scoping the sandbox-side companion #2508.
8. Compute driver and gateway interceptor spans, with trace context propagation.
Both are gateway-side surfaces — the gateway is the client in each case — so they belong here rather than in #2508.
Compute drivers. This is the missing piece of the
CreateSandboxdecomposition in the problem statement above. Without driver spans, "sandbox creation took 4.2s" bottoms out at "the driver took 3s," with no visibility into whether the time went to image pull, scheduling, or readiness wait.Note this is a moving target: drivers are expected to move from in-process backends to external gRPC services over
proto/compute_driver.proto, while remaining in-tree. Today only the VM driver is a separate process (crates/openshell-server/src/compute/vm.rs); Docker, Podman, and Kubernetes are in-process. Once that migration happens, every driver call becomes a real process boundary andcompute_driver.protobecomes a trace-context propagation point.That has two consequences for how this sub-issue should be built:
- Do not design driver instrumentation as in-process spans that later need rewriting. Treat the driver contract as a propagation boundary from the start, even for backends that are currently in-process, so the same instrumentation survives the migration.
- Trace context propagation over
compute_driver.protoshould be settled as part of that contract's evolution rather than bolted on afterward. Worth coordinating with whoever owns the external-driver migration before implementing.
Gateway interceptors.
Evaluateonproto/gateway_interceptor.protoalready calls an out-of-process, operator-run service, so this is a propagation boundary today. The existingopenshell_gateway_interceptor_fail_closed_total,fail_open_total, andlatency_secondsmetrics (crates/openshell-gateway-interceptors/src/runtime.rs) are evidence the observability need was already felt here, and they give a ready-made baseline to validate span timings against.Both cases share a property that makes them high value: the operator controls the service on the other side, so propagating context yields a genuine cross-service trace rather than a span that dead-ends at the boundary. This is the same argument that makes supervisor middleware the strongest case in #2508.
Concrete motivator from #2516 / #2498 follow-up:
The rootless Podman E2E lane on main timed out after #2498 raised the production Podman graceful stop timeout to 45s. The tests had already passed; the remaining time was spent in gateway/sandbox teardown. #2516 fixes CI by overriding the Podman E2E harness timeout to 15s while preserving the production default.
This is a good example of why compute-driver lifecycle spans belong in this observability work. For this class of issue, operators need to see teardown decomposed by phase, not just the outer sandbox delete duration or warning logs. Useful span attributes/events would include:
- driver.name
- sandbox.id
- operation = stop|remove|cleanup
- timeout_secs
- elapsed_ms
- outcome = graceful|forced|not_found|error
- backend-specific identifiers such as container/pod/vm id
For future external drivers, this also reinforces the need for trace context propagation across compute_driver.proto: the gateway can record the outer driver RPC span, while the driver records child spans for backend-specific phases such as Podman stop, Kubernetes pod deletion, or VM shutdown.
Reacted by krishicks- added 7 commits that reference this issue
on Jul 28, 2026 21 remaining items
Persona workflows from the observability RFC work (PRs #2628/#2629, now closed in favor of issues). These workflows motivate the driver instrumentation and deployment configuration tracked here:
Platform operator: debugging slow sandbox creation
An operator receives an alert that sandbox creation latency exceeds the SLO. They open Jaeger and search for
sandbox.createspans in the last hour. They find one taking 15 seconds. They drill into the child spans, expecting to see which compute driver operation was slow, but the trace ends at the gateway request span. The K8s driver'sprovision_podcall, the Docker driver'screate_container, the Podman driver's equivalent are all invisible because none of them have#[tracing::instrument]annotations.How this issue enables the workflow: The in-process driver spans appear as children of the gateway request span. The operator sees that
provision_podtook 14 of the 15 seconds, drills into the K8s API calls, and identifies that the node was under memory pressure. No new infrastructure is needed because these drivers already inherit the gateway's tracing subscriber.Platform integrator: wiring OpenShell into an existing monitoring stack
A platform integrator is deploying OpenShell on a managed K8s cluster that already runs Prometheus, Grafana, and Tempo. They need OpenShell to emit traces to their Tempo endpoint and expose Prometheus metrics for their existing dashboards and alerts. Today, the Helm chart has no OTLP configuration values, no ServiceMonitor template, and the TOML config reference does not document the OTLP section.
How this issue enables the workflow: The integrator sets
otlp.enabled=trueandotlp.endpoint=http://tempo:4317in the Helm values. They enable the ServiceMonitor and Prometheus starts scraping the gateway's/metricsendpoint. OpenShell traces appear in Tempo alongside their other services. The published docs atdocs/reference/gateway-config.mdxdocument the configuration surface.- added 10 commits that reference this issue
on Aug 6, 2026 This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.
- addedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Aug 23, 2026 - added a commit that references this issue
on Aug 26, 2026 - added a commit that references this issue
on Aug 27, 2026
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsPlanning
Problem Statement
OpenShell gives gateway operators three observability surfaces today, and none of them answer "what is my gateway doing right now, and why is it slow or failing?"
/metrics(Prometheus) exposes gRPC/HTTP request counters and duration histograms, a readiness gauge, and gateway-interceptor counters. It is scrape-only, disabled by default (--metrics-portdefaults to0), and has no Helm scrape wiring.openshell-core::telemetry) forwards aggregate usage to NVIDIA. It is deliberately coarse and is not readable by the operator running the gateway.What is missing is the connective tissue operators expect from any modern control plane: distributed traces, an OTLP push path for environments that do not scrape, and a turnkey collector wiring so this works without hand-assembling a stack.
Concretely, an operator investigating a slow
CreateSandboxtoday can see that the request took 4.2s from the request-duration histogram, and can find its log lines via the request ID from #932 — but cannot see where those 4.2 seconds went across compute-driver call, policy evaluation, supervisor session establishment, and readiness wait. That decomposition is exactly what a trace provides and no current surface does.This issue is the tracking parent for that work. It is a sub-issue of #1055, which remains the roadmap-level umbrella covering OCSF, log export, product telemetry, and dashboards.
Audience: the operator running an OpenShell gateway. Tracing for users of the SDKs who create sandboxes (#1818) is explicitly out of scope here.
Proposed Design
Deliver an OTel export surface for the gateway process, decomposed into sub-issues. PR #1270 built a working version of most of this and was closed unmerged on 2026-05-27; it is a valuable reference, but the implementation should be re-derived rather than rebased (see Agent Investigation for why).
Proposed sub-issues
Gateway OTLP trace export. Append a
tracing-opentelemetrylayer to the existing subscriber intracing_bus.rs, so the currenttower_http::trace::TraceLayerper-request span becomes the OTLP root without rewriting handlers. Resolve configuration from standardOTEL_*environment variables with a CLI/TOML fallback, honorOTEL_TRACES_SAMPLER, and flush the batch span processor on shutdown. Default off.Span coverage for gateway operations. Instrument the spans that make a trace useful rather than merely present: sandbox lifecycle transitions, compute-driver calls, policy evaluation, supervisor session establishment and relay claim, and credential-broker actions. Decide per operation whether it is a span, an event on a parent span, or neither.
Correlation identifiers. Carry stable fields across spans and align them with OCSF event correlation IDs so an operator can pivot from a trace to the security record for the same activity, and reuse the request ID from feat(server): add request-ID middleware for request correlation #932 rather than inventing a parallel identifier. This subsumes the emission half of Add OpenTelemetry trace correlation across gateway activity #1758.
OTLP metrics export. Offer OTLP push for the existing
metricsfamilies alongside the current Prometheus scrape endpoint, for operators whose collectors do not scrape. Per the discussion on feat(server): metrics instrumentation #909, scrape remains the default and push is opt-in.Kubernetes monitoring surface. Helm
ServiceMonitor(gated, off by default) targeting the existing named metrics port,OTEL_*projection into the StatefulSet, and values documentation. Also decide whether--metrics-portshould default to enabled when the chart wires up scraping.Local development stack. A
misetask installing a collector plus a trace backend and Grafana into the k3d dev cluster, so contributors can validate the pipeline end to end. Mirrors the existing Keycloak dev add-on pattern.Documentation. An operator-facing page under
docs/observability/covering enablement, configuration surface, what spans exist, and a reference collector setup — plus an update toarchitecture/gateway.md.Sub-issues are expected to land independently; the trace export path (1–3) is the critical path and the rest can follow.
Open design questions
OTEL_*env vars are the ecosystem convention, but OpenShell has a TOML gateway config (RFC 0003) that should probably own this. Precedence between the two needs a decision, anddocs/reference/gateway-config.mdxneeds updating either way.Alternatives Considered
Agent Investigation
Verified against
mainat 0d5e5c5.No OpenTelemetry exists in the repository.
grep -rni "opentelemetry|otlp"acrosscrates/,python/,deploy/, anddocs/returns a single hit, indeploy/sbom/resolve_licenses.py, unrelated to product code.What is present:
metricsfacade +metrics-exporter-prometheus,/metricsroute atcrates/openshell-server/src/http.rs:172, builder init atcrates/openshell-server/src/lib.rs:472.openshell_server_grpc_requests_total,openshell_server_grpc_request_duration_seconds,openshell_server_http_requests_total,openshell_server_http_request_duration_seconds(multiplex.rs), a readiness gauge (readiness.rs), and interceptor fail-open/fail-closed/latency (openshell-gateway-interceptors/src/runtime.rs). The broader catalog proposed in feat(server): metrics instrumentation #909 is unimplemented.--metrics-port/OPENSHELL_METRICS_PORTdefaults to0, meaning the dedicated metrics listener is off unless configured (cli.rs:68).Why PR #1270 should be re-derived rather than rebased: it pinned
opentelemetry0.29 andtracing-opentelemetry0.30, described in the PR body as the latest set compatible with the workspace'stonic 0.12+prost 0.13. The workspace is now ontonic 0.14/tonic-prost 0.14/prost 0.14, so that pin rationale no longer holds and the version selection has to be redone. The PR was also authored before the gateway TOML config (#1317), the gateway interceptors (#2005), and the supervisor middleware crates landed — all of which are plausible span sources. It was closed with no review comments explaining the decision, so the reason for abandonment is not recoverable from the PR itself and may be worth confirming with a maintainer before investing.Related issues:
state:triage-needed, no comments since 2026-06-04. Its emission and correlation requirements should be absorbed by sub-issues 1–3 and the issue closed as superseded, or retargeted as one of them.Deliberately adjacent, not children:
state:stalewith an unmerged POC branch.Checklist