Skip to content

feat(supervisor): stage gateway configuration snapshot delivery - #3244

Open
pimlock wants to merge 5 commits into
mainfrom
1731-config-update-stage-1/pimlock
Open

feat(supervisor): stage gateway configuration snapshot delivery#3244
pimlock wants to merge 5 commits into
mainfrom
1731-config-update-stage-1/pimlock

Conversation

@pimlock

@pimlock pimlock commented Sep 9, 2026

Copy link
Copy Markdown
Collaborator

Summary

Add Stage 1 of gateway-pushed configuration over ConnectSupervisor. The gateway sends complete sandbox configuration and provider-environment snapshots; supervisors receive them while polling remains authoritative.

Related Issue

Part of #1731. Replaces #2967. Stage 2 will apply snapshots and acknowledge revisions; Stage 3 will add durable completion semantics and remove polling.

Changes

  • Add bootstrap, snapshot, and acknowledgement contracts with gateway/supervisor protocol revision checks.
  • Share read-only snapshot builders between polling and push delivery.
  • Route committed updates through an async interface with ordered, coalesced delivery, bounded fanout, and payload limits. Remote-owner routing follows in feat(kubernetes): support HA gateway rebalancing #1868.
  • Initialize policy history atomically without overwriting existing apply results.
  • Update architecture, compatibility documentation, and generated bindings.

Testing

  • Pre-commit and full local CI
  • Go and TypeScript SDK CI
  • Docker conformance and all four live-policy-update E2E tests

Postgres-specific concurrency coverage requires OPENSHELL_TEST_POSTGRES_URL and was not run.

Checklist

  • Conventional Commits and DCO sign-off
  • Architecture docs updated

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown

The startup repair that creates version-one policy history for legacy
sandboxes propagated validation failures, so a single stored policy that
no longer passes current validation rules prevented the gateway from
starting. Skip such sandboxes with a warning and a completion summary so
they keep the pre-repair behavior where only their own configuration reads
report the failure. Store errors remain fatal.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Fleet-wide configuration changes spawned one snapshot build per connected
sandbox and component with no concurrency limit, so a global setting or
provider change issued every store query and credential-driver call at
once. Gate builds behind a semaphore sized from the database pool and
start the build deadline only once a permit is held.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Sandboxes keep their supervisor binary until they are recreated, so a
gateway upgrade meets supervisors that predate the handshake and report
revision zero. Rejecting them severs every running sandbox with no
automatic recovery. Accept revision zero for one release, log a warning
per session, and count them in
openshell_supervisor_protocol_legacy_sessions_total. The supervisor
mirrors the allowance for gateways that predate the handshake.

Add a shared ConnectSupervisor test harness and handler-level tests for
legacy acceptance and unknown-revision rejection. Move the skill
troubleshooting paragraph out of the numbered deployment list so the
list renders.

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Sep 9, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@pimlock

pimlock commented Sep 9, 2026

Copy link
Copy Markdown
Collaborator Author

/ok to test a9387c7

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant