Skip to content

Add an optional connect timeout for robot connections - #585

Open
dkar-sto wants to merge 1 commit into
UniversalRobots:masterfrom
dkar-sto:feat/connect-timeout
Open

dkar-sto wants to merge 1 commit into
UniversalRobots:masterfrom
dkar-sto:feat/connect-timeout

Conversation

@dkar-sto

Copy link
Copy Markdown
Contributor

A robot that is switched off doesn't answer connection requests at all, so every connection attempt waits until the operating system gives up. On Linux that takes about two minutes. We ran into this with an RTDE client that should notice within seconds that the robot is off and keep retrying in the background. #519 made the connect interruptible, but a single attempt still has no upper bound unless another thread calls disconnect().

This PR adds an opt-in timeout for a single connection attempt. The default is 0, which keeps the current behavior.

Changes

  • TCPSocket::setConnectTimeout() and getConnectTimeout(). The existing poll loop in openInterruptible() gives up at the deadline. One deadline covers all addresses the host resolves to. Name resolution isn't included. A failed socket is now closed right away, so a timed-out request can't still complete during the back-off before the next attempt.
  • RTDEClient, PrimaryClient and DashboardClient forward the timeout to their socket, so it also applies to automatic reconnects. For PolyScope X the dashboard client hands it to cpp-httplib instead. The 5 s default stays, and 0 falls back to httplib's 300 s because httplib can't turn the timeout off.
  • UrDriverConfiguration::socket_connect_timeout, applied before the RTDE and primary clients connect. I put it last in the struct so positional initializers keep compiling.
  • A short "Connection timeout" section in the RTDE, primary and dashboard client docs.

Why a setter instead of another parameter. I know setReconnectionTime() was deprecated in favor of passing values to connect(). The timeout also has to reach the automatic reconnect in URProducer, which calls stream_.reconnect() without arguments. A parameter would mean changing the virtual IProducer::setupProducer() and Pipeline::init(), which breaks custom producers. The setter leaves both alone. If you'd rather go another way, I'm happy to change it.

Compatibility. Existing code compiles unchanged. The ABI changes, though. TCPSocket gets a data member, DashboardClientImpl gets two virtual functions with no-op defaults (like setReceiveTimeout()), and UrDriverConfiguration and UrDriver grow.

Tests

The tests don't depend on an unroutable address. UnresponsiveServer in test_utils.h listens on loopback and fills its accept queue, so the kernel drops further connection requests without answering, just like a switched-off robot. Where the operating system refuses them instead (I expect Windows), the timing tests skip.

  • test_tcp_socket.cpp covers the timeout for one attempt and across retries, reconnect(), 0 still waiting for the OS, disconnect() still interrupting, refused connections not being delayed, and a timed-out request being closed before the back-off.
  • test_client_connect_timeout.cpp is new and needs no robot. It covers RTDEClient::init(), PrimaryClient::start(), the G5 and PolyScope X dashboard clients and UrDriver.

With the deadline disabled, the timing tests fail. Without the early close, the back-off test fails. I ran the unit tests on Linux with GCC, in Debug and with the CI flags (ASAN, coverage, warnings as errors), and all pass. I haven't run them on Windows or macOS.

🤖 Generated with Claude Code

A robot that is switched off neither accepts nor refuses connection
requests, so every connection attempt waited for the operating system's
connect timeout (about two minutes on Linux).

TCPSocket::setConnectTimeout() bounds a single connection attempt,
including all addresses the host resolves to. A timed-out socket is
closed before the wait for the next attempt. RTDEClient, PrimaryClient,
DashboardClient and UrDriverConfiguration expose the timeout; it also
applies to automatic reconnects. The default of zero keeps the current
behavior.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cursor

cursor Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

PR Summary

Medium Risk
Changes core connection/reconnect timing across TCPSocket and all major clients; default behavior is unchanged but ABI grows (new TCPSocket member, virtuals on DashboardClientImpl, UrDriverConfiguration field).

Overview
Adds an opt-in connect timeout so a single TCP connection attempt can fail fast when the robot is off (instead of waiting ~2 minutes for the OS). Default 0 keeps today’s OS-level behavior.

TCPSocket gets setConnectTimeout() / getConnectTimeout(). The interruptible connect loop in openInterruptible() now stops at a per-attempt deadline (shared across resolved addresses, not DNS). Timed-out sockets are closed immediately so a stale SYN cannot succeed during retry back-off.

RTDEClient, PrimaryClient, and DashboardClient expose the same API and forward to their transport (TCP for G5; PolyScope X maps to cpp-httplib with a 5 s default and 0 → 300 s fallback). UrDriverConfiguration::socket_connect_timeout wires the limit into primary/RTDE setup on driver init.

Docs add a Connection timeout section for RTDE, primary, and dashboard clients. New UnresponsiveServer tests and expanded test_tcp_socket / test_client_connect_timeout cover timeouts, reconnect, and driver init without a real robot.

Reviewed by Cursor Bugbot for commit 210bd5b. Bugbot is set up for automated code reviews on this repo. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant