Conversation
A robot that is switched off neither accepts nor refuses connection requests, so every connection attempt waited for the operating system's connect timeout (about two minutes on Linux). TCPSocket::setConnectTimeout() bounds a single connection attempt, including all addresses the host resolves to. A timed-out socket is closed before the wait for the next attempt. RTDEClient, PrimaryClient, DashboardClient and UrDriverConfiguration expose the timeout; it also applies to automatic reconnects. The default of zero keeps the current behavior. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PR SummaryMedium Risk Overview
Docs add a Connection timeout section for RTDE, primary, and dashboard clients. New Reviewed by Cursor Bugbot for commit 210bd5b. Bugbot is set up for automated code reviews on this repo. Configure here. |
A robot that is switched off doesn't answer connection requests at all, so every connection attempt waits until the operating system gives up. On Linux that takes about two minutes. We ran into this with an RTDE client that should notice within seconds that the robot is off and keep retrying in the background. #519 made the connect interruptible, but a single attempt still has no upper bound unless another thread calls
disconnect().This PR adds an opt-in timeout for a single connection attempt. The default is 0, which keeps the current behavior.
Changes
TCPSocket::setConnectTimeout()andgetConnectTimeout(). The existing poll loop inopenInterruptible()gives up at the deadline. One deadline covers all addresses the host resolves to. Name resolution isn't included. A failed socket is now closed right away, so a timed-out request can't still complete during the back-off before the next attempt.RTDEClient,PrimaryClientandDashboardClientforward the timeout to their socket, so it also applies to automatic reconnects. For PolyScope X the dashboard client hands it to cpp-httplib instead. The 5 s default stays, and 0 falls back to httplib's 300 s because httplib can't turn the timeout off.UrDriverConfiguration::socket_connect_timeout, applied before the RTDE and primary clients connect. I put it last in the struct so positional initializers keep compiling.Why a setter instead of another parameter. I know
setReconnectionTime()was deprecated in favor of passing values toconnect(). The timeout also has to reach the automatic reconnect inURProducer, which callsstream_.reconnect()without arguments. A parameter would mean changing the virtualIProducer::setupProducer()andPipeline::init(), which breaks custom producers. The setter leaves both alone. If you'd rather go another way, I'm happy to change it.Compatibility. Existing code compiles unchanged. The ABI changes, though.
TCPSocketgets a data member,DashboardClientImplgets two virtual functions with no-op defaults (likesetReceiveTimeout()), andUrDriverConfigurationandUrDrivergrow.Tests
The tests don't depend on an unroutable address.
UnresponsiveServerintest_utils.hlistens on loopback and fills its accept queue, so the kernel drops further connection requests without answering, just like a switched-off robot. Where the operating system refuses them instead (I expect Windows), the timing tests skip.test_tcp_socket.cppcovers the timeout for one attempt and across retries,reconnect(), 0 still waiting for the OS,disconnect()still interrupting, refused connections not being delayed, and a timed-out request being closed before the back-off.test_client_connect_timeout.cppis new and needs no robot. It coversRTDEClient::init(),PrimaryClient::start(), the G5 and PolyScope X dashboard clients andUrDriver.With the deadline disabled, the timing tests fail. Without the early close, the back-off test fails. I ran the unit tests on Linux with GCC, in Debug and with the CI flags (ASAN, coverage, warnings as errors), and all pass. I haven't run them on Windows or macOS.
🤖 Generated with Claude Code