Problem
ProcessMonitor.update_pid_monitor_map (asap-tools/experiments/classes/process_monitor.py:162) calls p.as_dict(...) without handling psutil.NoSuchProcess. If any monitored process exits mid-run, the exception escapes run(), the monitor stops, and it never sends its data back. remote_monitor.py then times out waiting and skips writing monitor_output.json, so the samples for every process in that run are lost, not only the dead one.
Observed
Experiment prom_test_2_10x, sketchdb mode: the query engine container (sketchdb-queryengine-rust) died about 35 min before the run ended.
sketchdb/remote_monitor.out:
Error in monitor process: process no longer exists (pid=112517)
...
psutil.NoSuchProcess: process no longer exists (pid=112517)
sketchdb/remote_monitor_output/remote_monitor.log:
ERROR | __main__:main:517 - Process monitor did not return data before timeout; not writing monitor_output.json (leave file absent)
As a result, compare_costs.py --all_experiment_modes reported only baseline costs.
Suggested fix
- Catch
psutil.NoSuchProcess in the sampling loop (for top-level and child handles), log it, and stop sampling that PID instead of crashing.
- Keep the other PIDs' series aligned:
compare_costs.py asserts that all PIDs have equal-length cpu_percent/memory_info, so either pad the dead PID's series (for example with 0 or None) or record the death index and handle it downstream.
- Consider flagging the run in
monitor_output.json (for example "exited_at_sample": N) so analysis scripts can warn that a component died mid-run.
Problem
ProcessMonitor.update_pid_monitor_map(asap-tools/experiments/classes/process_monitor.py:162) callsp.as_dict(...)without handlingpsutil.NoSuchProcess. If any monitored process exits mid-run, the exception escapesrun(), the monitor stops, and it never sends its data back.remote_monitor.pythen times out waiting and skips writingmonitor_output.json, so the samples for every process in that run are lost, not only the dead one.Observed
Experiment
prom_test_2_10x,sketchdbmode: the query engine container (sketchdb-queryengine-rust) died about 35 min before the run ended.sketchdb/remote_monitor.out:sketchdb/remote_monitor_output/remote_monitor.log:As a result,
compare_costs.py --all_experiment_modesreported only baseline costs.Suggested fix
psutil.NoSuchProcessin the sampling loop (for top-level and child handles), log it, and stop sampling that PID instead of crashing.compare_costs.pyasserts that all PIDs have equal-lengthcpu_percent/memory_info, so either pad the dead PID's series (for example with0orNone) or record the death index and handle it downstream.monitor_output.json(for example"exited_at_sample": N) so analysis scripts can warn that a component died mid-run.