Skip to content

ProcessMonitor loses all samples when one monitored process dies #740

Description

@milindsrivastava1997

Problem

ProcessMonitor.update_pid_monitor_map (asap-tools/experiments/classes/process_monitor.py:162) calls p.as_dict(...) without handling psutil.NoSuchProcess. If any monitored process exits mid-run, the exception escapes run(), the monitor stops, and it never sends its data back. remote_monitor.py then times out waiting and skips writing monitor_output.json, so the samples for every process in that run are lost, not only the dead one.

Observed

Experiment prom_test_2_10x, sketchdb mode: the query engine container (sketchdb-queryengine-rust) died about 35 min before the run ended.

sketchdb/remote_monitor.out:

Error in monitor process: process no longer exists (pid=112517)
...
psutil.NoSuchProcess: process no longer exists (pid=112517)

sketchdb/remote_monitor_output/remote_monitor.log:

ERROR | __main__:main:517 - Process monitor did not return data before timeout; not writing monitor_output.json (leave file absent)

As a result, compare_costs.py --all_experiment_modes reported only baseline costs.

Suggested fix

  • Catch psutil.NoSuchProcess in the sampling loop (for top-level and child handles), log it, and stop sampling that PID instead of crashing.
  • Keep the other PIDs' series aligned: compare_costs.py asserts that all PIDs have equal-length cpu_percent/memory_info, so either pad the dead PID's series (for example with 0 or None) or record the death index and handle it downstream.
  • Consider flagging the run in monitor_output.json (for example "exited_at_sample": N) so analysis scripts can warn that a component died mid-run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions