My two recent jobs (12097387 and 12107358) got a SIGTERM after running for 00:29:21 and 0:28:55, respectively, while they were scheduled for 12 and 24 hours, respectively.
Thank you in advance.
Steps to Reproduce
Launch a job on gpu048, with 8 GPUs for 12 or 24h
Expected Result
Job runs uninterrupted for the requested time.
Actual Result
srun: error: gpu048: task 0: Exited with exit code 1 3%|▎ | 15/498 [26:13<14:01:49, 104.57s/it]W0829 11:05:31.954000 1161860 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 1162140 closing signal SIGTERM
The error output is not really clear, I was wondering if it coul be possible the script request more ram very quickly and the trigger for OOM didn’t have time to dectect. Then the python fail and the job is in error :
We have updated a system parameter, and the storage issue should now be resolved.
Please let us know if you experience any further I/O performance or stability issues.