[baobab] gpu048 - SIGTERM without a reason

Primary informations

Username: falkiew2
Cluster: baobab

Description

Dear HPC Team,

My two recent jobs (12097387 and 12107358) got a SIGTERM after running for 00:29:21 and 0:28:55, respectively, while they were scheduled for 12 and 24 hours, respectively.

Thank you in advance.

Steps to Reproduce

Launch a job on gpu048, with 8 GPUs for 12 or 24h

Expected Result

Job runs uninterrupted for the requested time.

Actual Result

srun: error: gpu048: task 0: Exited with exit code 1
3%|▎ | 15/498 [26:13<14:01:49, 104.57s/it]W0829 11:05:31.954000 1161860 site-packages/torch/distributed/elastic/multiprocessing/api.py:897] Sending process 1162140 closing signal SIGTERM

Kind regards,

Maciej Falkiewicz

Dear Maciej,

The error output is not really clear, I was wondering if it coul be possible the script request more ram very quickly and the trigger for OOM didn’t have time to dectect. Then the python fail and the job is in error :

Memory Utilized: 46.26 GB
Memory Efficiency: 74.01% of 62.50 GB (62.50 GB/node)

Best regards,

Dear Maciej,

Since we migrate beegfs from v7 to v8, it seems it may have some issues. We are investigating with beegfs support.

#1-6A951454-1
[259786.704022] beegfs: beegfs_Flusher(4678): Remoting (write file): Error storage targetID: 4; Msg: Potential cache loss for open file handle. (Server crash detected.); FileHandle: 10B8C93F#1-6A951454-1
[259789.484423] beegfs: wandb-core(3461357): Remoting (write file): Error storage targetID: 3; Msg: Potential cache loss for open file handle. (Server crash detected.); FileHandle: 10BB1BAE#54-6A952450-1
[259789.485677] beegfs: wandb-core(3461349): Remoting (write file): Error storage targetID: 3; Msg: Potential cache loss for open file handle. (Server crash detected.); FileHandle: 10BB1BAE#54-6A952450-1
[259790.238209] beegfs: wandb-core(3399219): Remoting (write file): Error storage targetID: 4; Msg: Potential cache loss for open file handle. (Server crash detected.); FileHandle: 1656DA6D#4-6A951461-4
[259790.239379] beegfs: wandb-core(3448278): Remoting (write file): Error storage targetID: 4; Msg: Potential cache loss for open file handle. (Server crash detected.); FileHandle: 1656DA6D#4-6A951461-4
(baobab)-[root@gpu048 scratch]$

Best regards,

Dear Maciej,

We have updated a system parameter, and the storage issue should now be resolved.
Please let us know if you experience any further I/O performance or stability issues.

Best regards,