# Slurm lost many jobs, can AiiDA deal with that?

**URL:** <https://aiida.discourse.group/t/slurm-lost-many-jobs-can-aiida-deal-with-that/560>\
**Category:** General Usage\
**Created:** [February 22, 2025, 10:55am UTC](https://aiida.discourse.group/t/slurm-lost-many-jobs-can-aiida-deal-with-that/560 "2025-02-22T10:55:01Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![pie](https://avatars.discourse-cdn.com/v4/letter/p/f9ae1b/32.png) [@pie](https://aiida.discourse.group/u/pie)\
**Post date:** [February 22, 2025, 10:55am UTC](https://aiida.discourse.group/t/slurm-lost-many-jobs-can-aiida-deal-with-that/560/1 "2025-02-22T10:55:01Z")

</div>

Dear All,

while running AiiDA on a slurm based cluster, the HPC system had quite a few issues and many jobs went lost, i.e. something like

```
scontrol show jobid 123455

```

returns `slurm_load_jobs error: Invalid job id specified` but the ID is perfectly valid and is written in the stdout of a job that was actually successfully completed.

On the AiiDA side, all calculations are reported as `RUNNING`.

Generally I would just kill everything and resubmit, but since it’s quite a low of simulations that were all successful, I was wondering if there’s a way to force the workers to collect the results and move forward.

Hacky solutions are welcome.

Thanks!

---

<div class="post-metadata">

**Author:** ![giovannipizzi](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/giovannipizzi/32/17_2.png) [@giovannipizzi](https://aiida.discourse.group/u/giovannipizzi)\
**Post date:** [February 22, 2025, 2:27pm UTC](https://aiida.discourse.group/t/slurm-lost-many-jobs-can-aiida-deal-with-that/560/2 "2025-02-22T14:27:24Z")

</div>

Hi! What does squeue reports? Is the job 123455 still in the queue? Is squeue does not report it and squeue does not return an error, AiiDA should assume the job is done and retrieve it, I think. Is maybe the job still there in some weird state? (if so, you can try to kill those jobs using scancel, and AiiDA will then proceed to retrieve them, if they actually finished I think there should be no problem. You can try with 1 job first, to see if it’s working, if that’s what’s happening

---

<div class="post-metadata">

**Author:** ![pie](https://avatars.discourse-cdn.com/v4/letter/p/f9ae1b/32.png) [@pie](https://aiida.discourse.group/u/pie)\
**Post date:** [February 23, 2025, 7:57am UTC](https://aiida.discourse.group/t/slurm-lost-many-jobs-can-aiida-deal-with-that/560/3 "2025-02-23T07:57:33Z")

</div>

Here’s an example:

```
893835 5D ago ElkCalculation ⏵ Waiting Monitoring scheduler: job state RUNNING

```

I use `calcjob gotocomputer` and land here

```auto
$ cat _scheduler-stdout.txt
================================================================================
JobID = 6059487
Partition = standard96:shared, Nodelist = gcn2022
================================================================================
============ Job Information ===================================================
Submitted: 2025-02-18T08:02:20
Started: 2025-02-18T08:02:24
Ended: 2025-02-19T09:06:13
Elapsed: 1504 min, Limit: 2880 min, Difference: 1376 min
CPUs: 16, Nodes: 1
Estimated Consumption: 200.53 core-hours
================================================================================

```

My `squeue -u $USER` is empty, this is the relevant part of the job

```auto
$ cat _aiidasubmit.sh
#!/bin/bash
#SBATCH --no-requeue
#SBATCH --job-name="aiida-893835"

```

and finally

```auto
$ sacct -j 6059487 --format="JobID,JobName%30,Partition,Account,AllocCPUS,State,ExitCode"
JobID JobName Partition Account AllocCPUS State ExitCode
------------ ------------------------------ ---------- ---------- ---------- ---------- --------
6059487 aiida-893835 standard9+ bep00079 16 COMPLETED 0:0
6059487.bat+ batch bep00079 16 COMPLETED 0:0
6059487.ext+ extern bep00079 16 COMPLETED 0:0

```

Edit: I forgot to post process report which is full of errors like this:

```auto
| raise SchedulerError(
 | aiida.schedulers.scheduler.SchedulerError: squeue returned exit code 1 (_parse_joblist_output function)
 | stdout=''
 | stderr='Loading software stack: nhr-lmod
 | Found project directory, setting $PROJECT_DIR to '/projects/extern/nhr/nhr_be/bep00079/dir.project'
 | Found scratch directory, setting $WORK to '/mnt/lustre-grete/usr/u14590'
 | Found scratch directory, setting $TMPDIR to '/mnt/lustre-grete/tmp/u14590'
 | __________ _ _________  ____  _____________  ____
 | \ \ / / ____| | /____ / __\| \/ |____ | | ____ / __ \
 | \ \ /\ / /| | __| | | | | | | | \ / | |__ | | | | | |
 | \ \/ \/ / | __| | | | | | | | | |\/| |__ | | | | | | |
 | \ /\ / | | ____| |___ | | ___| |__ | | | | | | ____| | | |__ | |
 | _ \/ _\/ _| ______|______ \ _____\____ /|_| |_| ______|____ |_| __\____ /
 | | \ | | | | | __\____ / ____\ \ / /__ \ / ____ |
 | | \| | | __| | |__ ) | / __\ | |__ \ \ /\ / /| | | | | __
 | | . ` | __ | _ / / / _` | | | |_ | \ \/ \/ / | | | | | |_ |
 | | |\ | | | | | \ \ | | (_| | | | __| | \ /\ / | |__ | | |__| |
 | |_| \_|_| |_|_| \_\ \ \ __,_| \_____ | \/ \/ | _____/ \_____ |
 | \ ____ /
 |
 | Documentation https://docs.hpc.gwdg.de Support nhr-support@gwdg.de
 | slurm_load_jobs error: Connection reset by peer'

```

---

<div class="post-metadata">

**Author:** ![giovannipizzi](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/giovannipizzi/32/17_2.png) [@giovannipizzi](https://aiida.discourse.group/u/giovannipizzi)\
**Post date:** [February 27, 2025, 6:28am UTC](https://aiida.discourse.group/t/slurm-lost-many-jobs-can-aiida-deal-with-that/560/4 "2025-02-27T06:28:31Z")

</div>

Hi, that’s a bit strange.  
Interesting the message `slurm_load_jobs error: Connection reset by peer'` that probably happened during squeue. However you are saying that now squeue works fine? (you don’t get any error?)

If AiiDA is not picking it up, maybe the daemon lost the corresponding task?

You might try this first. Not sure if it will help, but let’s try first.

> [@Graceful kill - instruct a paused job to retrieve results](https://aiida.discourse.group/t/graceful-kill-instruct-a-paused-job-to-retrieve-results/177/5):
>
> This should fix it verdi daemon stop verdi devel rabbitmq tasks analyze --fix verdi daemon start # Now wait a bit (depends on how many active procs you have) but couple seconds verdi process play PK

Also good to confirm that `squeue` does not give any error.

---

<div class="post-metadata">

**Author:** ![giovannipizzi](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/giovannipizzi/32/17_2.png) [@giovannipizzi](https://aiida.discourse.group/u/giovannipizzi)\
**Post date:** [February 27, 2025, 6:35am UTC](https://aiida.discourse.group/t/slurm-lost-many-jobs-can-aiida-deal-with-that/560/5 "2025-02-27T06:35:59Z")

</div>

One more thing, try to run the squeue commands also with the option `--jobs=6059487,6059487` which is what aiida does - yes, the job ID twice, see comment here:

> <https://github.com/aiidateam/aiida-core/blob/f4c55f5f78cd7fde9a5b4a4e48cad2159fd666b2/src/aiida/schedulers/plugins/slurm.py#L203-L231>
