# WorkChain exception inside AiidaLab container

**URL:** <https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152>\
**Category:** General Usage\
**Tags:** aiidalab\
**Created:** [November 10, 2023, 11:46am UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152 "2023-11-10T11:46:39Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![AndresOrtegaGuerrero](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/andresortegaguerrero/32/187_2.png) [@AndresOrtegaGuerrero](https://aiida.discourse.group/u/AndresOrtegaGuerrero)\
**Post date:** [November 10, 2023, 11:46am UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/1 "2023-11-10T11:46:39Z")

</div>

Dear Community,

I am trying to run a workchain that has several steps, the workchain seems to work fine in with a small system in the localhost (i am working in the aiidalab container ) but when i tried a bigger system and using a cluster i get this error

```auto
2023-11-10 11:08:44 [6174 | REPORT]: [23510|DielectricWorkChain|on_except]: Traceback (most recent call last):
  File "/opt/conda/lib/python3.9/site-packages/plumpy/base/state_machine.py", line 324, in transition_to
    self._enter_next_state(new_state)
  File "/opt/conda/lib/python3.9/site-packages/plumpy/base/state_machine.py", line 388, in _enter_next_state
    self._fire_state_event(StateEventHook.ENTERED_STATE, last_state)
  File "/opt/conda/lib/python3.9/site-packages/plumpy/base/state_machine.py", line 300, in _fire_state_event
    callback(self, hook, state)
  File "/opt/conda/lib/python3.9/site-packages/plumpy/processes.py", line 331, in <lambda>
    lambda _s, _h, from_state: self.on_entered(cast(Optional[process_states.State], from_state)),
  File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/process.py", line 426, in on_entered
    super().on_entered(from_state)
  File "/opt/conda/lib/python3.9/site-packages/plumpy/processes.py", line 714, in on_entered
    self._communicator.broadcast_send(body=None, sender=self.pid, subject=subject)
  File "/opt/conda/lib/python3.9/site-packages/plumpy/communications.py", line 175, in broadcast_send
    return self._communicator.broadcast_send(body, sender, subject, correlation_id)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/threadcomms.py", line 258, in broadcast_send
    result = self._loop_scheduler.await_(
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 164, in await_
    return self.await_submit(awaitable).result(timeout=self.task_timeout)
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 446, in result
    return self.__get_result()
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 391, in __get_result
    raise self._exception
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 256, in __step
    result = coro.send(None)
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 178, in coro
    res = await awaitable
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 522, in broadcast_send
    result = await publisher.broadcast_send(body, sender, subject, correlation_id)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 66, in broadcast_send
    return await self.publish(message, routing_key=defaults.BROADCAST_TOPIC, mandatory=False)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/messages.py", line 209, in publish
    result = await self._exchange.publish(message, routing_key=routing_key, mandatory=mandatory)
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/exchange.py", line 233, in publish
    return await asyncio.wait_for(
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 442, in wait_for
    return await fut
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 508, in basic_publish
    async with self.lock:
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 90, in lock
    raise ChannelInvalidStateError("%r closed" % self)
aiormq.exceptions.ChannelInvalidStateError: <Channel: "4"> closed

```

I was wondering any of you could help me of what could be the sources, sometimes i need to do a verdi daemon restart for the workchain to run since it can remain in the created status for a while. I am using aiida-core 2.3.1

---

<div class="post-metadata">

**Author:** ![mbercx](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/mbercx/32/345_2.png) [@mbercx](https://aiida.discourse.group/u/mbercx)\
**Post date:** [November 10, 2023, 12:01pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/2 "2023-11-10T12:01:56Z")

</div>

Hi @AndresOrtegaGuerrero! Are you submitting a lot processes within a short time frame? The error seems similar to the one raised in this issue:

> <https://github.com/aiidateam/aiida-core/issues/5899>
>
> \### Describe the bug
> 
> I am not sure if the daemons getting overwhelmed is the …reason behind it. But when I launch ~200 calculations together, they get excepted throwing a \`aiormq.exceptions.ChannelInvalidStateError: \<Channel: "4"\> closed\` error. It is similar to this \[issue \](https://github.com/aiidateam/aiida-core/issues/5800) that I opened previously.
> 
> Following is the full error report 
> 
> \`\`\` 
> (aiida168) tthakur@theospc31:~$ verdi process report 918283
> 2023-02-03 18:05:18 \[564668 | REPORT\]: \[918283|LinDiffusionWorkChain|setup\]: launching WorkChain with pinball coefficients defined by \<813418\>
> 2023-02-03 18:05:18 \[564669 | REPORT\]: \[918283|LinDiffusionWorkChain|run\_process\]: launching ReplayMDWorkChain\<918287\>
> 2023-02-03 18:05:21 \[564673 | REPORT\]: \[918287|ReplayMDWorkChain|run\_process\]: launching FlipperCalculation\<918302\> iteration #1
> 2023-02-09 03:27:35 \[572361 | REPORT\]: \[918287|ReplayMDWorkChain|report\_error\_handled\]: FlipperCalculation\<918302\> failed with exit status 312: The stdout output file was incomplete probably because the calculation got interrupted.
> 2023-02-09 03:27:35 \[572362 | REPORT\]: \[918287|ReplayMDWorkChain|report\_error\_handled\]: Action taken: Restarting calculation...
> 2023-02-09 03:27:35 \[572363 | REPORT\]: \[918287|ReplayMDWorkChain|inspect\_process\]: FlipperCalculation\<918302\> failed but a handler dealt with the problem, restarting
> 2023-02-09 03:27:35 \[572364 | REPORT\]: \[918287|ReplayMDWorkChain|check\_energy\_fluctuations\]: FlipperCalculation\<918302\> \[check\_energy\_fluctuations\]: Total energy fluctuations = 0.004842710000957595 \< threshold (uuid: dd4cc43d-935f-4abe-abc9-d2646a108927 (pk: 918276) value: 180.0) OK
> 2023-02-09 03:27:35 \[572365 | REPORT\]: \[918287|ReplayMDWorkChain|update\_mdsteps\]: FlipperCalculation\<918302\> ran 109190 steps (109190 done - 890810 to go).
> 2023-02-09 03:31:09 \[572501 | REPORT\]: \[918287|ReplayMDWorkChain|run\_process\]: launching FlipperCalculation\<924759\> iteration #2
> 2023-02-14 21:24:05 \[637726 | ERROR\]: Traceback (most recent call last):
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiida/manage/external/rmq.py", line 208, in \_continue
> result = await super().\_continue(communicator, pid, nowait, tag)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/process\_comms.py", line 607, in \_continue
> proc = cast('Process', saved\_state.unbundle(self.\_load\_context))
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/persistence.py", line 60, in unbundle
> return Savable.load(self, load\_context)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/persistence.py", line 452, in load
> return load\_cls.recreate\_from(saved\_state, load\_context)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/processes.py", line 239, in recreate\_from
> call\_with\_super\_check(process.init)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/base/utils.py", line 29, in call\_with\_super\_check
> wrapped(\*args, \*\*kwargs)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiida/engine/processes/process.py", line 159, in init
> super().init()
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/base/utils.py", line 16, in wrapper
> wrapped(self, \*args, \*\*kwargs)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/processes.py", line 298, in init
> identifier = self.\_communicator.add\_rpc\_subscriber(self.message\_receive, identifier=str(self.pid))
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/communications.py", line 141, in add\_rpc\_subscriber
> return self.\_communicator.add\_rpc\_subscriber(converted, identifier)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/kiwipy/rmq/threadcomms.py", line 215, in add\_rpc\_subscriber
> return self.\_loop\_scheduler.await\_(
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/pytray/aiothreads.py", line 159, in await\_
> return self.await\_submit(awaitable).result(timeout=self.task\_timeout)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/concurrent/futures/\_base.py", line 446, in result
> return self.\_\_get\_result()
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/concurrent/futures/\_base.py", line 391, in \_\_get\_result
> raise self.\_exception
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/pytray/aiothreads.py", line 36, in done
> result = done\_future.result()
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/futures.py", line 201, in result
> raise self.\_exception
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/tasks.py", line 258, in \_\_step
> result = coro.throw(exc)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/pytray/aiothreads.py", line 178, in proxy
> return await awaitable
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 482, in add\_rpc\_subscriber
> identifier = await msg\_subscriber.add\_rpc\_subscriber(subscriber, identifier)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 123, in add\_rpc\_subscriber
> rpc\_queue = await self.\_channel.declare\_queue(exclusive=True, arguments=self.\_rmq\_queue\_arguments)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aio\_pika/robust\_channel.py", line 173, in declare\_queue
> queue = await super().declare\_queue(
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aio\_pika/channel.py", line 325, in declare\_queue
> await queue.declare(timeout=timeout)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aio\_pika/queue.py", line 92, in declare
> self.declaration\_result = await asyncio.wait\_for(
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/tasks.py", line 442, in wait\_for
> return await fut
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiormq/channel.py", line 703, in queue\_declare
> return await self.rpc(
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiormq/base.py", line 168, in wrap
> return await self.create\_task(func(self, \*args, \*\*kwargs))
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiormq/base.py", line 25, in \_\_inner
> return await self.task
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/futures.py", line 284, in \_\_await\_\_
> yield self # This tells Task to wait for completion.
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/tasks.py", line 328, in \_\_wakeup
> future.result()
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/futures.py", line 201, in result
> raise self.\_exception
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/tasks.py", line 256, in \_\_step
> result = coro.send(None)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiormq/channel.py", line 121, in rpc
> raise ChannelInvalidStateError("writer is None")
> aiormq.exceptions.ChannelInvalidStateError: writer is None
> 
> 2023-02-14 21:24:06 \[637729 | ERROR\]: Traceback (most recent call last):
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiida/manage/external/rmq.py", line 208, in \_continue
> result = await super().\_continue(communicator, pid, nowait, tag)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/process\_comms.py", line 613, in \_continue
> await proc.step\_until\_terminated()
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/processes.py", line 1230, in step\_until\_terminated
> await self.step()
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/processes.py", line 1216, in step
> self.transition\_to(next\_state)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/base/state\_machine.py", line 335, in transition\_to
> self.transition\_failed(initial\_state\_label, label, \*sys.exc\_info()\[1:\])
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/base/state\_machine.py", line 351, in transition\_failed
> raise exception.with\_traceback(trace)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/base/state\_machine.py", line 320, in transition\_to
> self.\_enter\_next\_state(new\_state)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/base/state\_machine.py", line 386, in \_enter\_next\_state
> self.\_fire\_state\_event(StateEventHook.ENTERED\_STATE, last\_state)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/base/state\_machine.py", line 299, in \_fire\_state\_event
> callback(self, hook, state)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/processes.py", line 326, in \<lambda\>
> lambda \_s, \_h, from\_state: self.on\_entered(cast(Optional\[process\_states.State\], from\_state)),
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiida/engine/processes/process.py", line 390, in on\_entered
> super().on\_entered(from\_state)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/processes.py", line 700, in on\_entered
> self.\_communicator.broadcast\_send(body=None, sender=self.pid, subject=subject)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/plumpy/communications.py", line 175, in broadcast\_send
> return self.\_communicator.broadcast\_send(body, sender, subject, correlation\_id)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/kiwipy/rmq/threadcomms.py", line 258, in broadcast\_send
> result = self.\_loop\_scheduler.await\_(
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/pytray/aiothreads.py", line 159, in await\_
> return self.await\_submit(awaitable).result(timeout=self.task\_timeout)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/concurrent/futures/\_base.py", line 446, in result
> return self.\_\_get\_result()
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/concurrent/futures/\_base.py", line 391, in \_\_get\_result
> raise self.\_exception
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/pytray/aiothreads.py", line 36, in done
> result = done\_future.result()
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/futures.py", line 201, in result
> raise self.\_exception
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/tasks.py", line 256, in \_\_step
> result = coro.send(None)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/pytray/aiothreads.py", line 178, in proxy
> return await awaitable
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 522, in broadcast\_send
> result = await publisher.broadcast\_send(body, sender, subject, correlation\_id)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 66, in broadcast\_send
> return await self.publish(message, routing\_key=defaults.BROADCAST\_TOPIC, mandatory=False)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/kiwipy/rmq/messages.py", line 209, in publish
> result = await self.\_exchange.publish(message, routing\_key=routing\_key, mandatory=mandatory)
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aio\_pika/exchange.py", line 233, in publish
> return await asyncio.wait\_for(
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/asyncio/tasks.py", line 442, in wait\_for
> return await fut
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiormq/channel.py", line 508, in basic\_publish
> async with self.lock:
> File "/home/tthakur/miniconda3/envs/aiida168/lib/python3.9/site-packages/aiormq/channel.py", line 90, in lock
> raise ChannelInvalidStateError("%r closed" % self)
> aiormq.exceptions.ChannelInvalidStateError: \<Channel: "4"\> closed
> \`\`\`
> 
> \### Steps to reproduce
> 
> Steps to reproduce the behavior:
> 
> 1. Launch a lot of WorkChains (\>100) withing a few hours.
> 2. Wait for all the processes to leave the \`Created\` state and start running properly.
> 3. Optionally restart the daemon, but this is not strictly required.
> 4. Most WorkChains will except at this point.
> 
> \### Expected behavior
> 
> Nothing should happen, the workchains should run normally and not get excepted.
> 
> \### Your environment
> 
> \- Operating system: Ubuntu 22.04
> \- Python version: 3.9.13
> \- aiida-core version: 1.6.8
> \- RabbitMQ: 3.7.28
> \- PostgreSQL: 14.5
> 
> \### Additional context
> 
> For some reason I am seeing this issue much more frequently now. It used to happen once in a blue moon only if I restarted the daemons, but last time it happened I didn't do anything, the WCs just got excepted after I left the machine alone over the weekend. My environment is still the same, only my aiida database has become bigger.

---

<div class="post-metadata">

**Author:** ![AndresOrtegaGuerrero](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/andresortegaguerrero/32/187_2.png) [@AndresOrtegaGuerrero](https://aiida.discourse.group/u/AndresOrtegaGuerrero)\
**Post date:** [November 10, 2023, 12:03pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/3 "2023-11-10T12:03:43Z")

</div>

Hi @mbercx , i think so , I am running the DielectricWorkChain from aiida-vibroscopy

---

<div class="post-metadata">

**Author:** ![sphuber](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/sphuber/32/6_2.png) [@sphuber](https://aiida.discourse.group/u/sphuber)\
**Post date:** [November 10, 2023, 2:33pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/4 "2023-11-10T14:33:41Z")

</div>

Thanks for the report @AndresOrtegaGuerrero . I had a look at the code and this exception can be thrown when the connection is temporarily unavailable. The code is trying to broadcast a state change of the process to all subscribers. Although useful, it is not critical if this broadcast fails. Listening processes have a polling mechanism as backup. So we definitely shouldn’t let this exception topple the entire process. I looked in the code of `plumpy` and we already catch a `ConnectionClosed` exception. I opened a PR that simply also catches this other exception. This should hopefully fix the issue. @mbercx could you please review the PR? [Catch `ChannelInvalidStateError` in process state change by sphuber · Pull Request #278 · aiidateam/plumpy · GitHub](https://github.com/aiidateam/plumpy/pull/278)

Once merged in, I will make a release of plumpy (there is also another feature already that I want to release). I will ping here when it is available so you can update and hopefully the problem is gone.

---

<div class="post-metadata">

**Author:** ![sphuber](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/sphuber/32/6_2.png) [@sphuber](https://aiida.discourse.group/u/sphuber)\
**Post date:** [November 10, 2023, 2:51pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/5 "2023-11-10T14:51:54Z")

</div>

Ok, the release is done. You can run `pip install plumpy==0.21.9` and then `verdi daemon restart --reset`. Then try to run your workchain again. Please let us know here how it goes.

---

<div class="post-metadata">

**Author:** ![AndresOrtegaGuerrero](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/andresortegaguerrero/32/187_2.png) [@AndresOrtegaGuerrero](https://aiida.discourse.group/u/AndresOrtegaGuerrero)\
**Post date:** [November 10, 2023, 6:39pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/6 "2023-11-10T18:39:00Z")

</div>

> [@sphuber](#):
>
> verdi daemon restart --reset

Thank you @sphuber , i tried what you suggested,  
this is the ouput from the logs of the deamon

```auto
    res = await coro()
  File "/opt/conda/lib/python3.9/site-packages/plumpy/process_comms.py", line 536, in __call__
    return await self._continue(communicator, **task.get(TASK_ARGS, {}))
  File "/opt/conda/lib/python3.9/site-packages/aiida/manage/external/rmq/launcher.py", line 87, in _continue
    return future.result()
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 439, in result
    return self.__get_result()
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 391, in __get_result
    raise self._exception
aiida.engine.exceptions.PastException: aiormq.exceptions.ChannelInvalidStateError: writer is None

```

the workchain again was excepted , and some childs of the workchain as well  
and some have this

```auto
25187 3h ago PwCalculation ⏸ Waiting Pausing after failed transport task: submit_calculation failed 5 times consecutively

```

---

<div class="post-metadata">

**Author:** ![sphuber](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/sphuber/32/6_2.png) [@sphuber](https://aiida.discourse.group/u/sphuber)\
**Post date:** [November 12, 2023, 12:39pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/7 "2023-11-12T12:39:17Z")

</div>

What are the timestamps of that exception though? Could they not be exceptions from before the change? Is the calculation with pk 25187 one you launched _after_ having installed the new version and restarted the daemon, or did that one already exist? What is the output of `verdi process report 25187`?

---

<div class="post-metadata">

**Author:** ![AndresOrtegaGuerrero](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/andresortegaguerrero/32/187_2.png) [@AndresOrtegaGuerrero](https://aiida.discourse.group/u/AndresOrtegaGuerrero)\
**Post date:** [November 12, 2023, 4:45pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/8 "2023-11-12T16:45:32Z")

</div>

Hi @sphuber , this is the last part of the output

```auto
+-> WARNING at 2023-11-10 16:55:22.169019+00:00
 | maximum attempts 5 of calling do_upload, exceeded
+-> ERROR at 2023-11-10 16:59:34.390073+00:00
 | Traceback (most recent call last):
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/utils.py", line 187, in exponential_backoff_retry
 | result = await coro()
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/calcjobs/tasks.py", line 146, in do_submit
 | return execmanager.submit_calculation(node, transport)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/daemon/execmanager.py", line 379, in submit_calculation
 | result = scheduler.submit_from_script(workdir, submit_script_filename)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/schedulers/scheduler.py", line 410, in submit_from_script
 | return self._parse_submit_output(*result)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/schedulers/plugins/slurm.py", line 430, in _parse_submit_output
 | raise SchedulerError(f'Error during submission, retval={retval}\nstdout={stdout}\nstderr={stderr}')
 | aiida.schedulers.scheduler.SchedulerError: Error during submission, retval=1
 | stdout=
 | stderr=sbatch: error: Unable to open file _aiidasubmit.sh
 | 
+-> ERROR at 2023-11-10 17:00:48.109295+00:00
 | Traceback (most recent call last):
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/utils.py", line 187, in exponential_backoff_retry
 | result = await coro()
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/calcjobs/tasks.py", line 146, in do_submit
 | return execmanager.submit_calculation(node, transport)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/daemon/execmanager.py", line 379, in submit_calculation
 | result = scheduler.submit_from_script(workdir, submit_script_filename)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/schedulers/scheduler.py", line 410, in submit_from_script
 | return self._parse_submit_output(*result)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/schedulers/plugins/slurm.py", line 430, in _parse_submit_output
 | raise SchedulerError(f'Error during submission, retval={retval}\nstdout={stdout}\nstderr={stderr}')
 | aiida.schedulers.scheduler.SchedulerError: Error during submission, retval=1
 | stdout=
 | stderr=sbatch: error: Unable to open file _aiidasubmit.sh
 | 
+-> ERROR at 2023-11-10 17:01:57.093854+00:00
 | Traceback (most recent call last):
 | File "/opt/conda/lib/python3.9/site-packages/paramiko/transport.py", line 2271, in _check_banner
 | buf = self.packetizer.readline(timeout)
 | File "/opt/conda/lib/python3.9/site-packages/paramiko/packet.py", line 380, in readline
 | buf += self._read_timeout(timeout)
 | File "/opt/conda/lib/python3.9/site-packages/paramiko/packet.py", line 609, in _read_timeout
 | raise EOFError()
 | EOFError
 | 
 | During handling of the above exception, another exception occurred:
 | 
 | Traceback (most recent call last):
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/utils.py", line 187, in exponential_backoff_retry
 | result = await coro()
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/calcjobs/tasks.py", line 145, in do_submit
 | transport = await cancellable.with_interrupt(request)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/utils.py", line 94, in with_interrupt
 | result = await next(wait_iter)
 | File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 611, in _wait_for_one
 | return f.result() # May raise f.exception().
 | File "/opt/conda/lib/python3.9/asyncio/futures.py", line 201, in result
 | raise self._exception
 | File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 258, in __step
 | result = coro.throw(exc)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/calcjobs/manager.py", line 180, in updating
 | await self._update_job_info()
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/calcjobs/manager.py", line 132, in _update_job_info
 | self._jobs_cache = await self._get_jobs_from_scheduler()
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/calcjobs/manager.py", line 98, in _get_jobs_from_scheduler
 | transport = await request
 | File "/opt/conda/lib/python3.9/asyncio/futures.py", line 284, in __await__
 | yield self # This tells Task to wait for completion.
 | File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 328, in __wakeup
 | future.result()
 | File "/opt/conda/lib/python3.9/asyncio/futures.py", line 201, in result
 | raise self._exception
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/transports.py", line 89, in do_open
 | transport.open()
 | File "/opt/conda/lib/python3.9/site-packages/aiida/transports/plugins/ssh.py", line 498, in open
 | proxy_client.connect(proxy['host'], **proxy_connargs)
 | File "/opt/conda/lib/python3.9/site-packages/paramiko/client.py", line 421, in connect
 | t.start_client(timeout=timeout)
 | File "/opt/conda/lib/python3.9/site-packages/paramiko/transport.py", line 699, in start_client
 | raise e
 | File "/opt/conda/lib/python3.9/site-packages/paramiko/transport.py", line 2094, in run
 | self._check_banner()
 | File "/opt/conda/lib/python3.9/site-packages/paramiko/transport.py", line 2275, in _check_banner
 | raise SSHException(
 | paramiko.ssh_exception.SSHException: Error reading SSH protocol banner
+-> ERROR at 2023-11-10 17:03:22.373291+00:00
 | Traceback (most recent call last):
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/utils.py", line 187, in exponential_backoff_retry
 | result = await coro()
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/calcjobs/tasks.py", line 146, in do_submit
 | return execmanager.submit_calculation(node, transport)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/daemon/execmanager.py", line 379, in submit_calculation
 | result = scheduler.submit_from_script(workdir, submit_script_filename)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/schedulers/scheduler.py", line 410, in submit_from_script
 | return self._parse_submit_output(*result)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/schedulers/plugins/slurm.py", line 430, in _parse_submit_output
 | raise SchedulerError(f'Error during submission, retval={retval}\nstdout={stdout}\nstderr={stderr}')
 | aiida.schedulers.scheduler.SchedulerError: Error during submission, retval=1
 | stdout=
 | stderr=sbatch: error: Unable to open file _aiidasubmit.sh
 | 
+-> ERROR at 2023-11-10 17:07:08.424598+00:00
 | Traceback (most recent call last):
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/utils.py", line 187, in exponential_backoff_retry
 | result = await coro()
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/calcjobs/tasks.py", line 146, in do_submit
 | return execmanager.submit_calculation(node, transport)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/engine/daemon/execmanager.py", line 379, in submit_calculation
 | result = scheduler.submit_from_script(workdir, submit_script_filename)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/schedulers/scheduler.py", line 410, in submit_from_script
 | return self._parse_submit_output(*result)
 | File "/opt/conda/lib/python3.9/site-packages/aiida/schedulers/plugins/slurm.py", line 430, in _parse_submit_output
 | raise SchedulerError(f'Error during submission, retval={retval}\nstdout={stdout}\nstderr={stderr}')
 | aiida.schedulers.scheduler.SchedulerError: Error during submission, retval=1
 | stdout=
 | stderr=sbatch: error: Unable to open file _aiidasubmit.sh
 | 
+-> WARNING at 2023-11-10 17:07:08.434079+00:00
 | maximum attempts 5 of calling do_submit, exceeded

```

---

<div class="post-metadata">

**Author:** ![sphuber](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/sphuber/32/6_2.png) [@sphuber](https://aiida.discourse.group/u/sphuber)\
**Post date:** [November 13, 2023, 8:36am UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/9 "2023-11-13T08:36:28Z")

</div>

I think there is a problem here that is not related but may have to do with the updating of the code for an existing calculation. For this calculation, apparently the `_aiidasubmit.sh` script was not created. There is no way to recover from this, and we should simply kill and delete this calculation.

Could you simply please try to launch a new calculation/workchain and see if that works?

---

<div class="post-metadata">

**Author:** ![AndresOrtegaGuerrero](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/andresortegaguerrero/32/187_2.png) [@AndresOrtegaGuerrero](https://aiida.discourse.group/u/AndresOrtegaGuerrero)\
**Post date:** [November 13, 2023, 10:18am UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/10 "2023-11-13T10:18:46Z")

</div>

Hi @sphuber , I just tried again the workchain got an except

```auto
2023-11-13 09:59:13 [7059 | WARNING]: Process<25288>: no connection available to broadcast state change from running to excepted
2023-11-13 09:59:13 [7060 | ERROR]: Traceback (most recent call last):
  File "/opt/conda/lib/python3.9/site-packages/plumpy/processes.py", line 888, in on_close
    cleanup()
  File "/opt/conda/lib/python3.9/site-packages/plumpy/communications.py", line 144, in remove_rpc_subscriber
    return self._communicator.remove_rpc_subscriber(identifier)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/threadcomms.py", line 221, in remove_rpc_subscriber
    return self._loop_scheduler.await_(self._communicator.remove_rpc_subscriber(identifier))
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 164, in await_
    return self.await_submit(awaitable).result(timeout=self.task_timeout)
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 446, in result
    return self.__get_result()
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 391, in __get_result
    raise self._exception
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 258, in __step
    result = coro.throw(exc)
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 178, in coro
    res = await awaitable
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 487, in remove_rpc_subscriber
    await msg_subscriber.remove_rpc_subscriber(identifier)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 140, in remove_rpc_subscriber
    await rpc_queue.cancel(identifier)
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/robust_queue.py", line 140, in cancel
    result = await super().cancel(consumer_tag, timeout, nowait)
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/queue.py", line 264, in cancel
    return await asyncio.wait_for(
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 442, in wait_for
    return await fut
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 395, in basic_cancel
    return await self.rpc(
  File "/opt/conda/lib/python3.9/site-packages/aiormq/base.py", line 168, in wrap
    return await self.create_task(func(self, *args, **kwargs))
  File "/opt/conda/lib/python3.9/site-packages/aiormq/base.py", line 25, in __inner
    return await self.task
  File "/opt/conda/lib/python3.9/asyncio/futures.py", line 284, in __await__
    yield self # This tells Task to wait for completion.
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 328, in __wakeup
    future.result()
  File "/opt/conda/lib/python3.9/asyncio/futures.py", line 201, in result
    raise self._exception
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 256, in __step
    result = coro.send(None)
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 121, in rpc
    raise ChannelInvalidStateError("writer is None")
aiormq.exceptions.ChannelInvalidStateError: writer is None

2023-11-13 09:59:23 [7061 | REPORT]: [25288|DielectricWorkChain|on_terminated]: cleaned remote folders of calculations: 25301 25432 25653

```

I will explore the options within the workchains to delay the submission of scf, and also to run serial to see if the issue is resolved

---

<div class="post-metadata">

**Author:** ![sphuber](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/sphuber/32/6_2.png) [@sphuber](https://aiida.discourse.group/u/sphuber)\
**Post date:** [November 13, 2023, 10:37am UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/11 "2023-11-13T10:37:12Z")

</div>

Hmm. Did you launch just a single workchain? How many processes does that spawn.

There seems to be something really off with your connection to RabbitMQ. Where is RabbitMQ running? Can you report the output of `verdi status`?

---

<div class="post-metadata">

**Author:** ![AndresOrtegaGuerrero](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/andresortegaguerrero/32/187_2.png) [@AndresOrtegaGuerrero](https://aiida.discourse.group/u/AndresOrtegaGuerrero)\
**Post date:** [November 13, 2023, 12:05pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/12 "2023-11-13T12:05:14Z")

</div>

Hi @sphuber ,

this my status, no issue on RabbitMQ

```auto
 ✔ version: AiiDA v2.3.1
 ✔ config: /home/jovyan/.aiida
 ✔ profile: default
 ✔ storage: Storage for 'default' [open] @ postgresql://aiida:***@localhost:5432/aiida_db / DiskObjectStoreRepository: 406d090665c941ef98807cc2109af721 | /home/jovyan/.aiida/repository/default/container
 ✔ rabbitmq: Connected to RabbitMQ v3.9.13 as amqp://guest:guest@127.0.0.1:5672?heartbeat=600
 ✔ daemon: Daemon is running with PID 253

```

The nature of the WorkChain usually launch several WorkChains

```auto
HarmonicWorkChain<25278> Finished [401] [2:inspect_processes]
    ├── generate_preprocess_data<25279> Finished [0]
    ├── PhononWorkChain<25284> Finished [0] [7:if_(should_run_phonopy)(1:inspect_phonopy)]
    │ ├── generate_preprocess_data<25289> Finished [0]
    │ ├── get_supercell<25296> Finished [0]
    │ ├── create_kpoints_from_distance<25298> Finished [0]
    │ ├── PwBaseWorkChain<25304> Finished [0] [3:results]
    │ │ └── PwCalculation<25307> Finished [0]
    │ ├── get_supercells_with_displacements<25321> Finished [0]
    │ ├── PwBaseWorkChain<25359> Finished [0] [3:results]
    │ │ └── PwCalculation<25435> Finished [0]
    │ ├── PwBaseWorkChain<25361> Finished [0] [3:results]
    │ │ └── PwCalculation<25438> Finished [0]
    │ ├── PwBaseWorkChain<25363> Finished [0] [3:results]
    │ │ └── PwCalculation<25441> Finished [0]
    │ ├── PwBaseWorkChain<25365> Finished [0] [3:results]
    │ │ └── PwCalculation<25444> Finished [0]
    │ ├── PwBaseWorkChain<25367> Finished [0] [3:results]
    │ │ └── PwCalculation<25447> Finished [0]
    │ ├── PwBaseWorkChain<25369> Finished [0] [3:results]
    │ │ └── PwCalculation<25450> Finished [0]
    │ ├── PwBaseWorkChain<25371> Finished [0] [3:results]
    │ │ └── PwCalculation<25453> Finished [0]
    │ ├── PwBaseWorkChain<25373> Finished [0] [3:results]
    │ │ └── PwCalculation<25456> Finished [0]
    │ ├── PwBaseWorkChain<25375> Finished [0] [3:results]
    │ │ └── PwCalculation<25459> Finished [0]
    │ ├── PwBaseWorkChain<25377> Finished [0] [3:results]
    │ │ └── PwCalculation<25462> Finished [0]
    │ ├── PwBaseWorkChain<25379> Finished [0] [3:results]
    │ │ └── PwCalculation<25465> Finished [0]
    │ ├── PwBaseWorkChain<25381> Finished [0] [3:results]
    │ │ └── PwCalculation<25468> Finished [0]
    │ ├── PwBaseWorkChain<25383> Finished [0] [3:results]
    │ │ └── PwCalculation<25471> Finished [0]
    │ ├── PwBaseWorkChain<25385> Finished [0] [3:results]
    │ │ └── PwCalculation<25474> Finished [0]
    │ ├── PwBaseWorkChain<25387> Finished [0] [3:results]
    │ │ └── PwCalculation<25477> Finished [0]
    │ ├── PwBaseWorkChain<25389> Finished [0] [3:results]
    │ │ └── PwCalculation<25480> Finished [0]
    │ ├── PwBaseWorkChain<25391> Finished [0] [3:results]
    │ │ └── PwCalculation<25483> Finished [0]
    │ ├── PwBaseWorkChain<25393> Finished [0] [3:results]
    │ │ └── PwCalculation<25486> Finished [0]
    │ ├── PwBaseWorkChain<25395> Finished [0] [3:results]
    │ │ └── PwCalculation<25489> Finished [0]
    │ ├── PwBaseWorkChain<25397> Finished [0] [3:results]
    │ │ └── PwCalculation<25492> Finished [0]
    │ ├── PwBaseWorkChain<25399> Finished [0] [3:results]
    │ │ └── PwCalculation<25495> Finished [0]
    │ ├── PwBaseWorkChain<25401> Finished [0] [3:results]
    │ │ └── PwCalculation<25498> Finished [0]
    │ ├── PwBaseWorkChain<25403> Finished [0] [3:results]
    │ │ └── PwCalculation<25501> Finished [0]
    │ ├── PwBaseWorkChain<25405> Finished [0] [3:results]
    │ │ └── PwCalculation<25504> Finished [0]
    │ ├── PwBaseWorkChain<25407> Finished [0] [3:results]
    │ │ └── PwCalculation<25507> Finished [0]
    │ ├── PwBaseWorkChain<25409> Finished [0] [3:results]
    │ │ └── PwCalculation<25510> Finished [0]
    │ ├── PwBaseWorkChain<25411> Finished [0] [3:results]
    │ │ └── PwCalculation<25513> Finished [0]
    │ ├── PwBaseWorkChain<25413> Finished [0] [3:results]
    │ │ └── PwCalculation<25516> Finished [0]
    │ ├── PwBaseWorkChain<25415> Finished [0] [3:results]
    │ │ └── PwCalculation<25519> Finished [0]
    │ ├── PwBaseWorkChain<25417> Finished [0] [3:results]
    │ │ └── PwCalculation<25522> Finished [0]
    │ ├── PwBaseWorkChain<25419> Finished [0] [3:results]
    │ │ └── PwCalculation<25525> Finished [0]
    │ ├── PwBaseWorkChain<25421> Finished [0] [3:results]
    │ │ └── PwCalculation<25528> Finished [0]
    │ ├── PwBaseWorkChain<25423> Finished [0] [3:results]
    │ │ └── PwCalculation<25531> Finished [0]
    │ ├── PwBaseWorkChain<25425> Finished [0] [3:results]
    │ │ └── PwCalculation<25534> Finished [0]
    │ ├── PwBaseWorkChain<25427> Finished [0] [3:results]
    │ │ └── PwCalculation<25537> Finished [0]
    │ ├── PwBaseWorkChain<25429> Finished [0] [3:results]
    │ │ └── PwCalculation<25540> Finished [0]
    │ ├── generate_phonopy_data<25740> Finished [0]
    │ └── PhonopyCalculation<25742> Finished [0]
    └── DielectricWorkChain<25288> Excepted [8:while_(should_run_electric_field_scfs)]
        ├── create_kpoints_from_distance<25290> Finished [0]
        ├── PwBaseWorkChain<25295> Finished [0] [3:results]
        │ └── PwCalculation<25301> Finished [0]
        ├── PwBaseWorkChain<25320> Finished [0] [3:results]
        │ └── PwCalculation<25432> Finished [0]
        ├── compute_critical_electric_field<25642> Finished [0]
        ├── get_accuracy_from_critical_field<25644> Finished [0]
        ├── get_electric_field_step<25646> Finished [0]
        ├── PwBaseWorkChain<25650> Finished [0] [3:results]
        │ └── PwCalculation<25653> Finished [0]
        └── PwBaseWorkChain<25752> Created

```

---

<div class="post-metadata">

**Author:** ![Xing](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/xing/32/10_2.png) [@Xing](https://aiida.discourse.group/u/Xing)\
**Post date:** [November 13, 2023, 1:04pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/13 "2023-11-13T13:04:39Z")

</div>

Hi @AndresOrtegaGuerrero , please use RabbitMQ version \< 3.8.15.  
Please check this [issue](https://github.com/aiidateam/aiida-core/issues/5105#:~:text=I%27ve%20just%20had%20the%20issue%20with%20the%20channel%20closed%20error%2C%20while%20running%20the%20RabbitMQ%20v3.9.13.), and the [doc](https://aiida.readthedocs.io/projects/aiida-core/en/latest/intro/troubleshooting.html#:~:text=RabbitMQ%20incompatibility).

---

<div class="post-metadata">

**Author:** ![sphuber](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/sphuber/32/6_2.png) [@sphuber](https://aiida.discourse.group/u/sphuber)\
**Post date:** [November 13, 2023, 7:04pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/14 "2023-11-13T19:04:31Z")

</div>

@Xing is right that in most cases it is preferred to use RabbitMQ \< 3.8.15, unless the server is configured as mentioned in AiiDA’s docs. Then it is fine to use more modern versions. But regardless of the version, the problem we see here should not be related to the RabbitMQ version.

The last exception you posted is similar in nature to the one I fixed, but just in a different place. I will try to patch that as well and try and find other locations where this could occur, hypothetically. But it is a bit difficult to spot them without actually running into the problem.

---

<div class="post-metadata">

**Author:** ![AndresOrtegaGuerrero](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/andresortegaguerrero/32/187_2.png) [@AndresOrtegaGuerrero](https://aiida.discourse.group/u/AndresOrtegaGuerrero)\
**Post date:** [November 14, 2023, 2:57pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/15 "2023-11-14T14:57:08Z")

</div>

@sphuber I try again with no luck

```auto
2023-11-14 14:27:25 [7279 | WARNING]: Process<26839>: no connection available to broadcast state change from running to excepted
2023-11-14 14:27:25 [7280 | ERROR]: Traceback (most recent call last):
  File "/opt/conda/lib/python3.9/site-packages/plumpy/processes.py", line 888, in on_close
    cleanup()
  File "/opt/conda/lib/python3.9/site-packages/plumpy/communications.py", line 144, in remove_rpc_subscriber
    return self._communicator.remove_rpc_subscriber(identifier)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/threadcomms.py", line 221, in remove_rpc_subscriber
    return self._loop_scheduler.await_(self._communicator.remove_rpc_subscriber(identifier))
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 164, in await_
    return self.await_submit(awaitable).result(timeout=self.task_timeout)
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 446, in result
    return self.__get_result()
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 391, in __get_result
    raise self._exception
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 258, in __step
    result = coro.throw(exc)
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 178, in coro
    res = await awaitable
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 487, in remove_rpc_subscriber
    await msg_subscriber.remove_rpc_subscriber(identifier)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 140, in remove_rpc_subscriber
    await rpc_queue.cancel(identifier)
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/robust_queue.py", line 140, in cancel
    result = await super().cancel(consumer_tag, timeout, nowait)
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/queue.py", line 264, in cancel
    return await asyncio.wait_for(
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 442, in wait_for
    return await fut
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 395, in basic_cancel
    return await self.rpc(
  File "/opt/conda/lib/python3.9/site-packages/aiormq/base.py", line 168, in wrap
    return await self.create_task(func(self, *args, **kwargs))
  File "/opt/conda/lib/python3.9/site-packages/aiormq/base.py", line 25, in __inner
    return await self.task
  File "/opt/conda/lib/python3.9/asyncio/futures.py", line 284, in __await__
    yield self # This tells Task to wait for completion.
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 328, in __wakeup
    future.result()
  File "/opt/conda/lib/python3.9/asyncio/futures.py", line 201, in result
    raise self._exception
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 256, in __step
    result = coro.send(None)
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 121, in rpc
    raise ChannelInvalidStateError("writer is None")
aiormq.exceptions.ChannelInvalidStateError: writer is None

```

should i downgrade the RabbitMQ ?

---

<div class="post-metadata">

**Author:** ![sphuber](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/sphuber/32/6_2.png) [@sphuber](https://aiida.discourse.group/u/sphuber)\
**Post date:** [November 14, 2023, 9:16pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/16 "2023-11-14T21:16:04Z")

</div>

> [@AndresOrtegaGuerrero](#):
>
> should i downgrade the RabbitMQ ?

As I mentioned before, I don’t think the version of RabbitMQ is the cause here. The problem is that at the end of the process, the `cleanup` method is called. This tries to remove itself as an rpc subscriber, which it needs to do over the connection to RabbitMQ. This connection fails, causing the exception. The thing I _don’t_ understand is that the `cleanup()` call is wrapped in a try-except block (see here [plumpy/src/plumpy/processes.py at ff5770f55da9974b693bed5e731c211ee47c39cd · aiidateam/plumpy · GitHub](https://github.com/aiidateam/plumpy/blob/ff5770f55da9974b693bed5e731c211ee47c39cd/src/plumpy/processes.py#L888) ). It should catch the exception and simply log it, but not let the process fall over as is happening in your case. I am not sure why this is not being caught. If I can figure that out, we can fix the problem, but I need more time to look at it.

---

<div class="post-metadata">

**Author:** ![sphuber](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/sphuber/32/6_2.png) [@sphuber](https://aiida.discourse.group/u/sphuber)\
**Post date:** [November 14, 2023, 9:22pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/17 "2023-11-14T21:22:05Z")

</div>

There is 1 other thing you can try in the meantime. I have an open branch that updates `aio-pika` and `aiormq` (the libraries that are used to connect to RabbitMQ) to newer versions that are supposed to be more stable. It would be great if you could give that branch a go.

Doing so is simple. You just need to do the following:

- activate your virtual environment (conda or other)
- `git clone https://github.com/sphuber/aiida-core`
- `cd aida-core`
- `git checkout fix/bump-engine-dependencies`
- `pip install -e .`
- `verdi daemon restart --reset`

Now you can relaunch a new workchain and see if that helps. It would be of great help to see if this reduces the problem since you seem to have a case that is so reproducible.

---

<div class="post-metadata">

**Author:** ![AndresOrtegaGuerrero](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/andresortegaguerrero/32/187_2.png) [@AndresOrtegaGuerrero](https://aiida.discourse.group/u/AndresOrtegaGuerrero)\
**Post date:** [November 15, 2023, 1:31pm UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/18 "2023-11-15T13:31:31Z")

</div>

@sphuber thanks for the help, so currently I am working in the aiidalab-launch container, and since i am working with the Qe App (latest version) the branch of the aiida-core you shared is incompatible. I am going to run locally to see if maybe it works, one thing i noticed is that is only in an specific workchain were i get the except `(https://github.com/bastonero/aiida-vibroscopy/blob/main/src/aiida_vibroscopy/workflows/dielectric/base.py)`

I will test locally if the workchain doesnt fail, in case is an issue within the container, though in the container i can run the workchain with pw.localhost with no problem

---

<div class="post-metadata">

**Author:** ![AndresOrtegaGuerrero](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/andresortegaguerrero/32/187_2.png) [@AndresOrtegaGuerrero](https://aiida.discourse.group/u/AndresOrtegaGuerrero)\
**Post date:** [November 17, 2023, 10:55am UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/19 "2023-11-17T10:55:54Z")

</div>

@sphuber , so I did two test, one is in my personal computer and submitting the job in verdi shell, and the other one using the aiidalab-container, in both i am using aiida-core AiiDA v2.4… In my personal computer the workchain completes, but in the container i have the same issue

```auto
"023-11-17 10:05:11 [10042 | REPORT]: [33693|PwBaseWorkChain|on_terminated]: remote folders will not be cleaned
2023-11-17 10:06:11 [10043 | WARNING]: Process<33660>: no connection available to broadcast state change from waiting to running
2023-11-17 10:06:11 [10044 | WARNING]: Process<33660>: no connection available to broadcast state change from running to running
2023-11-17 10:06:11 [10045 | REPORT]: [33660|DielectricWorkChain|on_except]: Traceback (most recent call last):
  File "/opt/conda/lib/python3.9/site-packages/plumpy/process_states.py", line 228, in execute
    result = self.run_fn(*self.args, **self.kwargs)
  File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/workchains/workchain.py", line 314, in _do_step
    finished, stepper_result = self._stepper.step()
  File "/opt/conda/lib/python3.9/site-packages/plumpy/workchains.py", line 295, in step
    finished, result = self._child_stepper.step()
  File "/opt/conda/lib/python3.9/site-packages/plumpy/workchains.py", line 538, in step
    finished, result = self._child_stepper.step()
  File "/opt/conda/lib/python3.9/site-packages/plumpy/workchains.py", line 295, in step
    finished, result = self._child_stepper.step()
  File "/opt/conda/lib/python3.9/site-packages/plumpy/workchains.py", line 246, in step
    return True, self._fn(self._workchain)
  File "/home/jovyan/aiida-vibroscopy/src/aiida_vibroscopy/workflows/dielectric/base.py", line 690, in run_electric_field_scfs
    node = self.submit(PwBaseWorkChain, **inputs)
  File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/process.py", line 544, in submit
    return self.runner.submit(process, **kwargs)
  File "/opt/conda/lib/python3.9/site-packages/aiida/engine/runners.py", line 183, in submit
    process_inited = self.instantiate_process(process, **inputs)
  File "/opt/conda/lib/python3.9/site-packages/aiida/engine/runners.py", line 169, in instantiate_process
    return instantiate_process(self, process, **inputs)
  File "/opt/conda/lib/python3.9/site-packages/aiida/engine/utils.py", line 64, in instantiate_process
    process = process_class(runner=runner, inputs=inputs)
  File "/opt/conda/lib/python3.9/site-packages/plumpy/base/state_machine.py", line 195, in __call__
    call_with_super_check(inst.init)
  File "/opt/conda/lib/python3.9/site-packages/plumpy/base/utils.py", line 31, in call_with_super_check
    wrapped(*args, **kwargs)
  File "/opt/conda/lib/python3.9/site-packages/aiida/engine/processes/process.py", line 188, in init
    super().init()
  File "/opt/conda/lib/python3.9/site-packages/plumpy/base/utils.py", line 16, in wrapper
    wrapped(self, *args, **kwargs)
  File "/opt/conda/lib/python3.9/site-packages/plumpy/processes.py", line 309, in init
    identifier = self._communicator.add_rpc_subscriber(self.message_receive, identifier=str(self.pid))
  File "/opt/conda/lib/python3.9/site-packages/plumpy/communications.py", line 141, in add_rpc_subscriber
    return self._communicator.add_rpc_subscriber(converted, identifier)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/threadcomms.py", line 215, in add_rpc_subscriber
    return self._loop_scheduler.await_(
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 164, in await_
    return self.await_submit(awaitable).result(timeout=self.task_timeout)
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 446, in result
    return self.__get_result()
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 391, in __get_result
    raise self._exception
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 258, in __step
    result = coro.throw(exc)
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 178, in coro
    res = await awaitable
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 483, in add_rpc_subscriber
    identifier = await msg_subscriber.add_rpc_subscriber(subscriber, identifier)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 124, in add_rpc_subscriber
    rpc_queue = await self._channel.declare_queue(exclusive=True, arguments=self._rmq_queue_arguments)
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/robust_channel.py", line 173, in declare_queue
    queue = await super().declare_queue(
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/channel.py", line 325, in declare_queue
    await queue.declare(timeout=timeout)
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/queue.py", line 92, in declare
    self.declaration_result = await asyncio.wait_for(
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 442, in wait_for
    return await fut
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 703, in queue_declare
    return await self.rpc(
  File "/opt/conda/lib/python3.9/site-packages/aiormq/base.py", line 168, in wrap
    return await self.create_task(func(self, *args, **kwargs))
  File "/opt/conda/lib/python3.9/site-packages/aiormq/base.py", line 25, in __inner
    return await self.task
  File "/opt/conda/lib/python3.9/asyncio/futures.py", line 284, in __await__
    yield self # This tells Task to wait for completion.
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 328, in __wakeup
    future.result()
  File "/opt/conda/lib/python3.9/asyncio/futures.py", line 201, in result
    raise self._exception
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 256, in __step
    result = coro.send(None)
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 121, in rpc
    raise ChannelInvalidStateError("writer is None")
aiormq.exceptions.ChannelInvalidStateError: writer is None

2023-11-17 10:06:11 [10046 | WARNING]: Process<33660>: no connection available to broadcast state change from running to excepted
2023-11-17 10:06:11 [10047 | ERROR]: Traceback (most recent call last):
  File "/opt/conda/lib/python3.9/site-packages/plumpy/processes.py", line 888, in on_close
    cleanup()
  File "/opt/conda/lib/python3.9/site-packages/plumpy/communications.py", line 144, in remove_rpc_subscriber
    return self._communicator.remove_rpc_subscriber(identifier)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/threadcomms.py", line 221, in remove_rpc_subscriber
    return self._loop_scheduler.await_(self._communicator.remove_rpc_subscriber(identifier))
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 164, in await_
    return self.await_submit(awaitable).result(timeout=self.task_timeout)
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 446, in result
    return self.__get_result()
  File "/opt/conda/lib/python3.9/concurrent/futures/_base.py", line 391, in __get_result
    raise self._exception
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 258, in __step
    result = coro.throw(exc)
  File "/opt/conda/lib/python3.9/site-packages/pytray/aiothreads.py", line 178, in coro
    res = await awaitable
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 488, in remove_rpc_subscriber
    await msg_subscriber.remove_rpc_subscriber(identifier)
  File "/opt/conda/lib/python3.9/site-packages/kiwipy/rmq/communicator.py", line 141, in remove_rpc_subscriber
    await rpc_queue.cancel(identifier)
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/robust_queue.py", line 140, in cancel
    result = await super().cancel(consumer_tag, timeout, nowait)
  File "/opt/conda/lib/python3.9/site-packages/aio_pika/queue.py", line 264, in cancel
    return await asyncio.wait_for(
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 442, in wait_for
    return await fut
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 395, in basic_cancel
    return await self.rpc(
  File "/opt/conda/lib/python3.9/site-packages/aiormq/base.py", line 168, in wrap
    return await self.create_task(func(self, *args, **kwargs))
  File "/opt/conda/lib/python3.9/site-packages/aiormq/base.py", line 25, in __inner
    return await self.task
  File "/opt/conda/lib/python3.9/asyncio/futures.py", line 284, in __await__
    yield self # This tells Task to wait for completion.
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 328, in __wakeup
    future.result()
  File "/opt/conda/lib/python3.9/asyncio/futures.py", line 201, in result
    raise self._exception
  File "/opt/conda/lib/python3.9/asyncio/tasks.py", line 256, in __step
    result = coro.send(None)
  File "/opt/conda/lib/python3.9/site-packages/aiormq/channel.py", line 121, in rpc
    raise ChannelInvalidStateError("writer is None")
aiormq.exceptions.ChannelInvalidStateError: writer is None

2023-11-17 10:06:22 [10048 | REPORT]: [33660|DielectricWorkChain|on_terminated]: cleaned remote folders of calculations: 33668 33679 33696"

```

---

<div class="post-metadata">

**Author:** ![jusong.yu](https://yyz2.discourse-cdn.com/free1/user_avatar/aiida.discourse.group/jusong.yu/32/3_2.png) [@jusong.yu](https://aiida.discourse.group/u/jusong.yu)\
**Post date:** [November 17, 2023, 11:27am UTC](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152/20 "2023-11-17T11:27:37Z")

</div>

Hi @AndresOrtegaGuerrero, the only difference I can see between running locally and inside the container is the version of rabbitmq, maybe you can did a test on switch the version of it. If I guess correctly, you have `3.9.13` means you use the arm64 image. Do you have an amd64 machine to test, since it will use a recommended version of rmq which is `3.8.15`.  
Or you can use rmq `3.9.13` in your local machine to see if the issue can be reproduced.

[Next page](https://aiida.discourse.group/t/workchain-exception-inside-aiidalab-container/152.md?page=2)
