"ChildFailedError" fine tuning Video-Llama and Video-ChatGPT

The simple answer is you are running distrubuted, and parent process is telling you that one of the child processes failed. It is not clear for which reason, but it could be:

  • In sufficient resources for the child process (GPU, GPU memory, CPU, memory)
  • Perhaps if this is a remote host it could be different python script, data, libraries, etc. If this is distributed across nodes. (or the Anaconda environment you are running on for the ‘parent process’)

It is mentioned in with similar issues:

And of course enable trackback on the child process/worker:
https://pytorch.org/docs/stable/elastic/errors.html