# Pytorch DataParallel two gpu Hang

**URL:** <https://deeptalk.lambda.ai/t/pytorch-dataparallel-two-gpu-hang/1886>\
**Category:** Technical Help\
**Created:** [September 23, 2020, 8:14pm UTC](https://deeptalk.lambda.ai/t/pytorch-dataparallel-two-gpu-hang/1886 "2020-09-23T20:14:51Z")\
**Posts on this page:** 1\
**Showing post:** 1

<div class="post-metadata">

**Author:** ![maskjp](https://avatars.discourse-cdn.com/v4/letter/m/e9a140/32.png) [@maskjp](https://deeptalk.lambda.ai/u/maskjp)\
**Post date:** [September 23, 2020, 8:14pm UTC](https://deeptalk.lambda.ai/t/pytorch-dataparallel-two-gpu-hang/1886/1 "2020-09-23T20:14:51Z")

</div>

Hi, I have two lambda dual with two TITAN RTX. One pc works fine. But another pc freeze when using Dataparallel.  
I got the not working one last mouth.

I tried the solution in this

> **[\[solved\] DataParallel Multiple V100s Hang](https://discuss.pytorch.org/t/solved-dataparallel-multiple-v100s-hang/24089)**
>
> I’m having some trouble getting multi-gpu working across several V100s. Here’s code: BATCH\_SIZE = 800 import torch import torchvision import torchvision.transforms as transforms from pathlib import Path transform = transforms.Compose( ...

But got an error message.

And I also try to disable iommu following the following methods. But I found that the iommu is already disabled by default.

> <https://github.com/pytorch/pytorch/issues/1637#issuecomment-338268158>
>
> I'm having trouble getting multi-gpu via \`DataParallel\` across two Tesla K80 GPU…s. The code I'm using is a modification of the MNIST example:
> 
> \`\`\`
> import torch
> import torch.nn as nn
> import torch.nn.functional as F
> import torch.optim as optim
> from torchvision import datasets, transforms
> from torch.autograd import Variable
> from data\_parallel import DataParallel
> 
> train\_loader = torch.utils.data.DataLoader(
> datasets.MNIST('../data', train=True, download=True,
> transform=transforms.Compose(\[
> transforms.ToTensor(),
> transforms.Normalize((0.1307,), (0.3081,))
> \])),
> batch\_size=256, shuffle=True, num\_workers=2, pin\_memory=True)
> 
> class Net(nn.Module):
> def \_\_init\_\_(self):
> super(Net, self).\_\_init\_\_()
> self.conv1 = nn.Conv2d(1, 10, kernel\_size=5)
> self.conv2 = nn.Conv2d(10, 20, kernel\_size=5)
> self.fc1 = nn.Linear(320, 50)
> self.fc2 = nn.Linear(50, 10)
> 
> def forward(self, x):
> x = F.relu(F.max\_pool2d(self.conv1(x), 2))
> x = F.relu(F.max\_pool2d(self.conv2(x), 2))
> x = x.view(-1, 320)
> x = F.relu(self.fc1(x))
> x = self.fc2(x)
> return F.log\_softmax(x)
> 
> model = DataParallel(Net())
> model.cuda()
> 
> optimizer = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)
> criterion = nn.NLLLoss().cuda()
> 
> model.train()
> for batch\_idx, (data, target) in enumerate(train\_loader):
> input\_var = Variable(data.cuda())
> target\_var = Variable(target.cuda())
> 
> print('Getting model output')
> output = model(input\_var)
> print('Got model output')
> 
> loss = criterion(output, target\_var)
> optimizer.zero\_grad()
> loss.backward()
> optimizer.step()
> 
> print('Finished')
> \`\`\`
> 
> This doesn't throw an error, but hangs after it prints "Getting model output" and never returns. I traced this down to the \`parallel\_apply\` spawning threads that then never finish. The line that hangs is \[here\](https://github.com/pytorch/pytorch/blob/e50a1f19b3dc735f0710929b97b0af384aafe09b/torch/nn/parallel/parallel\_apply.py#L25) where the threads are spawned using both GPU 0 and GPU 1, but never finish.
> 
> This is only a problem when \`CUDA\_VISIBLE\_DEVICES=0,1\` as both GPU0 and GPU1 work perfectly well individually.
> 
> Before running this, \`nvidia-smi\` shows
> 
> \`\`\`
> +------------------------------------------------------+
> | NVIDIA-SMI 352.68 Driver Version: 352.68 |
> |-------------------------------+----------------------+----------------------+
> | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
> | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
> |===============================+======================+======================|
> | 0 Tesla K80 Off | 0000:06:00.0 Off | 0 |
> | N/A 40C P0 57W / 149W | 55MiB / 11519MiB | 0% Default |
> +-------------------------------+----------------------+----------------------+
> | 1 Tesla K80 Off | 0000:07:00.0 Off | 0 |
> | N/A 35C P0 76W / 149W | 55MiB / 11519MiB | 99% Default |
> +-------------------------------+----------------------+----------------------+
> 
> +-----------------------------------------------------------------------------+
> | Processes: GPU Memory |
> | GPU PID Type Process name Usage |
> |=============================================================================|
> | No running processes found |
> +-----------------------------------------------------------------------------+
> \`\`\`
> 
> after running (while it hangs), \`nvidia-smi\` gives
> 
> \`\`\`
> +------------------------------------------------------+
> | NVIDIA-SMI 352.68 Driver Version: 352.68 |
> |-------------------------------+----------------------+----------------------+
> | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
> | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
> |===============================+======================+======================|
> | 0 Tesla K80 Off | 0000:06:00.0 Off | 0 |
> | N/A 42C P0 69W / 149W | 251MiB / 11519MiB | 99% Default |
> +-------------------------------+----------------------+----------------------+
> | 1 Tesla K80 Off | 0000:07:00.0 Off | 0 |
> | N/A 36C P0 90W / 149W | 249MiB / 11519MiB | 99% Default |
> +-------------------------------+----------------------+----------------------+
> 
> +-----------------------------------------------------------------------------+
> | Processes: GPU Memory |
> | GPU PID Type Process name Usage |
> |=============================================================================|
> | 0 4785 C python 194MiB |
> | 1 4785 C python 192MiB |
> +-----------------------------------------------------------------------------+
> \`\`\`
> 
> and \`top\` shows the main python process and the two python subprocesses. Wondering if this could be something similar to #554.
> 
> 
> Using \[this\](https://github.com/tensorflow/models/blob/master/tutorials/image/cifar10/cifar10\_multi\_gpu\_train.py) TensorFlow example, I get linear speedup using multiple GPUs as I change \`CUDA\_VISIBLE\_DEVICES\` so multiple K80s should certainly be viable.

Are there any other solutions?

Thanks!

---

_[View the full topic](https://deeptalk.lambda.ai/t/pytorch-dataparallel-two-gpu-hang/1886)._
