# Too many open files

**URL:** <https://deeptalk.lambda.ai/t/too-many-open-files/3795>\
**Category:** Technical Help\
**Created:** [June 18, 2023, 5:54pm UTC](https://deeptalk.lambda.ai/t/too-many-open-files/3795 "2023-06-18T17:54:33Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![pathos00011](https://avatars.discourse-cdn.com/v4/letter/p/ee59a6/32.png) [@pathos00011](https://deeptalk.lambda.ai/u/pathos00011)\
**Post date:** [June 18, 2023, 5:54pm UTC](https://deeptalk.lambda.ai/t/too-many-open-files/3795/1 "2023-06-18T17:54:33Z")

</div>

I seem to only get this while using an A100 with an attached filesystem. I have not done anything unusual that I know of. ulimit -n shows 1048576.

I have not run into this with any other instances before, the A6000 I was using previously with the filesystem had never ran into this problem.

Any ideas what might cause this?

---

<div class="post-metadata">

**Author:** ![mpapili](https://avatars.discourse-cdn.com/v4/letter/m/3e96dc/32.png) [@mpapili](https://deeptalk.lambda.ai/u/mpapili)\
**Post date:** [June 19, 2023, 10:13pm UTC](https://deeptalk.lambda.ai/t/too-many-open-files/3795/2 "2023-06-19T22:13:11Z")

</div>

The next time this happens can you capture the output of:

```auto
# print [num-handlers] [PID] for top 10 processes 
lsof -n | awk '{print $2}' | sort | uniq -c | sort -rn | head -n 10

```

and then the names of the processes attached to each PID (should be the second arg) with:

```auto
ps aux | grep -i <PID>

```

-once you found your likely culprit you can do `lsof -p <PID>` and we’ll be able to see what these files are exactly.

---

<div class="post-metadata">

**Author:** ![yapee23](https://avatars.discourse-cdn.com/v4/letter/y/65b543/32.png) [@yapee23](https://deeptalk.lambda.ai/u/yapee23)\
**Post date:** [July 28, 2023, 8:11am UTC](https://deeptalk.lambda.ai/t/too-many-open-files/3795/3 "2023-07-28T08:11:34Z")

</div>

Hi, did you find any solution for this? I am running into the same error and the instance fails every time I get close to finishing epoch 1. It fails before even finishing the first epoch during model training with this exact error “Too many open files.”

I am using the Lambda labs Persistent Storage beta as well.

---

<div class="post-metadata">

**Author:** ![yanos](https://avatars.discourse-cdn.com/v4/letter/y/a183cd/32.png) [@yanos](https://deeptalk.lambda.ai/u/yanos)\
**Post date:** [August 1, 2023, 12:11am UTC](https://deeptalk.lambda.ai/t/too-many-open-files/3795/4 "2023-08-01T00:11:21Z")

</div>

Hello!  
@pathos00011, @yapee23  
Can you send us an example of the command you run that causes this to happen, so we can reproduce this?

Best,  
Yanos

---

<div class="post-metadata">

**Author:** ![aigility](https://avatars.discourse-cdn.com/v4/letter/a/5fc32e/32.png) [@aigility](https://deeptalk.lambda.ai/u/aigility)\
**Post date:** [September 18, 2023, 9:52pm UTC](https://deeptalk.lambda.ai/t/too-many-open-files/3795/5 "2023-09-18T21:52:10Z")

</div>

I’ve run into this same error. Linking my current thread on it:

> [@"Too many open files" errors in many contexts](https://deeptalk.lambda.ai/t/too-many-open-files-errors-in-many-contexts/3963):
>
> I’m using a large dataset via the lambda cloud storage. This stores somewhere around 10 million files at around ~3TB of storage. I am doing some basic processing on these files, but can only process a certain number of file directories at time before running into “Too many open files” error. This error persists in all other applications and programs, and only goes away after restarting the instance. This has happened when using aws s3 sync, and doing some basic processing in some python progra…

Very easy to reproduce. Create a lambda cloud file system. Attach it to an instance. Put 2M dummy files into it (just doing that alone may have you run into the error). Then run  
`aws s3 sync ./local_file_dir/ s3://some-s3-file-bucket-test/`  
You’ll run into it before the folder is finished syncing.

I’m convinced this is a lambda cloud file system error, as `aws s3 sync` is engineered to scale and I have seen it in other code contexts multiple times.

Additionally, `lsof` seems unable to diagnose such open file handles. For example, the highest number of open file handles comes from jupyter-lab process when running `lsof -n | awk '{print $2}' | sort | uniq -c | sort -rn | head -n 10`, which I don’t think is the issue because (1) I’ve killed it and stopped it from spawning, which still doesn’t solve it, and (2) jupyter-lab shows a high number of open file handles, even when the instance is freshly restarted.

The only solution so far that’s worked for me is restarting the instance, but this makes lambda labs untenable for me for normal-scale machine learning work.

@mpapili @yanos

---

<div class="post-metadata">

**Author:** ![soheil](https://sea1.discourse-cdn.com/flex019/user_avatar/deeptalk.lambda.ai/soheil/32/638_2.png) [@soheil](https://deeptalk.lambda.ai/u/soheil)\
**Post date:** [October 5, 2023, 4:11am UTC](https://deeptalk.lambda.ai/t/too-many-open-files/3795/6 "2023-10-05T04:11:27Z")

</div>

I typically resolve this type of issue by either

Setting PAM Limits in `/etc/pam.d/common-session`

```auto
session required pam_limits.so

```

Or just a `ulimit -n unlimited`
