# VLLM cluster using VLLM production stack

**URL:** <https://deeptalk.lambda.ai/t/vllm-cluster-using-vllm-production-stack/4728>\
**Category:** Uncategorized\
**Created:** [August 18, 2025, 7:15am UTC](https://deeptalk.lambda.ai/t/vllm-cluster-using-vllm-production-stack/4728 "2025-08-18T07:15:20Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![morrow\_maki](https://avatars.discourse-cdn.com/v4/letter/m/ecb155/32.png) [@morrow\_maki](https://deeptalk.lambda.ai/u/morrow_maki)\
**Post date:** [August 18, 2025, 7:15am UTC](https://deeptalk.lambda.ai/t/vllm-cluster-using-vllm-production-stack/4728/1 "2025-08-18T07:15:20Z")

</div>

Hello community,

I’m looking forward to renting multiple lambdalabs machines (with H100 gpus) and deploying them on a cluster using the [vllm production stack](https://github.com/vllm-project/production-stack).

My goal is to be able to manually purchase more machines from lambda labs and connect them to the cluster as my infrastructure get more load (→ more GPUs and more VLLM instances).

Any one has achieved something similar to this and would be willing to give me a few hints ? I followed the vllm production stack guide but I only managed to make it work with minikube (a single worker). I’d like to make it work for multi worker

Cheers !
