Skip to main content
VM Clusters is in Beta. It’s ready for production pilots. For large or time-critical deployments, talk to sales.
With VM Clusters (also called VMaaS, short for VM as a Service), you rent a dedicated cluster of high-performance GPU nodes — 1 to 8 nodes, each with 8 GPUs. GPUs from NVIDIA’s Hopper to Blackwell generations (H100, H200, B200, GB200, B300, GB300) are available — for options that aren’t self-service yet, talk to sales. VESSL Cloud manages the GPU supply and provisioning; you get full root access on every node, with no other customers sharing your hardware.
VM Clusters page listing active, provisioning, in review, and confirmed clusters with GPU, nodes, and contract dates

Why VM Clusters

  • Multi-node distributed training — InfiniBand is provisioned by default, giving you high-bandwidth, low-latency RDMA networking between nodes for data- or model-parallel workloads.
  • Full root control — swap GPU drivers, load custom kernel modules, and configure the OS however you need.
  • Near bare-metal GPU performance — each VM gets its GPUs directly using GPU passthrough, keeping virtualization overhead minimal.
  • Bring your own stack — install Slurm, Kubernetes, security agents, or your own deployment scripts. Root access on every node.

What you get

Boot disk data is not durable. If a node is reprovisioned (for example, after a hardware failure), all data on its boot disk is lost. Keep training data, checkpoints, and any other important files on NFS shared storage, which is mounted across every node in the cluster.Data is deleted at contract expiry. Once a contract ends, all VMs are stopped and all associated data — including boot disks and NFS storage — is permanently deleted. There’s no grace period. We send reminder emails before expiry so you can back up or extend the contract in time.

Core features

  • NFS shared storage — a shared filesystem mounted across all nodes in the cluster.
  • Firewall rules — self-service inbound access control for the whole cluster.
  • InfiniBand by default — for node-to-node communication in multi-node training. A single-node cluster doesn’t use InfiniBand; that node uses all of its resources exclusively.
  • Health checks and Metrics — per-node health verdict with GPU, InfiniBand, and ECC checks, plus time-series charts for GPU and system metrics. Refreshed every minute.
  • Reboot — restart a VM yourself when you need to.

Workspace vs VM Cluster

If you only need a containerized environment for development or single-node work, a workspace is faster to start and cheaper. Choose a VM cluster when you need kernel or driver control, multi-node training with InfiniBand, or a reserved cluster. See Workspace vs VM Cluster for a full comparison.
VM Clusters is an Organization-level resource. Only Admin members can create a new cluster.

Order a VM cluster

Walk through the order form and what happens after you submit.