Skip to main content
Version: v1.4.x

Create an AI Virtual Machine

Create an AI Virtual Machine to get a GPU or CPU virtual machine with its own guest operating system on a Compute Pool. For what an AI Virtual Machine is and how it relates to a Compute Pool, refer to AI Virtual Machines.

Create an AI Virtual Machine​

Prerequisites​

  • A Compute Pool that runs AI Virtual Machines, and that your scope can use. The Compute Pool picker lists only dedicated, single-cluster Compute Pools built from a VMO Profile Bundle, and does not offer shared-cluster Compute Pools or Compute Pools that run apps and models.
  • The name of a machine pool on that Compute Pool.
  • Permission to create AI Virtual Machines in your scope.

Configure the AI Virtual Machine​

  1. From the left main menu, open AI Virtual Machines and select Create AI Virtual Machine. If no Compute Pool is available to your scope, a dialog explains why and the workflow does not open.

  2. On General Information, set the name and metadata:

    FieldDescription
    NameThe name of the virtual machine. It must be 3 to 45 characters, start with a lowercase letter, end with a lowercase letter or number, and contain only lowercase letters, numbers, and hyphens. It cannot change after the virtual machine is created.
    DescriptionAn optional description.
    Compute PoolThe Compute Pool that runs the virtual machine.
    MetadataOptional annotations and labels. Expand Metadata and add entries under Annotations or Labels.

    Below the fields, the sharing section shares the virtual machine with tenants or projects at lower scopes. It is optional. A scope you share the virtual machine with can view it, connect to it, and clone it, but cannot stop or delete it. Refer to Sharing Resources.

  3. Select Next.

  4. On Template, choose a template and set the sizing for this virtual machine. Refer to Set Sizing Overrides.

  5. Select Next.

  6. On Customizations, review the manifest PaletteAI rendered from your choices, and optionally edit it. Refer to Review and Validate the Manifest.

  7. Select Next.

  8. On Review, confirm the Overview, then select Create AI Virtual Machine.

Set Sizing Overrides​

This section applies to the Template step of the UI workflow. The step starts from the template you select and overrides its allocation for this virtual machine only. An override you leave alone keeps the template's value.

The Machine Pool field on this step selects which machine pool runs the virtual machine. On a multi-node Compute Pool, select a worker pool. On a single-node Compute Pool, the control plane machine pool is the only one available.

Under Template Overrides, set the following fields:

FieldDescription
CPU CoresThe number of vCPUs.
MemoryThe guest memory, with a unit selected beside the value.
Root Disk SizeThe root disk size, with a unit selected beside the value.
Enable GPURequests GPU devices. Leave it off for a CPU-only virtual machine.
GPU VariantThe GPU model, for example NVIDIA H100 | 80 GB. The variant determines the per-GPU memory.
GPU CountThe number of GPUs.
GPU MemoryRead-only. Reports the per-GPU memory of the selected variant for confirmation.
Multi-Instance GPU ProfileThe MIG profile to carve. Appears only on a MIG-capable machine pool, and must be set together with Instances per GPU.
Instances per GPUThe number of MIG instances per GPU, from 1 to 7. Must be set together with Multi-Instance GPU Profile.

PaletteAI supports the units KB/k, KiB/ki, MB/m, MiB/mi, GB/g, GiB/gi, TB/t, and TiB/ti. A decimal unit such as GB counts in powers of 1,000, and a binary unit such as GiB counts in powers of 1,024. Every value needs a unit and must be greater than zero.

For CPU Cores, Memory, and the GPU fields, the workflow shows the maximum the selected machine pool allows, which is the lower of the Compute Pool's per-VM limits and the machine pool's capacity, and it does not accept a value above that maximum. Root Disk Size has no ceiling here, because disk does not count against Compute Pool capacity. Instances per GPU is bounded by the hardware maximum of 7, not by the Compute Pool's Max MIG Instances.

Those bounds are applied when you create the virtual machine in the console. Read the ceiling for each machine pool from the Compute Pool's status.vmCapacityByMachinePool before you apply a manifest.

Review and Validate the Manifest​

This section applies to the Customizations step of the UI workflow. The step shows the KubeVirt VirtualMachine manifest that the PaletteAI controller rendered from the template you selected, your overrides, and the machine pool. Select Validate to check it for syntax errors before you submit it.

Two things are worth knowing before you edit the manifest:

  • Editing a sizing value directly in the manifest has no effect. CPU Cores, Memory, Root Disk Size, and the GPU and MIG fields come from the Template step, and those values win over anything the manifest says about them. Change sizing on the Template step instead.
  • Cloud-init user data and network data are what the guest runs at first boot. They live in Secrets that PaletteAI creates, and the manifest points at those Secrets rather than carrying their contents.

Verify the Creation​

  1. From the left main menu, open AI Virtual Machines, and find your virtual machine.
  2. Check the Status column. It reads Provisioning while the disks are prepared, Starting while the guest boots, and Running when the virtual machine is ready to use.

The virtual machine reaches Running shortly after you create it.

Next Steps​