AI Virtual Machines
An AI Virtual Machine is a complete virtual machine, with its own guest operating system, that runs on a Compute Pool built for virtualization. App Deployments and Model Deployments run containerized workloads. An AI Virtual Machine gives you the whole operating system instead, for workloads that need one: GPU drivers, desktop applications, or tooling that does not run in a container. PaletteAI builds each virtual machine on KubeVirt and manages its lifecycle through a Virtual Machine Orchestrator (VMO) stack installed on the Compute Pool's cluster.
The AIVirtualMachine custom resource represents each virtual machine on the hub cluster. For the complete manifest schema, refer to the AIVirtualMachine reference.
You create and manage AI Virtual Machines from the AI Virtual Machines page in the PaletteAI console. The page is available at the System, Tenant, and Project scopes, and it appears in the left main menu in the same group as Compute Pools. Advanced and GitOps users can also manage the resource declaratively with kubectl.
Compute Pools that Run AI Virtual Machines
A Compute Pool runs either containers or virtual machines, and you choose which when you create the Compute Pool. A Compute Pool that runs AI Virtual Machines is always dedicated and single-cluster, and it is built from the VMO Profile Bundle. The Compute Pool reports its kind in its status.infraKind field, which is vm for a Compute Pool that runs AI Virtual Machines and container for every other Compute Pool.
Each machine pool on a VM Compute Pool carries per-VM limits, set when the Compute Pool is created. The limits are Max CPU, Max Memory, Max Physical GPU, and Max MIG Instances. They cap what a single AI Virtual Machine may request. The create workflow applies a ceiling for CPU, memory, and GPUs, which is the lower of the limits you set and the machine pool's capacity, and it does not accept a request above it. Max MIG Instances is a separate per-virtual-machine limit and is not part of that ceiling. A machine pool with no limits set has no configured cap of its own, and the ceiling is then the machine pool's capacity.
For the Compute Pool creation workflow, refer to Create and Manage Compute Pools.
Templates
Every AI Virtual Machine starts from a VmTemplate. PaletteAI ships a curated set of built-in templates that are offered on every VM Compute Pool: Ubuntu 22.04 and Fedora 37. Built-in templates are read-only, and you cannot edit or clone them. Each one's root disk comes from a golden image on the Compute Pool, so the template carries no registry location and works the same in a connected and an air-gapped deployment.
None of the built-in templates include an NVIDIA guest driver, because PaletteAI does not redistribute it. A virtual machine that needs GPU acceleration installs the guest driver at first boot through its cloud-init user data, from a source you configure. The driver must match the branch of the vGPU Manager running on the host.
Machine Pool Placement
The machine pool you choose determines which nodes run the virtual machine. On a multi-node Compute Pool, select a worker pool. On a single-node Compute Pool, use the control plane machine pool, which is the only one available. A virtual machine that names a machine pool the Compute Pool does not have reports an error and never schedules.
Virtual Machine Lifecycle
Creating an AI Virtual Machine takes you through a workflow that ends with a review of the KubeVirt VirtualMachine manifest PaletteAI built from your choices. Submitting that review deploys the virtual machine. The manifest records what is deployed:
- PaletteAI never rebuilds a virtual machine from its template, so a later template change does not affect a deployed virtual machine.
- When you review the manifest before submitting, CPU, memory, GPU, MIG, and the machine pool are also set in the workflow's own fields. Those fields win, and a conflicting value in the manifest is ignored.
- Placement is fixed.
spec.computePoolRefandspec.machinePoolcannot change after the virtual machine is deployed. spec.virtualMachinecannot be cleared. To remove the virtual machine, delete the AI Virtual Machine.
After you submit, the state moves to Provisioning while the disks are prepared, then to Starting while the guest boots, and finally to Running. The other settled states are Paused, Stopped, and Terminating. A virtual machine in Unschedulable is waiting for capacity on its machine pool rather than failing: it starts on its own once capacity frees up. For the full state list, refer to the AIVirtualMachine reference.
Access
AI Virtual Machines are reached with virtctl, the KubeVirt command line tool, against the Compute Pool's cluster. virtctl proxies through the cluster's KubeVirt API, so you need no routable endpoint and no exposed port. The built-in templates create a default user with a well-known password so you can connect on first boot.
For the connection steps and the default credentials, refer to Access an AI Virtual Machine. For replacing those credentials with your own, refer to Change AI Virtual Machine Credentials.
Shared Access
An AI Virtual Machine can be shared with lower scopes through its spec.sharedWith field, using the same sharing model as other PaletteAI resources. A scope the virtual machine is shared with can view it, connect to it, and clone it. Stopping and deleting stay with the scope that owns the virtual machine.
A shared virtual machine runs on its owner's Compute Pool and counts toward the owner's usage, no matter which scope connects to it. For the sharing model, refer to Sharing Resources.
Advanced Management with VMO
PaletteAI manages the lifecycle operations that most AI VM workloads need. For advanced lifecycle management and configuration, use the VMO console on the Compute Pool's cluster. The Overview tab of an AI Virtual Machine carries a VMO link that opens that virtual machine directly, and the Compute Pool overview carries a link that opens the VMO dashboard for the cluster.
Next Steps
- Create an AI Virtual Machine: choose a template and sizing, review the generated manifest, and submit it.
- Access an AI Virtual Machine: connect over a serial console, SSH, or VNC.
- Create and Manage Compute Pools: stand up a Compute Pool that runs AI Virtual Machines.