This week in Las Vegas, Broadcom's VMware Explore floor is selling you a complete private-AI stack: VMware Cloud Foundation, NVIDIA GPUs, Tanzu, and a backup appliance from Veeam, Rubrik, Commvault, or Cohesity on the same booth row. The sentence in every briefing is a 2010s sentence wearing 2026 clothes: snapshot the GPU node, snapshot the Tanzu cluster, and the AI estate is covered.
It is not.
A VMDK, an AMI, a PersistentVolumeClaim snapshot: those are image-level recoverability. They bring disks back. They do not bring back the runnable LLM application that was serving traffic last Tuesday. Sign the Q3 private-AI plus backup budget on that confusion and you own a 2026 contract that restores a virtual machine and cannot reconstruct the system.
Two primitives vendors are collapsing
"Backup" on the show floor now means three things at once: hypervisor snapshot, cluster snapshot, and "AI-ready DR." Those are not the same object. Split them or you will buy the wrong one.
Image-level recoverability is what vSphere snapshots, Veeam VM backups, Rubrik SLA domains, and CSI volume snapshots already do well. You get a point-in-time copy of block devices. The guest OS boots. The GPU driver loads. That is real work, and the hypervisor is fine at it.
Operational-state reconstructability is different. It is the ability to stand the same LLM application back up: system prompt and policy pack, retrieval-index version, tool and API bindings, safety settings, secrets layout, and the exact runtime config that produced last week's behavior.
You can have the first and none of the second. Private AI can be fully snapshotted at the infrastructure layer and still be impossible to stand back up as the same system.
We already argued this in Why Backup Strategies Are the New AI Imperative. The Explore pitch is the counterfeit version of that argument. It uses the word "backup" for disks because disks are what those vendors already sell.
What a VMDK actually contains
A typical private-AI layout on VCF: a GPU-backed VM or Tanzu worker, vLLM or TensorRT-LLM serving from a PVC, a vector store on another volume, an inference gateway in front.
A hypervisor snapshot of that GPU node captures VMDK bytes, VMX hardware config (vCPU, vRAM, vGPU profile), and a crash-consistent or quiesced file-system image. Right artifact for host loss. Wrong artifact for the application.
It does not capture these as named, versioned, restorable units:
system_prompt.mdandpolicy_pack.yamlloaded into the serving process (often a ConfigMap edited in place, or a prompt store outside the VM)- FAISS, Milvus, or pgvector collection version, plus the embedding-model hash that built it
- Tool and API bindings: OpenAPI specs, allow-lists, which endpoint is prod this week
- Safety settings: filter thresholds, groundedness checks, PII redaction, output schemas
- Secrets layout: not the values, the map of env names, CSI secret-provider classes, and Secret objects
- Runtime config: max model length, tensor parallel size, LoRA IDs, tokenizer revision, model SHA, NIM digest
Restore that VMDK onto a new host and you get a machine, not last week's application. The index has moved. The prompt was hot-patched. The NIM tag floated. The vector-volume PVC sits at a different RPO than the serving VM, so the model was never evaluated against that index.
That is not a restore. That is a new system sharing some disk blocks with the old one.
The RFP line that will lock you in
Backup vendors will win the "do you have backup?" line this week because procurement already knows how to score it. SLA domain, immutable copy, 3-2-1, ransomware cleanroom: all real, all useful, all about VMs and PVCs.
Watch the language stretch. "AI-ready backup" means GPU VMs and Tanzu namespaces are in the job. "Application-consistent AI" means they quiesced the guest, which is consistent for the guest, not for the LLM application.
None of that is a lie about disks. It is a category error about the workload. SQL has a WAL and a documented restore sequence. An LLM application has a prompt pack, an index generation, a binding file, and a runtime flag set that never appear in the VMDK cut. Quiescing ext4 does not version policy_pack.yaml.
Ask one bake-off question and refuse to let it collapse into "we snapshot the cluster": can you restore a point-in-time, portable snapshot of the full AI operational state, independent of vSphere, Veeam, and the hypervisor, and prove it is the same system by replaying a golden query set?
If the answer is a VM powering on, score image-level recoverability. Do not score operational-state reconstructability.
What to do before the Q3 markup hardens
Keep VCF, NVIDIA, and the backup appliance. Hypervisor snapshots are the right tool for host loss, guest ransomware, and datacenter failover. Treat the LLM application as a second restore domain with its own artifacts and its own RPO.
Inventory operational state as named objects, not as "the VM":
- Prompts and policy packs. Version
system_prompt.md,policy_pack.yaml, and the safety-settings blob with a content hash. No prompt hash, no restore. - Retrieval indexes. Collection name, build job ID, embedding model plus revision, chunking config, source corpus snapshot ID. A PVC of
milvus-datawithout the build manifest is an orphan. - Tool and API bindings. Allow-list, OpenAPI fingerprints, destination map. Bindings change with no VM change.
- Secrets layout. Schema only: env names, provider class, mount paths. Re-hydrate keys from the vault after restore.
- Runtime config. Model ID, tokenizer revision, NIM/vLLM/Triton digest (not the tag), tensor parallel size, context window, LoRA IDs.
- Golden query pack. Ten to fifty prompts with expected retrieval hits and policy outcomes. That is how you prove sameness.
Schedule the operational-state snapshot on its own clock. A 24-hour VM RPO and a 5-minute index RPO are two backups. They will drift.
If the private-AI program exists so the model and corpus stay out of someone else's cloud, that still is not recoverability. We covered Why Your AI Needs More Than Just Cloud Solutions; the dual is now true on-prem. Sovereignty of the VMDK is not sovereignty of the application.
Do not let the floor redefine the word
vSphere, VCF, NVIDIA NIMs, Tanzu, Veeam, Rubrik: keep them. The failure is definitional. If "backup" in the RFP means "the cluster came back," you will hit the gap the first time last week's prompt pack, index, and bindings are what you actually need.
SaveState exists to take that second primitive seriously: a point-in-time, portable snapshot of the full AI operational state, independent of the hypervisor and independent of whoever sold you the VM job. Keep the VMDK. Add the thing the VMDK cannot see.
Before you mark up the Q3 line, run one restore test that is not "the VM booted." Restore the prompt pack, the index version, the bindings, the safety settings, the secrets layout, and the runtime digest, then run the golden query pack. If you cannot name those artifacts, you do not have an AI backup. You have a disk.
Keep the VMDK. Add the memory layer.
Pro is $9/month. Encrypted portable memory. Card today. No waitlist.
Subscribe to Pro — $9/mo Team is $29/monthAfter you pay, your API key is emailed. Then savestate login. Card today — no waitlist.