TL;DR
Treat the target as an exact AOS, AHV, Prism Central, NCC, LCM, firmware, and product-service combination. “AOS 7 and AHV 10” is not specific enough for a change record.
Validate the supported upgrade path and compatibility matrix before downloading software.
Upgrade Prism Central first only when the validated dependency plan requires it. Back it up before changing it.
Refresh the supported LCM framework, perform inventory, update NCC, run NCC, and complete an independent LCM precheck before the maintenance window.
Require green data resiliency, healthy cluster services, no active rebuild or unresolved critical alert, and enough CPU, memory, and storage headroom to tolerate one node in maintenance.
Use the order selected by LCM and the applicable release notes. A practical planning sequence is Prism Central if required, LCM, NCC, AOS, validation, AHV, validation, firmware, dependent services, and final evidence.
Do not describe rollback as a generic downgrade. Define safe stop points, Prism Central recovery, support-directed component recovery, and application-level failback separately.
The change is complete only after infrastructure health, workload transactions, data protection, monitoring, and audit evidence all pass.
What This Runbook Accomplishes
The objective is to move a production Nutanix AHV cluster from an approved source state to an approved AOS 7.x and AHV 10.x target while preserving data availability, workload service, manageability, and recovery options.
The runbook assumes:
The cluster is running Nutanix AHV.
Prism Element is available and the cluster may be registered to Prism Central.
LCM is the authorized lifecycle mechanism. Nutanix requires LCM for AOS upgrades to AOS 6.8 or later. [3]
The operator has administrative access, a tested support path, an approved maintenance window, and named application validators.
The source and target releases are supported on the exact server model and hardware configuration.
This is not a universal procedure for single-node clusters, two-node clusters, Metro Availability pairs, Nutanix Cloud Clusters, or environments with special device-backed workloads. Those designs require their version-specific Nutanix procedures in addition to this baseline. LCM supports some software operations on single-node clusters, but Nutanix warns about data-loss and service-disruption risk, so a normal rolling-upgrade assumption is inappropriate. [4]
Why AOS, AHV, Prism Central, and Firmware Need Separate Gates
The platform components have different functions and failure domains:
Component
Operational role
Upgrade concern
Prism Central
Central management and services across clusters
Version compatibility with registered Prism Element clusters, service dependencies, backup, and management-plane availability
LCM
Inventory, dependency evaluation, prechecks, software and firmware orchestration
Framework compatibility, current metadata, network or dark-site bundle access, and task history
NCC
Platform health assessment
Current compatible version, unresolved findings, and repeatable before-and-after results
AOS
Distributed storage and cluster services running through CVMs
CVM health, data resiliency, active rebuilds, service state, and rolling component behavior
AHV
Hypervisor on each node
Host evacuation, maintenance mode, VM mobility, device-backed workloads, capacity, and host restart behavior
Firmware
BIOS, BMC, NIC, HBA or controller, drive, and other platform firmware
Hardware model support, Redfish or BMC access, OEM bundles, reboot duration, and dependency order
Product services
Files, Objects, Flow, NDB, NKP, DR, and other services
Separate interoperability matrices, upgrade guides, and application-facing validation
LCM can calculate dependency-aware component order for selected updates, but that does not remove the need to validate the whole environment or define gates between major components. [5]
The workflow below shows the sequence that matters. A failed gate returns to remediation. It does not flow into the next component.
Define the Change Before Opening LCM
A change ticket should name the exact desired state. At minimum, record:
Field
Required value
Cluster
Name, UUID, site, node count, hardware model, and fault-tolerance design
Current stack
Prism Central, Prism Element or AOS, AHV, LCM, NCC, Foundation, firmware, and product-service versions
Target stack
Exact approved version for every component being changed
Upgrade path
Supported hop or sequence returned by Nutanix Upgrade Paths
Compatibility evidence
Saved result from the Compatibility and Interoperability Matrix
Release notes
Reviewed AOS, AHV, Prism Central, LCM, and hardware bundle notes
Service impact
Expected management, migration, VM, and application behavior
Stop conditions
Specific technical and application conditions that halt the workflow
Recovery owners
Nutanix platform, network, storage, backup, application, business, and support contacts
Validation
Infrastructure checks, application transactions, monitoring checks, and observation period
Avoid a target such as “latest AOS 7.” A valid target looks like an exact maintenance release paired with an exact AHV build and a compatible Prism Central release. The target must also account for the server model, disk and controller firmware, NIC firmware, GPU or vGPU stack, Files or Objects versions, backup integrations, and DR relationships.
Pre-Upgrade Readiness Gates
Validate the Supported Path and Product Compatibility
Use three separate checks because they answer different questions:
Lifecycle status: Is the proposed target still maintained and supported?
Upgrade path: Can the current version move directly to the target, or are intermediate hops required?
Interoperability: Are Prism Central, AOS, AHV, firmware, hardware, and dependent services supported together?
LCM showing an available update is useful evidence, but it is not a substitute for reviewing the target release notes and the formal compatibility data. This matters when the environment includes products that LCM does not fully model, such as backup software, guest drivers, third-party monitoring, GPU host drivers, security agents, or application certification requirements.
Record the result, date, and operator. Compatibility pages change over time, so an audit record should preserve what was approved for the change rather than relying on a future lookup.
Resolve Prism Central Dependencies
Prism Central and the registered Prism Element clusters must remain compatible during every transition state, not only after the final upgrade. Nutanix provides a PC-to-PE compatibility health check, and the software compatibility matrix remains the authoritative planning source. [6]
Use this decision sequence:
If Prism Central is already compatible with both the current and target AOS releases, retain it unless release notes or a required service say otherwise.
If the target AOS requires a newer Prism Central release, upgrade Prism Central before the managed cluster.
If Prism Central hosts Self-Service, Flow, DR, automation, or other services, validate each service against the proposed PC release.
If Prism Central itself runs on the cluster being upgraded, confirm that its VMs can tolerate host evacuation and that local Prism Element access remains available.
Do not unregister a cluster merely to bypass a compatibility finding. Registration can carry service, policy, and protection dependencies.
Back up Prism Central using its supported backup mechanism and verify that the backup is current. Nutanix documents One-Click Recovery for protected Prism Central instances, which makes the backup a real recovery control rather than a checkbox. [7]
Capture:
Prism Central version and deployment size
Registered clusters and connection state
Identity provider and RBAC health
Enabled Prism Central services
Backup type, destination, last successful timestamp, and recovery prerequisites
Current critical alerts and failed services
Refresh LCM and Perform Inventory
LCM inventory identifies the software and firmware entities visible to the framework and the bundles available to it. In a connected site, inventory can also update the LCM framework. In a dark site, the correct framework and payload bundles must be staged using the method supported by the installed LCM and platform versions. [8]
Perform inventory before the maintenance window, then verify:
Every expected node and component appears.
Current versions match the change record.
The proposed target versions are available or the correct offline bundles are staged.
No inventory operation is stuck or reporting stale data.
BMC or Redfish credentials work where firmware orchestration requires them.
DNS, NTP, proxy, firewall, and repository paths required by LCM are healthy.
LCM operations history does not contain an unresolved prior lifecycle task.
An unexpected missing component is a stop condition. Do not assume LCM will safely ignore hardware or software it cannot inventory.
Update NCC and Run the Full Health Check
NCC is a prerequisite, not a post-failure diagnostic. Nutanix explicitly instructs administrators to run NCC before an upgrade and resolve results other than INFO or PASS before proceeding. [9]
Update NCC to the latest version supported by the current stack, then run the complete check from Prism or from a CVM. The following read-only baseline is intentionally small so it can be repeated after each major stage:
# Run from one CVM in each Prism Element cluster as the nutanix user.
date -u
ncc –version
cluster status
ncc health_checks run_all
The commands record the UTC time, NCC version, cluster service state, and complete NCC result. Copy the raw output into the change evidence repository. Do not reduce the record to a screenshot of a green summary.
Successful execution means:
Cluster services expected for that release are up.
NCC completes without unresolved findings outside INFO or PASS.
Any accepted informational finding has a documented disposition.
The same commands can be rerun after AOS and AHV without relying on shell history.
NCC is point-in-time evidence. A passing result from the previous week does not authorize tonight’s change.
Verify CVM, Cluster, and Data Resiliency Health
Before any rolling activity, verify:
All CVMs are powered on and reachable.
Cluster services are healthy and stable.
Data Resiliency Status is green.
No node, disk, storage pool, or metadata rebuild is active.
No critical hardware, storage, network, CVM, or hypervisor alert is unresolved.
Replication and protection jobs are current.
Time synchronization and name resolution are healthy.
Cluster performance is within its normal baseline.
Nutanix procedures use green Data Resiliency Status and NCC as prerequisites before taking a node out of service. That same principle belongs in an upgrade gate. [10]
Do not proceed while the platform is already consuming fault-tolerance capacity. In an RF2 cluster, for example, multiple simultaneous node outages can exceed the design’s protection boundary. The exact allowable failure state depends on cluster design, failure domains, RF, and current data placement.
Prove Capacity for One Node in Maintenance
AHV upgrades place hosts into maintenance mode and process them sequentially. The cluster must be able to run the remaining workload while a node is unavailable. Nutanix documents maintenance mode as a required state for host and firmware operations. [11]
Validate all three capacity planes.
Compute capacity
Remaining hosts can absorb the CPU and memory demand of the evacuated node.
HA reservation and admission policies remain satisfied.
Memory overcommit, NUMA placement, and oversized VMs will not block evacuation.
Peak demand, not only current demand, fits within the remaining cluster.
Storage capacity
Free and rebuild-reserved capacity are healthy.
No container is near its critical threshold.
The cluster can tolerate normal upgrade I/O plus workload I/O.
Capacity-tier and data-placement behavior are stable.
Network capacity
Management, CVM, storage, and live-migration paths are healthy.
Uplinks, bonds, VLANs, MTU, and upstream switch ports are stable.
The network can absorb migration traffic without starving production flows.
There is no safe universal free-capacity percentage for every Nutanix cluster. Node count, RF, block awareness, storage media, failure domains, workload shape, and reserved rebuild capacity all affect the decision. Use the live resiliency and capacity state for the exact cluster.
Identify Workloads That May Not Evacuate Cleanly
Inventory VMs with placement or device constraints before the AHV stage:
Host affinity or anti-affinity rules
Large memory reservations
PCI passthrough, SR-IOV, GPU, or vGPU configurations
Special network functions or latency-sensitive appliances
VMs with mounted media or unusual device state
Workloads restricted by license, application clustering, or node identity
Prism Central or infrastructure services hosted on the cluster
AHV 10 includes configuration-specific live-migration rules. Some capabilities improved with AOS 7 and AHV 10, so it is equally unsafe to assume every device-backed VM can migrate or that none can. Validate each configuration against the AHV 10 documentation and current host-driver matrix. [12]
For each exception, define one approved treatment: live migrate, shut down and restart, move to another cluster, suspend the dependent service, or exclude the host from the change. “Handle during the window” is not a plan.
Validate Firmware and Hardware Compatibility
Firmware should be planned as its own risk domain even when LCM can orchestrate it in the same workflow. Record the current and target versions for:
BIOS or UEFI
BMC
Storage controller or HBA
NIC and adapter firmware
NVMe, SSD, and HDD firmware where applicable
GPU firmware and host drivers
Foundation and platform support components
Confirm the exact server model and component revisions are supported by the target AOS and AHV releases. OEM platforms may have different bundle, credential, and sequencing requirements. If LCM’s dependency resolver or the hardware release notes prescribe a different order than the generic sequence in this article, the validated vendor order wins.
Verify Backup, Replication, and Application Recovery
An infrastructure upgrade plan is incomplete without application recovery. Confirm:
Prism Central backup is current and recoverable.
Application backups completed successfully.
Protection domains, recovery plans, and replication schedules are healthy.
Metro, NearSync, asynchronous replication, and witness dependencies are explicitly reviewed where used.
Backup proxies, agents, and storage integrations support the target releases.
Application owners know the validation transaction and the point at which application failback is invoked.
Do not create an improvised rollback mechanism by snapshotting or manipulating CVMs, PCVMs, or AHV hosts outside a documented Nutanix procedure. Platform recovery should use supported product mechanisms or Nutanix Support direction.
Run LCM Prechecks as a Separate Dry Run
LCM can run prechecks independently of the upgrade, which helps separate readiness failures from failures that occur after software execution begins. [13]
Run the precheck early enough to remediate findings, then run it again close to the maintenance window. A valid precheck record includes:
Selected cluster and component
Source and target versions
Date and time
LCM task or operation identifier
Complete findings
Remediation owner and disposition
Rerun result
Do not waive a failed precheck because an earlier cluster passed or because the same issue was harmless in a lab. A failed precheck is a no-go until the exact finding is resolved or Nutanix Support provides written direction for that environment.
Recommended Upgrade Order
The following is a planning sequence, not a command to override LCM. LCM uses compatibility information and dependencies to select update order. Release notes, Upgrade Paths, the compatibility matrix, and the order presented by LCM take precedence.
Order
Stage
Gate before continuing
1
Prism Central, if the target plan requires it
PC services healthy, backup verified, registered clusters connected, identity and required services validated
2
LCM framework and inventory
All expected entities visible, metadata current, target bundles available, no stuck task
3
NCC
Supported NCC installed, full health check passes, raw result saved
4
AOS
Cluster services healthy, data resiliency green, no rebuild, capacity and application baseline captured
5
AOS validation
Versions correct, services stable, NCC clean, storage and application checks pass
6
AHV
Every host can enter maintenance, constrained VMs have an approved disposition, remaining capacity is sufficient
7
AHV validation
All hosts on target build, no host left in maintenance, VMs healthy, network and storage checks pass
8
Firmware
Model-specific compatibility and credentials validated, software stack stable, reboot expectations approved
9
Dependent services
Each service’s own upgrade path, matrix, backup, and validation plan approved
10
Final evidence
All infrastructure, application, monitoring, protection, and audit gates pass
For connected environments, selecting multiple updates may be appropriate when LCM presents and validates the complete dependency plan. For high-risk production changes, separate validation gates between AOS, AHV, and firmware make failure isolation and change control clearer.
For dark sites, follow Nutanix’s version-specific dark-site method and recommended order. Do not assume a bundle prepared for one LCM framework or catalog state is valid for another.
Execute the Upgrade in Controlled Stages
Upgrade Prism Central When Required
If the approved dependency plan includes Prism Central:
Confirm the most recent PC backup completed.
Capture PC health, services, alerts, connected clusters, identity, and version.
Start the supported Prism Central upgrade workflow.
Monitor the task until it reaches a completed state.
Confirm the new version and service health.
Validate login, RBAC, identity provider integration, registered clusters, and required applications.
Rerun the PC-to-PE compatibility check.
Do not begin the cluster upgrade while Prism Central is degraded unless the approved plan explicitly removes that dependency and Nutanix Support agrees with the recovery path.
Upgrade AOS
Before selecting AOS, verify the change record one final time. Confirm the exact target and review LCM’s proposed plan.
During the AOS stage:
Watch the LCM operation, Prism alerts, cluster services, CVM reachability, storage latency, and application telemetry.
Do not make unrelated storage, network, CVM, or cluster configuration changes.
Do not restart several CVMs or services in response to a slow step.
Do not start AHV or firmware work manually while AOS is still running.
Record the operation identifier and the start and completion time.
When LCM reports completion, pause. Rerun the baseline commands, verify Data Resiliency Status, review new alerts, confirm the target AOS version on all nodes, and run the agreed application smoke tests.
Only the approved AOS validation gate authorizes the AHV stage.
Upgrade AHV
The AHV stage is host-disruptive even when workloads remain available through migration. LCM typically handles hosts sequentially, including maintenance entry, workload evacuation, update, restart, and maintenance exit.
Monitor:
Which host is active in the workflow
VM migration and placement
CVM state on that node
Maintenance-mode entry and exit
Host management and storage connectivity
Network uplinks and virtual switches
Application latency and error rate
If a host cannot evacuate, do not force the cluster into a second simultaneous maintenance event. Identify the blocking VM or constraint, apply the approved exception plan, and rerun the supported precheck or task.
After AHV completes, verify every host reports the exact target build and no host remains in maintenance mode. Confirm all CVMs and user VMs are in their intended state.
Upgrade Firmware and Dependent Services
Begin firmware only after the AOS and AHV stack is stable, unless the dependency plan explicitly requires another sequence. Firmware operations may involve longer host restarts, BMC communication, and OEM-specific behavior.
Upgrade product services only with their own compatibility and recovery checks. A green AOS and AHV result does not prove Files, Objects, NDB, NKP, Flow, backup, DR, or third-party integrations are ready for their next version.
Stop Conditions and Failure Handling
The operator needs authority to stop the change before the window begins. Use explicit stop conditions such as:
Data Resiliency Status is not green.
A new critical hardware, storage, CVM, hypervisor, or network alert appears.
A cluster service is down or repeatedly restarting.
AOS completes on only part of the cluster and LCM does not present a supported continuation.
A host cannot enter or exit maintenance mode.
A VM cannot migrate and no approved exception exists.
Storage latency, packet loss, application error rate, or transaction time crosses the approved threshold.
Prism Central loses required services or cluster connectivity.
Replication, backup, or protection state becomes unhealthy.
The maintenance window no longer contains enough time for validation and recovery.
Use the response matrix below rather than improvising.
Failure point
Immediate response
Do not do
Inventory or precheck failure
Stop, preserve findings, correct the dependency, rerun inventory and prechecks
Start the upgrade because the cluster “looks healthy”
Package download or staging failure
Verify bundle integrity, repository access, proxy, DNS, certificates, and space; restage through the supported method
Repeatedly upload different bundles without confirming target metadata
Prism Central upgrade failure
Preserve task details, validate local clusters through Prism Element, collect PC evidence, use the supported PC recovery path
Unregister clusters or rebuild PC without dependency analysis
AOS step fails
Do not proceed to AHV; verify cluster services and resiliency, preserve LCM history, collect logs, engage Nutanix Support
Restart multiple CVMs, run an undocumented downgrade, or clear task state manually
AHV host is stuck
Stop at the safe boundary, identify VM or host constraint, validate CVM and host state, follow support guidance
Put another host into maintenance or power-cycle several nodes
Firmware update fails
Keep the stable software state, preserve BMC and LCM evidence, use model-specific recovery
Retry power or firmware operations blindly
Infrastructure is green but application test fails
Freeze further lifecycle work, invoke the application incident and recovery plan, compare pre-change telemetry
Close the change based only on Prism health
LCM’s Stop Update function chooses a safe point rather than interrupting an unsafe operation immediately. Nutanix also documents that if AOS and AHV are selected together and a stop is requested after the AOS update begins, LCM allows AOS to finish and cancels before AHV. [14]
That behavior reinforces an important distinction: stop, rollback, retry, restore, and failback are different operations.
Rollback Is a Recovery Model, Not a Downgrade Button
A credible plan defines several recovery levels.
Abort Before Change
If a readiness gate fails before software execution, cancel the change. No platform rollback is required. Preserve the evidence, remediate the finding, and reschedule.
Stop at a Safe LCM Boundary
If the change begins but risk rises, request the supported LCM stop operation. Expect LCM to stop at its next safe point. A stop may leave one component successfully upgraded while later components remain unchanged.
For example, a completed AOS upgrade is not automatically reversed because the AHV stage was canceled. Validate whether the resulting AOS and AHV combination is supported, stabilize it, and obtain support guidance before deciding the next action.
Retry or Resume the Failed Component
When LCM or the product guide presents a supported retry or resume operation, use it only after identifying the original failure and verifying that cluster health is stable. Preserve the original task identifier and evidence. Repeated retries without causal remediation can turn a contained problem into an outage.
Restore Prism Central
If Prism Central cannot be recovered in place, use the documented Prism Central backup and One-Click Recovery process. This recovers the management platform and important service configuration. It is not an AOS or AHV rollback.
Recover the Application Service
If infrastructure is healthy but the business service fails, use the application-specific recovery plan. That may mean reverting an application change, restarting a clustered service, failing traffic back, restoring data, or activating DR. The correct action depends on the application’s consistency and recovery design.
Escalate for Component-Level Recovery
Manual AOS downgrade, CVM image replacement, AHV reinstallation, task-database edits, or forced multi-node restarts should not appear as routine operator rollback steps. Those are support-directed recovery actions when applicable.
Post-Upgrade Validation
Validation should prove the desired service state, not merely the absence of a red banner.
Rebuild the LCM Inventory
Perform a fresh inventory and verify:
Every node and component reports the approved version.
No unexpected component remains on the source release.
Firmware inventory matches the change record.
No lifecycle operation remains running or failed.
The next recommended updates are understood but not accidentally included in the current change.
Export LCM operations history and inventory for the audit record. Nutanix provides an export workflow specifically for this evidence. [15]
Rerun Platform Health Checks
Repeat the same commands used for the baseline:
# Post-upgrade check from one CVM in each cluster.
date -u
ncc –version
cluster status
ncc health_checks run_all
Compare the results with the baseline. Resolve or formally disposition every new finding.
Then verify in Prism:
Data Resiliency Status is green.
All CVMs and hosts are healthy.
No node remains in maintenance mode.
No disk, node, or metadata rebuild is unexpectedly active.
Storage pools and containers have expected capacity.
Cluster latency, IOPS, throughput, CPU, and memory remain within normal range.
Alerts are reviewed, not simply acknowledged in bulk.
Validate Workload and Network Behavior
Confirm:
Expected VMs are powered on and placed correctly.
HA and affinity policies remain intact.
A representative live migration succeeds if the change plan authorizes an active test.
Virtual networks, uplinks, bonds, VLANs, MTU, routing, DNS, and load balancer paths work.
Device-backed and exception workloads operate as designed.
Guest time, network, and storage behavior are stable.
Validate Management, Protection, and Integrations
Confirm:
Prism Central sees the cluster and reports the correct versions.
SSO, directory integration, RBAC, certificates, and API access work.
Backup jobs, snapshots, protection domains, recovery plans, and replication schedules resume.
Monitoring, logging, SIEM, ticketing, and alert forwarding receive current data.
Files, Objects, NDB, NKP, Flow, and other installed services pass their own checks.
Hardware management and BMC connectivity remain healthy.
Validate the Application, Not Only the VM
Application owners should execute the same transactions captured before the upgrade. Examples include:
User authentication
API request and response
Database read and write
File creation and retrieval
Message publish and consume
Batch or scheduled job completion
North-south and east-west application flows
A powered-on VM is not evidence that the service is healthy.
Post-Upgrade Evidence Package
The following table can be copied into the change record.
Evidence
Before
After
Owner
Result
Prism Central version and health
Attached
Attached
Platform
Pass or fail
LCM inventory export
Attached
Attached
Platform
Pass or fail
AOS and AHV versions by node
Attached
Attached
Platform
Pass or fail
Firmware versions by node
Attached
Attached
Hardware
Pass or fail
NCC version and full result
Attached
Attached
Platform
Pass or fail
cluster status output
Attached
Attached
Platform
Pass or fail
Data Resiliency Status
Attached
Attached
Platform
Pass or fail
Capacity and performance baseline
Attached
Attached
Operations
Pass or fail
Alerts and event review
Attached
Attached
Operations
Pass or fail
Backup and replication state
Attached
Attached
Data protection
Pass or fail
Network and migration validation
Attached
Attached
Network and platform
Pass or fail
Application transactions
Attached
Attached
Application owner
Pass or fail
LCM operation IDs and timestamps
Not applicable
Attached
Change owner
Complete
Exceptions and support case
Documented
Documented
Change owner
Closed or tracked
Keep the change open through the approved observation period. A technically completed LCM task is a milestone, not the definition of done.
Common Upgrade Mistakes
Treating a Major Version as the Target
“AOS 7” and “AHV 10” are search terms. Production changes need exact maintenance releases and builds.
Assuming Prism Central Is Always First or Always Optional
The correct answer depends on PC-to-PE compatibility and enabled services. Validate the transition state.
Running NCC Only After Something Fails
NCC belongs before the change, after AOS, after AHV, and after final remediation when the risk warrants it.
Combining AOS, AHV, and Firmware Without Gates
LCM may support a combined dependency-aware workflow, but operational gates make ownership and failure isolation clearer.
Equating Rolling with Risk-Free
Rolling operations reduce planned interruption. They do not eliminate workload, migration, capacity, network, hardware, or application risk.
Calling Every Recovery Action a Rollback
Safe stop, retry, Prism Central restore, application failback, and support-directed component repair solve different problems.
Closing on Green Infrastructure Alone
The final gate is the application transaction plus protection, monitoring, and evidence, not the color of the Prism dashboard.
Frequently Asked Questions
Should Prism Central always be upgraded before AOS?
No. Upgrade Prism Central first when the validated source-to-target compatibility plan, Upgrade Paths, release notes, or an enabled service requires it. If the installed PC release supports both the current and target cluster versions, a PC upgrade may not be necessary for that change.
Can AOS and AHV be selected together in LCM?
LCM can present a dependency-aware combined plan. For production control, review the exact plan and retain a validation gate after AOS. If a stop is requested after AOS begins, LCM may complete AOS and cancel before AHV at a safe boundary.
Does a rolling Nutanix upgrade guarantee zero downtime?
No. Rolling orchestration reduces planned service interruption, but constrained VMs, capacity pressure, migration failures, network faults, firmware behavior, or application dependencies can still cause impact.
Can an AOS or AHV upgrade be rolled back with one click?
Do not plan on a universal downgrade button. Define pre-change abort, safe LCM stop, supported retry or resume, Prism Central recovery, application failback, and Nutanix Support escalation as separate controls.
Should firmware be upgraded in the same maintenance window?
Only when the compatibility plan, available time, hardware procedure, and recovery plan support it. Firmware has a different failure domain and often deserves a separate gate or separate change window.
Conclusion
The safest Nutanix upgrade is not the one with the fewest clicks. It is the one with the clearest dependencies, the strongest readiness gates, and the most objective evidence.
For AOS 7.x and AHV 10.x, that means resolving an exact supported target, protecting Prism Central, refreshing LCM inventory, updating and running NCC, proving CVM and cluster health, preserving node-maintenance capacity, validating hardware and workload mobility, and separating AOS, AHV, firmware, and product services with explicit gates.
Rollback should be treated as a recovery model rather than a promise of simple downgrade. When the change defines safe stop points, support escalation, management-plane recovery, application failback, and measurable validation before execution begins, LCM becomes what it should be: the orchestration engine inside a disciplined lifecycle process.
Official References
Some Nutanix Support & Insights pages may require an authenticated Nutanix account.
Nutanix Software End of Life Information
Nutanix Compatibility and Interoperability Matrix
Prism 7.0: AOS Upgrade
LCM 3.3: LCM Limitations
LCM: Life Cycle Manager Overview
NCC Health Check: PC and PE Version Compatibility
Prism Central Backup, Restore, and Migration
LCM 3.3: LCM Inventory
Prism 7.0: Running NCC from Prism Element
Nutanix Node Shutdown Precheck
AHV 10.3: Node Maintenance Mode
AHV 10.0: Live Migration Restrictions
LCM: Performing Pre-Upgrade Checks
LCM 3.3: Stop Firmware and Software Upgrades
LCM 3.3: Exporting Operations History and Inventory
Nutanix CVM Troubleshooting Command Reference: Stargate, Curator, Cassandra, Genesis, and Cluster Health
TL;DR A Nutanix CVM incident should be investigated as an evidence chain, not as a race to restart services. Begin with cluster…
Next PostHCX 9.1 Migration Runbook: Network Profiles, Compute Profiles, Service Mesh, and First VM MoveTL;DR An HCX 9.1 migration should not begin with the Migrate Virtual Machines button. It should begin with a dependency-ordered runbook that proves routing, firewall policy, IP pools, site pairing,…
The post Nutanix AOS 7 and AHV 10 Upgrade Runbook: LCM Prechecks, Order, Rollback, and Validation appeared first on Digital Thought Disruption.

