VCF 9.1 → 9.1.1: “Failed to upload NSX bundle” in the NSX Upgrade Coordinator

Today I wrote a field report from a VMware Cloud Foundation 9.1 to 9.1.1 maintenance upgrade: the order that matters, the error that stopped the NSX Manager upgrade, how to prove the bundle was not the problem, and what got the upgrade moving again.
Abstract
- In VCF 9.1.1 the NSX Manager upgrade is started from VCF Operations (Build → Lifecycle), not from the NSX UI – and only after the Fleet Lifecycle components and SDDC Manager are already on 9.1.1.
- During the NSX step the upgrade failed with “NSX Upgrade Coordinator – Failed to upload NSX bundle”. The SDDC Manager log showed an HTTP 500 returned by one NSX Manager node on POST /api/v1/upgrade/bundles.
- The bundle on SDDC Manager was intact (size and SHA-256 matched), disk space was fine, and the NSX cluster was STABLE – so the problem was on the NSX Manager side.
- A rolling reboot of the three NSX Managers, one at a time and with a cluster-status check in between, cleared it. The retried upgrade ran through.
Prerequisites: Everything above NSX must already be on 9.1.1
Before touching NSX, be clear about where you are in the upgrade chain. VCF 9.1.1 is a maintenance release, and the patch order is dependency-driven. The NSX Manager step can only start once the management plane “above” it is on 9.1.1:
- Fleet Lifecycle (the lifecycle services that drive everything else) – patched first.
- VCF services runtime and the other VCF Management components in the order the release notes dictate (VCF Operations is not patched in parallel with others; Identity Broker and Salt RaaS only after the services runtime; Migration Service Engine only after VCF Automation; patching a Software Depot blocks other patches).
- SDDC Manager – must be on 9.1.1 before any workload- or management-domain component is upgraded.
Only then follow the core domain order: SDDC Manager → NSX Manager → vCenter → ESX → NSX Edge/finalize. Do not use “Upgrade All” before you have complied with these dependencies. In our environment, the Software Depot sync had run, Fleet Lifecycle and the services runtime were upgraded, and SDDC Manager was on 9.1.1:

Tip: Right after the services runtime update, VCF Operations may show an “upstream connect error or disconnect/reset before headers” message. In our case this was transient while the nodes rotated and it disappeared on its own after a few minutes. Don’t start fixing things while the platform is still settling.

Confirmation that the prerequisite is met – VCF Operations → Build → Lifecycle → VCF Instances → <instance> → <management domain> → Component Versions. SDDC Manager is “On Target” (9.1.1.0); NSX, vCenter and ESX still show “Version Drift”, i.e. they are the next in line:

Pre-Flight: Set up the backups BEFORE! you start – and start them again right before NSX
Do not start the upgrade without working backups of both SDDC Manager and NSX. Both must be configured (target, credentials, schedule) before the first component is patched. A configured schedule is not enough on its own – the last backup might be hours old or, as in our SDDC Manager case, automatic backups may be disabled altogether.
Recommended routine – two backup runs per component:
- First backup start: configure the backup target and trigger a manual backup before you patch the very first component (before Fleet Lifecycle / SDDC Manager). This also proves that the target is reachable and the credentials work.
- Second backup start: immediately before the NSX Manager upgrade, trigger a fresh manual backup of SDDC Manager and of NSX again, and check that both finish as “Successful”. The state you want to be able to roll back to is the one right before NSX is touched – SDDC Manager is already on 9.1.1 at that point, and the schedule may not have run since.
Where to do it:
- SDDC Manager: Operate → SDDC Manager → Backup Settings (SFTP target, then trigger the backup manually). In our case the last backup was successful at 12:56 PM, taken manually because automatic backup was disabled.
- NSX Manager: System → Backup & Restore. Configure the SFTP target and a schedule (every 4 hours in our case), then start a manual backup and wait for “Last backup successful”.
Also check the NSX cluster health – three managers, cluster STABLE, no unexplained alarms. Our only alarms were remote-logging and license-expiry related, which do not block the upgrade.



Where to start the NSX upgrade: from VCF Operations, not from the NSX UI. If you open System → Upgrade in NSX, you get this warning: “This NSX Manager is managed by VCF Operations. Upgrading NSX through this workflow may impact VCF Operations.” Take it seriously and drive the upgrade from Build → Lifecycle → VCF Instances → <instance> → <domain> → Upgrades.

The Problem
The NSX upgrade plan was started from VCF Operations. Shortly after the NSX Upgrade Coordinator step began, the task failed with:
NSX Upgrade Coordinator - Failed to upload NSX bundle.
Review the LCM log files ... (127.0.0.1:/var/log/vmware/vcf/lcm ...)
The message tells you where to look, but not what is wrong. “Failed to upload” can mean several quite different things, so the first job is to narrow it down.
Diagnosis
On the SDDC Manager appliance the lifecycle logs live in /var/log/vmware/vcf/lcm/. Note that there is no plain lcm.log – the useful ones are:
- lcm-debug.log – detailed trace, where the real error is
- lcm-activity.log – API calls, good for the timeline
- lcm.err – errors; in our case only a harmless Spring shutdown trace and syslog noise
Filtering lcm-debug.log for errors and warnings around the time of the failure:
cd /var/log/vmware/vcf/lcm
grep -E "ERROR|WARN" lcm-debug.log | grep "<time window of the failure>"
This returns a lot of noise. None of the following is related to the failure:
ERROR ... UpgradeServiceImpl getResourceName ...] Resource type is unknown NSX_T_PARALLEL_CLUSTER
ERROR ... BundleManifestHandler registerDedupBundles ...] Error occurred filtering supportedSoftwareType
ERROR ... VCenterBundleMetadataValidator ...] Error verifying signature for upgradeInfo file
/nfs/vmware/vcf/nfs-mount/bundle/depot/local/tmpDir/upgrade_info.xml error INVALID_SIGNATURE
The relevant block starts with the upload failure of the NSX upgrade element (timestamps and IDs shortened, host names replaced):
2026-10-03T11:43:49.611Z ERROR vcf_lcm ... upgradeId="<upgrade-id>"
resourceType="NSX_T_PARALLEL_CLUSTER" ... class="c.v.e.s.l.p.i.nsxt.NsxtPrecheckUtil"
method="uploadPayload"] Payload upload failed:
org.springframework.web.client.HttpServerErrorException$InternalServerError
2026-10-03T11:43:49.611Z ERROR vcf_lcm ... class="c.v.e.s.l.p.i.nsxt.NsxtUpgradeUtil"
method="handleNsxtExceptions"] Handling NSX Exception
org.springframework.web.client.HttpServerErrorException$InternalServerError:
500 Internal Server Error on POST request for
"https://<nsx-manager-02>/api/v1/upgrade/bundles": [no body]
The stack trace right below it shows that the call went through the NSX retry helper, i.e. it had already been retried before giving up:
at NsxtBundleUploadOperations.uploadUpgradeBundle(:72)
at NsxtPrecheckUtil.lambda$uploadPayload$11(:1074)
at NsxtTransientRetryHelper.execute(:56)
at NsxtPrecheckUtil.uploadPayload(:1070)
at NsxtUcUpgradeStageRunner.doUpgradeStage(:413)
at NsxtParallelClusterPrimitiveImpl.runUpgrade(:702)
A few lines further down, the orchestrator records the consequence – the failing stage and the failed element:
DEBUG ... NsxtUpgradeUtil buildUpgradeError] Setting Upgrade Error for stage
NSX_UPGRADE_STAGE_SET_UPGRADE_PAYLOAD
ERROR ... NsxtParallelClusterPrimitiveImpl runUpgrade] Set the current element
<nsx-vip>:UpgradeCoordinator as COMPLETED_WITH_FAILURE
ERROR ... NsxtParallelClusterPrimitiveImpl runUpgrade] Setting the batch upgrade status
as COMPLETED_WITH_FAILURE
In plain words: SDDC Manager issues POST /api/v1/upgrade/bundles to NSX Manager node 2 in the stage NSX_UPGRADE_STAGE_SET_UPGRADE_PAYLOAD, the NSX transient-retry helper retries it, and the NSX Manager still answers with an HTTP 500 and no body. That is the whole error the UI condensed into “Failed to upload NSX bundle”.
So the upload does reach NSX. It is NSX that rejects it with an internal server error.
Before blaming NSX, make sure the source is good. The bundle is stored on SDDC Manager below /nfs/vmware/vcf/nfs-mount/bundle/<bundle-id>/:
ls -l /nfs/vmware/vcf/nfs-mount/bundle/<bundle-id>/<bundle-id>/
sha256sum /nfs/vmware/vcf/nfs-mount/bundle/<bundle-id>/<bundle-id>/*.mub
df -h
- The main NSX bundle (.mub, about 6.4 GB) and the pre-check bundle (about 1.15 GB) were present with the expected sizes.
- The SHA-256 of the .mub matched the checksum from the bundle manifest.
- The bundle metadata reported a download status of SUCCESS, and there was ample free space on SDDC Manager.
Bundle intact, space available – the failure was not a download or corruption problem.
Now check the NSX cluster. Log in to an NSX Manager as admin (SSH/console) and check the cluster:
get cluster status
Overall Status was STABLE and all members of all groups were UP, and the UI agreed. Nothing visibly broken, yet one node returned a 500 for the bundle upload.
The Fix: Rolling reboot of the NSX Managers
A known pattern for failing NSX bundle uploads is a rolling restart of the NSX Manager nodes – Broadcom KB 430635 describes bundle-upload failures that are resolved by rebooting the NSX Managers one after another, linked to an issue in the JDK (JDK-8330017) that can show up when the managers have been running for a long time. Caveat: that KB was written for NSX 9.0.0, and I did not prove that this specific JDK issue is what I hit on 9.1.0.x. I applied the same remediation because the symptom matched and the risk is low when done correctly.
Do it strictly one node at a time:
- Make sure the SDDC Manager and NSX backups from section 2 exist and that you triggered the second (fresh) backup run right before.
- On NSX Manager 1: run get cluster status – it must be STABLE with all members UP.
- Reboot that node (from the NSX CLI: reboot). The CLI broadcasts “The system will reboot now!”.
- Wait until the node is back and run get cluster status again from a node. Only continue when the cluster is STABLE and all members are UP again.
- Repeat for NSX Manager 2, then NSX Manager 3.
Never reboot the next manager while the cluster is still re-forming – with three managers you would lose quorum.
In our case: manager 3 was checked at 13:01 UTC (STABLE, all UP), rebooted, and about ten minutes later the cluster status at 13:10 UTC was again STABLE with every service UP on all three nodes.
With the cluster healthy again, retry the failed task from VCF Operations (Build → Lifecycle → … → Upgrades), re-running the precheck if prompted. This time the bundle upload to NSX succeeded and the NSX Manager upgrade started and ran.
Honest note: the rolling reboot fixed the symptom in my environment. I did not capture the NSX-side logs that would prove the exact root cause, so treat KB 430635 as the best matching explanation, not as a confirmed diagnosis. If the error survives the reboots, collect an NSX support bundle and open a case with Broadcom.
That’s it from this post, if you have any questions, I will love to help but as I said, consider to open a Support Ticket at Broadcom.