Recovering a VxRail Manager with nearly every pod in CrashLoopBackOff

The VxRail Manager VM had been rebooted and the VxRail pages in the vSphere Client no longer loaded. VxRail Manager runs its services as Kubernetes pods on the appliance itself, so the first look was at one of the pods that was obviously unhappy.
vxm01:~ # kubectl describe pod do-vxrail-system
Name: do-vxrail-system-xxxxxxxxxx-aaaaa
Namespace: helium
Status: Running
Containers:
do-vxrail-system:
Image: do-main/do-vxrail-system:2.14.26
Port: 5000/TCP
State: Waiting
Reason: CrashLoopBackOff
Last State: Terminated
Reason: Error
Exit Code: 1
Started: <today> 08:37:26 +0000
Finished: <today> 08:37:26 +0000
Ready: False
Restart Count: 12
Limits:
memory: 1024M
Liveness: http-get http://:5000/do-vxrail-system/health-check delay=0s timeout=60s period=30sTwo details matter here. The container starts and finishes within the same second, and the exit code is 1. That rules out an out-of-memory kill (which would show OOMKilled) and a failing liveness probe (which needs minutes, not milliseconds). The application itself gives up immediately at startup, so the reason has to be in its log.
More than 30 pods were crash-looping. The handful that were running had all restarted at the same moment, which is simply the reboot. When that many services fail together, the pod you started with is a casualty, not the cause. Deleting pods one by one at this point would have achieved nothing.
vxm01:~ # kubectl get pods -n helium
NAME READY STATUS RESTARTS AGE
api-gateway-xxxxxxxxx-aaaaa 0/1 Unknown 1 509d
auth-service-xxxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 23 (45s ago) 509d
cacheservice-xxxxxxxxxx-aaaaa 1/1 Running 2 (66m ago) 509d
cms-service-xxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 23 (36s ago) 509d
compute-service-xxxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 24 (42s ago) 509d
do-cluster-xxxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 22 (53s ago) 509d
do-host-xxxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 22 (38s ago) 509d
do-vxrail-system-xxxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 21 (63s ago) 509d
...
lockbox-xxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 19 (2m38s ago) 509d
logging-xxxxxxxxx-aaaaa 1/1 Running 2 (47m ago) 509d
property-collector-xxxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 18 (2m37s ago) 509d
serviceregistry-xxxxxxxxx-aaaaa 1/1 Running 2 (47m ago) 509d
workflow-engine-xxxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 17 (2m13s ago) 509d
The –previous flag returns the log of the last terminated container instead of the one currently waiting. Addressing the deployment instead of the pod avoids typing the generated pod name.
vxm01:~ # kubectl logs -n helium deploy/do-vxrail-system --previous --tail=50
REGISTERED:true
Registering do-vxrail-system.helium.svc.cluster.local.
port:5000
curl: (7) Failed to connect to api-gateway port 8080 after 1 ms: Couldn't connect to server
The service tries to register itself at api-gateway on port 8080 and exits when that fails. The refusal after 1 ms is informative: the name resolves, but nothing is listening behind the service. That points away from DNS and toward a missing backend pod. The api-gateway pod is stuck in Unknown:
vxm01:~ # kubectl get pods -n helium | grep api-gateway
api-gateway-xxxxxxxxx-aaaaa 0/1 Unknown 1 509d
vxm01:~ # kubectl get endpoints -n helium api-gateway
NAME ENDPOINTS AGE
api-gateway 509d
The endpoints column is empty, which matches the refused connection. A describe of the pod showed all volumes and ConfigMaps in place and, at the bottom, Events: <none>. No image pull error, no failed mount. The kubelet was not even trying to start the pod any more; it had been left in Unknown across the reboot.
With no concrete error to fix, the reasonable move is to delete the pod and let the ReplicaSet create a fresh one. We took a snapshot of the VxRail Manager VM first.
vxm01:~ # kubectl delete pod -n helium api-gateway-xxxxxxxxx-aaaaa
pod "api-gateway-xxxxxxxxx-aaaaa" deleted
vxm01:~ # kubectl get pods -n helium | grep api-gateway
api-gateway-xxxxxxxxx-bbbbb 1/1 Running 0 25s
vxm01:~ # kubectl get endpoints -n helium api-gateway
NAME ENDPOINTS AGE
api-gateway 10.42.0.239:8080 509d
If the delete hangs. A pod in Unknown can stay in Terminating. In that case kubectl delete pod … –force –grace-period=0 removes it.
Immediately afterwards, the list of failing pods looked exactly the same. That is expected. After more than 20 restarts Kubernetes waits five minutes between attempts, and the last attempts had all happened before the new gateway existed. Nothing needed to be done except wait.
vxm01:~ # kubectl get pods -n helium | grep -v -E "Running|Completed"
NAME READY STATUS RESTARTS AGE
auth-service-xxxxxxxxxx-aaaaa 0/1 Error 27 (5m30s ago) 509d
compute-service-xxxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 28 (15s ago) 509d
infra-config-service-xxxxx-aaaaa 0/1 CrashLoopBackOff 21 (4m56s ago) 509d
(a few minutes later)
NAME READY STATUS RESTARTS AGE
compute-service-xxxxxxxxxx-aaaaa 0/1 CrashLoopBackOff 29 (40s ago) 509d
From more than 30 down to one. It would have been easy to stop here, delete the last pod and call it done. But compute-service had already failed several times with the gateway up, so it had a reason of its own. The straggler points at another service:
vxm01:~ # kubectl logs -n helium deploy/compute-service --previous --tail=50
...
compute-service [INFO] db.go tryToOpenDB() (110): successfully opened DB
compute-service [INFO] db.go initDB() (64): Datasource connected: postgres://10.42.0.38:5432/vxrail
compute-service [INFO] mq.go tryToConnectToMQ() (249): mq service connection success
compute-service [INFO] mq.go tryToWaitPropertyCollectorService() (182): http://property-collector:5000/propertycollector/health-check
compute-service [WARNING] mq.go tryToWaitPropertyCollectorService() (190): Get "http://property-collector:5000/propertycollector/health-check": context deadline exceeded (Client.Timeout exceeded while awaiting headers)
compute-service [WARNING] mq.go tryToWaitPropertyCollectorService() (191): property collector service seems not ready yet
compute-service [INFO] mq.go tryToWaitPropertyCollectorService() (192): sleep 10 seconds for 1 time
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x40 pc=0x...]
goroutine 1 [running]:
vxrail/compute-svc/compute/cmd/app.(*App).tryToWaitPropertyCollectorService(...)
/compute-svc/compute/cmd/app/mq.go:202 +0x26e
Database and message queue connect fine. The service then waits for the health check of property-collector, which does not answer within five seconds, and instead of retrying it dies with a nil pointer panic. The panic is the loud part of the log, but it is only a consequence. The real question is why property-collector does not answer, even though Kubernetes lists it as 1/1 Running.
A direct call from another pod confirmed it. The command ran and returned nothing, because -s hides the timeout message:
vxm01:~ # kubectl exec -n helium deploy/do-vxrail-system -- curl -s -m 15 http://property-collector:5000/propertycollector/health-check
vxm01:~ #
The root cause in the property-collector log:
vxm01:~ # kubectl logs -n helium deploy/property-collector --tail=50
Traceback (most recent call last):
File ".../do_common/connector/vsphere/soap_client.py", line 111, in _update_shared_session
self._make_fresh_new_conn()
File ".../do_common/connector/vsphere/do_smart_connect.py", line 61, in DoSmartConnect
return SmartConnect(**kwargs)
File ".../pyVim/connect.py", line 985, in SmartConnect
...
File "/usr/lib64/python3.11/ssl.py", line 1382, in do_handshake
self._sslobj.do_handshake()
ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: unable to get local issuer certificate (_ssl.c:1006)
property-collector [ERROR] vc_credential.py _connect() (68): SSLCertVerificationError(1, '[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: unable to get local issuer certificate (_ssl.c:1006)')
property-collector [WARNING] utils.py call() (99): Retryed 1 times, raise Exception('Could not connect to host vcenter01.example.local:443 with username svc-vxrail@vsphere.local') error
The last line reads like a credential problem, but the line above it says otherwise. unable to get local issuer certificate means the client does not have the CA that signed the server certificate. The VxRail Manager did not trust the vCenter certificate. The same check can be made with openssl against the trust store directory VxRail Manager uses. All of these commands are read-only.
vxm01:~ # openssl s_client -connect vcenter01.example.local:443 -verify_return_error -brief \
-CApath /var/lib/vmware-marvin/trust/lin -verify_hostname vcenter01.example.local </dev/null
depth=0 C = XX, O = Example Corp, OU = IT, CN = vcenter01.example.local
verify error:num=20:unable to get local issuer certificate
...:error:0A000086:SSL routines:tls_post_process_server_certificate:certificate verify failed:...
vxm01:~ # openssl s_client -connect vcenter01.example.local:443 </dev/null 2>/dev/null | openssl x509 -noout -issuer -dates
issuer=C = XX, O = Example Corp, CN = Example Internal Issuing CA
notBefore=Jun 5 11:08:47 2025 GMT
notAfter=Jun 5 11:08:47 2027 GMT
vxm01:~ # ls -l /var/lib/vmware-marvin/trust/lin/
-rw-r--r-- 1 tcserver pivotal 1489 May 13 2025 aaaaaaaa.0
-rw-r--r-- 1 tcserver pivotal 779 May 13 2025 aaaaaaaa.r0
Three outputs, one story. The trust store holds a single CA from the day the cluster was deployed. About three weeks later the vCenter machine certificate was replaced with one issued by the internal company CA. Nobody imported that chain into the VxRail Manager, and nothing complained, because the running services kept their established sessions. Sixteen months later a reboot forced a new handshake, and the missing trust finally surfaced.
Dell documents the import in KB 000077894. From VxRail 7.0.480 on, the script is already on the appliance at /mystic/ssl/cert_util.py and must be run as root.
vxm01:~ # python3 /mystic/ssl/cert_util.py --help
Running script with version 2024.07.11
usage: cert_util.py [-h] [-r] [-i]
-r, --regencert Regenerate VxRail Manager self signed cert.
-i, --vc_import Import vCenter Root CA certificates to VxRail Manager.
Mind the option. -i imports the vCenter CA certificates, which is what is needed here. -r regenerates the VxRail Manager’s own certificate, a different operation that was not the problem.
The script cleans out the existing trust store entries before writing the new ones, so we copied the directory first.
vxm01:~ # cp -a /var/lib/vmware-marvin/trust /root/trust-backup-$(date +%F)
vxm01:~ # python3 /mystic/ssl/cert_util.py -i
Running script with version 2024.07.11
Downloaded root CA certificate zip from vcenter01.example.local
Verify certificate against vCenter vcenter01.example.local
Found certificates ['certs/lin/aaaaaaaa.r0', 'certs/lin/bbbbbbbb.0', 'certs/lin/cccccccc.0', ...] that can verify server certificate
Clean up existing certificates in /var/lib/vmware-marvin/trust/
- Removing /var/lib/vmware-marvin/trust/lin/aaaaaaaa.r0
- Removing /var/lib/vmware-marvin/trust/lin/aaaaaaaa.0
Processing file [2/4]: certs/lin/bbbbbbbb.0
Root CA certificate /tmp/certs/lin is saved at /var/lib/vmware-marvin/trust/.
...
Delete saved CRL info in cacheservice...
Restarting vmware-marvin service...
The line to look for is Found certificates […] that can verify server certificate. If it is missing, vCenter itself does not publish the complete chain (for example a missing intermediate CA) and the CA certificates have to be imported by hand.
vxm01:~ # ls -l /var/lib/vmware-marvin/trust/lin/
-rw-r--r-- 1 tcserver pivotal 3073 Oct 5 10:07 bbbbbbbb.0
-rw-r--r-- 1 tcserver pivotal 2130 Oct 5 10:07 dddddddd.0
-rw-r--r-- 1 tcserver pivotal 1489 Oct 5 10:07 aaaaaaaa.0
-rw-r--r-- 1 tcserver pivotal 779 Oct 5 10:07 aaaaaaaa.r0
-rw-r--r-- 1 tcserver pivotal 2264 Oct 5 10:07 cccccccc.0
-rw-r--r-- 1 root root 1702 Oct 5 10:07 cccccccc.r0
vxm01:~ # openssl s_client -connect vcenter01.example.local:443 -verify_return_error -brief \
-CApath /var/lib/vmware-marvin/trust/lin -verify_hostname vcenter01.example.local </dev/null
CONNECTION ESTABLISHED
Protocol version: TLSv1.2
Peer certificate: C = XX, O = Example Corp, OU = IT, CN = vcenter01.example.local
Verification: OK
Verified peername: vcenter01.example.local
DONE
Now Restart the two pods. compute-service waits for property-collector at startup, so the collector should be up and answering before the other one is restarted.
vxm01:~ # kubectl delete pod -n helium property-collector-xxxxxxxxxx-aaaaa
pod "property-collector-xxxxxxxxxx-aaaaa" deleted
vxm01:~ # kubectl delete pod -n helium compute-service-xxxxxxxxxx-aaaaa
pod "compute-service-xxxxxxxxxx-aaaaa" deleted
vxm01:~ # kubectl get pods -n helium | grep -E "property-collector|compute-service"
property-collector-xxxxxxxxxx-bbbbb 1/1 Running 0 85s
compute-service-xxxxxxxxxx-bbbbb 1/1 Running 0 78s
vxm01:~ # kubectl get pods -n helium | grep -v -E "Running|Completed"
NAME READY STATUS RESTARTS AGE
An empty list. Every pod in the namespace was running again and the VxRail Page inside of vSphere was visible. That’s it from this Blog Post, if you have any questions use the comment section below or contact Dell Support.