VxVerify: “DO host queries failed”

We wanted a VxVerify health check before updating the environment to the latest 8.0.x. The package was unpacked on the appliance and we started it the way we had done on older clusters.
vxm01:/home/mystic/tmp/vxv # ls -l
-rw-r--r-- 1 root root 53248 Oct 1 17:34 rule.db
-rwxrwxrwx 1 mystic root 923023 Oct 5 07:12 vxverify_3-61-002.pyc
-rwxrwxrwx 1 root root 8524 Oct 1 18:32 vxverify.sh
-rw-r--r-- 1 root root 499 Oct 1 17:34 vxverify.sha256
vxm01:/home/mystic/tmp/vxv # python vxverify_3-61-002.pyc -r root
Enter VCSA root credentials password:
Running VxVerify 3.61.002 healthcheck on VxRail 8.0.330.
In case of program errors consult article www.dell.com/support/kbdoc/000066460.
Step 1 of 10: Querying Lockbox and Config services for Management credentials
Step 0: Preliminary container tests in progress
Step 1: Running VxRM and VC API tests
Step 1: Querying DO-host and Lockbox services for Node credentials
DO host queries failed. See vxv.log for details
Rerun vxv as root or use sudo method as described in KB 21527
The message points in two directions, and both were wrong for us. “DO host queries failed” suggests that the do-host microservice or the ESXi nodes have a problem. “Rerun as root” suggests a permission problem, but the prompt already showed a root shell. Because the appliance had just been through an outage, we took the message at face value first and looked at the do-host pod.
vxm01:~ # kubectl logs -n helium deploy/do-host --tail=40
do-host [INFO] linzhi_port_cache.py is_server_reachable() (72): The host name esx04.example.local
do-host [INFO] linzhi_port_cache.py is_server_reachable() (74): connect successfully
do-host [INFO] linzhi_dataloader.py batch_load_fn() (72): The request url is https://esx04.example.local:39090/rest/ps/private/v1/settings/disk/led
...
do-host [INFO] linzhi_dataloader.py fetch_async() (102): fetch_async: the response from linzhi is <ClientResponse(https://esx04.example.local:39090/...) [200 OK]>
do-host [INFO] linzhi_dataloader.py fetch_async() (102): fetch_async: the response from linzhi is <ClientResponse(https://esx03.example.local:39090/...) [200 OK]>
do-host [INFO] linzhi_dataloader.py fetch_async() (102): fetch_async: the response from linzhi is <ClientResponse(https://esx02.example.local:39090/...) [200 OK]>
do-host [INFO] linzhi_dataloader.py fetch_async() (102): fetch_async: the response from linzhi is <ClientResponse(https://esx01.example.local:39090/...) [200 OK]>
All four nodes answered with 200 OK. The service was doing its job. Whatever stopped VxVerify, it was not the hosts. VxVerify writes its log to /tmp/vxv/vxv.log, regardless of the directory it is started from. A filter for error lines is enough.
vxm01:~ # grep -n -i -E "error|fail|timeout|certificate|ssl" /tmp/vxv/vxv.log | tail -40
603: 11:26:15-ERROR [config-service] <class 'NameError'> for key system_ntp failed: Traceback (most recent call last):
612: NameError: name 'aiohttp' is not defined
621: 11:26:15-ERROR [config-service] <class 'NameError'> for key system_sso_tenant failed: Traceback (most recent call last):
630: NameError: name 'aiohttp' is not defined
...
783: 11:26:15-ERROR [config-service] <class 'NameError'> for key witness_vm_host failed: Traceback (most recent call last):
792: NameError: name 'aiohttp' is not defined
806: 11:26:15-ERROR [do_host_q] <class 'NameError'> for query {configuredHosts {
828: NameError: name 'aiohttp' is not defined
837: 11:26:15-WARNING Failed do-host: {}
838: 11:26:15-WARNING DO host queries failed. See vxv.log for details
Three things stand out. Every query fails, not only the one to do-host. They all fail within the same second. And the error is a Python NameError, not a timeout, an HTTP status or a TLS error. The tool never reached the network. Its HTTP client library, aiohttp, was not available to the interpreter it was running in, so each query died before it was sent. The do-host query merely happened to be the one whose failure is printed on screen.
python is a link to Python 3.6. Python 3.11 is installed next to it but is only used when called by its full name. A compiled .pyc file is tied to the interpreter version it was built for, which is why VxVerify ships in two builds: 3.x for Python 3.6 and 4.x for Python 3.11. Dell KB 000021527 documents the 4.x build with python3.11 for VxRail 8.0.x.
So the 3.x file did start under python, because the versions matched, and then failed inside because this appliance release no longer provides the modules it expects for Python 3.6.
vxm01:~ # ls -l /usr/bin/python*
lrwxrwxrwx 1 root root 7 Nov 28 2023 /usr/bin/python -> python3
lrwxrwxrwx 1 root root 9 Feb 10 2025 /usr/bin/python3 -> python3.6
-rwxr-xr-x 1 root root 6288 Feb 10 2025 /usr/bin/python3.11
-rwxr-xr-x 2 root root 10560 Feb 10 2025 /usr/bin/python3.6
...
We got the 4.x build from the package and started it the same way as before.
vxm01:/home/mystic/tmp/vxv # python vxverify_4-61-002.pyc -r root
RuntimeError: Bad magic number in .pyc file
This error is at least an honest one. “Bad magic number” means the file was compiled for a different Python version than the one trying to load it. The 4.x build under Python 3.6 cannot work.
The Fix
We got the 4.x build from the package and started it the same way as before.
vxm01:/home/mystic/tmp/vxv # python vxverify_4-61-002.pyc -r root
RuntimeError: Bad magic number in .pyc file
This error is at least an honest one. “Bad magic number” means the file was compiled for a different Python version than the one trying to load it. The 4.x build under Python 3.6 cannot work.
A check that misled us. Before trying this we ran python3.11 -c “import aiohttp”, and it failed with ModuleNotFoundError, just as it did for Python 3.6. From that test alone, the 4.x build looked like it would fail too. It did not. We did not investigate how the 4.x build gets its dependencies, so the takeaway is only this: a manual import test in the system interpreter does not predict whether the tool will run. Try the documented invocation first.
That’s it from this Blog post, if you have any questions use the comment section below or contact Dell Support.