07.10.2026

VxVerify: “DO host queries failed”

By H. Cemre Günay

We wanted a VxVerify health check before updating the environment to the latest 8.0.x. The package was unpacked on the appliance and we started it the way we had done on older clusters.


vxm01:/home/mystic/tmp/vxv # ls -l
-rw-r--r-- 1 root   root   53248 Oct  1 17:34 rule.db
-rwxrwxrwx 1 mystic root  923023 Oct  5 07:12 vxverify_3-61-002.pyc
-rwxrwxrwx 1 root   root    8524 Oct  1 18:32 vxverify.sh
-rw-r--r-- 1 root   root     499 Oct  1 17:34 vxverify.sha256
 
vxm01:/home/mystic/tmp/vxv # python vxverify_3-61-002.pyc -r root
Enter VCSA root credentials password:
Running VxVerify 3.61.002 healthcheck on VxRail 8.0.330.
In case of program errors consult article www.dell.com/support/kbdoc/000066460.
Step 1 of 10: Querying Lockbox and Config services for Management credentials
Step 0: Preliminary container tests in progress
Step 1: Running VxRM and VC API tests
Step 1: Querying DO-host and Lockbox services for Node credentials
DO host queries failed. See vxv.log for details
Rerun vxv as root or use sudo method as described in KB 21527

The message points in two directions, and both were wrong for us. “DO host queries failed” suggests that the do-host microservice or the ESXi nodes have a problem. “Rerun as root” suggests a permission problem, but the prompt already showed a root shell. Because the appliance had just been through an outage, we took the message at face value first and looked at the do-host pod.


vxm01:~ # kubectl logs -n helium deploy/do-host --tail=40
do-host [INFO] linzhi_port_cache.py is_server_reachable() (72): The host name esx04.example.local
do-host [INFO] linzhi_port_cache.py is_server_reachable() (74): connect successfully
do-host [INFO] linzhi_dataloader.py batch_load_fn() (72): The request url is https://esx04.example.local:39090/rest/ps/private/v1/settings/disk/led
...
do-host [INFO] linzhi_dataloader.py fetch_async() (102): fetch_async: the response from linzhi is <ClientResponse(https://esx04.example.local:39090/...) [200 OK]>
do-host [INFO] linzhi_dataloader.py fetch_async() (102): fetch_async: the response from linzhi is <ClientResponse(https://esx03.example.local:39090/...) [200 OK]>
do-host [INFO] linzhi_dataloader.py fetch_async() (102): fetch_async: the response from linzhi is <ClientResponse(https://esx02.example.local:39090/...) [200 OK]>
do-host [INFO] linzhi_dataloader.py fetch_async() (102): fetch_async: the response from linzhi is <ClientResponse(https://esx01.example.local:39090/...) [200 OK]>

All four nodes answered with 200 OK. The service was doing its job. Whatever stopped VxVerify, it was not the hosts. VxVerify writes its log to /tmp/vxv/vxv.log, regardless of the directory it is started from. A filter for error lines is enough.


vxm01:~ # grep -n -i -E "error|fail|timeout|certificate|ssl" /tmp/vxv/vxv.log | tail -40
603: 11:26:15-ERROR    [config-service] <class 'NameError'> for key system_ntp failed: Traceback (most recent call last):
612: NameError: name 'aiohttp' is not defined
621: 11:26:15-ERROR    [config-service] <class 'NameError'> for key system_sso_tenant failed: Traceback (most recent call last):
630: NameError: name 'aiohttp' is not defined
...
783: 11:26:15-ERROR    [config-service] <class 'NameError'> for key witness_vm_host failed: Traceback (most recent call last):
792: NameError: name 'aiohttp' is not defined
806: 11:26:15-ERROR    [do_host_q] <class 'NameError'> for query {configuredHosts {
828: NameError: name 'aiohttp' is not defined
837: 11:26:15-WARNING  Failed do-host: {}
838: 11:26:15-WARNING  DO host queries failed. See vxv.log for details

Three things stand out. Every query fails, not only the one to do-host. They all fail within the same second. And the error is a Python NameError, not a timeout, an HTTP status or a TLS error. The tool never reached the network. Its HTTP client library, aiohttp, was not available to the interpreter it was running in, so each query died before it was sent. The do-host query merely happened to be the one whose failure is printed on screen.

python is a link to Python 3.6. Python 3.11 is installed next to it but is only used when called by its full name. A compiled .pyc file is tied to the interpreter version it was built for, which is why VxVerify ships in two builds: 3.x for Python 3.6 and 4.x for Python 3.11. Dell KB 000021527 documents the 4.x build with python3.11 for VxRail 8.0.x.

So the 3.x file did start under python, because the versions matched, and then failed inside because this appliance release no longer provides the modules it expects for Python 3.6.


vxm01:~ # ls -l /usr/bin/python*
lrwxrwxrwx 1 root root     7 Nov 28  2023 /usr/bin/python -> python3
lrwxrwxrwx 1 root root     9 Feb 10  2025 /usr/bin/python3 -> python3.6
-rwxr-xr-x 1 root root  6288 Feb 10  2025 /usr/bin/python3.11
-rwxr-xr-x 2 root root 10560 Feb 10  2025 /usr/bin/python3.6
...

We got the 4.x build from the package and started it the same way as before.


vxm01:/home/mystic/tmp/vxv # python vxverify_4-61-002.pyc -r root
RuntimeError: Bad magic number in .pyc file

This error is at least an honest one. “Bad magic number” means the file was compiled for a different Python version than the one trying to load it. The 4.x build under Python 3.6 cannot work.

The Fix

We got the 4.x build from the package and started it the same way as before.


vxm01:/home/mystic/tmp/vxv # python vxverify_4-61-002.pyc -r root
RuntimeError: Bad magic number in .pyc file

This error is at least an honest one. “Bad magic number” means the file was compiled for a different Python version than the one trying to load it. The 4.x build under Python 3.6 cannot work.

That’s it from this Blog post, if you have any questions use the comment section below or contact Dell Support.