Slackbot
09/04/2023, 2:25 AMJian Shen Yap
09/04/2023, 6:12 AMJanson Liew
09/04/2023, 6:52 AMJian Shen Yap
09/04/2023, 6:52 AMJiang
09/04/2023, 7:04 AMJanson Liew
09/04/2023, 7:05 AMJiang
09/04/2023, 7:05 AMJanson Liew
09/04/2023, 7:08 AMexecve("/usr/bin/nvidia-smi", ["nvidia-smi"], 0x7fff5a6004e0 /* 82 vars */) = -1 EACCES (Permission denied)
strace: exec: Permission denied
+++ exited with 1 +++
This is the strace output for nvidia-smi.
The next strace is from a modified bento with root user access (Same image, just with root access). And the strace for nvidia-smi is as follows:Jiang
09/04/2023, 7:18 AMls -l /usr/bin/nvidia*Jiang
09/04/2023, 7:21 AMdf?Jiang
09/04/2023, 7:23 AM/usr/local/nvidia/bin/nvidia-smi ?Janson Liew
09/04/2023, 7:27 AMJiang
09/04/2023, 7:28 AM/usr/local/nvidia/bin/nvidia-smiJiang
09/04/2023, 7:29 AMJiang
09/04/2023, 7:30 AMJiang
09/04/2023, 7:31 AMJiang
09/04/2023, 7:33 AMrm /usr/bin/nvidia-smi
nvidia-smiJiang
09/04/2023, 7:34 AMJanson Liew
09/04/2023, 7:36 AMJiang
09/04/2023, 7:36 AMJiang
09/04/2023, 7:37 AMJiang
09/04/2023, 7:41 AM/usr/bin/nvidia-smi instead of /usr/local/nvidia .
If your container runtime also acts like that, it should just workJiang
09/04/2023, 7:45 AMJiang
09/04/2023, 7:46 AMJanson Liew
09/04/2023, 7:47 AMJiang
09/04/2023, 7:47 AMJiang
09/04/2023, 7:47 AMJanson Liew
09/04/2023, 7:49 AMJanson Liew
09/04/2023, 7:50 AMJiang
09/04/2023, 7:51 AMJiang
09/04/2023, 7:53 AMdriver.enabled=true if the system image of nodes didn't have a driver pre-installed.Janson Liew
09/04/2023, 8:04 AMJanson Liew
09/04/2023, 8:04 AMJiang
09/04/2023, 8:06 AMJiang
09/04/2023, 8:10 AMCUDA version in the nvidia-smi first.
If you don't want to touch GPU operators, appending /usr/local/nvidia/lib or /usr/local/nvidia/lib32 or /usr/local/nvidia/lib64 to
LD_LIBRARY_PATH may workaround it. Just a personal hint.Janson Liew
09/04/2023, 8:29 AMJanson Liew
09/04/2023, 10:02 AMafter our analysis, the current issue is that the GPU plugin version used in your production cluster is 1.2.15, and the plugin version configures the permission of the /opt/cloud/cce/nvidia directory on the node to 755, which causes a permission issue in your production cluster. The GPU plugin version used in your test cluster is 1.2.28, which configures the permission of the plugin directory to 650, causing errors in your business.
The current solution is that you can log in to the node and give this directory 755 permission by executing the command "chmod 755 -R /opt/cloud/cce/nvidia".
Hi @Jiang, Huawei Cloud support has provided me this solution, and this does solve my issue of bentoml not able to run on 1.25 cluster.
To explain more, /opt/cloud/cce/nvidia is the directory where the driver is installed on the node using their GPU add-on. From the message, it seems to suggest that setting driver directory to 650 is an intended behaviour and not a bug.
Just want to know what do you think of this? Also, does bentocloud does this kind of workaround to your nodes? Personally I feel like they have misconfigured the GPU add-on 😅Jiang
09/04/2023, 10:07 AMcurrent solution, from my pov)
Anyway, we just use gpu-opeator in BentoCloud. It works as promised, thus we don't need to workaroundJanson Liew
09/04/2023, 10:14 AMJiang
09/04/2023, 10:15 AMJanson Liew
09/04/2023, 10:17 AMJiang
09/04/2023, 10:23 AMJiang
09/04/2023, 10:24 AMJanson Liew
09/04/2023, 10:34 AM