This message was deleted.
# ask-for-help
s
This message was deleted.
🏁 1
j
are you using yatai or just pure BentoML ?
j
Pure bentoml, hosted on managed kubernetes cluster by Huawei Cloud
j
@Jiang have we encounter things like this before?
j
Yes we do have BentoCloud served on both 1.25 and 1.26. Never encountered an issue like that. Would you minding sharing your detailed error stack around "Permission Denied" ?
j
I do have some strace output for nvidia-smi. Is this what you need?
j
Yes
j
Copy code
execve("/usr/bin/nvidia-smi", ["nvidia-smi"], 0x7fff5a6004e0 /* 82 vars */) = -1 EACCES (Permission denied)
strace: exec: Permission denied
+++ exited with 1 +++
This is the strace output for nvidia-smi. The next strace is from a modified bento with root user access (Same image, just with root access). And the strace for nvidia-smi is as follows:
j
What's the permission of nvidia files?
ls -l /usr/bin/nvidia*
The result of
df
?
/usr/local/nvidia/bin/nvidia-smi
?
j
If bentoml user will get permission denied error. Need to use root to get this output
j
directly run
/usr/local/nvidia/bin/nvidia-smi
that's it
There're two things we know from here.
1. The image itself has a nvidia-smi, even CUDA toolkit installed. With NVIDIA enabled container runtime, it is not necessary. NVIDIA container toolkit will inject nvidia-smi and libs to the container automatically. A nvidia-smi shipped with the image may affect the behavior
Copy code
rm /usr/bin/nvidia-smi
nvidia-smi
test this script?
j
Run this in the node or in the pod?
j
In the pod, just for test.
Or fully remove CUDA/NVIDIA things from the image and retry.
2. The nvidia container runtime may not be correctly installed or configured. My nvidia docker will mount nvidia-smi to
/usr/bin/nvidia-smi
instead of
/usr/local/nvidia
. If your container runtime also acts like that, it should just work
If you can fix 2, then we can just ignore 1
How is the gpu-operator installed? @Janson Liew
j
Failed to remove nvidia-smi
j
It's okay. Just for an extra test.
Did you use custom base image in the bento?
j
As for how GPU operator is installed, the GPU operator is installed by a add-on provided by Huawei Cloud. Effectively we just provide the add-on with Nvidia-driver link, and it will install the driver
Not custom base image, the base img is python 3.9 cuda 11.6.2
j
I'm not familiar with Huawei Cloud addons. I'll suggest install gpu-opeator via NVIDIA's doc by helm.
Drivers could also be installed via helm values
driver.enabled=true
if the system image of nodes didn't have a driver pre-installed.
j
Hmm, I would prefer not to use Helm chart. Last time I use helm chart on HWC, there are a few blockers in terms of their implementation
Anyways, to summarize our discussion just now, the only thing left to check is the GPU driver installation, which might be the root issue of this? Since you guys have been running BentoCloud on 1.25 and 1.26 without issue, it should not be issue with cluster version transition.
j
Yeah it is an issue about the driver and NVIDIA libs.
You should ensure that you can see
CUDA version
in the
nvidia-smi
first. If you don't want to touch GPU operators, appending
/usr/local/nvidia/lib
or
/usr/local/nvidia/lib32
or
/usr/local/nvidia/lib64
to
LD_LIBRARY_PATH
may workaround it. Just a personal hint.
j
I see, thanks for the personal tips and your insights! I will raise a ticket to Huawei Cloud about this issue with all the evidence, hopefully it can pinpoint the issue
🍻 1
Copy code
after our analysis, the current issue is that the GPU plugin version used in your production cluster is 1.2.15, and the plugin version configures the permission of the /opt/cloud/cce/nvidia directory on the node to 755, which causes a permission issue in your production cluster. The GPU plugin version used in your test cluster is 1.2.28, which configures the permission of the plugin directory to 650, causing errors in your business.
The current solution is that you can log in to the node and give this directory 755 permission by executing the command "chmod 755 -R /opt/cloud/cce/nvidia".
Hi @Jiang, Huawei Cloud support has provided me this solution, and this does solve my issue of bentoml not able to run on 1.25 cluster. To explain more,
/opt/cloud/cce/nvidia
is the directory where the driver is installed on the node using their GPU add-on. From the message, it seems to suggest that setting driver directory to 650 is an intended behaviour and not a bug. Just want to know what do you think of this? Also, does bentocloud does this kind of workaround to your nodes? Personally I feel like they have misconfigured the GPU add-on 😅
j
Good to know you found a solution. (It seems like a bug in add-on 1.2.28, thus you will need a
current solution
, from my pov) Anyway, we just use gpu-opeator in BentoCloud. It works as promised, thus we don't need to workaround
j
I see, good to know that. Hopefully they will fix it in the near future. Again, thanks for your help!
j
Glad it helps 🍻
j
Speaking of BentoCloud, actually I have sent an email to contact@bentoml.com to enquire about bentocloud on behalf of my organization back at July but I have yet to receive a response. Can I know who to contact regarding this?
j
Sure, I will ping my colleagues to confirm. BentoCloud is currently undergoing significant feature iterations, and perhaps my colleagues are planning to get in touch with you and other customers in queue after the major updates are ready.
Can you please tell me the name of the organization?
j
The organization name is EM2AI