This message was deleted.
# ask-for-help
s
This message was deleted.
a
hmm, this is odd. It thinks the device is “busy or unavailable.” I tested like this before, but it may have been when using a different cuda version?
Copy code
bentoml@embeddings-6d7948b77-b2ljx:~/bento$ nvidia-smi
Thu Sep 28 16:40:40 2023
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 470.182.03   Driver Version: 470.182.03   CUDA Version: 11.6     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                               |                      |               MIG M. |
|===============================+======================+======================|
|   0  Tesla K80           On   | 00000000:00:1E.0 Off |                    0 |
| N/A   26C    P8    29W / 149W |      0MiB / 11441MiB |      0%      Default |
|                               |                      |                  N/A |
+-------------------------------+----------------------+----------------------+

+-----------------------------------------------------------------------------+
| Processes:                                                                  |
|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |
|        ID   ID                                                   Usage      |
|=============================================================================|
|  No running processes found                                                 |
+-----------------------------------------------------------------------------+
bentoml@embeddings-6d7948b77-b2ljx:~/bento$ python3
Python 3.11.5 (main, Aug 25 2023, 13:19:53) [GCC 9.4.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> torch
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
NameError: name 'torch' is not defined
>>> import torch
>>> torch.cuda.is_available()
True
>>> from sentence_transformers import SentenceTransformer
/usr/lib/python3/dist-packages/requests/__init__.py:89: RequestsDependencyWarning: urllib3 (1.26.16) or chardet (3.0.4) doesn't match a supported version!
  warnings.warn("urllib3 ({}) or chardet ({}) doesn't match a supported "

>>>
>>> model = SentenceTransformer("thenlper/gte-base", cache_folder="./src/models")
>>> model.encode(["food is good"])
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/usr/local/lib/python3.11/dist-packages/sentence_transformers/SentenceTransformer.py", line 153, in encode
    <http://self.to|self.to>(device)
  File "/usr/local/lib/python3.11/dist-packages/torch/nn/modules/module.py", line 1145, in to
    return self._apply(convert)
           ^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/torch/nn/modules/module.py", line 797, in _apply
    module._apply(fn)
  File "/usr/local/lib/python3.11/dist-packages/torch/nn/modules/module.py", line 797, in _apply
    module._apply(fn)
  File "/usr/local/lib/python3.11/dist-packages/torch/nn/modules/module.py", line 797, in _apply
    module._apply(fn)
  [Previous line repeated 1 more time]
  File "/usr/local/lib/python3.11/dist-packages/torch/nn/modules/module.py", line 820, in _apply
    param_applied = fn(param)
                    ^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/torch/nn/modules/module.py", line 1143, in convert
    return <http://t.to|t.to>(device, dtype if t.is_floating_point() or t.is_complex() else None, non_blocking)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
RuntimeError: CUDA error: CUDA-capable device(s) is/are busy or unavailable
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
hmm, I think perhaps the cuda 11.6 is incompatible with the k80. I’d need to downgrade to 11.4 according to this discussion https://www.reddit.com/r/MLQuestions/comments/137aex3/setting_up_cuda_for_tesla_k80/
No, I don’t think that’s it. Running now on a T4, and I am able to successfully utilize the gpu with a script, but the runner doesn’t seem to be using gpu still. I think I’ve configured something incorrectly in the bentoml_configuration.yaml causing the gpu to not be assigned to the runner
yup, that was it, I had cpus assigned to the runners in the bentoml_configuration, but not gpus 🤦