This message was deleted.
# announcements
s
This message was deleted.
b
Hey @Nicolas Jaccard can you post your bento service definition and more context?
n
Yep!
Copy code
import bentoml
from bentoml.adapters import FileInput
from bentoml.frameworks.fastai import FastaiModelArtifact

@bentoml.env(requirements_txt_file='./requirements.txt')
@bentoml.artifacts([FastaiModelArtifact('learner')])
class DRGradingService2(bentoml.BentoService):
    @bentoml.api(input=FileInput(), batch=False)
    def predict(self, file):       
        grade, _, scores = self.artifacts.learner.predict(file.read())
        return scores.tolist()
This is the service definition
When calling
predict
from a notebook (using the same learner), I get a different output (not very different but still different)
b
@Nicolas Jaccard I see. We are tracking a similiar issue here: https://github.com/bentoml/BentoML/issues/1491 From looking at yours and the issue's bento service. I think the fastai's learner is different between saving and loading. BentoML is using
model.export()
and
fastai.basics.load_learner
to save/load artifact. Is that still the recommend methods for fastai? https://github.com/bentoml/BentoML/blob/50f7debcb8ff4a73c745dff8e5ed43d094e7ff71/bentoml/frameworks/fastai.py#L183
n
@Bo I think it is. I will try to
export
and re-load using
load_learner
, and see if this makes a difference!
b
Great. Let me know how it goes. I would love to get this issue solved.
n
So weirdly enough, it seems OK now, I restarted my environment and I have been unable to reproduce the error. Will keep an eye out!
b
were you working in a notebook environment?
n
Got an issue with
containerize
though (wanted to test in a Docker env to see if the same issue was happening). The image created crashes on startup with
TypeError: expected str, bytes or os.PathLike object, not NoneType
I was, yes!
b
Can you pull out the container log for us to debug?
n
Yep here it is:
Copy code
[2021-05-07 17:28:37,908] INFO - Starting BentoML proxy in production mode..
[2021-05-07 17:28:37,909] INFO - Starting BentoML API server in production mode..
[2021-05-07 17:28:38 +0000] [16] [INFO] Starting gunicorn 20.1.0
[2021-05-07 17:28:38,100] INFO - Running micro batch service on :5000
[2021-05-07 17:28:38 +0000] [16] [INFO] Listening at: <http://0.0.0.0:54905> (16)
[2021-05-07 17:28:38 +0000] [16] [INFO] Using worker: sync
[2021-05-07 17:28:38 +0000] [1] [INFO] Starting gunicorn 20.1.0
[2021-05-07 17:28:38 +0000] [1] [INFO] Listening at: <http://0.0.0.0:5000> (1)
[2021-05-07 17:28:38 +0000] [1] [INFO] Using worker: aiohttp.worker.GunicornWebWorker
[2021-05-07 17:28:38 +0000] [18] [INFO] Booting worker with pid: 18
[2021-05-07 17:28:38 +0000] [17] [INFO] Booting worker with pid: 17
[2021-05-07 17:28:38,140] INFO - Your system nofile limit is 1024, which means each instance of microbatch service is able to hold this number of connections at same time. You can increase the number of file descriptors for the server process, or launch more microbatch instances to accept more concurrent connection.
/opt/conda/lib/python3.6/site-packages/torch/cuda/__init__.py:52: UserWarning: CUDA initialization: Found no NVIDIA driver on your system. Please check that you have an NVIDIA GPU and installed a driver from <http://www.nvidia.com/Download/index.aspx> (Triggered internally at  /pytorch/c10/cuda/CUDAFunctions.cpp:100.)
  return torch._C._cuda_getDeviceCount() > 0
[2021-05-07 17:28:40 +0000] [17] [ERROR] Exception in worker process
Traceback (most recent call last):
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 589, in spawn_worker
    worker.init_process()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/workers/base.py", line 134, in init_process
    self.load_wsgi()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/workers/base.py", line 146, in load_wsgi
    self.wsgi = self.app.wsgi()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/app/base.py", line 67, in wsgi
    self.callable = self.load()
  File "/opt/conda/lib/python3.6/site-packages/bentoml/server/gunicorn_server.py", line 115, in load
    bento_service = load_from_dir(self.bento_service_bundle_path)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/saved_bundle/loader.py", line 107, in wrapper
    return func(bundle_path, *args)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/saved_bundle/loader.py", line 269, in load_from_dir
    svc = svc_cls()
  File "/opt/conda/lib/python3.6/site-packages/bentoml/service/__init__.py", line 488, in __init__
    self._config_artifacts()
  File "/opt/conda/lib/python3.6/site-packages/bentoml/service/__init__.py", line 557, in _config_artifacts
    self.artifacts.load_all(artifacts_path)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/service/artifacts/__init__.py", line 275, in load_all
    artifact.load(path)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/service/artifacts/__init__.py", line 158, in wrapped_load
    ret = original(*args, **kwargs)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/frameworks/fastai.py", line 183, in load
    model = fastai2_module.basics.load_learner(path + '/' + self._file_name)
  File "/opt/conda/lib/python3.6/site-packages/fastai/learner.py", line 375, in load_learner
    res = torch.load(fname, map_location='cpu' if cpu else None, pickle_module=pickle_module)
  File "/opt/conda/lib/python3.6/site-packages/torch/serialization.py", line 594, in load
    return _load(opened_zipfile, map_location, pickle_module, **pickle_load_args)
  File "/opt/conda/lib/python3.6/site-packages/torch/serialization.py", line 853, in _load
    result = unpickler.load()
ModuleNotFoundError: No module named 'timm'
[2021-05-07 17:28:40 +0000] [17] [INFO] Worker exiting (pid: 17)
[2021-05-07 17:28:40 +0000] [17] [WARNING] Exception during worker exit:
Traceback (most recent call last):
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 589, in spawn_worker
    worker.init_process()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/workers/base.py", line 134, in init_process
    self.load_wsgi()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/workers/base.py", line 146, in load_wsgi
    self.wsgi = self.app.wsgi()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/app/base.py", line 67, in wsgi
    self.callable = self.load()
  File "/opt/conda/lib/python3.6/site-packages/bentoml/server/gunicorn_server.py", line 115, in load
    bento_service = load_from_dir(self.bento_service_bundle_path)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/saved_bundle/loader.py", line 107, in wrapper
    return func(bundle_path, *args)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/saved_bundle/loader.py", line 269, in load_from_dir
    svc = svc_cls()
  File "/opt/conda/lib/python3.6/site-packages/bentoml/service/__init__.py", line 488, in __init__
    self._config_artifacts()
  File "/opt/conda/lib/python3.6/site-packages/bentoml/service/__init__.py", line 557, in _config_artifacts
    self.artifacts.load_all(artifacts_path)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/service/artifacts/__init__.py", line 275, in load_all
    artifact.load(path)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/service/artifacts/__init__.py", line 158, in wrapped_load
    ret = original(*args, **kwargs)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/frameworks/fastai.py", line 183, in load
    model = fastai2_module.basics.load_learner(path + '/' + self._file_name)
  File "/opt/conda/lib/python3.6/site-packages/fastai/learner.py", line 375, in load_learner
    res = torch.load(fname, map_location='cpu' if cpu else None, pickle_module=pickle_module)
  File "/opt/conda/lib/python3.6/site-packages/torch/serialization.py", line 594, in load
    return _load(opened_zipfile, map_location, pickle_module, **pickle_load_args)
  File "/opt/conda/lib/python3.6/site-packages/torch/serialization.py", line 853, in _load
    result = unpickler.load()
ModuleNotFoundError: No module named 'timm'

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 602, in spawn_worker
    sys.exit(self.WORKER_BOOT_ERROR)
SystemExit: 3

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 608, in spawn_worker
    self.cfg.worker_exit(self, worker)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/server/gunicorn_config.py", line 8, in worker_exit
    multiprocess.mark_process_dead(worker.pid)
  File "/opt/conda/lib/python3.6/site-packages/prometheus_client/multiprocess.py", line 161, in mark_process_dead
    for f in glob.glob(os.path.join(path, 'gauge_livesum_{0}.db'.format(pid))):
  File "/opt/conda/lib/python3.6/posixpath.py", line 80, in join
    a = os.fspath(a)
TypeError: expected str, bytes or os.PathLike object, not NoneType

Process Process-1:
Traceback (most recent call last):
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 209, in run
    self.sleep()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 357, in sleep
    ready = select.select([self.PIPE[0]], [], [], 1.0)
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 242, in handle_chld
    self.reap_workers()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 525, in reap_workers
    raise HaltServer(reason, self.WORKER_BOOT_ERROR)
gunicorn.errors.HaltServer: <HaltServer 'Worker failed to boot.' 3>

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 642, in kill_worker
    os.kill(pid, sig)
ProcessLookupError: [Errno 3] No such process

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/opt/conda/lib/python3.6/multiprocessing/process.py", line 258, in _bootstrap
    self.run()
  File "/opt/conda/lib/python3.6/multiprocessing/process.py", line 93, in run
    self._target(*self._args, **self._kwargs)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/server/__init__.py", line 192, in _start_prod_server
    gunicorn_app.run()
  File "/opt/conda/lib/python3.6/site-packages/bentoml/server/gunicorn_server.py", line 123, in run
    super(GunicornBentoServer, self).run()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/app/base.py", line 231, in run
    super().run()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/app/base.py", line 72, in run
    Arbiter(self).run()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 229, in run
    self.halt(reason=inst.reason, exit_status=inst.exit_status)
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 342, in halt
    self.stop()
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 390, in stop
    self.kill_workers(sig)
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 632, in kill_workers
    self.kill_worker(pid, sig)
  File "/opt/conda/lib/python3.6/site-packages/gunicorn/arbiter.py", line 648, in kill_worker
    self.cfg.worker_exit(self, worker)
  File "/opt/conda/lib/python3.6/site-packages/bentoml/server/gunicorn_config.py", line 8, in worker_exit
    multiprocess.mark_process_dead(worker.pid)
  File "/opt/conda/lib/python3.6/site-packages/prometheus_client/multiprocess.py", line 161, in mark_process_dead
    for f in glob.glob(os.path.join(path, 'gauge_livesum_{0}.db'.format(pid))):
  File "/opt/conda/lib/python3.6/posixpath.py", line 80, in join
    a = os.fspath(a)
TypeError: expected str, bytes or os.PathLike object, not NoneType
Just noticed that one of the error is due to
timm
not being installed, I can fix that
Not the NoneType one though
b
I see, can you try have
timm
installed and build the container to try again?
n
@Bo Finally had a chance to give it another go - it works when the
timm
dependency is correctly added to the service!
b
Great. Glad it works out for you. Let me know if I can help with anything else
Would love to learn more about your usecases and setup
n
@Bo We want to use BentoML to replace a bunch of custom tools that we use for model deployment. Our entire stack is based on k8s (currently hosted on AWS EKS but we are pretty much cloud agnostic). I am keen to explore the use of BentoML to handle all our deployment needs. In particular, I am keen to see what kind of performance we can get on Lambda. The only worry I have is regarding GPU support (not needed right now, but we have a few models that do require GPU, mostly object detection / segmentation)
j
Hello. I'm running into a very similar issue. When I use my trained model with bentoML, the weights are totally different from when I use the model directly from fastai. The inference is totally wrong. Also, I am using fastaiv1 while I see everyone else reporting this issue is using fasaiv2.
b
@Juan Acevedo I see. I think this issue related very close to how fastai is save/load models. Did you try to restart your notebook and see if still have different result?
j
Good morning @Bo yes I did try that. Same results. I also containerized the app and same results. I'm going to keep investigating. Thanks