This message was deleted.
# ask-for-help
s
This message was deleted.
e
I have some questions: • Is there a way that
bentoml.<flavor>.save_model(model)
can document the library versions used to create that model in the
model.yaml
file? The reason I ask is that I find myself having to document all the dependencies of a model in the
bentofile.yaml
, which feels like a strange non-separation of concerns. So the next question: • Is
bentofile.yaml
meant to describe the dependencies of the service only? Or the dependencies of both the service and all models that the service serves? If it's the latter, this requires us to know a lot about our models in advance of writing the
bentofile.yaml
. I feel that this prevents us from truly leveraging BentoML to create interfaces/abstractions. For example: what if I could write a service that does image classification with no knowledge about the underlying model? That would allow our data scientists to publish an image classifier using pytorch one day, tensorflow the next, and sklearn the next, all without having to modify the service. • Is
service.py
being used as a config file of sorts? The fact that
bentoml build ...
actually executes the
service.py
file has been causing us some grief. Our service endpoints require preprocessing logic, which require some fancy libraries like
opencv
and even C-binaries like
ffmpeg
to be installed. In order to execute the file, the
import
statements have to execute... and in order for those to execute, we have to have
opencv
and
ffmpeg
installed at the point that we run
bentoml build ...
. I'm probably being a noob, but our CI pipeline is pip installing the same requirements twice to build our docker images 🤣. a. install
opencv
and
ffmpeg
when we run
bentoml build
b. (not sure) does
bentoml build
turn around and install
opencv
again while it's locking the python requirements? c. install
opencv
and
ffmpeg
when we run
bentoml containerize
(because
docker build ...
installs
opencv
into a layer)
I kind of wish we could document our service name and model tags in a YAML file so
bentoml build ...
didn't have to execute
service.py
. That would save us this grief of needing to install certain dependencies before building bentos. Another option: I really like how
flask
and
FastAPI
allow you to define a "factory method" for creating your
svc
object. Here's an example of that from my open source project: https://github.com/rootski-io/rootski/blob/trunk/rootski_api/src/rootski/main/main.py#L27-L80 To launch the app, we run
Copy code
CMD gunicorn "rootski.main.main:create_default_app()" \
    --worker-class uvicorn.workers.UvicornWorker \
    --workers ${ROOTSKI__NUM_WORKERS} \
    --bind ${ROOTSKI__HOST}:${ROOTSKI__PORT} \
    --reload
You can see that
gunicorn
invokes
create_default_app()
on startup. This is cool for a few reasons: 1. Our
app
is not a global variable. This means we can create different versions of
app
for testing vs production. 2. Our singleton
Config
instance is not a global variable--supports [1] 3. Enables calling
import
only when the application is being run in production (if you put the
import
statements inside the
create_default_app()
function, which does feel ugly, but hey, it could work) If I could put my
import cv2
statement inside of a
create_default_app()
sort of function, I wouldn't have to install
opencv-python
and
ffmpeg
to execute
bentoml build ...
. That seems like a plus to me unless I'm missing something.
Typing this out helped me: maybe I could get around all this by doing no imports in
service.py
and instead importing
service:svc
into another file where I actually declare the endpoints. I think that would work except the case where you define custom
Runnables
that may need to import ML-specific libraries 😕
Last thought: it would be nice if
bentoml build ...
were smart enough to union the python/C-binary requirements of the service and all referenced models so that we didn't have to do that ourselves in advance in the
bentofile.yaml
. Would that violate an assumption of BentoML I haven't considered?
I imagine
bentoml build
would throw a warning or even an error message if it saw that a Model uses
sklearn
v2 and the
bentofile.yaml
declares
sklearn
v1. Or maybe a warning if the python versions don't match by a major version.
s
We actually do generally save the framework version into context, but I think that checking and automatically including model dependencies when packaging a bento is something we should absolutely do.
❤️ 1