This message was deleted.
# ask-for-help
s
This message was deleted.
j
Note if I run the example in the docs (with some small modifications to explicitly test my own model that's saved locally) I get the same runner not initialized error. I'm testing a hugging face pipeline model that's been saved locally
Copy code
class IrisFeatures(BaseModel):
    sepal_len: float
    sepal_width: float
    petal_len: float
    petal_width: float


runner = bentoml.transformers.get("model_name").to_runner()

svc = bentoml.Service("iris_fastapi_demo", runners=[runner])


@svc.api(input=JSON(pydantic_model=IrisFeatures), output=NumpyNdarray())
def predict_bentoml(input_data: IrisFeatures) -> np.ndarray:
    input_df = pd.DataFrame([input_data.dict()])
    return runner.predict.run(input_df)


fastapi_app = FastAPI()
svc.mount_asgi_app(fastapi_app)


# For demo purpose, here's an identical inference endpoint implemented via FastAPI
@fastapi_app.post("/predict_fastapi")
def predict():
    results = runner.run("TESTING 123")
    return {"prediction": results.tolist()[0]}


# BentoML Runner's async API is recommended for async endpoints
@fastapi_app.post("/predict_fastapi_async")
async def predict_async(features: IrisFeatures):
    input_df = pd.DataFrame([features.dict()])
    results = await runner.predict.async_run(input_df)
    return {"prediction": results.tolist()[0]}
j
Hey @Joe Mifsud, thanks for testing this feature out! the runner is initialized via
Copy code
svc = bentoml.Service("iris_fastapi_demo", runners=[iris_clf_runner])
on your first code example, its missing will take a look on your second one!
j
hey @Jian Shen Yap I tried it with
svc = bentoml.Service("iris_fastapi_demo", runners=[iris_clf_runner])
as well but it didn't seem to work, though I'm starting the fastapi application directly (not via bento), could that be the issue?
j
ah yeah, you have to start it with
bentoml serve
j
ok so does this require packaging the bento / etc? Or can I use bentoml serve to kick off my server similar to how I do now? This is how we currently kick off in docker:
Copy code
CMD ["uvicorn", "main:app", "--proxy-headers", "--host", "0.0.0.0", "--port", "8000"]
Asking because we'd really like to do this incrementally / not touch whats known to work in our systems
j
i think we have a few options here, 1. use the bentoml way of deploying, that means going down the
bentoml serve
and
bentoml containerize
route with the code above. you can specify the ports and other things as you would able to do on a normal uvicorn server. 2. another way to just use the runner as a standalone server.
bentoml start-runner-server
and you can configure your current application to forward the prediction request to the runner server. Either this requires you to package the bento anyways!
we'd really like to do this incrementally / not touch whats known to work in our systems
Right, that's usually a good way to do things! I would suggest spinning a bentoml service using the recommended way (choice 1) and run some e2e test to ensure there's no regression on your application
a
Hey it is often anti-pattern to include both save_model within the service definition.
j
what is the serve command doing under the hood? could we replicate that when our server boots?
@Jian Shen Yap @Aaron Pham would love to hear your thoughts on this!