Slackbot
08/30/2023, 8:10 PMJoe Mifsud
08/30/2023, 8:11 PMclass IrisFeatures(BaseModel):
sepal_len: float
sepal_width: float
petal_len: float
petal_width: float
runner = bentoml.transformers.get("model_name").to_runner()
svc = bentoml.Service("iris_fastapi_demo", runners=[runner])
@svc.api(input=JSON(pydantic_model=IrisFeatures), output=NumpyNdarray())
def predict_bentoml(input_data: IrisFeatures) -> np.ndarray:
input_df = pd.DataFrame([input_data.dict()])
return runner.predict.run(input_df)
fastapi_app = FastAPI()
svc.mount_asgi_app(fastapi_app)
# For demo purpose, here's an identical inference endpoint implemented via FastAPI
@fastapi_app.post("/predict_fastapi")
def predict():
results = runner.run("TESTING 123")
return {"prediction": results.tolist()[0]}
# BentoML Runner's async API is recommended for async endpoints
@fastapi_app.post("/predict_fastapi_async")
async def predict_async(features: IrisFeatures):
input_df = pd.DataFrame([features.dict()])
results = await runner.predict.async_run(input_df)
return {"prediction": results.tolist()[0]}Jian Shen Yap
08/31/2023, 1:53 AMsvc = bentoml.Service("iris_fastapi_demo", runners=[iris_clf_runner])
on your first code example, its missing
will take a look on your second one!Joe Mifsud
08/31/2023, 1:54 AMsvc = bentoml.Service("iris_fastapi_demo", runners=[iris_clf_runner]) as well but it didn't seem to work, though I'm starting the fastapi application directly (not via bento), could that be the issue?Jian Shen Yap
08/31/2023, 1:54 AMbentoml serveJoe Mifsud
08/31/2023, 2:00 AMCMD ["uvicorn", "main:app", "--proxy-headers", "--host", "0.0.0.0", "--port", "8000"]Joe Mifsud
08/31/2023, 2:00 AMJian Shen Yap
08/31/2023, 3:28 AMbentoml serve and bentoml containerize route with the code above. you can specify the ports and other things as you would able to do on a normal uvicorn server.
2. another way to just use the runner as a standalone server. bentoml start-runner-server and you can configure your current application to forward the prediction request to the runner server.
Either this requires you to package the bento anyways!
we'd really like to do this incrementally / not touch whats known to work in our systemsRight, that's usually a good way to do things! I would suggest spinning a bentoml service using the recommended way (choice 1) and run some e2e test to ensure there's no regression on your application
Aaron Pham
08/31/2023, 12:53 PMJoe Mifsud
08/31/2023, 2:02 PMJoe Mifsud
09/05/2023, 7:12 PM