This message was deleted.
# ask-for-help
s
This message was deleted.
j
I am attempting to use an endpoint that uploads a large file through the swaggerui page associated with the container. It seems that files larger than about 250-300mb don't get uploaded to the container. I setup remote debugging and was able to pinpoint what I believe to be the entry point for the server receiving the file, in bentoml/_internal/server/service_app.py, the asynch function api_func at line 302. I am running bentoml v1.0.5. I don't have this issue when running the container locally, only on an azure ml compute instance. Additionally, when running a different kind of container and uploading a file to it (like mltooling/mlworkspace), there is no issue with file uploads, only with the bentoml swagger ui page. Any ideas?
b
ingress might be limiting the upload size
what errors do u get?
j
Sometimes, there is no error and other times there is either an entity too large error or a timeout error
b
Copy code
<http://nginx.ingress.kubernetes.io/proxy-body-size|nginx.ingress.kubernetes.io/proxy-body-size>: 8m
Yup this one 👆 causes the entity too large
j
Is there a place that can be set?
Also this isnt a kubernetes instance
b
what's your setup like?
j
I have a containerized bento running on an azure Machine Learning compute instance. service.py looks like:
Copy code
#import os
#os.environ['CUDA_VISIBLE_DEVICES'] = ''
from warnings import catch_warnings
import bentoml
import logging
import shutil
from pathlib import Path
from datetime import datetime
import numpy as np
import openslide
import torch
from fastai.vision.all import *
import gdown

from icevision.models.checkpoint import model_from_checkpoint
import icevision.tfms.albumentations as A

from <http://bentoml.io|bentoml.io> import NumpyNdarray, Image, Multipart, JSON, Text, File

from bentoml_extensions.io_descriptors.wsi_file import WSIFile

from bentoml_extensions.environment import show_install
from MicroDetectFramework.slide_access import SlideContainer
from MicroDetectFramework.Segmentation.tissue_segmentation import TissueSegmentation

from PIL import ImageFile
ImageFile.LOAD_TRUNCATED_IMAGES = True
import PIL.Image
PIL.Image.MAX_IMAGE_PIXELS = int(1024 * 1024 * 1024) * 3 # 2GB max upload size


import hydra
from hydra import compose, initialize



hydra.core.global_hydra.GlobalHydra.instance().clear()
initialize(version_base=None, config_path="MitoticFigures/deployment/config")
cfg = compose(config_name="config.yaml", overrides=[])

# Logging
Path(cfg.deployment.logging.folder).mkdir(parents=True, exist_ok=True)
logger = logging.getLogger("ObjectDetection")
logger.setLevel(<http://logging.INFO|logging.INFO>)

# Stream to the console
sh = logging.StreamHandler()
sh.setLevel(getattr(logging, cfg.deployment.logging.level))
sh.setFormatter(logging.Formatter(cfg.deployment.logging.format))
logger.addHandler(sh)

# Stream to a file
# TimedRotatingFileHandler
fh = logging.handlers.TimedRotatingFileHandler(filename=Path(cfg.deployment.logging.folder) / cfg.deployment.logging.file,
                        when="midnight", backupCount=cfg.deployment.logging.backupCount)

fh.setLevel(getattr(logging, cfg.deployment.logging.level))
fh.setFormatter(logging.Formatter(cfg.deployment.logging.format))
logger.addHandler(fh)



<http://logger.info|logger.info>("starting server")

object_detection_runner = bentoml.pytorch.get(f"{cfg.object_detection.name}:{cfg.object_detection.tag}").to_runner()
second_stage_runner = bentoml.pytorch.get(f"{cfg.second_stage.name}:{cfg.second_stage.tag}").to_runner() #{cfg.second_stage.tag}
svc = bentoml.Service("gestalt-micro-detection-service", runners=[second_stage_runner, object_detection_runner]) #object_detection_runner, second_stage_runner

# Unfortuantely, icevision needs to load some meta information from the model.
checkpoint_path = cfg.object_detection.model_path
checkpoint_and_model = model_from_checkpoint(checkpoint_path)
model = checkpoint_and_model["model"] # free memory
model.eval()

model_type = checkpoint_and_model["model_type"]
backbone = checkpoint_and_model["backbone"]
class_map = checkpoint_and_model["class_map"]
patch_size = checkpoint_and_model["img_size"]
device = 'cuda' if torch.cuda.is_available() else 'cpu'
<http://model.to|model.to>(device)


valid_tfms = A.Adapter([*A.resize_and_pad(patch_size), A.Normalize(mean=cfg.object_detection.mean, std=cfg.object_detection.std)])
valid_tfms_second_stage = A.Compose([*A.resize_and_pad(cfg.second_stage.img_size), A.Normalize(mean=cfg.second_stage.mean, std=cfg.second_stage.std)])



@svc.api(input=Text(), output=JSON())
def enviroment(show_nvidia_smi:str="True"):

    enviroment = show_install(show_nvidia_smi=True)
    enviroment["config"] = str(cfg)

    return enviroment

@svc.api(input=Text(), output=JSON())
def start_debugger(text:str="None"):





@svc.api(input=Multipart(image=WSIFile(), meta=JSON()), output=JSON())
def predict_wsi(image, meta:JSON): #
    ....

 


@svc.api(input=Multipart(meta=JSON(), source_path=Text()), output=JSON())
def predict_path(source_path:str, meta): 
    ....

@svc.api(input=Multipart(image=Image(), meta=JSON()), output=JSON())
def predict_cell(image: Image, meta:JSON):
    ....
b
When you deploy this, is there a webserver in front of the BentoML service? I don't have zero experience with Azure, but basically it seems like the endpoint limiting the size of your file
another way could be instead of uploading the file you would point it to an object storage and the service would fetch the file from there
So bottom line it's not a Bento limitation. You'll need to figure out how to get around the Entity Too Large limitation in Azure
j
I posted this question on the microsoft forums:https://learn.microsoft.com/en-us/answers/questions/1090211/azure-ml-compute-instance-times-out-during-file-up.html The answer is somewhat confusing, since I am able to upload files over this supposed 512mb limit to the mltooling/ml-workspace container I deployed in a docker container on the same compute, and the files I have been trying to upload to the bento container have been less than that size. overall very confusing, but thanks for looking into it
b
woah they reply really quickly
what are you uploading that's so large? can it be compressed first then uncompressed in the service?
j
@Benjamin Tan We are uploading whole slide image files, which are .tiff and .svs. Sure, we could compress them and then decompress, but that would only get some of the files under the threshold, and wouldn't be sustainable because of that. I dug into the mltooling/mlworkspace code and found that the reason it is able to take large uploads on a container running on a az ml compute is because they set the maximum request body size in their nginx configuration to 10gb. How is the Bentoml Swagger ui page set up in relation to this? I noticed a deprecated configuration option for the max request size in the bentoml source code, why was that deprecated? Is any size accepted by default? I don't think Bentoml is using nginx, but I could be mistaken. Any clarification on how that all works and if there is any way to modify it (maybe through the swagger ui bundle that bentoml uses?) would be helpful, since it seems the folks at mltooling were able to get around this by setting a larger max request size with the setting "client_max_body_size" (https://github.com/ml-tooling/ml-workspace/blob/main/resources/nginx/nginx.conf)