This message was deleted.
# ask-for-help
s
This message was deleted.
🏁 1
s
Hi Paulo, yes, and that is the goal of adaptive batching feature in BentoML. Clients can make requests in parallel and the requests arrive around the same time can be batched. See more, https://docs.bentoml.org/en/latest/guides/batching.html
p
Thanks for the answer @Sean. Do you know why I have this error? Some requests come with no response when I try to run it with --production. But when I run with --reload the entries are batched and all requests are answered
s
Hi Paulo, could you please clarify what you meant by “the entries are batched”. The
--reload
parameter is designed for development environment. In this mode, adaptive batching (linked above) is disabled, so no batching is expected. To investigate why some requests come with no response, could you please share your code and reproducing steps?
p
@Sean, I'm trying to make concurrent requests to my predict method to test its efficiency. I tried in several ways and in all of them not all requests return the correct answer.
Copy code
import httpx
import asyncio
import time

async def post_async(data):

    async with httpx.AsyncClient() as client:
        return await <http://client.post|client.post>('<http://localhost:3000/model_binary/predict>', data=json.dumps(data))


async def launch():

    resps = await asyncio.gather(*map(post_async, all_context_tokens[:200]))
    data = [resp.text for resp in resps]
    print(data)
    

tm1 = time.perf_counter()

asyncio.run(launch())

tm2 = time.perf_counter()
print(f'Total time elapsed: {tm2-tm1:0.2f} seconds')
Here is my service code:
Copy code
@svc.api(input=NumpyNdarray(), output=NumpyNdarray(), route='/model_binary/predict')
async def model_binary_predict(input) :
    response = await binary_runner.predict.async_run([input])
    return response[0]
and here is my custom runner code:
Copy code
class BinaryModelRunnable(bentoml.Runnable):
    
    SUPPORTED_RESOURCES = ("<http://nvidia.com/gpu|nvidia.com/gpu>", "cpu")
    SUPPORTS_CPU_MULTI_THREADING = True

    def __init__(self):
        """
        Starts the class by loading the corresponding models
        """
        self.binary_model = bentoml.pytorch.load_model(binary_model)
        <http://self.binary_model.to|self.binary_model.to>(device)    
    

    @bentoml.Runnable.method(batchable=True, batch_dim=0 )
    def predict(self, input):
        
        input = input[0]
    
        all_ids = torch.tensor(
            [np.array(input[0])]
        )
        attention_mask = torch.tensor(
            [np.array(input[1])]
        )
        all_ids = <http://all_ids.to|all_ids.to>(torch.device('cuda'), dtype = torch.long)
        attention_mask = <http://attention_mask.to|attention_mask.to>(torch.device('cuda') , dtype = torch.long)

        prediction = []

        with torch.no_grad():
            output = self.binary_model(all_ids, attention_mask)
            prediction.extend(torch.sigmoid(output).cpu().detach().numpy().tolist())
        
        return np.array(prediction) 
        # return np.array([])
please help me, I've been stuck in this problem for a long time