This message was deleted.
# ask-for-help
s
This message was deleted.
c
@sauyon @Aaron Pham could you help take a look?
s
This looks like the client is trying to write to a closed server connection, I think it has something to do with the way starlette handles connections. I had to work around this in the client by just creating a new session for every request.
s
@sauyon thank you for your reply. what client were you referring to? is the workaround something i had to implement on my end? any guidance is much appreciated.
s
Ah, sorry, I should have been clearer. I think this is an issue with the
aiohttp
session we're using to communicate with the runner server. Let me see if I can write something up to fix it really quickly.
s
much appreciate it! in the mean time, i will try to run the same jmeter warm up locally to see if i can reproduce it with
--production
.
I managed to recreate it locally with
--production
. My JMeter setting is 1 request per second for 60 users. 1 out of 60 threw the error above. I have many runner and api server instances on my local. But, on dev, I only have 1 api server and 1 runner because the service doesn’t need a lot of cpu and memory. XGBoost is pretty efficient in memory and prediction. I have 3 pods to match the prod environment. The local test confirmed that the error was not cause by lack of resources.
j
@sauyon
just creating a new session for every request.
That will re-open the tcp connection everytime, which may drag the server down.
I think we better make it clear why the server closed the connection first.
From my experience, it usually happens when the CPU is overloaded. Allocation more cluster resources should do the trick.
s
Yeah I'll probably just catch the error and retry n times.
s
okay. let me bump the CPU resources and report back.
The connection reset error went away after I bumped the CPU. It is unfortunate that the new architecture and tech stack utilizes more CPU. But, it is not too terrible. Thanks everyone for your help! The CPU bump helped to alleviate the problem. But, I am still getting connection reset once in a while. This doesn’t happen in BentoML 0.13.
s
This is more of a bug with the way we handle a closed TCP connection, we should really just retry after reconnecting if it's been closed.
s
that will be great! your fix also may help with the connection problem where the service is accessed for the first time after deployment.
s
Hm, do you have a link to the slack thread for that issue? I think I vaguely remember you reporting something but I don't quite remember what the exact context was.
s
hmm, i didn’t report it. but, i did come across similar thread related to warm-up problem. you are right that it is a bug because my xgboost inferences had 1 failed request out of 3600 requests from my load test. I was shooting for 12 requests per second. bentoml 0.13 can handle it without any failure. i confirmed from Datadog that it had enough CPU.
i am not sure if i am comfortable deploying the upgraded bentoml to production. but, i am stuck right now because if i need to deploy any updated model, i need to roll back to bentoml 0.13. Any thought @Chaoyu @Bo?
s
Ah, yeah, that was something else I was meaning to see if we could fix. I should really get around to putting up my proposal...
b
@Shihgian Lee let me dm you in our channel.
👍 1