Lars Kæraa Lücke
01/29/2024, 4:53 PM❯ docker run --memory=4g --cpus=2 -p 3000:3000 ann_bert_service:hwthbwf6wcew3nxt
tokenizer_config.json: 100%|██████████| 28.0/28.0 [00:00<00:00, 1.94kB/s]
vocab.txt: 100%|██████████| 872k/872k [00:00<00:00, 2.19MB/s]
tokenizer.json: 100%|██████████| 1.72M/1.72M [00:00<00:00, 3.55MB/s]
config.json: 100%|██████████| 625/625 [00:00<00:00, 179kB/s]
model.safetensors: 100%|██████████| 672M/672M [00:25<00:00, 26.2MB/s]
2024-01-29T16:47:12+0000 [INFO] [cli] Using the default model signature for pickable model ({'__call__': ModelSignature(batchable=False, batch_dim=(0, 0), input_spec=None, output_spec=None)}) for model "ann_model".
2024-01-29T16:47:14+0000 [INFO] [cli] Using the default model signature for pickable model ({'__call__': ModelSignature(batchable=False, batch_dim=(0, 0), input_spec=None, output_spec=None)}) for model "feature_extractor".
2024-01-29T16:47:27+0000 [INFO] [cli] Service loaded from Bento directory: bentoml.Service(tag="ann_bert_service:hwthbwf6wcew3nxt", path="/home/bentoml/bento/")
2024-01-29T16:47:29+0000 [INFO] [cli] Prometheus metrics for HTTP BentoServer from "/home/bentoml/bento" can be accessed at <http://localhost:3000/metrics>.
2024-01-29T16:47:30+0000 [INFO] [cli] Starting production HTTP BentoServer from "/home/bentoml/bento" listening on <http://0.0.0.0:3000> (Press CTRL+C to quit)
I see docker stats running into at least 1.7Gib OO 2Gib allocated. For Manjaro it appears to hug so many resources.
Anyone familiar with this? If not, I will update with more details later -- possibly on Github.