Slackbot
08/08/2023, 9:45 AMMirdul Swarup
08/08/2023, 1:27 PMJian Shen Yap
08/08/2023, 3:44 PMMirdul Swarup
08/08/2023, 3:46 PMopenllm in /usr/local/lib/python3.10/dist-packages (0.2.17)Mirdul Swarup
08/08/2023, 3:48 PMMirdul Swarup
08/08/2023, 6:57 PMChaoyu
08/08/2023, 8:18 PMChaoyu
08/08/2023, 9:26 PMMirdul Swarup
08/08/2023, 9:28 PMAaron Pham
08/09/2023, 12:10 AMChaoyu
08/09/2023, 12:10 AMAaron Pham
08/09/2023, 12:11 AMtorch.cuda.device_count()?Chaoyu
08/09/2023, 12:14 AMAaron Pham
08/09/2023, 12:17 AMAaron Pham
08/09/2023, 12:17 AMAaron Pham
08/09/2023, 12:17 AMllm = openllm.Runner("opt", model_id="facebook/opt-125m", serialisation="legacy", init_local=True)
llm("What is the difference between a duck and a goose?")Chaoyu
08/09/2023, 12:33 AMChaoyu
08/09/2023, 12:33 AMRuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument index in method wrapper_CUDA__index_select)Chaoyu
08/09/2023, 12:34 AMIn [2]: import openllm
In [3]: llm = openllm.Runner("opt", model_id="facebook/opt-125m", serialisation="legacy", init_local=True)
...:
...: llm("What is the difference between a duck and a goose?")
2023-08-09 00:33:34.180240: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX512F AVX512_VNNI
To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags.
2023-08-09 00:33:34.294262: I tensorflow/core/util/port.cc:104] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
--------------------------------------------------------------------------
WARNING: No preset parameters were found for the device that Open MPI
detected:
Local host: 192-9-143-60
Device name: mlx5_0
Device vendor ID: 0x02c9
Device vendor part ID: 4126
Default device parameters will be used, which may result in lower
performance. You can edit any of the files specified by the
btl_openib_device_param_files MCA parameter to set values for your
device.
NOTE: You can turn off this warning by setting the MCA parameter
btl_openib_warn_no_device_params_found to 0.
--------------------------------------------------------------------------
--------------------------------------------------------------------------
No OpenFabrics connection schemes reported that they were able to be
used on a specific port. As such, the openib BTL (OpenFabrics
support) will be disabled for this port.
Local host: 192-9-143-60
Local device: mlx5_0
Local port: 1
CPCs attempted: udcm
--------------------------------------------------------------------------
/home/ubuntu/.local/lib/python3.8/site-packages/transformers/generation/utils.py:1468: UserWarning: You are calling .generate() with the `input_ids` being on a device type different than your model's device. `input_ids` is on cuda, whereas the model is on cpu. You may experience unexpected behaviors or slower generation. Please make sure that you have put `input_ids` to the correct device by calling for example input_ids = <http://input_ids.to|input_ids.to>('cpu') before running `.generate()`.
warnings.warn(
---------------------------------------------------------------------------
RuntimeError Traceback (most recent call last)
<ipython-input-3-1458a16db2be> in <module>
1 llm = openllm.Runner("opt", model_id="facebook/opt-125m", serialisation="legacy", init_local=True)
2
----> 3 llm("What is the difference between a duck and a goose?")
~/.local/lib/python3.8/site-packages/openllm/_llm.py in _wrapped_generate_run(__self, prompt, **kwargs)
1159 """
1160 prompt, generate_kwargs, postprocess_kwargs = self.sanitize_parameters(prompt, **kwargs)
-> 1161 return self.postprocess_generate(prompt, __self.generate.run(prompt, **generate_kwargs), **postprocess_kwargs)
1162
1163 def _wrapped_embeddings_run(__self: LLMRunner[M, T], prompt: str | list[str]) -> LLMEmbeddings:
~/.local/lib/python3.8/site-packages/bentoml/_internal/runner/runner.py in run(self, *args, **kwargs)
50
51 def run(self, *args: P.args, **kwargs: P.kwargs) -> R:
---> 52 return self.runner._runner_handle.run_method(self, *args, **kwargs)
53
54 async def async_run(self, *args: P.args, **kwargs: P.kwargs) -> R:
~/.local/lib/python3.8/site-packages/bentoml/_internal/runner/runner_handle/local.py in run_method(self, _LocalRunnerRef__bentoml_method, *args, **kwargs)
46 )
47
---> 48 return getattr(self._runnable, __bentoml_method.name)(*args, **kwargs)
49
50 async def async_run_method(
~/.local/lib/python3.8/site-packages/bentoml/_internal/runner/runnable.py in method(*args, **kwargs)
138 def __get__(self, obj: T, _: t.Type[T] | None = None) -> t.Callable[P, R]:
139 def method(*args: P.args, **kwargs: P.kwargs) -> R:
--> 140 return self.func(obj, *args, **kwargs)
141
142 return method
~/.local/lib/python3.8/site-packages/openllm/_llm.py in generate(_Runnable__self, prompt, **attrs)
1123 adapter_name = attrs.pop("adapter_name", None)
1124 if adapter_name is not None: __self.set_adapter(adapter_name)
-> 1125 return self.generate(prompt, **attrs)
1126
1127 @bentoml.Runnable.method(**method_signature(generate_sig))
~/.local/lib/python3.8/site-packages/openllm/models/opt/modeling_opt.py in generate(self, prompt, **attrs)
48 def generate(self, prompt: str, **attrs: t.Any) -> list[str]:
49 with torch.inference_mode():
---> 50 return self.tokenizer.batch_decode(self.model.generate(**self.tokenizer(prompt, return_tensors="pt").to(self.device), do_sample=True, generation_config=self.config.model_construct_env(**attrs).to_generation_config()), skip_special_tokens=True)
~/.local/lib/python3.8/site-packages/torch/utils/_contextlib.py in decorate_context(*args, **kwargs)
113 def decorate_context(*args, **kwargs):
114 with ctx_factory():
--> 115 return func(*args, **kwargs)
116
117 return decorate_context
~/.local/lib/python3.8/site-packages/transformers/generation/utils.py in generate(self, inputs, generation_config, logits_processor, stopping_criteria, prefix_allowed_tokens_fn, synced_gpus, assistant_model, streamer, **kwargs)
1586
1587 # 13. run sample
-> 1588 return self.sample(
1589 input_ids,
1590 logits_processor=logits_processor,
~/.local/lib/python3.8/site-packages/transformers/generation/utils.py in sample(self, input_ids, logits_processor, stopping_criteria, logits_warper, max_length, pad_token_id, eos_token_id, output_attentions, output_hidden_states, output_scores, return_dict_in_generate, synced_gpus, streamer, **model_kwargs)
2640
2641 # forward pass to get next token
-> 2642 outputs = self(
2643 **model_inputs,
2644 return_dict=True,
~/.local/lib/python3.8/site-packages/torch/nn/modules/module.py in _call_impl(self, *args, **kwargs)
1499 or _global_backward_pre_hooks or _global_backward_hooks
1500 or _global_forward_hooks or _global_forward_pre_hooks):
-> 1501 return forward_call(*args, **kwargs)
1502 # Do not call functions when jit is used
1503 full_backward_hooks, non_full_backward_hooks = [], []
~/.local/lib/python3.8/site-packages/transformers/models/opt/modeling_opt.py in forward(self, input_ids, attention_mask, head_mask, past_key_values, inputs_embeds, labels, use_cache, output_attentions, output_hidden_states, return_dict)
942
943 # decoder outputs consists of (dec_features, layer_state, dec_hidden, dec_attn)
--> 944 outputs = self.model.decoder(
945 input_ids=input_ids,
946 attention_mask=attention_mask,
~/.local/lib/python3.8/site-packages/torch/nn/modules/module.py in _call_impl(self, *args, **kwargs)
1499 or _global_backward_pre_hooks or _global_backward_hooks
1500 or _global_forward_hooks or _global_forward_pre_hooks):
-> 1501 return forward_call(*args, **kwargs)
1502 # Do not call functions when jit is used
1503 full_backward_hooks, non_full_backward_hooks = [], []
~/.local/lib/python3.8/site-packages/transformers/models/opt/modeling_opt.py in forward(self, input_ids, attention_mask, head_mask, past_key_values, inputs_embeds, use_cache, output_attentions, output_hidden_states, return_dict)
633
634 if inputs_embeds is None:
--> 635 inputs_embeds = self.embed_tokens(input_ids)
636
637 batch_size, seq_length = input_shape
~/.local/lib/python3.8/site-packages/torch/nn/modules/module.py in _call_impl(self, *args, **kwargs)
1499 or _global_backward_pre_hooks or _global_backward_hooks
1500 or _global_forward_hooks or _global_forward_pre_hooks):
-> 1501 return forward_call(*args, **kwargs)
1502 # Do not call functions when jit is used
1503 full_backward_hooks, non_full_backward_hooks = [], []
~/.local/lib/python3.8/site-packages/torch/nn/modules/sparse.py in forward(self, input)
160
161 def forward(self, input: Tensor) -> Tensor:
--> 162 return F.embedding(
163 input, self.weight, self.padding_idx, self.max_norm,
164 self.norm_type, self.scale_grad_by_freq, self.sparse)
~/.local/lib/python3.8/site-packages/torch/nn/functional.py in embedding(input, weight, padding_idx, max_norm, norm_type, scale_grad_by_freq, sparse)
2208 # remove once script supports set_grad_enabled
2209 _no_grad_embedding_renorm_(weight, input, max_norm, norm_type)
-> 2210 return torch.embedding(weight, input, padding_idx, scale_grad_by_freq, sparse)
2211
2212
RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument index in method wrapper_CUDA__index_select)Aaron Pham
08/09/2023, 1:43 AMMirdul Swarup
08/09/2023, 9:22 AM!openllm start opt --model-id "facebook/opt-125m" --serialisation legacy the model gets loaded but when given a prompt, I always receive an 500 error code which points that the model got loaded on CPU.Mirdul Swarup
08/09/2023, 4:38 PMChaoyu
08/09/2023, 7:19 PMChaoyu
08/09/2023, 7:19 PMMirdul Swarup
08/09/2023, 8:02 PMMirdul Swarup
08/09/2023, 8:02 PMMirdul Swarup
08/09/2023, 8:03 PMMirdul Swarup
08/09/2023, 8:03 PMChaoyu
08/09/2023, 8:58 PMMirdul Swarup
08/09/2023, 11:33 PMinput_ids being on a device type different than your model’s device. input_ids is on cuda, whereas the model is on cpu. You may experience unexpected behaviors or slower generation. Please make sure that you have put input_ids to the correct device by calling for example input_ids = input_ids.to(‘cpu’) before running .generate().
warnings.warn(
2023-08-09T233147+0000 [ERROR] [clillm mpt runner1] Exception in ASGI application
Traceback (most recent call last):
File “/usr/local/lib/python3.10/dist-packages/uvicorn/protocols/http/h11_impl.py”, line 408, in run_asgi
result = await app( # type: ignore[func-returns-value]
File “/usr/local/lib/python3.10/dist-packages/uvicorn/middleware/proxy_headers.py”, line 84, in call
return await self.app(scope, receive, send)
File “/usr/local/lib/python3.10/dist-packages/starlette/applications.py”, line 122, in call
await self.middleware_stack(scope, receive, send)
File “/usr/local/lib/python3.10/dist-packages/starlette/middleware/errors.py”, line 184, in call
raise exc
File “/usr/local/lib/python3.10/dist-packages/starlette/middleware/errors.py”, line 162, in call
await self.app(scope, receive, _send)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/http/traffic.py”, line 26, in call
await self.app(scope, receive, send)
File “/usr/local/lib/python3.10/dist-packages/opentelemetry/instrumentation/asgi/__init__.py”, line 580, in call
await self.app(scope, otel_receive, otel_send)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/http/instruments.py”, line 252, in call
await self.app(scope, receive, wrapped_send)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/http/access.py”, line 126, in call
await self.app(scope, receive, wrapped_send)
File “/usr/local/lib/python3.10/dist-packages/starlette/middleware/exceptions.py”, line 62, in call
await wrap_app_handling_exceptions(self.app, conn)(scope, receive, send)
File “/usr/local/lib/python3.10/dist-packages/starlette/_exception_handler.py”, line 57, in wrapped_app
raise exc
File “/usr/local/lib/python3.10/dist-packages/starlette/_exception_handler.py”, line 46, in wrapped_app
await app(scope, receive, sender)
File “/usr/local/lib/python3.10/dist-packages/starlette/routing.py”, line 727, in call
await route.handle(scope, receive, send)
File “/usr/local/lib/python3.10/dist-packages/starlette/routing.py”, line 285, in handle
await self.app(scope, receive, send)
File “/usr/local/lib/python3.10/dist-packages/starlette/routing.py”, line 74, in app
await wrap_app_handling_exceptions(app, request)(scope, receive, send)
File “/usr/local/lib/python3.10/dist-packages/starlette/_exception_handler.py”, line 57, in wrapped_app
raise exc
File “/usr/local/lib/python3.10/dist-packages/starlette/_exception_handler.py”, line 46, in wrapped_app
await app(scope, receive, sender)
File “/usr/local/lib/python3.10/dist-packages/starlette/routing.py”, line 69, in app
response = await func(request)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/runner_app.py”, line 273, in _request_handler
payload = await infer(params)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/marshal/dispatcher.py”, line 182, in _func
raise r
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/marshal/dispatcher.py”, line 377, in outbound_call
outputs = await self.callback(
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/runner_app.py”, line 253, in infer_single
ret = await runner_method.async_run(*params.args, **params.kwargs)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runner.py”, line 55, in async_run
return await self.runner._runner_handle.async_run_method(self, *args, **kwargs)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runner_handle/local.py”, line 59, in async_run_method
return await anyio.to_thread.run_sync(
File “/usr/local/lib/python3.10/dist-packages/anyio/to_thread.py”, line 33, in run_sync
return await get_asynclib().run_sync_in_worker_thread(
File “/usr/local/lib/python3.10/dist-packages/anyio/_backends/_asyncio.py”, line 877, in run_sync_in_worker_thread
return await future
File “/usr/local/lib/python3.10/dist-packages/anyio/_backends/_asyncio.py”, line 807, in run
result = context.run(func, *args)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runnable.py”, line 140, in method
return self.func(obj, *args, **kwargs)
File “/usr/local/lib/python3.10/dist-packages/openllm/_llm.py”, line 1125, in generate
return self.generate(prompt, **attrs)
File “/usr/local/lib/python3.10/dist-packages/openllm/models/mpt/modeling_mpt.py”, line 90, in generate
generated_tensors = self.model.generate(**inputs, **attrs)
File “/usr/local/lib/python3.10/dist-packages/torch/utils/_contextlib.py”, line 115, in decorate_context
return func(*args, **kwargs)
File “/usr/local/lib/python3.10/dist-packages/transformers/generation/utils.py”, line 1538, in generate
return self.greedy_search(
File “/usr/local/lib/python3.10/dist-packages/transformers/generation/utils.py”, line 2362, in greedy_search
outputs = self(
File “/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py”, line 1501, in _call_impl
return forward_call(*args, **kwargs)
File “/root/.cache/huggingface/modules/transformers_modules/72e5f594ce36f9cabfa2a9fd8f58b491eb467ee7/modeling_mpt.py”, line 270, in forward
outputs = self.transformer(input_ids=input_ids, past_key_values=past_key_values, attention_mask=attention_mask, prefix_mask=prefix_mask, sequence_id=sequence_id, return_dict=return_dict, output_attentions=output_attentions, output_hidden_states=output_hidden_states, use_cache=use_cache)
File “/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py”, line 1501, in _call_impl
return forward_call(*args, **kwargs)
File “/root/.cache/huggingface/modules/transformers_modules/72e5f594ce36f9cabfa2a9fd8f58b491eb467ee7/modeling_mpt.py”, line 168, in forward
tok_emb = self.wte(input_ids)
File “/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py”, line 1501, in _call_impl
return forward_call(*args, **kwargs)
File “/root/.cache/huggingface/modules/transformers_modules/72e5f594ce36f9cabfa2a9fd8f58b491eb467ee7/custom_embedding.py”, line 11, in forward
return super().forward(input)
File “/usr/local/lib/python3.10/dist-packages/torch/nn/modules/sparse.py”, line 162, in forward
return F.embedding(
File “/usr/local/lib/python3.10/dist-packages/torch/nn/functional.py”, line 2210, in embedding
return torch.embedding(weight, input, padding_idx, scale_grad_by_freq, sparse)
RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument index in method wrapper_CUDA__index_select)
2023-08-09T23:31:47+0000 [ERROR] [cli:llm-mpt-service:2] Exception on /v1/generate [POST] (trace=7e5ca5643ff8d332dcdf527989ead501,span=b936e421df547ed2,sampled=1,service.name=llm-mpt-service)
Traceback (most recent call last):
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/http_app.py”, line 341, in api_func
output = await api.func(*args)
File “/usr/local/lib/python3.10/dist-packages/openllm/_service.py”, line 45, in generate_v1
responses = await runner.generate.async_run(qa_inputs.prompt, **config)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runner.py”, line 55, in async_run
return await self.runner._runner_handle.async_run_method(self, *args, **kwargs)
File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runner_handle/remote.py”, line 242, in async_run_method
raise RemoteException(
bentoml.exceptions.RemoteException: An unexpected exception occurred in remote runner llm-mpt-runner: [500] Internal Server Error
2023-08-09T23:31:47+0000 [INFO] [cli:llm-mpt-service:2] 66.135.13.190:40454 (scheme=http,method=POST,path=/v1/generate,type=application/json,length=762) (status=500,type=application/json,length=2) 3008.876ms (trace=7e5ca5643ff8d332dcdf527989ead501,span=b936e421df547ed2,sampled=1,service.name=llm-mpt-service)Mirdul Swarup
08/09/2023, 11:33 PMMirdul Swarup
08/09/2023, 11:34 PMMirdul Swarup
08/09/2023, 11:35 PMimport openllm
client = openllm.client.HTTPClient('<http://e.0.0.0nt.query>('Explain to me the difference between "further" and "farther"')Mirdul Swarup
08/10/2023, 4:53 PMChaoyu
08/10/2023, 6:58 PMMirdul Swarup
08/10/2023, 7:03 PMAaron Pham
08/10/2023, 7:03 PMMirdul Swarup
08/10/2023, 7:53 PMAaron Pham
08/10/2023, 7:53 PMMirdul Swarup
08/10/2023, 9:19 PMopenllm.client.HTTPClient initialisation throws an error saying openllm_client module not found. this persists on CLI as well as the manual python console.Aaron Pham
08/10/2023, 9:33 PMpython -c "import importlib.metadata, orjson;print(orjson.dumps([i.as_posix() for i in importlib.metadata.distribution('openllm').files], option=orjson.OPT_INDENT_2).decode());"Aaron Pham
08/10/2023, 9:37 PMAaron Pham
08/10/2023, 9:40 PMMirdul Swarup
08/10/2023, 9:47 PMfrom langchain.llms import OpenLLM would return ModuleNotFoundError: No module named 'langchain.llms'; 'langchain' is not a package on a server but not on google collab?Aaron Pham
08/10/2023, 9:49 PMAaron Pham
08/10/2023, 9:50 PMAaron Pham
08/10/2023, 9:50 PMMirdul Swarup
08/10/2023, 9:57 PMAaron Pham
08/10/2023, 9:58 PMMirdul Swarup
08/10/2023, 9:58 PMMirdul Swarup
08/10/2023, 9:58 PMMirdul Swarup
08/10/2023, 10:36 PMAaron Pham
08/11/2023, 12:13 PMMirdul Swarup
08/11/2023, 4:42 PMChaoyu
08/11/2023, 9:23 PM