This message was deleted.
# ask-for-help
s
This message was deleted.
m
message has been deleted
j
Hi @Mirdul Swarup, could you share the OpenLLM version and the python version that you are running on?
m
openllm in /usr/local/lib/python3.10/dist-packages (0.2.17)
Both latest
any solution to this ? any reason why this is happening ?
c
I can reproduce the same issue
cc @Aaron Pham any insight?
m
@Aaron Pham has updated the GitHub repo something related to CUDA GPU, but it’s not working
a
This is on one GPU correct?
c
Yes for me
a
What is the output of
torch.cuda.device_count()
?
c
1
a
There must been a regression with Langchain. Since they recently did a lot of refactor
Please try the following
Copy code
llm = openllm.Runner("opt", model_id="facebook/opt-125m", serialisation="legacy", init_local=True)

llm("What is the difference between a duck and a goose?")
c
Same error
Copy code
RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument index in method wrapper_CUDA__index_select)
Full log:
Copy code
In [2]: import openllm

In [3]: llm = openllm.Runner("opt", model_id="facebook/opt-125m", serialisation="legacy", init_local=True)
   ...:
   ...: llm("What is the difference between a duck and a goose?")
2023-08-09 00:33:34.180240: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations:  AVX512F AVX512_VNNI
To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags.
2023-08-09 00:33:34.294262: I tensorflow/core/util/port.cc:104] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
--------------------------------------------------------------------------
WARNING: No preset parameters were found for the device that Open MPI
detected:

  Local host:            192-9-143-60
  Device name:           mlx5_0
  Device vendor ID:      0x02c9
  Device vendor part ID: 4126

Default device parameters will be used, which may result in lower
performance.  You can edit any of the files specified by the
btl_openib_device_param_files MCA parameter to set values for your
device.

NOTE: You can turn off this warning by setting the MCA parameter
      btl_openib_warn_no_device_params_found to 0.
--------------------------------------------------------------------------
--------------------------------------------------------------------------
No OpenFabrics connection schemes reported that they were able to be
used on a specific port.  As such, the openib BTL (OpenFabrics
support) will be disabled for this port.

  Local host:           192-9-143-60
  Local device:         mlx5_0
  Local port:           1
  CPCs attempted:       udcm
--------------------------------------------------------------------------
/home/ubuntu/.local/lib/python3.8/site-packages/transformers/generation/utils.py:1468: UserWarning: You are calling .generate() with the `input_ids` being on a device type different than your model's device. `input_ids` is on cuda, whereas the model is on cpu. You may experience unexpected behaviors or slower generation. Please make sure that you have put `input_ids` to the correct device by calling for example input_ids = <http://input_ids.to|input_ids.to>('cpu') before running `.generate()`.
  warnings.warn(
---------------------------------------------------------------------------
RuntimeError                              Traceback (most recent call last)
<ipython-input-3-1458a16db2be> in <module>
      1 llm = openllm.Runner("opt", model_id="facebook/opt-125m", serialisation="legacy", init_local=True)
      2
----> 3 llm("What is the difference between a duck and a goose?")

~/.local/lib/python3.8/site-packages/openllm/_llm.py in _wrapped_generate_run(__self, prompt, **kwargs)
   1159     """
   1160     prompt, generate_kwargs, postprocess_kwargs = self.sanitize_parameters(prompt, **kwargs)
-> 1161     return self.postprocess_generate(prompt, __self.generate.run(prompt, **generate_kwargs), **postprocess_kwargs)
   1162
   1163   def _wrapped_embeddings_run(__self: LLMRunner[M, T], prompt: str | list[str]) -> LLMEmbeddings:

~/.local/lib/python3.8/site-packages/bentoml/_internal/runner/runner.py in run(self, *args, **kwargs)
     50
     51     def run(self, *args: P.args, **kwargs: P.kwargs) -> R:
---> 52         return self.runner._runner_handle.run_method(self, *args, **kwargs)
     53
     54     async def async_run(self, *args: P.args, **kwargs: P.kwargs) -> R:

~/.local/lib/python3.8/site-packages/bentoml/_internal/runner/runner_handle/local.py in run_method(self, _LocalRunnerRef__bentoml_method, *args, **kwargs)
     46                 )
     47
---> 48         return getattr(self._runnable, __bentoml_method.name)(*args, **kwargs)
     49
     50     async def async_run_method(

~/.local/lib/python3.8/site-packages/bentoml/_internal/runner/runnable.py in method(*args, **kwargs)
    138     def __get__(self, obj: T, _: t.Type[T] | None = None) -> t.Callable[P, R]:
    139         def method(*args: P.args, **kwargs: P.kwargs) -> R:
--> 140             return self.func(obj, *args, **kwargs)
    141
    142         return method

~/.local/lib/python3.8/site-packages/openllm/_llm.py in generate(_Runnable__self, prompt, **attrs)
   1123       adapter_name = attrs.pop("adapter_name", None)
   1124       if adapter_name is not None: __self.set_adapter(adapter_name)
-> 1125       return self.generate(prompt, **attrs)
   1126
   1127     @bentoml.Runnable.method(**method_signature(generate_sig))

~/.local/lib/python3.8/site-packages/openllm/models/opt/modeling_opt.py in generate(self, prompt, **attrs)
     48   def generate(self, prompt: str, **attrs: t.Any) -> list[str]:
     49     with torch.inference_mode():
---> 50       return self.tokenizer.batch_decode(self.model.generate(**self.tokenizer(prompt, return_tensors="pt").to(self.device), do_sample=True, generation_config=self.config.model_construct_env(**attrs).to_generation_config()), skip_special_tokens=True)

~/.local/lib/python3.8/site-packages/torch/utils/_contextlib.py in decorate_context(*args, **kwargs)
    113     def decorate_context(*args, **kwargs):
    114         with ctx_factory():
--> 115             return func(*args, **kwargs)
    116
    117     return decorate_context

~/.local/lib/python3.8/site-packages/transformers/generation/utils.py in generate(self, inputs, generation_config, logits_processor, stopping_criteria, prefix_allowed_tokens_fn, synced_gpus, assistant_model, streamer, **kwargs)
   1586
   1587             # 13. run sample
-> 1588             return self.sample(
   1589                 input_ids,
   1590                 logits_processor=logits_processor,

~/.local/lib/python3.8/site-packages/transformers/generation/utils.py in sample(self, input_ids, logits_processor, stopping_criteria, logits_warper, max_length, pad_token_id, eos_token_id, output_attentions, output_hidden_states, output_scores, return_dict_in_generate, synced_gpus, streamer, **model_kwargs)
   2640
   2641             # forward pass to get next token
-> 2642             outputs = self(
   2643                 **model_inputs,
   2644                 return_dict=True,

~/.local/lib/python3.8/site-packages/torch/nn/modules/module.py in _call_impl(self, *args, **kwargs)
   1499                 or _global_backward_pre_hooks or _global_backward_hooks
   1500                 or _global_forward_hooks or _global_forward_pre_hooks):
-> 1501             return forward_call(*args, **kwargs)
   1502         # Do not call functions when jit is used
   1503         full_backward_hooks, non_full_backward_hooks = [], []

~/.local/lib/python3.8/site-packages/transformers/models/opt/modeling_opt.py in forward(self, input_ids, attention_mask, head_mask, past_key_values, inputs_embeds, labels, use_cache, output_attentions, output_hidden_states, return_dict)
    942
    943         # decoder outputs consists of (dec_features, layer_state, dec_hidden, dec_attn)
--> 944         outputs = self.model.decoder(
    945             input_ids=input_ids,
    946             attention_mask=attention_mask,

~/.local/lib/python3.8/site-packages/torch/nn/modules/module.py in _call_impl(self, *args, **kwargs)
   1499                 or _global_backward_pre_hooks or _global_backward_hooks
   1500                 or _global_forward_hooks or _global_forward_pre_hooks):
-> 1501             return forward_call(*args, **kwargs)
   1502         # Do not call functions when jit is used
   1503         full_backward_hooks, non_full_backward_hooks = [], []

~/.local/lib/python3.8/site-packages/transformers/models/opt/modeling_opt.py in forward(self, input_ids, attention_mask, head_mask, past_key_values, inputs_embeds, use_cache, output_attentions, output_hidden_states, return_dict)
    633
    634         if inputs_embeds is None:
--> 635             inputs_embeds = self.embed_tokens(input_ids)
    636
    637         batch_size, seq_length = input_shape

~/.local/lib/python3.8/site-packages/torch/nn/modules/module.py in _call_impl(self, *args, **kwargs)
   1499                 or _global_backward_pre_hooks or _global_backward_hooks
   1500                 or _global_forward_hooks or _global_forward_pre_hooks):
-> 1501             return forward_call(*args, **kwargs)
   1502         # Do not call functions when jit is used
   1503         full_backward_hooks, non_full_backward_hooks = [], []

~/.local/lib/python3.8/site-packages/torch/nn/modules/sparse.py in forward(self, input)
    160
    161     def forward(self, input: Tensor) -> Tensor:
--> 162         return F.embedding(
    163             input, self.weight, self.padding_idx, self.max_norm,
    164             self.norm_type, self.scale_grad_by_freq, self.sparse)

~/.local/lib/python3.8/site-packages/torch/nn/functional.py in embedding(input, weight, padding_idx, max_norm, norm_type, scale_grad_by_freq, sparse)
   2208         # remove once script supports set_grad_enabled
   2209         _no_grad_embedding_renorm_(weight, input, max_norm, norm_type)
-> 2210     return torch.embedding(weight, input, padding_idx, scale_grad_by_freq, sparse)
   2211
   2212

RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument index in method wrapper_CUDA__index_select)
a
Ah it seems that opt has a separate loading logics. I’m doing a new release to fix this behaviour
👍 1
🚀 1
m
The error is not related to langchain as when i am running
!openllm start opt --model-id "facebook/opt-125m" --serialisation legacy
the model gets loaded but when given a prompt, I always receive an 500 error code which points that the model got loaded on CPU.
Ran the same commands with MPT still facing the same issue.
c
Yes it is an issue with the OPT models support in OpenLLM
Could you share the error message for the 500 error?
m
this is on the webapp
this is on the console
i tried dolly-v2
still the same issue
c
Could you include the full error trace above that line in the console as well? That should include the error information from the runner processes
m
2023-08-09T233141+0000 [INFO] [clillm mpt runner1] _ (scheme=http,method=GET,path=/readyz,type=,length=) (status=200,type=text/plain; charset=utf-8,length=1) 0.370ms (trace=2fcec5a6c6e2879ddf4ed2603521a0f6,span=dc1052a86ec5a3e5,sampled=1,service.name=llm-mpt-runner) 2023-08-09T233141+0000 [INFO] [clillm mpt service1] 66.135.13.190:40400 (scheme=http,method=GET,path=/readyz,type=,length=) (status=200,type=text/plain; charset=utf-8,length=1) 71.900ms (trace=2fcec5a6c6e2879ddf4ed2603521a0f6,span=107520b75f0ee72f,sampled=1,service.name=llm-mpt-service) 2023-08-09T233141+0000 [INFO] [clillm mpt runner1] _ (scheme=http,method=GET,path=/readyz,type=,length=) (status=200,type=text/plain; charset=utf-8,length=1) 0.306ms (trace=56e5f9f3f08156d8f74bf48fc7e19701,span=17dfb62518dff56f,sampled=1,service.name=llm-mpt-runner) 2023-08-09T233141+0000 [INFO] [clillm mpt service2] 66.135.13.190:40408 (scheme=http,method=GET,path=/readyz,type=,length=) (status=200,type=text/plain; charset=utf-8,length=1) 48.832ms (trace=56e5f9f3f08156d8f74bf48fc7e19701,span=8dcce142134b6d3c,sampled=1,service.name=llm-mpt-service) 2023-08-09T233141+0000 [INFO] [clillm mpt service1] 66.135.13.190:40410 (scheme=http,method=GET,path=/docs.json,type=,length=) (status=200,type=application/json,length=7294) 12.244ms (trace=10e0d3ae52e568db78725fb530b11a00,span=10aac2b8c75ff0a2,sampled=1,service.name=llm-mpt-service) 2023-08-09T233141+0000 [INFO] [clillm mpt service1] 66.135.13.190:40426 (scheme=http,method=POST,path=/v1/metadata,type=text/plain; charset=utf-8,length=0) (status=200,type=application/json,length=917) 2.298ms (trace=ed9350e4cb5b65551caeaa7decb6d4db,span=5adb931cfbc224a2,sampled=1,service.name=llm-mpt-service) 2023-08-09T233141+0000 [INFO] [clillm mpt service2] 66.135.13.190:40438 (scheme=http,method=POST,path=/v1/metadata,type=text/plain; charset=utf-8,length=0) (status=200,type=application/json,length=917) 2.189ms (trace=b40563ae4274782482e87063d8896820,span=423ba8f1d8589c5b,sampled=1,service.name=llm-mpt-service) /usr/local/lib/python3.10/dist-packages/transformers/generation/utils.py1468 UserWarning: You are calling .generate() with the
input_ids
being on a device type different than your model’s device.
input_ids
is on cuda, whereas the model is on cpu. You may experience unexpected behaviors or slower generation. Please make sure that you have put
input_ids
to the correct device by calling for example input_ids = input_ids.to(‘cpu’) before running
.generate()
. warnings.warn( 2023-08-09T233147+0000 [ERROR] [clillm mpt runner1] Exception in ASGI application Traceback (most recent call last): File “/usr/local/lib/python3.10/dist-packages/uvicorn/protocols/http/h11_impl.py”, line 408, in run_asgi result = await app( # type: ignore[func-returns-value] File “/usr/local/lib/python3.10/dist-packages/uvicorn/middleware/proxy_headers.py”, line 84, in call return await self.app(scope, receive, send) File “/usr/local/lib/python3.10/dist-packages/starlette/applications.py”, line 122, in call await self.middleware_stack(scope, receive, send) File “/usr/local/lib/python3.10/dist-packages/starlette/middleware/errors.py”, line 184, in call raise exc File “/usr/local/lib/python3.10/dist-packages/starlette/middleware/errors.py”, line 162, in call await self.app(scope, receive, _send) File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/http/traffic.py”, line 26, in call await self.app(scope, receive, send) File “/usr/local/lib/python3.10/dist-packages/opentelemetry/instrumentation/asgi/__init__.py”, line 580, in call await self.app(scope, otel_receive, otel_send) File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/http/instruments.py”, line 252, in call await self.app(scope, receive, wrapped_send) File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/http/access.py”, line 126, in call await self.app(scope, receive, wrapped_send) File “/usr/local/lib/python3.10/dist-packages/starlette/middleware/exceptions.py”, line 62, in call await wrap_app_handling_exceptions(self.app, conn)(scope, receive, send) File “/usr/local/lib/python3.10/dist-packages/starlette/_exception_handler.py”, line 57, in wrapped_app raise exc File “/usr/local/lib/python3.10/dist-packages/starlette/_exception_handler.py”, line 46, in wrapped_app await app(scope, receive, sender) File “/usr/local/lib/python3.10/dist-packages/starlette/routing.py”, line 727, in call await route.handle(scope, receive, send) File “/usr/local/lib/python3.10/dist-packages/starlette/routing.py”, line 285, in handle await self.app(scope, receive, send) File “/usr/local/lib/python3.10/dist-packages/starlette/routing.py”, line 74, in app await wrap_app_handling_exceptions(app, request)(scope, receive, send) File “/usr/local/lib/python3.10/dist-packages/starlette/_exception_handler.py”, line 57, in wrapped_app raise exc File “/usr/local/lib/python3.10/dist-packages/starlette/_exception_handler.py”, line 46, in wrapped_app await app(scope, receive, sender) File “/usr/local/lib/python3.10/dist-packages/starlette/routing.py”, line 69, in app response = await func(request) File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/runner_app.py”, line 273, in _request_handler payload = await infer(params) File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/marshal/dispatcher.py”, line 182, in _func raise r File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/marshal/dispatcher.py”, line 377, in outbound_call outputs = await self.callback( File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/runner_app.py”, line 253, in infer_single ret = await runner_method.async_run(*params.args, **params.kwargs) File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runner.py”, line 55, in async_run return await self.runner._runner_handle.async_run_method(self, *args, **kwargs) File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runner_handle/local.py”, line 59, in async_run_method return await anyio.to_thread.run_sync( File “/usr/local/lib/python3.10/dist-packages/anyio/to_thread.py”, line 33, in run_sync return await get_asynclib().run_sync_in_worker_thread( File “/usr/local/lib/python3.10/dist-packages/anyio/_backends/_asyncio.py”, line 877, in run_sync_in_worker_thread return await future File “/usr/local/lib/python3.10/dist-packages/anyio/_backends/_asyncio.py”, line 807, in run result = context.run(func, *args) File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runnable.py”, line 140, in method return self.func(obj, *args, **kwargs) File “/usr/local/lib/python3.10/dist-packages/openllm/_llm.py”, line 1125, in generate return self.generate(prompt, **attrs) File “/usr/local/lib/python3.10/dist-packages/openllm/models/mpt/modeling_mpt.py”, line 90, in generate generated_tensors = self.model.generate(**inputs, **attrs) File “/usr/local/lib/python3.10/dist-packages/torch/utils/_contextlib.py”, line 115, in decorate_context return func(*args, **kwargs) File “/usr/local/lib/python3.10/dist-packages/transformers/generation/utils.py”, line 1538, in generate return self.greedy_search( File “/usr/local/lib/python3.10/dist-packages/transformers/generation/utils.py”, line 2362, in greedy_search outputs = self( File “/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py”, line 1501, in _call_impl return forward_call(*args, **kwargs) File “/root/.cache/huggingface/modules/transformers_modules/72e5f594ce36f9cabfa2a9fd8f58b491eb467ee7/modeling_mpt.py”, line 270, in forward outputs = self.transformer(input_ids=input_ids, past_key_values=past_key_values, attention_mask=attention_mask, prefix_mask=prefix_mask, sequence_id=sequence_id, return_dict=return_dict, output_attentions=output_attentions, output_hidden_states=output_hidden_states, use_cache=use_cache) File “/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py”, line 1501, in _call_impl return forward_call(*args, **kwargs) File “/root/.cache/huggingface/modules/transformers_modules/72e5f594ce36f9cabfa2a9fd8f58b491eb467ee7/modeling_mpt.py”, line 168, in forward tok_emb = self.wte(input_ids) File “/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py”, line 1501, in _call_impl return forward_call(*args, **kwargs) File “/root/.cache/huggingface/modules/transformers_modules/72e5f594ce36f9cabfa2a9fd8f58b491eb467ee7/custom_embedding.py”, line 11, in forward return super().forward(input) File “/usr/local/lib/python3.10/dist-packages/torch/nn/modules/sparse.py”, line 162, in forward return F.embedding( File “/usr/local/lib/python3.10/dist-packages/torch/nn/functional.py”, line 2210, in embedding return torch.embedding(weight, input, padding_idx, scale_grad_by_freq, sparse) RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0!
Copy code
(when checking argument for argument index in method wrapper_CUDA__index_select)
2023-08-09T23:31:47+0000 [ERROR] [cli:llm-mpt-service:2] Exception on /v1/generate [POST] (trace=7e5ca5643ff8d332dcdf527989ead501,span=b936e421df547ed2,sampled=1,service.name=llm-mpt-service)
Traceback (most recent call last):
  File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/server/http_app.py”, line 341, in api_func
    output = await api.func(*args)
  File “/usr/local/lib/python3.10/dist-packages/openllm/_service.py”, line 45, in generate_v1
    responses = await runner.generate.async_run(qa_inputs.prompt, **config)
  File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runner.py”, line 55, in async_run
    return await self.runner._runner_handle.async_run_method(self, *args, **kwargs)
  File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/runner/runner_handle/remote.py”, line 242, in async_run_method
    raise RemoteException(
bentoml.exceptions.RemoteException: An unexpected exception occurred in remote runner llm-mpt-runner: [500] Internal Server Error
2023-08-09T23:31:47+0000 [INFO] [cli:llm-mpt-service:2] 66.135.13.190:40454 (scheme=http,method=POST,path=/v1/generate,type=application/json,length=762) (status=500,type=application/json,length=2) 3008.876ms (trace=7e5ca5643ff8d332dcdf527989ead501,span=b936e421df547ed2,sampled=1,service.name=llm-mpt-service)
here is the whole error
' File “/root/query.py”, line 3, in <module> client.query(‘Explain to me the difference between “further” and “farther”‘) File “/usr/local/lib/python3.10/dist-packages/openllm_client/runtimes/base.py”, line 206, in query else: result = self.call(“generate”, inputs.model_dump()) File “/usr/local/lib/python3.10/dist-packages/openllm_client/runtimes/base.py”, line 150, in call return self._cached.call(f”{name}_{self._api_version}“, *args, **attrs) File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/client/__init__.py”, line 53, in call return self._sync_call( File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/client/__init__.py”, line 133, in _sync_call return asyncio.run(self._call(inp, _bentoml_api=_bentoml_api, **kwargs)) File “/usr/lib/python3.10/asyncio/runners.py”, line 44, in run return loop.run_until_complete(main) File “/usr/lib/python3.10/asyncio/base_events.py”, line 646, in run_until_complete return future.result() File “/usr/local/lib/python3.10/dist-packages/bentoml/_internal/client/http.py”, line 171, in _call raise BentoMLException( bentoml.exceptions.BentoMLException: Error making request: 500: b’“”' '
this is when i run the query command:
import openllm
client = openllm.client.HTTPClient('<http://e.0.0.0nt.query>('Explain to me the difference between "further" and "farther"')
I have given the full console log @Chaoyu
c
Thanks @Mirdul Swarup, we will look into this issue today
m
thanks a lot, it would be really great if this issue can be solved as soon as possible. As a developer advocate i am excited to document and build on OpenLLM.
a
can you try again with opt on the latest version?
m
Sure
a
sorry on main
m
openllm.client.HTTPClient
initialisation throws an error saying
openllm_client
module not found. this persists on CLI as well as the manual python console.
a
can you send me the output of the following
Copy code
python -c "import importlib.metadata, orjson;print(orjson.dumps([i.as_posix() for i in importlib.metadata.distribution('openllm').files], option=orjson.OPT_INDENT_2).decode());"
Hmm it seems like compiled wheels won’t include arbitrary include files for some reason
Will release a patch for this promptly
m
can you also help me think of reasons why
from langchain.llms import OpenLLM
would return
ModuleNotFoundError: No module named 'langchain.llms'; 'langchain' is not a package
on a server but not on google collab?
a
sure, did you install langchain on the server?
within the bentofile.yaml, you need to add langchain as a python dependency
m
Yes I installed the Langchain on the server
a
Can you send the stacktrace then?
m
How is it that the same thing works on colab but not on a server
Sure I will send it
Also I guess GPU error is sorted, so thanks for taking care of that ✅
a
can you try with 0.2.22 to see you still see the client issue?
m
yes so the client issue is also resolved and GPU error is also resolved. The langchain error was from my side due to the file naming. Thanks a lot for being persistent and patient from the last 2 weeks. Looking forward to document and build on OpenLLM. @Chaoyu @Aaron Pham 🙌
🙌 1
c
Thank you so much for your patience and feedback @Mirdul Swarup!
🙌 1