This message was deleted.
# ask-for-help
s
This message was deleted.
c
Hi @Bertrand Thia - thank you for reporting this issue. What’s your BentoML version and OpenLLM version?
cc @Aaron Pham
m
Hi is your error solved ?
b
openllm, version 0.2.14.dev5 bentoml, version 1.1.1 Please let me know if there's anything else i can do to help!
m
is this openLLM version completely functional
?
can you tell me the exact commads you are using to query the results on openLLM
b
Not sure. Guess I’ll try to install 2.13 and see if this error is fixed πŸ€”
The only commands I used are above. Just trying to instantiate the llm returned an error
m
Actually I tried the latest version I got the safetensor error and when I tried older versions I got a 500 error
c
cc @Aaron Pham
b
When I try version 0.2.13 I get a new error which is:
Copy code
>>> llm = OpenLLM(model_name="dolly-v2", model_id='databricks/dolly-v2-12b')

Traceback (most recent call last):
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/huggingface_hub/utils/_errors.py", line 261, in hf_raise_for_status
    response.raise_for_status()
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/requests/models.py", line 1021, in raise_for_status
    raise HTTPError(http_error_msg, response=self)
requests.exceptions.HTTPError: 404 Client Error: Not Found for url: <https://huggingface.co/databricks/dolly-v2-3b/resolve/19308160448536e378e3db21a73a751579ee7fdd/config.json>

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/transformers/utils/hub.py", line 417, in cached_file
    resolved_file = hf_hub_download(
                    ^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/huggingface_hub/utils/_validators.py", line 118, in _inner_fn
    return fn(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/huggingface_hub/file_download.py", line 1195, in hf_hub_download
    metadata = get_hf_file_metadata(
               ^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/huggingface_hub/utils/_validators.py", line 118, in _inner_fn
    return fn(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/huggingface_hub/file_download.py", line 1541, in get_hf_file_metadata
    hf_raise_for_status(r)
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/huggingface_hub/utils/_errors.py", line 271, in hf_raise_for_status
    raise EntryNotFoundError(message, response) from e
huggingface_hub.utils._errors.EntryNotFoundError: 404 Client Error. (Request ID: Root=1-64cd1feb-71f594f81aa8c1ce240f8490;76c322e3-65f7-4f6d-8a3f-03ca33f4467a)

Entry Not Found for url: <https://huggingface.co/databricks/dolly-v2-3b/resolve/19308160448536e378e3db21a73a751579ee7fdd/config.json>.

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/_llm.py", line 612, in from_pretrained
    _tag = cls.generate_tag(model_id, model_version)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/_llm.py", line 661, in generate_tag
    def generate_tag(cls, *param_decls: t.Any, **attrs: t.Any) -> bentoml.Tag: return bentoml.Tag.from_taglike(cls._generate_tag_str(*param_decls, **attrs))
                                                                                                               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/utils/__init__.py", line 223, in <lambda>
    return lambda *args, **kwargs: f1(f2(*args, **kwargs))
                                      ^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/_llm.py", line 655, in _generate_tag_str
    _config = transformers.AutoConfig.from_pretrained(model_id, trust_remote_code=cls.config_class.__openllm_trust_remote_code__, revision=first_not_none(model_version, default="main"))
              ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/transformers/models/auto/configuration_auto.py", line 983, in from_pretrained
    config_dict, unused_kwargs = PretrainedConfig.get_config_dict(pretrained_model_name_or_path, **kwargs)
                                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/transformers/configuration_utils.py", line 617, in get_config_dict
    config_dict, kwargs = cls._get_config_dict(pretrained_model_name_or_path, **kwargs)
                          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/transformers/configuration_utils.py", line 672, in _get_config_dict
    resolved_config_file = cached_file(
                           ^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/transformers/utils/hub.py", line 463, in cached_file
    raise EnvironmentError(
OSError: databricks/dolly-v2-3b does not appear to have a file named config.json. Checkout '<https://huggingface.co/databricks/dolly-v2-3b/19308160448536e378e3db21a73a751579ee7fdd>' for available files.

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/cli/entrypoint.py", line 237, in wrapper
    return func(*args, **attrs)
           ^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/cli/entrypoint.py", line 215, in wrapper
    return_value = func(*args, **attrs)
                   ^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/click/decorators.py", line 33, in new_func
    return f(get_current_context(), *args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/cli/entrypoint.py", line 197, in wrapper
    return f(*args, **attrs)
           ^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/cli/entrypoint.py", line 459, in import_command
    llm = infer_auto_class(impl).for_model(model_name, llm_config=llm_config, model_version=model_version, ensure_available=False, serialisation=serialisation_format)
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/models/auto/factory.py", line 51, in for_model
    llm = cls.infer_class_from_name(model).from_pretrained(model_id, model_version=model_version, llm_config=llm_config, **attrs)
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/_llm.py", line 614, in from_pretrained
    except Exception as err: raise OpenLLMException(f"Failed to generate a valid tag for {cfg_cls.__openllm_start_name__} with 'model_id={model_id}' (lookup to see its traceback):\n{err}") from err
                             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
openllm.exceptions.OpenLLMException: Failed to generate a valid tag for dolly-v2 with 'model_id=databricks/dolly-v2-3b' (lookup to see its traceback):
databricks/dolly-v2-3b does not appear to have a file named config.json. Checkout '<https://huggingface.co/databricks/dolly-v2-3b/19308160448536e378e3db21a73a751579ee7fdd>' for available files.

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/langchain/llms/openllm.py", line 172, in __init__
    runner = openllm.Runner(
             ^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/_llm.py", line 1023, in Runner
    runner = infer_auto_class(implementation).create_runner(model_name, llm_config=llm_config, ensure_available=ensure_available if ensure_available is not None else init_local, **attrs)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/models/auto/factory.py", line 71, in create_runner
    return cls.for_model(model, model_id=model_id, **attrs).to_runner(**runner_attrs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/models/auto/factory.py", line 52, in for_model
    if ensure_available: llm.ensure_model_id_exists()
                         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/_llm.py", line 804, in ensure_model_id_exists
    def ensure_model_id_exists(self) -> bentoml.Model: return import_model(self.config["start_name"], model_id=self.model_id, model_version=self._model_version, runtime=self.runtime, implementation=self.__llm_implementation__, quantize=self._quantize_method, serialisation_format=self._serialisation_format)
                                                              ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/cli/entrypoint.py", line 664, in _import_model
    return import_command.main(args=args, standalone_mode=False)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/click/core.py", line 1078, in main
    rv = self.invoke(ctx)
         ^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/click/core.py", line 1434, in invoke
    return ctx.invoke(self.callback, **ctx.params)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/click/core.py", line 783, in invoke
    return __callback(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/anaconda3/envs/description-std/lib/python3.11/site-packages/openllm/cli/entrypoint.py", line 239, in wrapper
    raise click.ClickException(click.style(f"[{group.name}] '{command_name}' failed: " + err.message, fg="red")) from err
click.exceptions.ClickException: [openllm] 'import' failed: Failed to generate a valid tag for dolly-v2 with 'model_id=databricks/dolly-v2-3b' (lookup to see its traceback):
databricks/dolly-v2-3b does not appear to have a file named config.json. Checkout '<https://huggingface.co/databricks/dolly-v2-3b/19308160448536e378e3db21a73a751579ee7fdd>' for available files.
It's interesting that it's checking for a URL for
dolly-v2-3b
while i am trying to instantiate
dolly-v2-12b
a
This has to do with most of these model doesn’t have safe tensors format
m
so how do you suppose we can use the models
b
Is there any fix? I was just trying to replicate the instructions in the README
a
Copy code
from langchain.llms import OpenLLM

llm = OpenLLM(model_name="mpt", model_id="mosaicml/mpt-30b-chat", serialisation="legacy")
πŸ‘€ 2
m
okay, so this is for a very specific finetuned model, for future preferences if users want to use more models. how can they do it ?
also is there a list of models where we can know which models do not have the safe tensors format.
a
You can just see the huggingface repo whether they have safetensors format or not
m
alright, also i would suggest you guys to change the example code because that dosnt work as dolly-V2 dosnt have the safe tensor format
and also the command you gave is for langchain, how i set this when i am starting the model
c
Thanks so much for the feedback!
m
because this error comes up when we are downloading the model in initial stages
a
Copy code
openllm import ... --serialisation legacy | safetensors
c
@Aaron Pham we should make sure the default serialization for each model works
πŸ‘ 1
m
openllm start --help
should also include
--serialisation
option description
a
there is a doc for
openllm start
for serialisation
m
ill try to find that, thank you πŸ™‚
b
hey, I am still getting errors:
Copy code
>>> from langchain.llms import OpenLLM
>>> llm = OpenLLM(model_name="dolly-v2", model_id="databricks/dolly-v2-12b", temperature=0, serialisation="legacy")
'NoneType' object has no attribute 'cadam32bit_grad_fp32'
Downloading pytorch_model.bin: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 23.8G/23.8G [08:36<00:00, 46.1MB/s]
zsh: killed     python
Every time i try to instantiate the model like this, it's killing python for some reason. Is there anything i can do to fix this please?
Copy code
>>> from langchain.llms import OpenLLM
>>> llm = OpenLLM(model_name="dolly-v2", model_id="databricks/dolly-v2-12b", temperature=0, serialisation="legacy")
'NoneType' object has no attribute 'cadam32bit_grad_fp32'

zsh: killed     python
a
are you running this on macos?
πŸ‘ 1
I don’t think you can load a 12b on macos for now, ofc it will oom as it will try to load the whole model into memory
maybe if you have a ec2 machine with GPU you can try it
but on mac i suggest with the smaller model
b
oh, i see. Will try this on the cloud then, thanks
a
We are working on ggml integration for CPU usage atm, and it is being worked on by a community member. Running on CPU is currently not our focus right now, unless there is a specific usecase for this
πŸ‘ 3
m
--help
does not include other options, also, it fails to find the model ID for some reason, and
--model-id
does not work either.
a
Openllm start llama β€”help
πŸ™Œ 1
Start is a subcommand, and each of the model in itself is a command
model_id is a command for eqch model
πŸ™ 2
m
fair enough, started to work, moving to langchain in a bit.
πŸ‘ 1
I am getting this 500 internal server error from quiet a few days, is there any way to solve it
also when i am using langchain my model is segmented in CPU and GPU and i am getting an error. RuntimeError: Expected all tensors to be on the same device, but found at least two devices, cpu and cuda:0! (when checking argument for argument index in method wrapper_CUDA__index_select)
from langchain.llms import OpenLLM
>>
>> llm = OpenLLM(model_name=β€œopt”, model_id=β€˜facebook/opt-125m’, serialisation=β€œlegacy”)
>> llm(β€œWhat is the difference between a duck and a goose? And why there are so many Goose in Canada?β€œ)
i ran this code
a
What is the GPU you have?
m
a100 80 gb
a
when you do
nvidia-smi
what is the GPU usage?
m
so i actually closed the session and i dont have that stat
i am running the model again
nvidia-smi reports 0 processes on gpu
ran this
a
Thanks, investigating atm
quick hack right now is
Copy code
<http://llm.runner.llm.model.to|llm.runner.llm.model.to>('cuda')
this should place the model to GPU
I will make a patch quickly to do GPU placement correctly
m
alright i will try this and keep you posted
a
sorry for the trouble but thanks for battle test this
m
no worries, thank you for being so cooperative, i make sure to test everything throughly before documenting it. love the interface you have made
Also please let me know once the patch is done, it would be a great help
πŸ‘ 1
still having this error which is causing 500 error code.
cannot test anything beyond this
Trying
openllm start opt --serialisation legacy --model-id facebook/opt-125m
on
openllm==0.2.17.dev8