This message was deleted.
# ask-for-help
s
This message was deleted.
👍 1
j
Hi @Daniel Holler, token streaming will be implemented in OpenLLM as soon as the SSE PR is merged. I believe it is within a week or two away!
d
Hey Jian, that is wonderful news. I just saw you're also the creator of the PR, so I just want to let you know that this is great to hear from you, and this will be an extremely useful feature! Much appreciated!
j
Thanks for your kind words Daniel! Token streaming is an absolute must for LLMs. I will drop you a ping here once there's any updates. and on the hand, what is the use case that you are currently working on?
d
Appreciate it Jian, would be happy to hear about updates from you! We're looking to use OpenLLM to deploy Llama v2 together with an array of different adapters that it should switch between depending on the incoming request - and then stream the responses back 🙂 It'll be interesting to see if the vLLM implementation works together with multiple adapters, and what the benchmarks look like in various scenarios!
h
I implemented a streaming response myself, using zeromq. In the API server, I sent the address bound by the zeromq server to the runner, who used this address to connect and sent the streaming results to the apiserver through this connection.😂
🙌 1
d
Congrats on getting the PR merged @Jian Shen Yap!
j
Thanks @Daniel Holler. I have seen OpenLLM internally having token streaming working, it should be out quite soon!