In python, is there a way to execute something bet...
# ask-ai
n
In python, is there a way to execute something between the sentence tokenizer step and when it gets sent to the provider? We're trying to implement a translation step and it's hard to translate on small chunks - we need it to be sentence sized chunks first.
t
Thanks for reaching out, Stephen. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Absolutely, you can add a step between sentence tokenization and sending the data to the provider to include a translation process. Here are a couple of friendly ways to approach this: • Built-in Translation: If you’re working with the Gladia STT plugin, you have the option to enable real-time translation by setting
translation_enabled=True
and specifying your target languages with
translation_target_languages
. This way, the translation happens smoothly within the STT provider itself. • Custom Translation Pipeline: Alternatively, you can create your own translation step. For example, after breaking the STT output into sentences using a sentence tokenizer, you can send each sentence to a custom translation function or class that leverages an LLM for translation. Please feel free to reach out if you’d like me to share detailed examples or if you need any help implementing either of these methods—I’m here to support you! Sources: Gladia integration guide | LiveKit Docs | livekit-plugins/livekit-plugins-gladia/README.md | VoicePipelineAgent | LiveKit Docs | Gladia integration guide | LiveKit Docs
n
I'm looking for translating after LLM and before TTS. How about that? This would be tokenizing the stream coming from the LLM
t
I’m glad you asked! To translate text after LLM processing and before TTS, a helpful way is to use the
tts_node
method within a custom
Agent
class. This lets you gently intercept the text stream coming from the LLM, apply the translation you need, and then send the translated text on to the TTS engine for synthesis. Here’s a friendly overview of how you can approach this: • Override the
tts_node
method in your custom agent. • Use an external translation API (like Google Translate) to translate the text chunks streamed from the LLM. • Keep the streaming smooth by handling
AsyncIterable[str]
, so you avoid any lag. • Then, pass the translated text stream to the default TTS node for synthesis. If you’d like, I’d be happy to walk you through a detailed code example to help you implement this. Just let me know! Sources: Docs::Agents | livekit-plugins/livekit-plugins-gladia/livekit/plugins/gladia/stt.py | chain-of-thought-tts/agent.py | Agents 0.x migration guide | LiveKit Docs
n
Is there a way to use a sentence tokenizer to increase the chunk sizes so that the translation context is large enough to be useful?
t
I'd love to better understand how you'd like to adjust the chunk sizes. Could you share a bit more about your preferences? For example: • Are you hoping to combine multiple sentences into larger chunks before translation? • Would you like to set a maximum or minimum size for each chunk? • Are you currently using a specific tokenizer or method? Knowing these details will really help me guide you in the best way possible.