I'm using the web embed voice agent + python agent...
# ask-ai
n
I'm using the web embed voice agent + python agent server sdk. I want to send screenshots from the web embed side to the python agent server, but I only want them on vad turn end detection. What's the best setup for this
t
Thanks for reaching out, rohit. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
Thanks for reaching out! Could you please share a bit more about how you're currently capturing and handling screenshots on the web embed side? Also, are you hoping for some guidance on detecting VAD turn ends, transmitting the screenshots, or maybe both? I’m here to help you through it! Sources: livekit-agents/livekit/agents/voice/agent_session.py | OpenAI Realtime API integration guide | LiveKit Docs | Turn detection and interruptions | LiveKit Docs | How does end-of-utterance detection work in conversations? | complex-agents/turn-taking/agent.py
n
looking for guidance on the best way to transmit the screenshot from the web embed side with low latency. I want to ideally pre-send the screenshot when the user starts talking and have it be processed with the users text when VAD detects the users turn ended
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!