in livekit, what client callback can we listen to ...
# ask-ai
n
in livekit, what client callback can we listen to to hook in taking a screenshot and passing it through
t
Thanks for reaching out, rohit. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I understand you're looking to capture screenshots using LiveKit, and I’m here to help! While LiveKit doesn’t offer a single client callback explicitly named for taking screenshots, different SDKs provide great ways to capture video frames or images depending on the platform you’re working with: • Android: The LiveKit SDK works with CameraX, utilizing the
ImageCapture
use case. You can use the
takePicture
method along with
ImageCapture.OnImageSavedCallback
to handle the captured image. • iOS: The
RPScreenRecorder
offers a sample buffer callback that lets you receive video frames during screen sharing. • Unity: You can use
ScreenCapture.CaptureScreenshotIntoRenderTexture
to capture the screen content into a render texture for further use. • Node-based SDKs: The
CaptureVideoFrameRequest
in the FFI interface indicates video frame capture capabilities. If you’re working on a particular platform, please feel free to share, and I’d be happy to guide you through the specific steps to implement screenshot functionality. I’m here to make this as smooth as possible for you! Sources: sample-app-compose/src/main/java/io/livekit/android/composesample/CallActiv... | Sources/LiveKit/Track/Capturers/MacOSScreenCapturer.swift | Runtime/Scripts/CameraVideoSource.cs | packages/livekit-rtc/src/proto/ffi_pb.ts | livekit-plugins/livekit-plugins-google/livekit/plugins/google/beta/realtime...
n
I have the screnshot, but how do I send it with the sentence being said over the voice call
t
Thanks for reaching out! Could you please share a bit more about how you’re planning to send the screenshot and the sentence? For instance: • Are you hoping to include them in a chat message? • Would you like to embed the screenshot within the audio stream? • Are you working with a particular SDK or platform for this? The more details you can provide, the better I can assist you. I’m here to help!
n
include them in the chat message (but it's a voice call), so perhaps include them at the STT -> LLM stage
t
I'm glad you're exploring how to include a screenshot and a transcribed sentence in a chat message during a LiveKit voice call! You can seamlessly integrate them at the STT (Speech-to-Text) to LLM (Language Model) stage. Here’s a friendly step-by-step on how to do this: 1. Capture the Screenshot: Start by using the video track to grab the latest frame. You can then encode it as a base64 image or save it for later use. 2. Obtain the Transcription: Next, use the STT system to get the transcribed sentence. This usually comes through the
on_user_turn_completed
callback. 3. Combine and Send: Finally, in the
before_llm_cb
callback, you can add both the screenshot and the transcription into the chat context before the LLM processes the input. Here’s an example:
Copy code
async def before_llm_cb(assistant: VoicePipelineAgent, chat_ctx: llm.ChatContext):
    transcript = "Hello, I need help with this."  # From STT
    latest_image = await get_latest_image(room)  # Capture latest frame

    if latest_image:
        chat_ctx.add_message(
            role="user",
            content=[
                transcript,
                ImageContent(image=latest_image)  # Can be URL, base64, or VideoFrame
            ]
        )
Just a quick tip: make sure your LLM supports multimodal input so it can handle both text and images smoothly. If you have any questions or want to dive deeper into the implementation, I’m here to help! Sources: Vision Agent Quickstart | LiveKit Docs | livekit-agents/livekit/agents/voice/agent_activity.py | Vision | LiveKit Docs
n
is it better to use a toolcall to take a screenshot, or every time a message starts we send over a screenshot and only analyze it if the user completes a sentence with text? or something else?
t
I don't have the answer you're looking for. You could also try asking your question: • in one of the other Slack channels or • to https://deepwiki.com/livekit/livekit_composite which is trained on all LiveKit source code If you find the answer, please post it here to help others!