AudioFrame format for custom TTS — float32 or int1...
# ask-ai
n
AudioFrame format for custom TTS — float32 or int16? Body: I'm integrating a custom TTS engine (NVIDIA Riva) with a LiveKit agent, and I'm generating audio frames using
rtc.AudioFrame
. I'm trying to determine the correct audio format to pass in
AudioFrame.data
. Here’s what I’ve tried: • When I send
float32
PCM (normalized to
[-1.0, 1.0]
), the audio plays but sounds distorted — like a corrupted radio. • When I send
int16
PCM, resampled to 48000Hz and sized correctly, I hear nothing at all. My question is: What audio format does LiveKit expect in
AudioFrame.data
? Should it be
float32
PCM or
int16
, and are there any alignment or endian requirements?
Appreciate any insights — thanks!
t
Thanks for reaching out, Mohamed. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
LiveKit expects the
AudioFrame.data
to be in 16-bit signed integer (int16) PCM format, not
float32
. Here are the specific requirements: • Data Type: 16-bit signed integers (
int16
) • Channel Layout: Interleaved by channel • Endianness: Little-endian (
pcm_s16le
format) • Alignment: The data length must be a multiple of 2 bytes (size of
int16
), and the total size should be at least
num_channels * samples_per_channel * sizeof(int16)
. If you are starting with
float32
PCM data, you need to convert it to
int16
. For example, in Python:
Copy code
audio_data = (float_data * 32767).astype(np.int16).tobytes()
Ensure that the data is properly resampled to 48000Hz and meets the alignment requirements. Let me know if you need further clarification! Sources: livekit-rtc/livekit/rtc/_proto/audio_frame_pb2.py | Track egress | LiveKit Docs | packages/livekit-rtc/src/audio_frame.ts | terminal-client/client.py | agents/src/audio.ts
n
2025-07-28 144815,875 - INFO riva-tts - Yielding audio frame: 1245 samples (int16), data size: 2490 bytes {"room": "xirk-anz5"} 2025-07-28 144815,876 - DEBUG riva-tts - Text buffer: ' What' {"room": "xirk-anz5"} 2025-07-28 144815,876 - DEBUG riva-tts - Text buffer: ' What’s' {"room": "xirk-anz5"} 2025-07-28 144815,878 - DEBUG riva-tts - Text buffer: ' What’s up' {"room": "xirk-anz5"} 2025-07-28 144815,879 - DEBUG riva-tts - Text buffer: ' What’s up?' {"room": "xirk-anz5"} 2025-07-28 144815,879 - INFO riva-tts - Processing text for TTS: 'What’s up?' {"room": "xirk-anz5"} 2025-07-28 144815,879 - DEBUG riva-tts - Synthesizing text: What’s up?... {"room": "xirk-anz5"} 2025-07-28 144816,568 - DEBUG riva-tts - Received audio chunk: 73562 bytes {"room": "xirk-anz5"} 2025-07-28 144816,568 - DEBUG riva-tts - Converted to numpy: 36781 samples (int16) {"room": "xirk-anz5"} 2025-07-28 144816,569 - DEBUG riva-tts - No resampling needed: 48000Hz matches target 48000Hz {"room": "xirk-anz5"} 2025-07-28 144816,569 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,569 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,570 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,570 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,570 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,570 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,570 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,571 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,571 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,571 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,571 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,572 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,572 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,572 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,572 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,573 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,573 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,573 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,573 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,573 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,573 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,575 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,575 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,575 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,575 - INFO riva-tts - Yielding audio frame: 1440 samples (int16), data size: 2880 bytes {"room": "xirk-anz5"} 2025-07-28 144816,576 - INFO riva-tts - Yielding audio frame: 781 samples (int16), data size: 1562 bytes {"room": "xirk-anz5"} 2025-07-28 144816,576 - WARNING livekit.plugins.silero - inference is slower than realtime {"room": "xirk-anz5", "delay": 0.6654774000118486} 2025-07-28 144816,578 - INFO livekit.agents - LLM metrics {"room": "xirk-anz5", "ttft": 0.77, "prompt_tokens": 52, "prompt_cached_tokens": 0, "completion_tokens": 8, "tokens_per_second": 3.68}
t
Hey there! I totally get how frustrating audio issues can be. Let me help you break down what might be causing those pesky audio problems. • Data Size and Alignment: Think of this like making sure all your puzzle pieces fit just right. Double-check that your data length is a clean multiple of 2 bytes and matches exactly what you're expecting:
num_channels * samples_per_channel * sizeof(int16)
. When things aren't lined up perfectly, playback can get wonky. • Improper Data Conversion: Converting audio data can be tricky! If you're switching from
float32
, make sure you're scaling those values to fit nicely in the
int16
range (-32768 to 32767). Here's a quick trick that might help:
audio_data = (float_data * 32767).astype(np.int16).tobytes()
• Channel Interleaving: For multi-channel audio, it's like making sure all your audio channels are playing nice together. If you're working with mono audio, just make sure it's duplicated correctly. • Value Range: Keep an eye on those audio values! If they go beyond the
int16
range, you might end up with some not-so-great distortion. From what I can see in the logs, your audio frames look like they're on the right track. But let's do a quick check: 1. Verify that your audio data is little-endian (pcm_s16le format). 2. Make sure everything is aligned just right and matches the frame size you're expecting. 3. Do a quick scan for any extra transformations that might be sneaking in and causing trouble. Need a hand working through this? I'm here to help! Just let me know what you find, and we'll solve this together. 😊 Sources: terminal-client/client.py | packages/livekit-rtc/src/audio_frame.ts | realtime/openai-realtime-pitch-shift.py | examples/video-stream/audio_wave.py
n
No audio output when using custom TTS with AudioFrame (int16 PCM, 48kHz, mono) Body: Hi everyone, I'm building a custom voice agent using LiveKit Agents and NVIDIA Riva TTS. My TTS pipeline receives float32 PCM from Riva, which I: • Resample to 48000Hz if needed • Normalize to [-1.0, 1.0] • Convert to
int16
using: •
audio_bytes = (audio_float * 32767).astype(np.int16).tobytes()
I then yield it like this:
Copy code
python
Copy code
yield rtc.AudioFrame(
    data=audio_bytes,
    sample_rate=48000,
    num_channels=1,
    samples_per_channel=len(audio_float)
)
I confirmed: • Audio plays fine when saved to WAV ( •
<http://scipy.io|scipy.io>.wavfile.write
) • The
AudioFrame
length is always a multiple of 2 (aligned) •
sample_rate
,
num_channels
, and
samples_per_channel
are correct • Logs show audio frames being yielded correctly Still, I hear nothing in the LiveKit room. I even tried skipping the int16 step and sending float32 (normalized), but same result. Questions: 1. Does
rtc.AudioFrame
in the Python SDK require float32 or int16? 2. Is there a known issue with
AudioFrame
playback when using
data=tobytes()
of int16? 3. Is there any way to debug or monitor audio output in the room?
t
Here are answers to your questions and suggestions for debugging the issue: • AudioFrame Data Format: Hey there! Let's talk about the
rtc.AudioFrame
in the Python SDK. Just a heads up, it needs int16 PCM format, not float32. Think of it like making sure your audio puzzle pieces fit perfectly - the data needs to be interleaved by channel and properly aligned (so the length should be a multiple of 2 bytes). • Using tobytes() with int16: No worries about using
data=tobytes()
with int16 arrays! This is actually a super common approach. Here's a friendly example you can try:
audio_data = (float_data * 32767).astype(np.int16).tobytes()
• Debugging and Monitoring Audio Output: I've got some cool techniques to help you troubleshoot those audio issues: ◦ Audio Level Monitoring: Think of this like checking the volume on your audio. Here's a neat way to calculate dB levels:
Copy code
audio_samples = np.frombuffer(frame.data, dtype=np.int16)
rms = np.sqrt(np.mean(audio_samples.astype(np.float32) ** 2))
max_int16 = np.iinfo(np.int16).max
db_level = 20.0 * np.log10(rms / max_int16 + 1e-6)
◦ Frame Counting and Logging: Let's keep track of those audio frames! This helps make sure everything's moving smoothly:
Copy code
frames_received += 1
if frames_received <= 5:
    <http://logging.info|logging.info>(f"Received audio frame {frames_received}")
elif frames_received % 100 == 0:
    <http://logging.info|logging.info>(f"Received {frames_received} frames total")
◦ Participant Audio Tracking: We can also keep an eye on audio levels for specific participants in the room. ◦ Audio Stream Status Monitoring: It's like having a health check for your audio stream:
Copy code
def callback(outdata, frames, time_info, status):
    if status:
        logging.warning(f"Audio output status: {status}")
If the issue is still hanging around, let's double-check a few things: make sure the audio data is little-endian (pcm_s16le format) and that the sample_rate, num_channels, and samples_per_channel values are playing nice together. Need more help? I'm right here for you! Sources: meditation_assistant/audio_handler.py | stream_audio.py | livekit-rtc/livekit/rtc/_proto/audio_frame_pb2.pyi