brash-barista-66564
07/30/2025, 9:21 AM{
"message": "worker failed",
"level": "ERROR",
"name": "livekit.agents",
"exc_info": "Traceback (most recent call last):
...
File \"/usr/local/lib/python3.11/site-packages/livekit/agents/worker.py\", line 387, in run
await self._inference_executor.initialize()
File \"/usr/local/lib/python3.11/site-packages/livekit/agents/ipc/supervised_proc.py\", line 169, in initialize
init_res = await asyncio.wait_for(
^^^^^^^^^^^^^^^^^^^^^^^
File \"/usr/local/lib/python3.11/asyncio/tasks.py\", line 502, in wait_for
raise exceptions.TimeoutError() from exc
TimeoutError"
}
## Environment
- ***Production VM***: 1GB RAM, 1 shared CPU (also tested with 4GB RAM, 2 CPUs - same issue)
- ***Docker container***: ~1.9GB image size
- ***Python***: 3.11-slim
- ***LiveKit SDK***: 1.2.2
- ***Models***: Turn detector, VAD (Silero), custom TTS
## Agent Configuration
python
# voice_agent.py
if __name__ == "__main__":
cli.run_app(
WorkerOptions(
entrypoint_fnc=entrypoint,
prewarm_fnc=prewarm,
worker_type=WorkerType.ROOM,
),
)
## Agent Components
python
async def entrypoint(ctx: JobContext):
# Using turn detector
session = AgentSession(
llm=openai.LLM(model="gpt-4o"),
stt=deepgram.STT(model="nova-3"),
tts=custom_tts_instance,
turn_detection=EnglishModel(), # Turn detector model
vad=ctx.proc.userdata["vad"], # Silero VAD
)
## Current Workarounds
1. Running with dev command instead of start - works but not ideal for production
2. Setting num_idle_processes=0 - disables inference executor entirely
## Questions
1. ***What's the recommended approach for running agents with inference models on resource-constrained VMs?* Should we disable the inference executor or is there a better configuration?
2. ***Is there a way to increase the inference executor initialization timeout?* The default seems too short for VMs with slower CPU/disk.
3. ***Can we configure the inference executor to use fewer resources*** (e.g., single process instead of pool) while still maintaining the benefits of non-blocking inference?
4. ***Are there environment variables or settings*** that can help with inference executor initialization in production environments?
The inference executor seems important for production performance, so I'd prefer not to disable it entirely if possible. What's the best practice for this scenario?tall-airline-16756
07/30/2025, 9:21 AMtall-airline-16756
07/30/2025, 9:22 AMWorkerOptions(
initialize_process_timeout=20.0, # Increase initialization timeout
shutdown_process_timeout=60.0, # Graceful shutdown timeout
)
2. Optimize Resource Usage
To reduce resource consumption while maintaining non-blocking inference, configure the number of idle processes:
WorkerOptions(
num_idle_processes=1, # Use a single idle process
)
3. Memory Management
Set memory limits to prevent resource exhaustion:
WorkerOptions(
job_memory_warn_mb=500, # Warning threshold in MB
job_memory_limit_mb=0, # Hard limit (0 = disabled)
)
4. Monitor and Detect Unresponsive Executors
The system includes built-in monitoring for unresponsive executors. You can configure the following parameters:
• ping_timeout=60 - Timeout for ping responses
• high_ping_threshold=2.5 - Threshold for detecting unresponsive executors
• ping_interval=5 - Interval between ping checks
5. Provide User Feedback
For operations that may take longer due to resource constraints, provide verbal feedback to users. This can include status updates during long-running operations or dynamic updates based on the query and estimated remaining time.
These adjustments should help mitigate the timeout issues while optimizing performance for your production environment. Let me know if you need further assistance with implementing these configurations.
Sources: Handling RAG Delays in Voice Agents | LiveKit Docs | livekit-agents/livekit/agents/ipc/inference_proc_executor.py | fly.io/fly.toml