How should I handle rolling updates for self deplo...
# ask-ai
b
How should I handle rolling updates for self deployed agents while achieving zero-downtime in docker containers. I am specifically asking about keeping active sessions alive, while routing new traffic to the new container.
t
Thanks for reaching out, Hannes. I'm an AI assistant and I'm here to help. Give me a moment while I research this topic.
I’m glad you’re looking to achieve zero-downtime rolling updates for your self-deployed agents in Docker containers while keeping active sessions alive and smoothly routing new traffic to updated containers. Here are some friendly strategies to help you along: • Graceful Shutdown: It’s important that your agent can handle termination signals like SIGTERM gracefully. This way, it stops accepting new sessions but lets the current ones finish without interruption. For example, LiveKit agents automatically enter a draining mode when they’re terminated, which is super helpful. • Use Kubernetes: If you’re using Kubernetes, you can take advantage of its RollingUpdate strategy. Just make sure to set
terminationGracePeriodSeconds
to give sessions time to wrap up, and use readiness probes to ensure traffic only goes to containers that are ready and healthy. • Fly.io Blue-Green Deployment: Fly.io offers a neat blue-green deployment setup that keeps your old instances running during updates. By configuring
kill_timeout
, you allow sessions to drain properly before the old containers shut down. • Docker Compose: For simpler environments, you can manually manage updates by starting new containers first, waiting for their health checks to pass, and then stopping the old ones with a grace period to avoid dropping sessions. If you’d like, just share which deployment platform you’re using, and I’d be more than happy to provide detailed configuration examples—like Kubernetes YAML or Fly.io settings—tailored just for you. I’m here to help every step of the way! Sources: Worker lifecycle | LiveKit Docs | Deploy to Kubernetes | LiveKit Docs | fly.io/fly.toml | Deploying to custom environments | LiveKit Docs | Deploying to production | LiveKit Docs
b
How can I deregister an agent to stop accepting new connections while keeping active ones alive?
t
Hi there! To gracefully deregister an agent and prevent it from accepting new connections while still allowing active ones to finish, you can send a
SIGTERM
signal to the agent's process. This will activate LiveKit's graceful shutdown mechanism, which works like this: • The agent stops taking on new jobs but keeps managing any active sessions until they’re done. • After all active sessions are complete, the agent will shut down fully. Here’s how you can do it: 1. First, find the agent's process ID (PID). 2. Then, send a
SIGTERM
signal using
kill <worker-pid>
, or if you’re on Kubernetes, use
kubectl delete pod <agent-pod-name> --grace-period=60
. If you’re running a self-hosted setup, you also have the option to adjust how long the agent waits for active sessions to finish by setting the
drain_timeout
parameter in the agent’s configuration. Feel free to reach out if you’d like me to walk you through that or if you have any other questions! Sources: Distributed Setup | LiveKit Docs | Agent deployment CLI reference | LiveKit Docs | agents/src/voice/agent_activity.ts | Worker options | LiveKit Docs | Deploying to production | LiveKit Docs
b
How do I implement this in a docker container? Currently, the agent disconnects when I stop the agents container in an active session
t
To help you gracefully shut down LiveKit agents in Docker containers while keeping your active sessions safe and sound, here’s a friendly step-by-step guide: 1. Enable Graceful Shutdown: When you stop the container (like with
docker stop
), the agent will get a
SIGTERM
signal. This gently puts the LiveKit agent into draining mode, so it won’t accept new connections but will let your current sessions finish nicely. 2. Set Drain Timeout: You can customize how long the agent waits for active sessions to wrap up by setting the
drain_timeout
in the agent’s
WorkerOptions
. For example:
const opts = new WorkerOptions({ drainTimeout: 3600 }); // 1 hour
3. Configure Container Grace Period: It’s important to align your Docker or orchestration platform’s termination grace period with the drain timeout. For instance, in Kubernetes:
Copy code
spec:
  terminationGracePeriodSeconds: 600
4. Use Rolling Deployments: This approach lets you bring up new containers alongside the old ones, so new traffic is smoothly directed while the old containers finish up active sessions without rushing. If you’d like, I’m here to help with specific examples tailored to your deployment setup or any other questions you might have. Just let me know! Sources: README.md | Operational Guidance | livekit/livekit_composite | DeepWiki | Deploying to production | LiveKit Docs | Inside a session | LiveKit Docs | Distributed Setup | LiveKit Docs
b
I am running in docker container, and when I stop the container while calling, the session stops without any logs
t
I understand how frustrating it can be when sessions stop without any logs after stopping a Docker container. This usually happens because the container shuts down too quickly, not giving the LiveKit agent enough time to close things gracefully or send out those final logs. Here are some steps that can really help: 1. Increase Shutdown Grace Period: It’s a good idea to give your container more time to shut down properly. For example, if you’re using Docker Compose, you can set
stop_grace_period: 600s
to allow up to 10 minutes for active sessions to finish smoothly. 2. Enable Unbuffered Logging: Adding
ENV PYTHONUNBUFFERED=1
to your Dockerfile makes sure logs are sent right away to stdout/stderr, so nothing gets held up. 3. Set Log Level to DEBUG: For more detailed insights, you can configure the agent to log at the DEBUG level by including
--log-level=DEBUG
in your Docker
CMD
. 4. Forward Logs Externally: To keep track of logs even if the container stops unexpectedly, consider forwarding them to external services like Datadog or CloudWatch. Following these steps should help keep your sessions running smoothly and make sure you don’t miss any important logs during shutdown. If you’d like, I’m here to help you with setting these up or answer any other questions you have! Sources: Operational Guidance | livekit/livekit_composite | DeepWiki | livekit-plugins/livekit-plugins-ultravox/livekit/plugins/ultravox/realtime/... | python-agent-example-app/Dockerfile | Dockerfile | Log collection | LiveKit Docs
b
Do I need to listen for SIGTERM in the container somehow?
t
Hey there! Just to reassure you, LiveKit agents don’t need any special custom code to handle
SIGTERM
signals when running inside Docker containers. The LiveKit framework takes care of this automatically when you start the agent using the official CLI interface (like with
cli.runApp
). So, when your container gets a
SIGTERM
signal (such as when you run
docker stop
), the agent smoothly switches into a "draining" mode. This means it stops taking new jobs but keeps managing any active sessions until they’re all wrapped up. This helpful behavior is built right into the framework. To make sure everything runs smoothly, here are a couple of things to keep in mind: • Set a sufficient container grace period: Make sure Docker or your orchestrator allows enough time for those active sessions to finish. For example, you can use
--stop-timeout
in Docker or
terminationGracePeriodSeconds
in Kubernetes. • Match the drain timeout: Double-check that the
drain_timeout
setting in your agent’s configuration aligns with the container’s grace period. If you’d like, I’m here to help you set these up just right for your environment—feel free to reach out anytime! Sources: README.md | Worker options | LiveKit Docs | Deployment and scaling | LiveKit Docs | livekit-agents/livekit/agents/cli/_run.py | Operational Guidance | livekit/livekit_composite | DeepWiki