Slackbot
11/14/2025, 5:00 PMKoorous Vargha
11/15/2025, 5:52 PMZac Policzer
11/17/2025, 5:09 PMSlackbot
11/21/2025, 5:00 PMKoorous Vargha
11/21/2025, 5:31 PMKoorous Vargha
11/21/2025, 11:34 PMFelix GV
11/25/2025, 1:59 PMSlackbot
11/28/2025, 5:00 PMFelix GV
12/09/2025, 7:52 PMSlackbot
12/12/2025, 5:00 PMFelix GV
12/16/2025, 2:24 AMFelix GV
12/16/2025, 2:25 AMSlackbot
12/26/2025, 5:00 PMSlackbot
02/20/2026, 5:00 PMSlackbot
03/06/2026, 5:00 PMFelix GV
03/18/2026, 7:22 PMSlackbot
03/20/2026, 4:00 PMSlackbot
04/03/2026, 4:00 PMSlackbot
04/17/2026, 4:00 PMZac Policzer
04/20/2026, 4:30 PMSlackbot
05/01/2026, 4:00 PMFelix GV
05/13/2026, 6:07 PMSlackbot
05/15/2026, 4:00 PMKoorous Vargha
05/21/2026, 5:31 PMSergey Makagonov
05/27/2026, 5:00 PM03:48:28 ERROR VenicePushJob: Kill check monitor detected that push job for store: ..., version: 19 has been killed. Status: ERROR
OfflinePushMonitor flips an in-flight push to terminal ERROR when it briefly sees a replica go OFFLINE.
To mitigate this, I added retries of the VenicePushJob inside of our Whatnot-specific Spark Job, and it helped.
I was also looking at node_removable API on the controller side and considered adding it into preStop hook of servers, - but it looks like it's mostly for decommissioning servers and making sure that the stores remain healthy (it can return WILL_LOSE_DATA, WILL_TRIGGER_LOAD_REBALANCE).
2. Partition skew mitigation
We used replication factor of 3 on our load-test setup with 10 venice servers. There were cases when some servers could have all 3 partitions of present for a particular store - not all active, but still, results in 33% of all reads for such store to be served from this particular server.
3. Per-store RCU quota and per-router quota
Looks like the RCU quota is a combination of per-store and per-router quota (hav a write-up on this). I'm curious, what is the max.read.capacity config that you use in production for routers? And what are the practical maximum per-store RCUs limits do you set?Sergey Makagonov
05/28/2026, 8:53 PMvenice-writer-router , alongside venice-router, venice-controller, etc- that will be a Netty HTTP server on top of the venice-producer client
• pros:
◦ not that hard to implement - we can reuse a bunch of code from router
◦ doesn't have to handle complex use-cases like write compute right away - we can start with simple PUT/DELETE for the entire key(s)
3. proxy solution, shipped as a DaemonSet
• can be extended further to support fast client and Da-Vinci client for reads. It's very likely we'll explore faster clients down the road.
Questions to the community:
• I heard from @Ankit that for their C++ service, they used go-proxy (overseeing the Java Da-vinci and Producer), and the writes and reads went through this proxy. Curious if there is any chance this part could be open-sourced? It should be generic enough, doesn't even have to compile, we can pick it up and make it work. Even design docs/proposal would be helpful.
• how common is the RT-only pattern for Venice, where writes happen outside of streaming/batch path?
cc @Koorous Vargha and @Zac Policzer would appreciate your input and recommendations on the direction.Slackbot
05/29/2026, 4:00 PMSlackbot
06/12/2026, 4:00 PMSlackbot
06/26/2026, 4:00 PMSlackbot
07/10/2026, 4:00 PM