https://venicedb.org logo
Join Slack
Powered by
# general
  • k

    Kacper Bielecki

    09/05/2025, 9:29 AM
    VPJ.log
    VPJ.log
  • k

    Kacper Bielecki

    09/05/2025, 9:30 AM
    Hi, we are slowly but steadily continue with testing Venice. We are trying now with the bigger setup and trying to push 4TB store but we fail. We use 0.4.613 version with the k8s setup through EKS. We are pushing with 100 partitions and 12 venice-servers. It seems that some partitions fail to allocate replicas. The root cause seems to be the failure of opening RocksDB metadata database on some venice-server pods which manifests in such exceptions:
    Copy code
    Caused by: org.rocksdb.RocksDBException: While lock file: /opt/venice/rocksdb/rocksdb/client-id-to-email-hash_v26/client-id-to-email-hash_v26_1000000000/LOCK: Resource temporarily unavailable
    I am attaching VPJ, leader controler and one of the venice-server logs. It seems that venice-server tries to open the metadata database from multiple threads and the locking failures are not retried. Is it a known issue? Are we doing anything wrong? 🤔
    k
    d
    +2
    • 5
    • 26
  • f

    Felix GV

    10/21/2025, 2:09 PM
    The 2nd edition of Designing Data-Intensive Applications mentions Venice!
    🚀 6
    • 1
    • 1
  • f

    Felix GV

    10/21/2025, 4:15 PM
    For those in the Bay Area, this could be an interesting meetup: https://luma.com/e7feg2i6
    🙌 2
    n
    • 2
    • 2
  • a

    Amre Shakim

    10/23/2025, 10:53 PM
    Has anyone looked into integrating Venice with Snowflake as a batch data source or batch ingestion layer? Since we already use Snowflake for transformations, I’m wondering what the main challenges or limitations might be.
    k
    n
    f
    • 4
    • 20
  • f

    Felix GV

    11/19/2025, 5:46 PM
    This is quite similar to how large values get chunked inside the Venice server. If instead of appending a flag to the key, Venice used a separate table (column family) to store chunks, then it would essentially be the same. https://www.linkedin.com/posts/ben-dicken-78797a73_postgres-uses-toast-to-store-large-variable-sized-activity-7396908837485727744-dG76?utm_source=share&utm_medium=member_ios&rcm=ACoAAAEk238BXlhz1s5hXc96bKIJJ-eWwXUnlas
    😮 1
  • g

    Gaojie Liu

    11/19/2025, 5:58 PM
    @Felix GV The latest Venice blog post is out today: https://www.linkedin.com/blog/engineering/infrastructure/evolution-of-the-venice-ingestion-pipeline
    🎉 6
  • f

    Felix GV

    11/19/2025, 7:00 PM
    Great to see that post finally come to light! I just re-read it and it’s even better than the first time I did! Congrats @Gaojie Liu and team!
    ➕ 2
  • z

    Zac Policzer

    11/19/2025, 7:39 PM
    Just posted on the VeniceDB LinkedIn page.
  • z

    Zac Policzer

    11/19/2025, 7:40 PM
    idk why, but using a social media account and posting as VeniceDB feels like this:
    😂 4
  • z

    Zac Policzer

    11/19/2025, 7:49 PM
    Blue sky post up as well: https://bsky.app/profile/venicedb.org/post/3m5yzjby7tk2t
    🙌 1
    f
    • 2
    • 5
  • m

    Minh Nguyen

    11/25/2025, 3:31 AM
    🎥 Proud to share our two latest ASQ videos! • Data Integrity Validation by @Lei Lu:

    https://youtu.be/sYytwZ4WJJw▾

    • Stateful CDC Client by @Koorous Vargha:

    https://youtu.be/6vvtmijdwUI▾

    Check them out and let us know your thoughts! 🚀
    🚀 2
    🎉 1
    👀 1
    venice black on white 1
    f
    • 2
    • 1
  • k

    Koorous Vargha

    12/08/2025, 5:52 PM
    https://www.confluent.io/blog/ibm-to-acquire-confluent/
    z
    • 2
    • 3
  • p

    Pavan

    12/09/2025, 4:53 AM
    Is the RocksDB implementation located in
    com.linkedin.davinci.store.rocksdb
    considered the canonical implementation for local storage across both the Venice Server and the Da Vinci Client?
    z
    f
    • 3
    • 5
  • s

    Sushant Mane

    02/13/2026, 8:57 PM
    Thanks to @Xun Yin for the reviews! PR is merged now! 🙏
    d
    • 2
    • 2
  • z

    Zac Policzer

    02/15/2026, 3:26 AM
    Man. Search on LinkedIn is awful. This semantic search thing isn't all it's cracked up to be.
    😂 1
    k
    f
    • 3
    • 12
  • k

    Koorous Vargha

    02/17/2026, 5:00 PM
    🚀 Venice Documentation Redesigned 🚀 We've completely overhauled the Venice documentation with a focus on making it easier to get started and find what you need. What's New: • Modern design with dark/light mode and improved navigation • New client guides with code examples (Thin Client, Fast Client, CDC) • Enhanced documentation for Da Vinci Client and Stream Processor • Better organized content for easier discovery Check it out: https://venicedb.org/
    🙌 3
    🚀 2
    🎉 2
    z
    f
    +2
    • 5
    • 9
  • z

    Zac Policzer

    03/02/2026, 2:05 AM
    For folks who might be interested (and by that, I mean venice devs), I have added a TON of new functionality to the heap dump parsing tool me and @Kai-Sern Lim made. It's now pretty fast and tracks all Java objects I've used. It's also got a very interesting property. If you parse the heapdump to a directory of parquet files, your favorite local LLM is pretty good at querying it with pyarrow. https://github.com/ZacAttack/HeapDumpStarDiver
    🙌 2
    ❤️ 2
    e
    f
    s
    • 4
    • 42
  • c

    Craig Alfieri

    03/11/2026, 9:01 PM
    Hi Venice-folks 👋... My name's Craig, I work at Antithesis and have worked with Felix a while back, and with some of the other teams at LinkedIn currently. We're hosting an industry conference in April, in Washington D.C. called BugBash, focusing on Systems and Building Reliable Applications. I had a recent customer who had their travel budget cut, and had to give back their three passes for BugBash. If anyone in the Venice community (Limit 3), is interested in attending with a no-cost pass, please DM, and I'd be happy to get you a passcode to register/attend. Below are some more details on the conference, the sessions, and speakers have been carefully curated to ensure a very thought-provoking two days. HMU directly if this piques interest. Best, _____ BUGBASH OVERVIEW 🪲 What's the Conference about: 😅 What does it take to build reliable software? BugBash is an industry conference hosted by Antithesis. The conference covers multiple aspects of building software that stands the test of time, including Testing, Formal Methods, Site Reliability Engineering, Observability, and Infrastructure. (Conference: Link) ⌚*️ When/Where? April 23rd-24th* in Washington D.C., there will also be talks on the 22nd that run adjacent to the conference What's the set-up: ~250 attendees, 2 days of official Conference, with a single track agenda (Makes for optimal hallway discussion). The day before, there will be adjacent seminar talks somewhat unconf-like. The evening before, there's a happy hour/social event, and a dinner on Day 2. A good write-up of last year's inaugural BugBash from an attendee at Datadog: https://concerningquality.com/bug-bash-2025/ 📣 Who are the speakers? 1. Will Wilson, Founder & CEO, Antithesis 2. Deb Chachra, Olin College, author of How Infrastructure Works 3. Ben Eggers, Member of Technical Staff @ OpenAI 4. Frank McSherry, CTO @ Materialize 5. Peter Alvaro, UC Santa Cruz / Father of Fault Injection 6. Chaitanya Bandari, Distributed Systems Engineer @ TigerBeetle 7. Brian Potter, author of Construction Physics and The Origins of Efficiency 8. Corwin Coburn, Senior Staff SWE at Google, Uber Tech Lead, Parallel File Systems 9. Ankush Desai, Principal Scientist, Systems Correctness, Snowflake 10. Ron Minsky, Co-head of Technology, Jane Street 11. Matt Barrett, Founder & CEO, Adaptive 12. Steve Klabnik, SWE, East River Source Control ⚡*️Confirmed lightning talks:* · Pierre Zemb - CleverCloud · Anish Agarwal - CEO Traversal · Fernanda Graciolli + Nada Amin - Midspiral · Jenny Qu - World-class Hacker 🎩 Formal methods: · Jacopo Tagliabue Bauplan · Fernanda Graciolli + Nada Amin / Midspiral · Tony Zhang / Basilisk 🪖 Hard problems in distributed systems testing: · Kyle Kingsbury - Jepsen · Marco Primi - Antithesis Red Team · Mark Logan - Mysten Labs · Who will be there: 85% Systems Engineers, 10% Academia, and Antithesis entire team of Product/Engineering will be in attendance.
    😮 3
    🎉 1
  • d

    Dmytro Prokhorenkov

    03/19/2026, 5:25 PM
    Hi all, We are still in the process of deploying the change for the fast-client. I would like to ask for some guidance on how I can debug if R2/D2 routing is configured properly, and my app will be able to reach the store with the fast-client. I'll add my current config files to the thread.
    z
    x
    • 3
    • 17
  • z

    Zac Policzer

    03/25/2026, 9:40 PM
    @Efe Gencer the man the myth the legend!
    👀 1
    😄 1
    👋 3
    🚀 1
    e
    s
    • 3
    • 4
  • s

    Sergey Makagonov

    03/30/2026, 4:34 PM
    Good morning @Felix GV @Koorous Vargha @Zac Policzer, I was looking at the realtime-ness and hybrid stores setup for Venice. It all made sense at first sight (new batch version is created, RT events are replayed up certain rewind time), but then the fundamental question is: how to handle the batch lag problem in practice. Consider the following example when batch and RT write the same field, where it seems like there is no clean resolution:
    Copy code
    1:00pm  VPJ1 pushes purchase_count = 3
    1:10pm  RT update: purchase_count = 4
    1:15pm  RT update: purchase_count = 5
    1:30pm  RT update: purchase_count = 6
    1:35pm  VPJ2 pushes, but batch data lags by 15min: purchase_count = 5
    As VPJ2's new version is being built, RT events from 110 130 are also being applied to it. But DCR resolves conflicts by timestamp: • Batch sets purchase_count = 5 (with timestamp 1:35pm) • RT event from 1:30pm (purchase_count = 6) has an earlier timestamp • Timestamp-based resolution: batch wins. Value regresses to 5. So far, my understanding is that achieving good consistency with Batch + RT updates is possible only if RT updates separate set of fields through write-compute. Am I missing something? I put up a doc on Venice Hybrid store considerations as a reference. cc @Hubert Puszklewicz @Amre Shakim
    ✅ 1
    z
    k
    f
    • 4
    • 29
  • f

    Felix GV

    04/01/2026, 9:21 PM
    @Jia congrats on the new gig! You absolutely wanted to work for the company that created Venice, didn’t you? 😂😁🤷‍♂️🤪
    🎉 4
    m
    j
    • 3
    • 6
  • f

    Felix GV

    04/02/2026, 11:11 AM
    And congrats @Koorous Vargha on the promo!
    👏 9
    🎉 6
    k
    • 2
    • 1
  • z

    Zac Policzer

    04/21/2026, 8:45 PM
    Somewhat random but @Sergey Makagonov and @Hubert Puszklewicz, you folks should talk to Ankit Patnaik, who I have come to learn today is a whatnot employee. He actually did a significant davinci integration at LinkedIn some time back that I worked with him on.
    😮 1
    🙌 4
    h
    a
    • 3
    • 7
  • s

    Sergey Makagonov

    05/05/2026, 4:00 PM
    Hi @Koorous Vargha, in earlier meetings, you mentioned about capacity: • biggest cluster: 300 servers, 280 routers (roughly 1:1 ratio). Trend is toward fewer, beefier boxes (vertical scaling of hardware). • Server hardware: 256 GB RAM, 32 cores, 6.4 TB NVMe SSD. • Router hardware: shared boxes with multiple router pods. Each pod: 24 CPUs, 36 GB RAM. Could you share what you configure for Controller and Zookeeper? I assume ZK doesn't need NVMe, EBS gp3 should suffice, but curious how the setup looks like in terms of number of instances, cores, RAM and disk. Thank you
    k
    a
    • 3
    • 10
  • z

    Zac Policzer

    05/19/2026, 4:21 PM
    Launching Zibra Labs publicly today. This is probably not the most appropriate channel to do this kind of endorsing, but I've been working with an awful lot of you for a long time now and I'm not above asking friends for a little help on the engagement algorithm 😅 (likes/comments/reposts appreciated)
    🚀 10
    m
    f
    +3
    • 6
    • 19
  • f

    Felix GV

    06/07/2026, 10:57 PM
    Would anyone be interested in an ultra wide Venice mouse pad? (More pics in thread)
    😮 1
    s
    h
    • 3
    • 7
  • s

    Sergey Makagonov

    06/12/2026, 8:47 PM
    Hi all. RocksDB supports ordered iteration and range scans natively. Has the Venice team ever considered exposing range queries (fetch all keys within a
    [lo, hi]
    range) as a first-class read API on top of it? What we understand about why it's not there today: - Stores are hash-partitioned on the full key, so logically-adjacent keys scatter across partitions. There's no bounded home for a range, unlike point/batch-get. - Keys are opaque (Avro-serialized) bytes, and Avro's encoding isn't bytewise-order-preserving, so even a within-partition RocksDB iterator wouldn't map to a logical range. - A general unbounded scan would break the predictable-tail-latency / RCU quota model. Curious whether it's ever been discussed or prototyped, and if there's a design doc or RFC. If it was ruled out, I'd love to understand the main reasons. Thanks! cc @Felix GV @Zac Policzer @Koorous Vargha @Hubert Puszklewicz @Amre Shakim fyi
    f
    k
    z
    • 4
    • 23
  • k

    Koorous Vargha

    07/20/2026, 4:00 PM
    📣 <!everyone> Venice is moving from Slack to Discord! Why are we moving? 🌍 Wider open-source adoption Discord is widely used in the open-source community, making it easier for contributors and newcomers to join Venice. ♾️ Free, unlimited message retention Important conversations and shared knowledge will remain accessible without Slack’s retention limits. 📚 Existing Slack messages have been migrated. Going forward, please use Discord for new conversations. 👉 Join the Venice Discord: discord.gg/…
    s
    • 2
    • 1