Rarely do I agree with every single tweet in a 18-...
# general
c
Rarely do I agree with every single tweet in a 18-tweet thread, but: https://x.com/jaykreps/status/1796992396548567490
👍 1
One day, hopefully, LittleHorse’s core engine can be donated to Apache.
But we have a Loooonnnggg way to go before that’s even a relevant conversation
c
I agree with a lot of what Jay said here too, my main difference of opinion is the bit about the core needing to be OSS. I don't think that's always true - for example, it's definitely not true for Snowflake. open protocols are great (see: Kafka), and we chose to open-source Bento rather than keeping it closed and only embedded in our product because we were happy with the way the Benthos project was being stewarded prior to the change, and we thought something should live on in its place. (Using Redpanda's project was obviously a nonstarter, as explained here.) What happened with Redpanda Connect™️, the project formerly known as Benthos could only happen because it had a sole maintainer and minimal governance structure. For Bento, we're actively exploring governance models that would make this situation extremely difficult/impossible. We would be happy to use the old project if it was still independent, had a strong governance model, was not controlled by a vendor, etc.
🙌 1
c
Interesting analysis. One of the first things JayK said was along the lines of how Open Source / Proprietary is not as much about morality as people would make it seem to be. When a company "takes something proprietary" they are not changing the license on anything existing. Rather, they are simply saying:
From now on I won't be paying my engineers to contribute to code under the same OSS license as before; all new contributions will belong to our company.
On a moral standpoint I don't think that's wrong. From the business perspective, who am I to judge? JayK has some good ideas. Like WarpStream, I think Confluent also is an implementation of the Kafka API rather than being purely Apache Kafka. Kora is quite differentiated from OSS Kafka. Nothing wrong with that, but Confluent would do well to ensure that the OSS AK project continues moving forward so that it remains the industry standard. In my somewhat naive opinion, it's much better to be "the best of many implementations of the industry standard" rather than "the only implementation of something that's not standard"
You could argue that the core of WarpStream is OSS just as much as the core of Confluent is OSS. That is, there's a pretty damn good implementation of the core API (the Kafka API) in OSS.
c
ha - I get your point - except that we don't run any Kafka, anywhere in WarpStream. Confluent is still very much running "Kafka" - even if it's their internal fork of it, with a ton of special sauce on top. But now we're splitting hairs. I generally agree w/ what you were saying. The open protocol is a huge deal. The decision to fork actually had very little to do with morality, it just happens to also be moral high ground, given the license changes on the original project and trademark threats, etc. 🙂 It was more that we did not want our customers to have to worry about the licenses of what they're running when they run the WarpStream Agents. Several people reached out to us about this exact issue the day the announcement was made.
c
Correct—btw, what's the produce latency your users observe with S31Z? I know it's 100-500ms with regular S3, but what about 1Z?
c
we'll put out a blog post with benchmarks at some point, but it's about 2/3 less than S3 Standard. It's not extremely low, like sub-100ms p99 E2E Produce>Consume, because in our implementation we don't sacrifice MZ durability. We have you create three directory buckets in 3 different AZs and the Agent writes to a quorum of those buckets before it can give an ack to the producer
c
Does it end up saving much on cost then for non-spikey workloads? Writing to 3 buckets, wherein writes are somewhat expensive, seems like it might add up. But obviously you can scale it to zero, which isn't possible with Apache Kafka.
c
it is more expensive because you make more S3 PUTs. But we also save some money by immediately compacting files from S31Z to S3 Standard in the background. it all depends on what your Kafka cluster looks like. the way we phrase it is S31Z gives you a big knob to turn between cost and latency, and because it's all within the S3 ecosystem, it's basically a single interface for your latency sensitive and non-latency sensitive stuff
🙌 1
c
Thanks! Very interesting. Definitely good for the overall Kafka ecosystem, where there's now a spectrum of cost/latency from Kafka -> WS S31Z -> WS S3
🙏 1
m
The morality argument bugs the crap out of me. There’s nothing moral about being “shamed” into giving things away.
đź’Ż 1
If a person/company/club/whatever decides to give something away, great! Gold star! But saying it wrong to not do that is silly false equivalence.
đź’Ż 1
It’s yours. Do with it what you want. If the reverse was true; every time someone builds something on Kafka or builds a connector they should be forced to OSs it.
c
There are licenses such that the reverse is also true. Selfishly, I am thankful that Kafka does not use those licenses (:
m
Which is great…if you produce something or buy the folks that do…do as you like. But your not a better person or better company for choosing either path
In many ways I agree connect is showing its age, and lack of ability to mature. I’m super stoked that the best alternative is going to be around .
c
I’m not familiar enough with Connect to know the problems with it. What are they?
m
Lack of resource isolation. The api is difficult to deal with any errors. The distributed herder(controller) isn’t great and causes a bunch of rebalance like issues. The security model(like resource isolation) is really not made for multi tenant systems.
đź‘€ 1
Don’t get me wrong, the criticism comes from a good place. I’ve used and loved connect, it’s just got some honestly horrible worts.
âś… 1
c
I would argue that those shortcomings apply equally to producers/consumers in general. Kafka lacks multi-tenancy (only weak isolation with prefix-based ACL's). Furthermore, quotas in Kafka are per-broker not per-cluster, which is annoying. I actually talked about some of the gymnastics we do using Strimzi + Cruise Control to get around that issue at the StrimziCon (recording not out yet)
m
Yep!
c
Mind you, I'm quite possibly one of the biggest Apache Kafka OSS fans in the world, so I too say that with love 🙂
m
Having dealt with streaming systems for the better part of 20 years, i can only think of one that solved those reasonably well. And it’s dead now.
đź‘€ 1
The one that did was painful. But did work