Interesting and relevant: <https://www.linkedin.co...
# general
c
How is “cluster” defined here? Is it like a service that you submit jobs to? I have never been a fan of the Flink or Kafka Connect architecture where there’s something running and you submit work to it. About rolling restarts of Kafka and statefulsets…just write a decent Kubernetes operator (or use Strimzi) and that problem goes away.
g
yeah, I think this was the definition
and if I understood Ryan correctly, this is exactly his point
c
Ah. Wouldn’t be surprised if Confluent completely rewrites the Flink implementation but keeps the API.
I mush prefer the kafka streams architecture and Operator pattern 😉
Thanks for sharing. Very interesting. Strimzi’s KafkaConnect CRD’s give a sort of declarative API over Connect which somewhat hides the actual cluster-style implementation.
g
interesting, I never tried it, but it sounds like something that will be super popular
m
The interesting question from this article is what are the medium sized enterprises doing? Or the less sophisticated? Not everyone is twitter or LinkedIn. What about someone that doesn’t need N+ clusters, but does need a few or needs more scale than a 500 person small business needs? Also the security and governance aspects…are info sec teams keeping up?
💡 1
💯 1
d
There's an analogy here where we start with fruit that has the seeds on the inside, but as we get better at breeding and managing the fruit we want the seeds on the outside so they're accessible to us. That said, the use of a remote controlled plane as the blog picture is a high grade pun and I'm here for it.
😂 1
r
Gwen, I'm not sure if you were around as much at the time, but Ryanne came to Cloudera through the Hortonworks merger and wrote the Mirror Maker 2 implementation while there. And he worked on Brooklin at LinkedIn. It wouldn't surprise me if he's been thinking about this sort of thing for a while.
c
This discussion about Kafka Connect just came up in an internal Technical Strategy meeting at LH... In particular, we want to build a similar project called "LittleHorse Connect" and the debate centered around whether we should build in any workload orchestration tools at all into the LH Connect framework. It seems from Ryanne's article that his main criticism of Kafka Connect is that the Java toolkit is useful (the SPI), but the concept of having a cluster with some orchestration capabilities that are quasi-competitive with K8s is a bit of a downside, because now you have to run that orchestrator on top of K8s rather than using native K8s stuff to assign workloads onto resources. So we are leaning towards making LH Connect "orchestrator-less", so that the users can use whatever orchestrator they want (which in practice will almost always be K8s, so our "paid product" will be a K8s Operator that orchestrates it on K8s)
👍 2
g
Also seems in-line with Kafka Streams and LittleHorse thinking? Going with apps, not clusters?
c
Yeah, for sure. Kafka Streams is a library for building apps, Flink is a cluster to which you submit jobs. Someone who knows the space well said earlier today in another slack:
I've never seen a small or medium company successfully self-manage Flink, so I don't know why there's so much buzz around it
On the flipside, it's hard to monetize a library like Kafka Streams, which might explain why there is so much more commercial investment into Flink (there is a server-side component, and monetizing a server-side thing is very simple). LH has no opinions about how your clients are deployed, but it also has a server-side component, which allows us to have a cloud service
m
Meroxa sorta did this already with conduit
They used Golang but same sort of ideas