This message was deleted.
# general
s
This message was deleted.
l
Somehow in pulsar i hope that works.
As we have learned from the SRE book, "hope is not a strategy" 🙂 In Pulsar, you can scale the storage layer (bookkeeper/"bookies") and the serving layer (brokers) separately. In addition, since Pulsar has per-message acknowledgements, this provides different ways of consuming messages in applications and scaling applications when compared to Kafka. For example, Pulsar has "shared" and "key_shared" subscription types . Performance and scalability has many aspects. In many cases, you should understand what the bottleneck is so that you can make conscious decisions. There's some advice buried in this thread: https://apache-pulsar.slack.com/archives/C5Z4T36F7/p1672932396402779?thread_ts=1672929210.109419&cid=C5Z4T36F7
🙌 1
m
I understand that scaling is more than just adding replicas... But... if i add another node in kafka, without configuration it does nothing... So how's pulsar there? in general? 🙂
brokers are stateless... so i expect them to perform better
3 are better than 1
l
That's true, that it's easier to scale out brokers.
âś… 1
m
maybe it breaks down to the question: do i need a developer to change stuff in code to scale?
l
maybe it breaks down to the question: do i need a developer to change stuff in code to scale?
it's not usual. However, if the implementation is very dependent on low latency, this could break under high load. Meeting strict latency requirement might require resource overprovisioning and this has an impact on cost.
m
ok got it
l
In Pulsar, in many cases, a lot of read operations can by-pass reading from BookKeeper and from disk when the consumers are in "tailing read" mode and all messages can be passed to the consumers from caches. On a very busy system, it's possible to get into a state that it's not easy to get the system running after a total crash if there's not enough resources to handle a lot of cache misses. When operating at very high scale, it's necessary to have a plan to handle this.
m
I would say we're in a state right now which extreme low throughput.... it's more about what system provides the easier way to scale up in production later...
l
yes, there are a lot of options for scaling.
🙌 1
m
thank you 🙂
d
Just to add to Lari’s great posts, Pulsar’s stateless brokers helps scalability in two ways. The first is the ease at which one can add brokers to the cluster and have them immediately start serving clients. The second less obvious one, is Pulsar’s automatic load balancing capability, which helps prevent individual broker “hot-spotting” scenarios where one broker has been assigned all the “busy” topics and is therefore overloaded.