This message was deleted.
# general
s
This message was deleted.
x
It’s from PageCache to ledger.
g
@Xiaogang Wen could you please explain more? I am specifically looking at the batch writes that happen o to the ledgers. small use case - we have journals enabled and also
journalSyncData
as
true
. We produce to a topic at 3MBps.. the journal observes continuous 3MBps write, but in ledger, the writes are only one every minute - so the complete 180Mb is written at once every minute. I want to control either this 1 minute or this 180Mb size
y
what’s the goal here? is that you have super fast ledger disk and the 1min or 180Mb is too small such that the io frequence is high?
g
No, that we are being io limited as the disk supports lesser write speed than 180MBps. If the flush happens more frequently i.e smaller sized batches, we would be fine
y
small size means you will have more io interruptions and more context switches. you are using more cpu to balance your slow io. I don’t get it why you do this. This will not improve your io rate from pulsar perspective.
g
When we do 180 MB in one minute, it takes over 3 seconds of induced io wait time by disk layer due to its write limitations
y
from the numbers of journal and ledger io rate, your pulsar throughput is limited by the journal disk 3MBps. brokers send back ack to producer once messages is persisted to journal disk.
g
If we do in 3 batches, 60 each, it would hardly take few ms
Agreed, about your last message, but we are uselessly wasting io resources being throttled unnecessarily
y
btw MB and Mb are different. 180MB/s is pretty slow. for journal, 3MB/s is way too slow you should consider swap journal and ledger disks if your goal is improving the performance.
g
Please assume it's bytes everywhere in my statements
Mobile autocorrect
3Mbps is just an example. It's not that the journal can't do more
Just to be clear, assume that our ledgers won't allow us to do more than 60MBps of write. So dumping 180 MB in one shot every minute takea 3seconds of io, while doing 60MB thrice a minute would take 10-15ms overall
y
this looks like io blocking. how many disks do you have? how many write threads do you have?
g
I don't think that changes anything. We have QOS limit on disk to not allow more than 60MBps
y
the ledger flush process is not just entry log, there are index flushing (to rockdb) too. the slowness may happen in index side.
g
And iostats clearly show an attempt of 180MB once a minute
The 60 limit is overall irrespective of io threads. There is a different iops limit which we aren't breaching
x
@Girish Sharma I can not find a in-depth description for the process of flushing an entry to ledger disk/file but this one by Hang in Chinese may give you some idea on it. Entries in ledger memtable are sorted and flushed to PageCache first. Then
flushEntrylogBytes
controls how frequent the entries are flushed from PageCache to ledger log file. The popular article from Jack Vanlightly also has a brief diagram on it. If using SSD for ledger disk you may use smaller
flushEntrylogBytes
as SSD is more friendly to small I/O ops.
g
@Xiaogang Wen Thank you for the references. I will go through them. So about smaller
flushEntryLogBytes
- i tried this, but i didn't seem to see any impact. In my small test case that I have been explaining above, the
flushEntryLogBytes
value is 10MB actually
x
Yes please check that blog post and if you still need further assistance on it please open a ZD ticket so that we can have the engineering team to check.