This message was deleted.
# general
s
This message was deleted.
j
Strange. All acked consumed messages are supposed to be deleted by default
fyi, in Pulsar, TTL and retention each solve different problems. Maybe u would you like to check the documentation for reference. https://pulsar.apache.org/docs/cookbooks-retention-expiry/
s
Thanks. Sure will re-look into this.
l
In the above doc, there's a single bullet point referring to BookKeeper garbage collection. That's what eventually releases disk space when ledgers get compacted and possibly deleted too. A Pulsar topic is like a linked list of "segments" which are Bookkeeper ledgers. The ledger would have to be closed before BookKeeper Garbage Collection can happen. The ledger is closed when the topic segment roll over happens. This is the essential part of the doc:
• Segment rollover period: basically, the segment rollover period is how often a new segment is created. Once a new segment is created, the old segment will be deleted. By default, this happens either when you have written 50,000 entries (messages) or have waited 240 minutes. You can tune this in your broker.
• Entry log rollover period: multiple ledgers in BookKeeper are interleaved into an entry log. In order for a ledger that has been deleted, the entry log must all be rolled over. The entry log rollover period is configurable but is purely based on the entry log size. For details, see here. Once the entry log is rolled over, the entry log can be garbage collected.
• Garbage collection interval: because entry logs have interleaved ledgers, to free up space, the entry logs need to be rewritten. The garbage collection interval is how often BookKeeper performs garbage collection. which is related to minor compaction and major compaction of entry logs. For details, see here.
Bookkeeper compaction is described here: https://bookkeeper.apache.org/docs/getting-started/concepts#data-compaction There are multiple tunable parameters for rollover and multiple tunable parameters for BookKeeper compaction / garbage collection.
👍 1
🎯 1
m
you could test these settings as per your requirement
Copy code
managedLedgerMaxEntriesPerLedger=10000
managedLedgerMinLedgerRolloverTimeMinutes=5
managedLedgerMaxLedgerRolloverTimeMinutes=10
managedLedgerMaxSizePerLedgerMbytes=512
Or Throttle the compaction bytes rate to 10MB/s on bookkeeper.conf to get faster cleanup
Copy code
isThrottleByBytes=true
compactionRateByBytes=10485760
s
Also found the following line to be insightful if you are trying to figure out whats occupying the disks
Copy code
If you do not have any retention period and you never have much of a backlog, the upper limit for retained messages, which are acknowledged, equals the Pulsar segment rollover period + entry log rollover period + (garbage collection interval * garbage collection ratios).