This message was deleted.
# troubleshooting
s
This message was deleted.
m
Hi Igor. This is a good place to read about tuning. Here's the Broker documentation.
Are you able to share anything about your use case?
i
I can share since it’s POC we ingesting approx 8Bil events from kafka with rollup, keeping 24 hrs of 5 mins buckets. Schema is time bucket + 4 long fields as part of dimensions: accountId, campaignId, adId, publisherId) and then anothery 4 long fields of kpis(1-4) meanwhile it’s only 1 server. pretty strong one(40 cores, 256GB memory) everything is running there as ‘large’ setup + some memory/buffer increases for broker(e.g.) I tried to run client that takes 100 Top accountIds(computed by ordering sum(adIds) desc) and then executes under diff number of threads(10 to 100) query with PreparedStatement that samples accountId and then aggregates data to hourly granularity for this account, something like:
Copy code
"select TIME_FLOOR(__time, 'PT1H') as _hour, " +
                "       SUM(kpi1) as kpi1, " +
                "       SUM(kpi2) as kpi2, " +
                "       SUM(kpi3) as kpi3, " +
                "       SUM(kpi4) as kpi4 " +
                "from raw_events " +
                "where __time >= CURRENT_TIMESTAMP - INTERVAL '24' HOUR " +
                "and accountId= ? " +
                "group by 1 " +
                "order by 1";
and then measured latency of query execution. Basically we are pleased with latency, however looking at heap of broker it’s not clear to me why it increased and continues to keep something in memory even though test finishes and full GC is triggered. So either it a)some misconfig from my side b) some missing code piece that release something or c) some leak(less probable)
s
Hi @Igor Berman.Take a look at heap sizing for the Broker. It points to 3 major sources of heap usage: • Partial unmerged query results.. these should not persist on heap after queries are completed. • The segment timeline and cached segment metadata do persist on heap throughout the broker's normal operation. Its size will depend on the number of segments and number of columns in the segments. How many segments did you end up with for the 4 billion rows? Another important optimization strategy is segment sizing which aims to create segments of approx 5 million rows / 300-700 MB each.
i
I’ve reduced retention lately to be 1/4 of what it was before. currently I have 929,631,452 total rows(after rollup), 90 segments, compaction strategy of up to 20Mil rows per segment. the Datasource is fully compacted(and was during the test), segments sizes are between 50MB(5Mil rows) and 140MB(20Mil rows)
ok, thanks for the pointer @Sergio Ferragut I’ll try to play with it with different number of segments to see if it correlates
s
Sounds good. Also, do you see that grow over time or it is stabilizing with a stable data set? If streaming continues, brokers also keep a record of any real-time segments being built.
i
it remains flat after test is done, but during test, when i go from 10 concurrent client threads to 100 concurrent threads it’s increasing
s
That makes sense while the queries are running. Does it stay high after the workload completes?
i
yes, it remains high
g
are you using JDBC (Avatica)?
if so, I think you discovered a regression in 24.0: these objects are leaked during JDBC connections. You won't see a similar leak for the JSON-over-HTTP API, and you won't see it in the prior release. You also won't see it in the next release, since we'll fix it! PR: https://github.com/apache/druid/pull/13259
🙌 1
i
oh! great to know. we indeed using jdbc(Avatica)!
g
OK, I am pretty sure you are running into that leak then. Sorry, in the meantime the best workaround is to restart your Broker, which will clear out those objects. Or, you can apply the patch in the linked PR if you are building from source
Thanks for including the visualvm screenshot, it helped track down the problem
i
Thanks for fast response! You building great product!
g
glad you like it so far 😄
i
@Gian Merlino last question 🙂 when do you think 24.0.1 will be released? (just to understand if we can wait for it or will need to build ourselves)
g
Can be tough to nail down a specific day in advance, since the voting process for each RC is minimum 72 hours even for a small release, by policy. However, I would guess release would go out in the next 1-2 weeks. Once an RC is up for voting, you can monitor the dev list for progress
👍 1
i
fair enough. I’ll monitor dev list thanks!
c
@Igor Berman, i was wondering if v24.0.1 fixes your problem? We are facing a similar problem with an increasing heap that remains high.
i
@Cl A Us yes, we see that it is better in v24.0.1. Having said that, we still see some heap problems from time to time(but they are not frequent) which version you are using?
c
Thanks for your quick reply. we are using v24.0.1. We migrated from v.0.22.1 and we are not sure if it got worse after the update.
i
I understood as well that it’s connected to sql api, if you using native api, than this memory leak(and it’s fix) is not relevant
c
Thanks for the hint. We are using the avatica sql driver to query for druid datasources. 😞
@Igor Berman, we reproduced the problem in our integration stage. It was not solved with v24.0.1 but v25.0.0 did it. The memory leak is gone with that version. https://apachedruidworkspace.slack.com/archives/C0309C9L90D/p1675739962904279?thread_ts=1674643179.664379&cid=C0309C9L90D
🙌 1
i
thanks a lot @Cl A Us for the update! will need to upgrade then, we do have occasional need for broker restarts in production so it may explain this please tell me how it’s going with 25.0.0
@Michael Taranov @Ori Amichay fyi
✅ 1