This message was deleted.
# troubleshooting
s
This message was deleted.
g
judging by the log messages mentioning a variety of different time ranges, I think you are running into an issue that can happen when you read historical data from a stream that is splayed over a wide array of time chunks. each one gets its own segment (set of files on disk & in-memory structures); the total # can become too high to handle
there's a few approaches you could take to solve it: 1) you could load your historical data in batch mode instead. I would recommend SQL-based batch ingestion: https://druid.apache.org/docs/latest/tutorials/tutorial-msq-extern.html. then you could hook up kafka to the same table to load additional data from the stream 2) you could try sizing up from
nano
to
micro
or
small
or even larger, to see if these larger sizes would be able to handle your situation. (assuming your server supports these larger sizes.) 3) you could set
maxTotalRows
(https://druid.apache.org/docs/latest/development/extensions-core/kafka-supervisor-reference.html) lower, such that handoff happens more often, which limits the number of segments in flight. since you observe that things are OK until about 500K rows, you could set it to something conservative like 250K (half that). if you do this then make sure to run a compaction afterwards, since that will improve query performance quite a bit. (lots of small handoffs is not a recipe for good query perf!)
m
Interesting, that makes a lot of sense since that is exactly how my data is - would changing "segmentGranularity": "week" help at all? I seem to recall picking this option at random but there may be historically overwhelmingly more records per week than now
I gather from your response that it's "too much" rather than "too fast" and slowing my ingestion rate would not accomplish anything
I think I understand more about why this happens, a segment is "finished" if it is now the next week from the start of the first message of that week, so it starts another segment based on the amount of data it encounters in that time, you're suggesting I additionally limit that to a row number, so if either is hit, it starts a new segment
Is there ever a benefit to having more segments?
I think there's about a million from 10 years of data then it's very slow, so there would not be too many handoffs, and it's rare that I would search for a time range
I suppose more segments may help with parallelisation, in which case I could set it to 250k and 1 year, so they are at least even, and all the tasks search the same amount (since I search it all at once most likely)
Trying with that ^
g
would changing "segmentGranularity": "week" help at all?
kind of: coarsening it up to "year" would mean fewer segments in flight. however, i didn't recommend this because coarsening up like that tends to cause other problems down the road with e.g. compaction. (you can't compact a time chunk that is receiving active ingestion; with "year" that's an entire year worth of data!)
for these reasons, for most situations i recommend
hour
or
day
or, sometimes,
month
Is there ever a benefit to having more segments?
you got it: it's the parallelism, mainly.