Hey everyone :wave: I’m experiencing an issue whic...
# troubleshooting
l
Hey everyone 👋 I’m experiencing an issue which makes it seem like the RealtimeToOfflineSegmentsTask is breaking my segments. Just wanted to check here before I raise an issue.. The minion task to convert the segments is taking a very long time, which causes the task to timeout. But while it has timed out, it looks like it still partially completes, and deleting segments and replacing them with segments in deep store with a strange prefix. Example segment: table_name__11__312__20220111T0936Za96a67f8-23e8-4213-b89b-9b199f3ffb21 In the actual table, all segments which end up like this show up as bad in error state, and can not be reloaded. Any guidance?
Fetching logs..
Untitled.txt
So what happens is that the mapper stage begins, ending on this log line..
Copy code
2022-01-12T12:40:28Z Initialized mapper with 43 record readers
Sometimes with up to 90 record readers. This stage then takes 35 minutes up to 1 hour. Once that’s done, it starts the usual reduction and sorting phase. (Even though the task should have been cancelled by now due to hitting timeout) It then starts destroying segments, and then immediately finishes with a timeout after destroying segments. I’m not sure exactly what happens here, but I’m wondering if this is causing an issue where the task is not properly cleaning up, causing loss of data. The logs above show the final moments where the task times out after issuing delete segments
Is this a bug?
r
which version of pinot are you using and do you have any raw columns?
l
No raw columns
version is 0.9.3
at least I don’t think I have raw columns
r
no dictionary columns?
just trying to figure out why it's taking so long, there was a known regression in 0.9.0 which will be fixed in 0.10.0, but it doesn't sound like it should matter here
l
Yes I have a bunch of no dictionary columns
sorry
So you’re saying the no dict columns is messing it up?
r
potentially
do you have a staging environment you can simulate this load in?
oh no
l
We don’t have customers in this yet
r
ok, so maybe try deploying from master here https://hub.docker.com/r/apachepinot/pinot/tags
if there is no customer impact
l
Ah you mean like that
r
try it, if it's still a problem, come back here. It will probably be this evening until someone very familiar with the realtime to offline flow will be around to help
l
Is there an issue link for that regression?
r
l
Thanks @Richard Startin
I’ve opted to remove the no dictionary columns instead of using the snapshot build to see if that does it
Hey @Richard Startin, does this also affect varLengthDictionaryColumns?
r
no
if dictionarizing the columns does help, then setting "rawIndexWriterVersion" : 4 should help for any string/bytes columns
it packs values more aggressively into chunks so the number of mmap/munmap calls will be very low (unintended high frequency of these syscalls for small chunk sizes was what caused the regression)
l
How do I set rawIndexWriterVersion? Is this backwards compatible, and is it present in 0.9.3?
Removing the non dict columns helped
But it’s still taking a very long time
Copy code
2022-01-12T14:04:33Z Initialized mapper with 39 record readers, output dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-2cf7ebb5-1166-4b03-9145-fb13fcb23aeb/workingDir/mapper_output, timeHandler: class org.apache.pinot.core.segment.processing.timehandler.EpochTimeHandler, partitioners: class org.apache.pinot.core.segment.processing.partitioner.TableConfigPartitioner
2022-01-12T14:51:10Z Beginning reduce phase on partitions: [0_12, 0_24, 0_3]
So that took 47 mins
In total this job was just below the threshold for timeouts
r
how many rows in the segment?
l
10,000,000
Here’s the task config
Copy code
"taskTypeConfigsMap": {
      "RealtimeToOfflineSegmentsTask": {
        "bucketTimePeriod": "3h",
        "bufferTimePeriod": "2d",
        "mergeType": "dedup",
        "maxNumRecordsPerSegment": 10000000,
        "roundBucketTimePeriod": "1h"
      }
    }
It created a total of 3 segments with 10M rows each
r
ok, it seems like a lot, could we profile it?
l
actually the segment is mega tiny
Copy code
segment.total.docs = 23513
So I believe this is happening because the “start of time” for my ingestion, there are some segments which contain very large time ranges
r
something seems wrong
l
Yeah I’ll try my best to explain the state of my segments
So the first 14 days of segments in my table are very dispersed, because of late data. This means that some segments contain very large time ranges. Since the RealtimeToOfflineSegmentsTask processes time buckets, it actually has to download some very large segments, but only extract partial data from it
So once the task has gone past those 14 days, segments should stabilise
Does that make sense?
So perhaps it doesn’t like that scenario
r
yes, makes sense
l
Can you think of any scenario in which this would slow down my tasks?
r
I'm not sure, if this is because there are very large segments to process there should be something in a profile of the minion, although a profile of 50 minutes is going to be a lot of data
are you familiar with running async-profiler as a native agent?
we should be able to figure out what it's doing that way, because I'm not very familiar with this workflow
l
I’m not familiar with that
The segments are in the 300MB range
m
My guess is that it is reading 14 days worth of data to generate 1 300MB segment. And it is all single threaded
r
what you could do is log on to the pod after you've seen the log statement about initialising the mapper, and then run
Copy code
jcmd <pid> JFR.start duration=300s filename=minion.jfr settings=profile
that will create a 5 minute profile of whatever it's doing before the sort phase and should pinpoint the theory it's downloading and sifting through a lot of data
l
Great I’ll do just that
Does this profile contain sensitive information?
I am not familiar
r
check the environment variables on the pod - if there's anything in there you wouldn't want someone outside your org to see, don't send it, or open it in JMC and send screenshots of the method profiling tab
no data is released
l
👍 thanks
r
there's also a setting to disable env var capture
I never use it so I need to look it up, but it would make it completely safe to share
it's going to be a pain to disable because you need a .jfc file, let's just hope there are no sensitive environment variables...
l
👍 It’s going to take me a little while before I can get around to doing it but I will let you know
Thanks for all the help
m
Is the look back window 14 days @Lars-Kristian Svenøy
l
It’s a 2 day buffer, 3 hr time window
but the segments from the “start of time” may contain data within a 14 day range interspersed
Let me get you a recent log file
Untitled.txt
So this log file shows the sequence of events in the minion from segment download to completion
Here’s another interesting logline
Copy code
2022-01-12T14:19:37Z Start sorting on numRows: 75094464, numSortFields: 9
2022-01-12T15:45:24Z Finish sorting in 5147498ms
1.42 hours to sort 75m rows
r
I suspected the sort
that's something concrete we can investigate
this is the relevant code:
Copy code
_fileReader = fileReader;
    int numRows = fileReader.getNumRows();
    _startRowId = 0;
    _endRowId = numRows;
    if (fileReader.getNumSortFields() > 0) {
      _sortedRowIds = new int[numRows];
      for (int i = 0; i < numRows; i++) {
        _sortedRowIds[i] = i;
      }
      Arrays
          .quickSort(0, _endRowId, (i1, i2) -> _fileReader.compare(_sortedRowIds[i1], _sortedRowIds[i2]), (i1, i2) -> {
            int temp = _sortedRowIds[i1];
            _sortedRowIds[i1] = _sortedRowIds[i2];
            _sortedRowIds[i2] = temp;
          });
    } else {
      _sortedRowIds = null;
    }
and there could be two factors: 1. you don't have enough heap and the GC is going crazy (that 5 min JFR profile after you see the sort start) would confirm that 2. the comparator is expensive when there are 9 sort fields
looking at what the comparator calls into I suspect there are some quick wins here, but checking your GC/heap settings are suitable for a 75M element sort would be handy too
@Lars-Kristian Svenøy do you really need a 9 dimensional sort?
l
Why is it doing a 9 dimensional sort?
r
Copy code
numSortFields: 9
this comes from your tableIndexingConfig.sortedColumn setting
l
So I’m just a little confused, because I only have one sorted column (I think you’re only allowed to have one anyway)
r
yes, I wouldn't expect this to be 9 either
ok sorry, I made a mistake, for ROLLUP it uses all non metric columns
and for DEDUP it's all fields
are you using DEDUP because you expect there to be duplicates, or just in case?
l
Yeah there can be duplicates
right so in this case, the dedup is making this slow as well I guess
why would it need to sort on all fields to do a dedup
r
that's the question I've raised with the rest of my team (sorry this has been a drawn out thread, I'm new to this particular component)
l
I’m running that profile now as well
r
Mayank also mentioned your data volume is high, so you might be stressing this component in a way it hasn't been before
l
Yes indeed
r
my suggestion is that we use your use case as guinea pig, I think we should be using a bloom filter for deduplication
sorting a large data set 9 ways is not going to end well, and it's just fortunate you don't have 100 dimensions...
do you have an id field that dedup could use?
l
😄
Yes I do
I actually wanted dedup to support the primaryidColumns
that would be excellent
r
a bloom filter on the primary key should work well
l
I store two ids which are composite together, and then I store days since epoch
I want to use those to reduce it to 1 record per day
[id_1, id_2, daysSinceEpoch]
r
can you tolerate duplicates for the time being? There's a gap which needs to be closed, in the meantime, you could resolve it by using concat and accept duplicates, or accept the performance impact of deduplication
l
Yes
I can
But I am not entirely sure dedup is the only reason why this is slow
Because it’s not only the sort that is slow, but it’s a start
r
yes, you mentioned the stage before dedup was slow too
l
I can disable dedup for now and we’ll see where that takes us
So far I’ve disabled my no dict columns as well
and also I’ve reduced the time buckets by a lot
r
if I can get a profile from that stage, I can find something concrete to do about it, just the dedup is such an obvious problem it's easy to suggest a path forwards
l
it’s actually to the point now where the segments are way too small
👍 Let me change my config right now
r
what data type are your no dict columns?
are any of them numeric?
l
Yes
long, double, string, int, int, string, int, bool, int, double, int, string, int, long, string, long, double, long, string
r
for any string raw columns, you could set "rawIndexWriterVersion" to 4 which doesn't have the problem mentioned, and packs more strings into each chunk within the segment
l
And this works for 0.9.3?
r
yes
l
I could do that, I think I’ll take it one step of a time first and go back there eventually. I’d rather just not have the indexes for the time being
But thanks for pointing me in that direction
FYI I raised another ticket about not hardcoding task timeouts as well
👍 1
It would have been useful, as some of my tasks take just a little longer than 1 hour
r
getting the profile from the first slow stage should be the next step, as well as temporarily using concat in preference to dedup
l
It also seems like the task actually completes most of the way, and just times out when it’s done which is annoying 🙂
Kind of like if you were to give me a deadline of 4 weeks, I finished it after 8 weeks and you told me to scrap it 😄
Yes, I’ll get that profile and the concat
What version of mission control do I need?
latest?
r
JMC 8 should be fine
if you don't have any sensitive environment variables it would be good for us to have a copy of it, but if you need to take screen shots having the flamegraph view in the hot methods tab would be great
l
Yeah I can’t send it over
r
also flame graps in the TLAB allocations tab by method
l
image.png
This is after
Copy code
2022-01-13T08:49:30Z Initialized mapper with 52 record readers, output dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-5d1a14ea-4e90-44d0-9350-89942f0001e7/workingDir/mapper_output, timeHandler: class org.apache.pinot.core.segment.processing.timehandler.EpochTimeHandler, partitioners: class org.apache.pinot.core.segment.processing.partitioner.TableConfigPartitioner
Where do I find the hot methods tab?
r
it's called method profiling
getting the flame graphs would be helpful, let me send an example
l
image.png
image.png
r
Screenshot 2022-01-13 at 09.49.54.png
is that the top method? Usually the top one is hidden and you need to scroll up (JMC bug 😞 )
this is going to be a pain, but can I have the screenshot of the stack trace for each of the top 5 profiled methods (select each method) and for the top 5 tlab allocations too? I think we can find some quick wins from that.
l
Yes
Give me a sec please I’m going into standup
Top 5 Methods
Top 5 TLAB
This is during
Copy code
2022-01-13T09:53:44Z Initialized mapper with 45 record readers, output dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/mapper_output, timeHandler: class org.apache.pinot.core.segment.processing.timehandler.EpochTimeHandler, partitioners: class org.apache.pinot.core.segment.processing.partitioner.TableConfigPartitioner
r
thanks
l
Confirmed there is no hidden top method
This isn’t the heaviest minion task run
I will get you a better profile later
r
we could get rid of getUnpaddedString and related frames by changing the string columns to raw but use V4
the GenericRow frames are things I have seen before and almost changed this a couple of months ago, so I might go ahead with that in the next few days
moving to raw V4 would also get rid of the thread locals
l
Is there a plan to default to v4?
Why aren’t we defaulting to it
Could you show me an example of how to use v4?
r
1. it's new 2. we need consensus to make it the default 3. there may end up being a v5 to support very large files according to some feedback
yes
Copy code
"fieldConfigList": [
    {
      "name": "col1",
      "encodingType": "RAW",
      "properties": {
        "rawIndexWriterVersion": "4"
      }
    }
  ]
directly in your table config, as a sibling of tableIndexingConfig
my word, zookeeper client allocates a lot. That isn't something we can do much about though.
l
Great thank you very much
Really appreciate all this help
Recording another profile for this
Copy code
2022-01-13T10:10:09Z Start sorting on numRows: 42757049, numSortFields: 9
It’s been stuck there for 25 minutes so far
Here’s the logs so far for that
Copy code
2022-01-13T09:53:44Z Beginning map phase on 45 record readers
2022-01-13T09:53:44Z Initialized mapper with 45 record readers, output dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/mapper_output, timeHandler: class org.apache.pinot.core.segment.processing.timehandler.EpochTimeHandler, partitioners: class org.apache.pinot.core.segment.processing.partitioner.TableConfigPartitioner
2022-01-13T10:08:47Z Beginning reduce phase on partitions: [0_0, 0_1, 0_10, 0_11, 0_12, 0_13, 0_14, 0_15, 0_2, 0_3, 0_4, 0_5, 0_6, 0_7, 0_8, 0_9]
2022-01-13T10:08:47Z Start reducing on partition: 0_0
2022-01-13T10:08:47Z Start sorting on numRows: 531663, numSortFields: 9
2022-01-13T10:08:51Z Finish sorting in 4563ms
2022-01-13T10:08:51Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_0
2022-01-13T10:08:52Z Finish creating dedup file in 655ms
2022-01-13T10:08:52Z Finish reducing in 5233ms
2022-01-13T10:08:52Z Start reducing on partition: 0_1
2022-01-13T10:08:52Z Start sorting on numRows: 170212, numSortFields: 9
2022-01-13T10:08:53Z Finish sorting in 1125ms
2022-01-13T10:08:53Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_1
2022-01-13T10:08:53Z Finish creating dedup file in 210ms
2022-01-13T10:08:53Z Finish reducing in 1340ms
2022-01-13T10:08:53Z Start reducing on partition: 0_10
2022-01-13T10:08:53Z Start sorting on numRows: 90414, numSortFields: 9
2022-01-13T10:08:54Z Finish sorting in 630ms
2022-01-13T10:08:54Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_10
2022-01-13T10:08:54Z Finish creating dedup file in 114ms
2022-01-13T10:08:54Z Finish reducing in 748ms
2022-01-13T10:08:54Z Start reducing on partition: 0_11
2022-01-13T10:08:54Z Start sorting on numRows: 427784, numSortFields: 9
2022-01-13T10:08:57Z Finish sorting in 2969ms
2022-01-13T10:08:57Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_11
2022-01-13T10:08:57Z Finish creating dedup file in 485ms
2022-01-13T10:08:57Z Finish reducing in 3465ms
2022-01-13T10:08:57Z Start reducing on partition: 0_12
2022-01-13T10:08:57Z Start sorting on numRows: 682524, numSortFields: 9
2022-01-13T10:09:02Z Finish sorting in 5152ms
2022-01-13T10:09:02Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_12
2022-01-13T10:09:03Z Finish creating dedup file in 715ms
2022-01-13T10:09:03Z Finish reducing in 5884ms
2022-01-13T10:09:03Z Start reducing on partition: 0_13
2022-01-13T10:09:03Z Start sorting on numRows: 2072511, numSortFields: 9
2022-01-13T10:09:22Z Finish sorting in 18428ms
2022-01-13T10:09:22Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_13
2022-01-13T10:09:24Z Finish creating dedup file in 2327ms
2022-01-13T10:09:24Z Finish reducing in 20806ms
2022-01-13T10:09:24Z Start reducing on partition: 0_14
2022-01-13T10:09:24Z Start sorting on numRows: 3382438, numSortFields: 9
2022-01-13T10:09:53Z Finish sorting in 28957ms
2022-01-13T10:09:53Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_14
2022-01-13T10:09:57Z Finish creating dedup file in 3919ms
2022-01-13T10:09:57Z Finish reducing in 32963ms
2022-01-13T10:09:57Z Start reducing on partition: 0_15
2022-01-13T10:09:57Z Start sorting on numRows: 142846, numSortFields: 9
2022-01-13T10:09:58Z Finish sorting in 957ms
2022-01-13T10:09:58Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_15
2022-01-13T10:09:58Z Finish creating dedup file in 175ms
2022-01-13T10:09:58Z Finish reducing in 1137ms
2022-01-13T10:09:58Z Start reducing on partition: 0_2
2022-01-13T10:09:58Z Start sorting on numRows: 127443, numSortFields: 9
2022-01-13T10:09:59Z Finish sorting in 820ms
2022-01-13T10:09:59Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_2
2022-01-13T10:09:59Z Finish creating dedup file in 154ms
2022-01-13T10:09:59Z Finish reducing in 978ms
2022-01-13T10:09:59Z Start reducing on partition: 0_3
2022-01-13T10:09:59Z Start sorting on numRows: 554682, numSortFields: 9
2022-01-13T10:10:03Z Finish sorting in 4233ms
2022-01-13T10:10:03Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_3
2022-01-13T10:10:04Z Finish creating dedup file in 670ms
2022-01-13T10:10:04Z Finish reducing in 4917ms
2022-01-13T10:10:04Z Start reducing on partition: 0_4
2022-01-13T10:10:04Z Start sorting on numRows: 251158, numSortFields: 9
2022-01-13T10:10:06Z Finish sorting in 2123ms
2022-01-13T10:10:06Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_4
2022-01-13T10:10:06Z Finish creating dedup file in 292ms
2022-01-13T10:10:06Z Finish reducing in 2423ms
2022-01-13T10:10:06Z Start reducing on partition: 0_5
2022-01-13T10:10:06Z Start sorting on numRows: 180554, numSortFields: 9
2022-01-13T10:10:08Z Finish sorting in 1343ms
2022-01-13T10:10:08Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_5
2022-01-13T10:10:08Z Finish creating dedup file in 227ms
2022-01-13T10:10:08Z Finish reducing in 1575ms
2022-01-13T10:10:08Z Start reducing on partition: 0_6
2022-01-13T10:10:08Z Start sorting on numRows: 110286, numSortFields: 9
2022-01-13T10:10:09Z Finish sorting in 712ms
2022-01-13T10:10:09Z Start creating dedup file under dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-a4085745-be54-4411-abfa-c0dfb8fa05a1/workingDir/reducer_output/0_6
2022-01-13T10:10:09Z Finish creating dedup file in 139ms
2022-01-13T10:10:09Z Finish reducing in 856ms
2022-01-13T10:10:09Z Start reducing on partition: 0_7
2022-01-13T10:10:09Z Start sorting on numRows: 42757049, numSortFields: 9
r
did you change to concat?
l
Not yet
Wanted to get a profile without
r
oh... I'm willing to go blind that the sort will show up there and nothing else...
l
😄 It seems likely
r
well, deserializing the values in the comparator will figure too
l
It seems really strange that dedup would need to sort on 9 dimensions
seems like a naive dedup approach
r
you're not wrong
it's a fairly new feature though
l
Would it not be more performant to store a simple hash of rows as they are seen, and then just discard if finding a match? Removes the ability for multi-threading, but pretty sure from what I’ve seen, there is no threading here
More multi threading where possible would probably help speed up the minion tasks drastically as well
Even with 4 cores, this minion task is currently using 0.25 of a core
I’ve checked IOPS, it isn’t using them all
r
I'd want to do that with a bloom filter to get a size bound
l
image.png
It barely uses CPU so there’s definitely something wrong
r
I've mentioned parallelism, but opinions differ on that
l
image.png
With 8GB heap ^
xms and xmx are set to 8
r
it's intended that the minion will process tasks concurrently, but I think it could use FJP if it sliced the task up
l
File system
image.png
huh maybe it is bottoming up on IO?
let me go check
r
the sort comparator does IO, so that should be expected
l
suspicious that it is capping out
r
but it's probably a symptom, not a cause
the problem is the sort, any resource it uses will get saturated
l
👍
The volume has 3000 IOPS, looking at monitoring it’s using 3200 read ops/s and 52 write ops/s
image.png
So it’s capping out on read ops
Is that right?
I’ve got the capture
Top 5 methods
Top 5 TLAB
So yes the sort is heavy, but look at the deserializer and databuffer copy
r
they come from the sort's comparator
l
yeah so removing dedup will speed that up?
r
I consider all of this a consequence of the sort
l
That makes sense
r
yes, by a factor of at least 9
l
So it’s actually sorting 42 M rows using quicksort * 9?
r
more if your sorted column is small or fixed width
l
that’s crazy
r
no let me grab the code
l
So latency is going to grow logarithmically with the amount of rows
That makes sense looking at the output from other sort operations
r
Copy code
_sortedRowIds = new int[numRows];
      for (int i = 0; i < numRows; i++) {
        _sortedRowIds[i] = i;
      }
      Arrays
          .quickSort(0, _endRowId, (i1, i2) -> _fileReader.compare(_sortedRowIds[i1], _sortedRowIds[i2]), (i1, i2) -> {
            int temp = _sortedRowIds[i1];
            _sortedRowIds[i1] = _sortedRowIds[i2];
            _sortedRowIds[i2] = temp;
          });
l
Yes right
that’s going to put a huge strain
I can see why my IOPS are through the roof considering that filereader call
Do you want me to raise an issue for this?
r
here's some low hanging fruit https://github.com/apache/pinot/pull/8013
Yes, please create an issue
can you do another run with v4 raw indexes? I would just change to concat to save waiting for the sort for now.
l
I’m running with concat right now
Would using v4 make any difference considering I’m not using raw indexes anymore, @Richard Startin
r
yes it would, because 2 of the hottest methods are related to the string dictionary
l
gotcha
do I still set field type to raw then?
r
so using V4 for the raw string columns should help (but only for string columns)
Copy code
"fieldConfigList": [
    {
      "name": "col1",
      "encodingType": "RAW",
      "properties": {
        "rawIndexWriterVersion": "4"
      }
    }
  ]
l
When you saw raw
Does that mean I need to use no dict columns for it?
What impact does it have to make it raw? for query performance
r
yes, no dict columns required also (it's a bit repetitive, I know)
l
Alright thanks
r
it means that you don't have a dictionary, so there's no hashing, and there's no indirection
l
Ok so with concat, one of these tables is now taking 23 minutes to execute the job
Need to check the heaviest tables too
r
when you have a dictionary, the values are all deduplicated in a dictionary, and within the column itself, the value is replaced by the maximum number of bits to encode an offset into the dictionary
l
Ok so even with concat, my largest table has been stuck on mapping for 30 mins
gotcha
I’ll try v4 then
No risk associated with it? 👁️
I’d prefer to avoid data loss
I’m OK with it crashing and burning
r
you won't lose data
l
Just not data loss, I don’t have backfill
great
r
well, wait
how large are your segments? There's a 4GB column limit with V4 (it was 2GB by default before V4, so you should be ok if you had raw columns before, just be aware of the limit)
l
They aren’t 4GB
right now they are super tiny
due to this defect
r
I might be able to do something about GenericRow being slow (the other 2 of the top 5) today, so mapping should speed up
I'll discuss using the FJP to divide and conquer when there are cores available this evening
l
👍
that would be a good idea
Hey Richard, the field config list
do I need that in both the OFFLINE and REALTIME table? @Richard Startin
r
just OFFLINE
l
gotcha
I’ll get right on it
do I keep the no dict cols in both realtime and offline then?
r
no dict cols in both, just the V4 config for offline
l
Thanks
Also one more question re. this
The recommender also told me to both my timestamp in the no dict col
Is that recommended?
r
yes
l
Great
Alright it’s going to take a while to write out all this json.. 🙂
r
the reasoning is that the dictionary is a glorified bidirectional hashmap, and when you have a dictionary you replace the value in the column with the offset into the dictionary, if you have a high cardinality fixed width type, it just adds indirection and increases size because there are almost as many dictionary codes as there are values
l
That makes sense
r
also, if you have large strings, hashing those strings to do dictionary lookups is costly. If they're basically unique (JSON messages) it's pointless
l
Makes sense
Also, the recommender told me to put something with a cardinality of 3 and avg length of 3 in the no dict column
I take it I ignore that?
r
I would ignore that
do you have more details? data type?
l
It’s basically an enum
a string
Only strings work for raw right?
Any rule of thumb for which vals should make it to no dict?
fixed set of values probably not yeah?
r
both work, but there's that regression with an unreleased fix which we need to work around. V4 only works for strings and bytes
l
Like if I have 315 distinct values for something, I don’t place it in no dict right?
r
I would make high cardinality column raw, make things you aggregate metric columns
high cardinality is millions
l
You can only aggregate numbers no?
You can’t make strings metric cols
r
if you have 315 distinct values, the dictionary code will be 7 bits, so you could fit 9 of them in a single long in the column itself, it really pays for the indirection
indeed
l
So you’re saying don’t use no dict for 315 distinct vals
r
but making numbers raw usually pays off, we just have this annoying regression to work around right now
l
Sorry just a little confused
r
315 distinct values should always have a dictionary
l
Yes gotcha
For some reason the recommender didn’t agree
r
I'd like to understand why, do you have case statements on that field in your queries?
l
Nope
Give me a few minutes
Can you use no dict cols on multi value fields? @Richard Startin
Also is there any point in having a dimension in both varLengthDict and invertedIndex?
Copy code
2022-01-13T11:13:37Z Beginning map phase on 64 record readers
2022-01-13T11:13:38Z Reflections took 186 ms to scan 24 urls, producing 6 keys and 219 values 
2022-01-13T11:13:38Z Initialized FunctionRegistry with 131 functions: [fromepochminutesbucket, arrayunionint, codepoint, mod, sha256, year, yearofweek, upper, ago, arraycontainsstring, arraydistinctstring, bytestohex, tojsonmapstr, trim, timezoneminute, sqrt, togeometry, normalize, fromepochdays, arraydistinctint, geotoh3, exp, stgeogfromwkb, stgeogfromtext, stgeomfromwkb, jsonpathlong, yow, toepochhoursrounded, lower, toutf8, concat, ceil, todatetime, jsonpathstring, substr, dayofyear, contains, jsonpatharray, arrayindexofint, fromepochhoursbucket, totimestamp, arrayindexofstring, minus, arrayunionstring, toepochhours, toepochdaysrounded, millisecond, fromepochhours, arrayreversestring, dow, doy, min, toepochsecondsrounded, strpos, jsonpath, tosphericalgeography, fromepochsecondsbucket, max, reverse, regexpextract, hammingdistance, stpoint, abs, timezonehour, stgeomfromtext, toepochseconds, arrayconcatint, quarter, md5, ln, toepochminutes, arraysortstring, replace, strrpos, jsonpathdouble, stastext, second, arraysortint, split, fromepochdaysbucket, lpad, day, toepochminutesrounded, strcmp, fromdatetime, fromepochseconds, arrayconcatstring, fromtimestamp, base64encode, ltrim, arraysliceint, chr, sha, plus, base64decode, month, arraycontainsint, toepochminutesbucket, startswith, week, jsonformat, sha512, arrayslicestring, fromepochminutes, remove, dayofmonth, times, hour, rpad, arrayremovestring, now, divide, bigdecimaltobytes, floor, toepochsecondsbucket, stasbinary, toepochdaysbucket, hextobytes, rtrim, length, toepochhoursbucket, bytestobigdecimal, toepochdays, arrayreverseint, datetrunc, minute, round, jsonpatharraydefaultempty, dayofweek, arrayremoveint, weekofyear] in 197ms
2022-01-13T11:13:38Z Initialized mapper with 64 record readers, output dir: /var/pinot/minion/data/RealtimeToOfflineSegmentsTask/tmp-bacd9aec-63bb-4926-8cbd-8b09d91af0b5/workingDir/mapper_output, timeHandler: class org.apache.pinot.core.segment.processing.timehandler.EpochTimeHandler, partitioners: class org.apache.pinot.core.segment.processing.partitioner.TableConfigPartitioner
2022-01-13T12:12:30Z Beginning reduce phase on partitions: [0_12, 0_13, 0_24, 0_3]
Largest table still struggling during mapping, I’ll try the raw indexes
all other tables are doing really well now
only 2 outliers
r
another profile during the processing of that large table would be good
don't worry about allocations, it doesn't seem to be the problem
I think we're probably justing seeing the cost of 42M * the cost of
GenericRow
being a bit slow now
l
Yeah makes sense
It would be nice to have configurable task timeouts
I’ll grab another profile once the task runs with the index changes
r
I was going to create an issue for that
l
I created one already
I’ve created a few issues the last few days
r
👍
l
I am grabbing a few more captures here. I’ll do one now, and then after the raw index change
I’ve added my new profiling runs on the issue
r
the raw index should definitely help looking at those profiles, half the time is in the string dictionary
l
That’s true, I am just not sure if I can use a raw index for some of the high cardinality strings I have
two of the strings are high cardinality and fairly large, but I need them
as they make up my composite id
r
why would you not want them to be raw? Do you have inverted indexes on them?
l
So one of them is my partition key
It’s always in a where clause
r
ok
l
The other one is also usually in a where clause
r
but the ones which used to be raw before we started digging in to this should be raw
l
yep
Really once it gets past those bad segments at the start of time, it shouldn’t struggle as much
I just don’t know if it ever will..
I wonder if I can hack whatever config the realtime task is using to make it skip ahead a bit
I wouldn’t mind doing that
r
is it coming from Kafka?
l
yep
The problem is the task is hanging behind by like 2 months right now, and the first 14 days are complete garbage
not garbage, but the segments are very large time ranges
r
@Neha Pawar should be able to help skip over the data
I have another PR I am about to submit before I get on with something else which should halve the time spent in
GenericRecord
operations, but getting dedup fixed is going to take at least a month because we'll need to write a design doc and discuss it, similar for any parallelisation (which is less likely to happen)
how did v4 go?
l
Stuck in the middle of something else atm, I’ll be able to check later @Richard Startin
n
You can manually bump up the start time ahead by 14 days, in zk browser, config/minion metadata
It's a hack for not. We should have a config to allow one to pick a start time
You could also try reducing the bucket time to even smaller(1h), or max rows to something smaller (1m). Once you get past the 14 days, bump it back up
l
I’ve tried doing the bucket change, but it’s not having an effect unfortunately
I’ll change the start time
thanks for the info
I am not seeing much of an impact with v4 unfortunately @Richard Startin. I’m guessing it’s because I don’t have raw indexes for some of my highest cardinality strings
👍 1
But overall performance has improved across the board for all my tables