When running the realtimeProvisioningHelper, we go...
# troubleshooting
m
When running the realtimeProvisioningHelper, we got a bunch of NAs. Any idea on how to troubleshoot this?
Copy code
RealtimeProvisioningHelper -tableConfigFile <tableConfig> -numPartitions 1 -pushFrequency null -numHosts 1,2,3,4 -numHours 1,2,3,4,56,12,18,24 -sampleCompletedSegmentDir <path-to-segment> -ingestionRate 1000 -maxUsableHostMemory 5120G -retentionHours 1
Note:

* Table retention and push frequency ignored for determining retentionHours since it is specified in command
* See <https://docs.pinot.apache.org/operators/operating-pinot/tuning/realtime>
Memory used per host (Active/Mapped)

numHosts --> 1 |2 |3 |4 |
numHours
1 --------> 8.1G/295.67G |4.05G/147.83G |4.05G/147.83G |4.05G/147.83G |
2 --------> NA |NA |NA |NA |
3 --------> NA |NA |NA |NA |
4 --------> NA |NA |NA |NA |
12 --------> NA |NA |NA |NA |
18 --------> NA |NA |NA |NA |
24 --------> NA |NA |NA |NA |
56 --------> NA |NA |NA |NA |

Optimal segment size

numHosts --> 1 |2 |3 |4 |
numHours
1 --------> 1.51G |1.51G |1.51G |1.51G |
2 --------> NA |NA |NA |NA |
3 --------> NA |NA |NA |NA |
4 --------> NA |NA |NA |NA |
12 --------> NA |NA |NA |NA |
18 --------> NA |NA |NA |NA |
24 --------> NA |NA |NA |NA |
56 --------> NA |NA |NA |NA |

Consuming memory

numHosts --> 1 |2 |3 |4 |
numHours
1 --------> 8.1G |4.05G |4.05G |4.05G |
2 --------> NA |NA |NA |NA |
3 --------> NA |NA |NA |NA |
4 --------> NA |NA |NA |NA |
12 --------> NA |NA |NA |NA |
18 --------> NA |NA |NA |NA |
24 --------> NA |NA |NA |NA |
56 --------> NA |NA |NA |NA |

Total number of segments queried per host (for all partitions)
numHosts --> 1 |2 |3 |4 |
numHours
1 --------> 2 |1 |1 |1 |
2 --------> NA |NA |NA |NA |
3 --------> NA |NA |NA |NA |
4 --------> NA |NA |NA |NA |
12 --------> NA |NA |NA |NA |
18 --------> NA |NA |NA |NA |
24 --------> NA |NA |NA |NA |
56 --------> NA |NA |NA |NA |
m
I think the problem is that you have 'retentionHours' set to
1
, so the command exits as soon as it's returned memory estimates for the first hour. Can you try change that value to 24:
Copy code
RealtimeProvisioningHelper -tableConfigFile <tableConfig> -numPartitions 1 -pushFrequency null -numHosts 1,2,3,4 -numHours 1,2,3,4,56,12,18,24 -sampleCompletedSegmentDir <path-to-segment> -ingestionRate 1000 -maxUsableHostMemory 5120G -retentionHours 24
m
thanks I will try. It’s odd though - shouldn’t retentionHours only be counted after numHours?
m
it builds up a 2D array of numHours x numHosts (containing all NA values to start with) and then iterates over that double loop, but then exits on this line: https://github.com/apache/pinot/blob/master/pinot-controller/src/main/java/org/apa[…]ntroller/recommender/realtime/provisioning/MemoryEstimator.java
m
Hmmm how retentionHours is counted in Pinot? From the time a segment is created or completed?
m
it seems to be based on
segment.end.time
which sounds more like it's the time the segment was completed rather than created. We can confirm with @Mayank though
m
Yeah that’s what I remember too. The definition of it seems to be different in the memory estimator tool
m
Retention is based on max time column value within segment. Also retention job kicks in periodically
m
how does it work with the RealtimeProvisioningHelper tool? It seems to read one retention value from the table config and another that's passed as an argument to the command
m
I’ve found the definition of retentionHours in this helper:
retentionHours
: This argument should specify how many hours of data will typically be queried on your table. It is assumed that these are the most recent hours. If
pushFrequency
is specified, then it is assumed that the older data will be served by the offline table, and the value is derived automatically. For example, if
pushFrequency
is
daily
, this value defaults to
72
. If
hourly
, then
24
. If
weekly
, then
8d
. If
monthly
, then
32d
. If neither
pushFrequency
nor
retentionHours
is specified, then this value is assumed to be the retention time of the realtime table (e.g. if the table is retained for 6 months, then it is assumed that most queries will retrieve all six months of data). As an example, if you have a realtime only table with a 21 day retention, and expect that 90% of your queries will be for the most recent 3 days, you can specify a
retentionHours
value of 72. This will help you configure a system that performs much better for most of your queries while taking a performance hit for those that occasionally query older data.
its meaning is different and confusing but okay 🙂
s
@Map Maybe its name is a bit confusing, but basically retentionHours parameter indicates for how many hours the realtime segments will be "actively" used for querying. In case of realtime only tables, the realtime segments will be active for the whole duration of their lifespan which is the table retention time. For hybrid tables, if there are offline segments overlapping with some realtime segments, the offline segments will be used for executing the queries. So these realtime segments won't be actively used for query execution, but they'll be around and won't get automatically deleted until they're older than the specified table retention time.
m
thanks for the further explanations. I would vote for something like “queryable hours” to avoid confusion 😄
s
@Sajjad Moradi if u can fix the doc to explain it better that will be nice. thanks.
s
Will do
m
probably makes sense to rename it to
queryableHours