Hey everyone. Quick question; When querying for a ...
# troubleshooting
l
Hey everyone. Quick question; When querying for a specific time range in Pinot, is it more efficient to use the primary time column defined in the segmentsConfig, or is it equivalent to using any other time column? From the docs it seems to indicate that the primary time column is only used for retention purposes, meaning that querying for another timestamp should be fine too. In my case, I am creating a copy of the primary timestamp, reducing the granularity of it, and calling it
daysSinceEpoch
, as I want to query for entities within certain days.
Copy code
"ingestionConfig": {
    "transformConfigs": [
      {
        "columnName": "daysSinceEpoch",
        "transformFunction": "toEpochDays(documentTimestamp)"
      }
    ],
...
Additionally, for the RealtimeToOfflineSegmentsTask, I am using this value for deduplication purposes. In the schema:
Copy code
"primaryKeyColumns": ["customerId", "machineId", "daysSinceEpoch"]
...
This is because for each event, I only want to keep the latest in a day. Here’s the RealtimeToOfflineSegmentsTask
Copy code
"RealtimeToOfflineSegmentsTask": {
        "bucketTimePeriod": "1d",
        "bufferTimePeriod": "2d",
        "mergeType": "dedup",
        "maxNumRecordsPerSegment": 10000000,
        "roundBucketTimePeriod": "1h"
      }
In the realtime table, I am also filtering out any events older than 14 days (Where documentTimestamp is the actual primary timeColumnName)
Copy code
"filterConfig": {
  "filterFunction": "Groovy({documentTimestamp < (new Date() - 14).getTime()}, documentTimestamp)"
},
Does that make sense?
n
you can use any time column. you’re right that primary time column is mainly used for things like retention
l
That’s great, thank you @Neha Pawar 🙂 I had assumed as much
n
you cannot really define your own primary keys for realtimeToOfflineSegments task dedup mode. It will dedup only if the entire row is same
l
Oh, it doesn’t use the primary key defined in the schema?
n
the primary Key columns field you see is for the upsertts feature. it doesnt take any effect for realtimeToOffline
l
aahh
Is there any reason why?
n
dedup is a relatively new feature in realtimeToOffline task. This version only does the full row dedup. We’d need to add a lot more config and code, to support the next level of smarter dedup
regarding filtering out events greater than 14d, you can just set table retention to 14d? any reason you’re using the filter function instead?
l
I sometimes get old events coming in through kafka which I don’t want to include in my segments
m
@Lars-Kristian Svenøy You can filter those rows at ingestion time: https://docs.pinot.apache.org/developers/advanced/ingestion-level-transformations#filtering
l
Yep that’s what I’m doing @Mayank 👍
👍 1