<https://docs.pinot.apache.org/basics/indexing/for...
# troubleshooting
x
https://docs.pinot.apache.org/basics/indexing/forward-index if i want to sort a column, should i sort within each partition (parquet file) or have it globally sorted across all partitions?
m
are you batch loading the data or loading it realtime (e.g. kafka)?
x
batch loading
sorry can i confirm my original question, should it be globally sorted across all segments (probably not for the whole table, but within a single day's worth of segments) or within segments is fine?
m
I thought across all segments (like how you did) but let's check with @Mayank
n
Each file becomes a segment. Hence, every file needs to be sorted
m
but do they need to be sorted across the files, or only within each file?
n
Only within each
m
ok cool, lemme update the docs to explain that. And does the offline ingestion job auto select a sorted column?
The column which is sorted in the data and hence will have a sorted index. This does not need to be specified for the offline table, as the segment generation job will automatically detect the sorted column in the data and create a sorted index for it.
n
if something is sorted in the input file used by the offline job, then the sort order will be maintained (it will be marked as “isSorted”). But for it to be used in queries, you have to configure it in the tableIndexConfig
m
Query planning happens for each segment individually at the lowest level. So if the data is sorted on a column within segment scope, query execution is able to take advantage of it @xtrntr