Hey Team, Have ingested the same data using `index...
# general
a
Hey Team, Have ingested the same data using
index-parallel
&
msqe
in two different tables. The Total data size in the table which is ingested through index-parallel is ~ 190mb although the total data size in the table which is i*ngested using the msqe is ~166 mb.* Do anyone knows why the data size is different in both the table, despite count of rows are same.
j
Hi ashish, This is likely due to the way the data is distributed across the segments, and how many columnar dictionaries have to be created, and the cardinality of each. Example -- If two segments have the same 1000 distinct values in a given column, then each segment will have a 1000 entry dictionary for that column ... so two identical dictionaries across two segments. However if you create larger segment sizes and the data from both of those segments ends up in one larger segment, then there would be only one 1000 entry dictionary. Space savings of 50% for the dictionary of that particular column in this case. So due to the optimizations made in segment storage structure, there are definitely space savings that can be achieved through economies of scale ... which it looks like you are seeing. Fyi MSQ is a very new engine and is known to do a better job of organizing and processing data ... so this doesn't surprise me 🙂
a
Thanks John 👍
Hi John, Does it impact on query performance also?
j
depending on the type of query and what part of the query execution we are looking at, the size and # rows of segment may not matter so much as the data organization within it. For example, filtering happens at both the segment pruning level and the row matching within the segment. For row matching within the segment the type and order of filter expressions might help. And if this is a group by query, then range partitioning could help to provide perfect rollup capabilities.