This message was deleted.
# troubleshooting
s
This message was deleted.
s
Yes. I believe that is how it works. The parquet file is downloaded to the Druid Indexer/MM and then parsed for the columns you are using. Subsequent processing (e.g.. rollup, sub-partitioning) will only use the columns that you have selected.
d
Is there a way to know the size of data druid processes after the columns are parsed?
s
I'm not sure, perhaps by looking at the temporary files it creates during the load. Just a thought.
g
There's some trickiness here with how you define "size of data"
Data takes up many different forms as it is processed, with various different kinds of compression, so we have to be really specific when we try to define "size"
Most of the time looking at the starting size (size of parquet files) or end size (size of druid segments) is enough
For intermediate sizes, if you use SQL-based ingestion, the report includes various counters showing the number of rows and size in bytes between each stage of the job. Note that the size in bytes is reported prior to compression, so it's not going to match how much is actually sent between servers or stored on disk (as those are compressed). But it's still useful for comparisons between jobs
When you run a SQL INSERT or REPLACE in the web console, the counters you see come from that report