Hello I am trying to migrate from Druid 24 to Drui...
# general
a
Hello I am trying to migrate from Druid 24 to Druid 29 to leverage MSQ for ingestion. Our pipeline are mainly batch pipelines which ingest yesterday's data by running the index_parallel jobs next day. In each of our job, we leverage Tuple sketch, as part of using MSQ for the same last year, we contributed to add sql function for the same. There is a large performance gap between MSQ JOB AND Index parallel job. For Index parallel it takes 40-60 mins to ingest 1 day of data around (100gb) - MM is used. But when I use MSQ it took 4 hours to ingest only 1hour of data, index parallel takes 7- 10 mins to do the same task. MSQ I am running on Indexers as compared to MM. Can I get some help where should I debug this.
k
How many Tasks are you running with MSQ.
I have a hunch that you are using just 1 worker task
a
I tried with 30/60 tasks
No, not 1task. It works very fast when I use Theta Sketch instead of Tuple Sketch
But then I don't have this problem in
index_parallel
k
Very surprising. Could you please share the controller task report
a
I will fetch it and share, will it be available in console or I need to check my my Infra team
Or is there an api to get task report.
k
It would be available on druid console
a
Ok
k
Go to tasks-> search for the task id, go to the report section.
👍 1
a
What are we looking in the logs.
l
it’d be good to see which stage consumes the most time. Are you ingesting a large amount of segments via INSERT?
a
Per day data is 100 gb, roughly 4-5gb per hour, 1000 partitions files
l
is it INSERT or REPLACE, and what are the number of segments you are ingesting?
a
INSERT
Final segment size is 5Mil Rows, so 2 segments per hour.
l
hmm, task reports/logs would help in this case. If there are a large number of segments, then INSERT has bad perf, but that doesn’t seem to be a problem in your case
a
I will get the logs for this sooner.