AR
03/15/2024, 2:45 AMLaksh Singla
03/15/2024, 10:16 AMsortMerge join. It doesn’t have the memory limit that the broadcast join suffers from (though under special circumstances it can still OOM, if the join keys are skewed, for a well distributed dataset it shouldn’t happen). This would be the best solution, if it works for your use case.
2. Yes, the limit applies to amount of data that can be materialized, which is same in both the cases. In the latest versions, MSQ doesn’t rebuild the broadcast tables, and therefore is more optimal in disk space usage.
3. You can try using the iceberg extensions, if you want to ingest the data using Spark. To export the data from the segment files to a location, you can use the newly introduced (experimentai) EXPORT functionality, however - it is experimental, and it is in 29.0.0, (hence requiring an update). In 24.0.2, I am not sure if there’s an easy way to export, you can run the SELECT query and export the data into a CSV.
4. No suggestions regarding this option
IMO, 1 would be the most ideal solution. It requires upgrading Druid version, though all the other answers also require upgrading.AR
03/15/2024, 4:20 PMsortMerge join is available in this version.
Can I upgrade directly from 24.0.2 - 27.0.0?
I looked at the release notes but didn't find any obvious blockers.
We are still on Java 8.
Will v27.0.0 run on Java 8 or is a higher version needed for it?
Regards,
AR.Laksh Singla
03/15/2024, 4:29 PMWill v27.0.0 run on Java 8 or is a higher version needed for it?Afaict all versions till now are supported on Java 8. Regarding compatibility, it’s difficult to make a blanket statement.
Laksh Singla
03/15/2024, 4:30 PMAR
03/15/2024, 5:06 PMLaksh Singla
03/15/2024, 6:31 PMwe don’t have a higher versionIt should be on Druid’s website. In either case, LMK how it goes!
Karan Kumar
03/18/2024, 6:45 AM