This message was deleted.
# general
s
This message was deleted.
b
Would you please share the link, where you see the description.
j
Hi Quoc, you may have older or erronrous information. Druid has two query engines now ... the "native" engine uses a fanout architecture and is limited to broadcast joins, which have limits in terms of secondary/intermediate table sizes ... But the new Multi-Stage-Query engine uses a peer-to-peer data shuffling architecture with sort-merge join that effectively removes those limits, providing a pretty strong table-join architecture in a shared-nothing environment. I have spent the better part of my four decade career in relational databases, and I have to say this is probably one of most advanced new architectures I've seen in a long time. Please let us know if you have more detailed questions on Druid product capabilities, and if you try it it out let us know how it goes. Thanks. John
a
Hello Jhon, So if we do ingest / load the data to datasource using MSQ, and after load data if we do SQL JOIN queries from client application to join multiple datasources & Lookup table to retrieve / report will work effectively ?
Jhon, Am using Apache Druid & Apache Kafka in our application. We deployed in Single-server (& Druid services are started using Auto config) My druid server has - 8 cpu & 64 gb ram As of now we using only Kafka streaming to ingest the data to druid datasource. Now my requirement is, to load the existing-old-bulk data using MSQ SQL queries (INSERT/REPLACE) to the existing datasource which is ingesting data using streaming index spec i.e. using Kafka stream) Is it possible to ingest / load old data to the datasource in-parallel while fresh/new data ingesting via kafka stream index spec (supervisor) ?
So in above case since am using Kafka stream to ingest, will it perform good if I use SQL JOINS ?
m
@Quoc Khanh @John Kowtko @Ashok Kumar Ragupathi I know that we implemented unions and other workarounds for joins at my last company. I wish we had experimented more with joins but there were several other factors there, including self-hosting. I would recommend Quoc and Ashok understand specifics about the Druid datasources being joined as that will make a huge difference. Often, doing joins at ETL is preferable. The business users will assert that they need "data munging" capabilities at Druid query time but often this requirement is better met with organized SQL in the data pipeline.
j
For streaming ingestion you can join in descriptive data via Lookups ... however this is single key/value pair only, so if you have several attributes to join in you will need several lookups. Lookups can also be used at query time, so your preference on whether the lookup dimension is changing or not to determine if you want to use the Lookup at ingestion or query time. Streaming ingestion is insert only (no overwrites) For batch ingestion you can use MSQ and join multiple sources together, including joins to existing datasources within Druid. This is SQL based so it is more like a regular SQL select statement that inserts its data into a destination datasource. This can be used in either insert or overwrite mode. If you want to stream in latest data while backfilling old data with batch ingestion jobs, that will work so long as you are working on different time intervals ... if you try to do both in the same time interval then by default the streaming ingestion will have priority and the batch ingestion job will abort.
a
Thank you John and Mike
m
Ya, @John Kowtko there were a fair number of people at my last company backfilling and upserting data using batch ingestion, while inserting new data via Kafka ingestion. Good to know that the batch ingestion will fail if there is contention.
j
Fyi the Lock priority will determine which conflicting job aborts ... the defaults are documented, but I believe they can be changed on any of the jobs: https://druid.apache.org/docs/latest/ingestion/tasks#lock-priority
👍 2