Hi all. I have a question about Druid data modelin...
# general
a
Hi all. I have a question about Druid data modeling for fast query execution. I have a datasource with fields 'UserId' and 'EventType'. Generally, when the value of 'EventType' is 'eventType1', the columns from 'eventType1Fields' are used; when it's 'eventType2', 'eventType2Fields' are utilized, and so forth for the 8 different event types. Many other fields remain the same across these events. The search is performed based on 'UserId' for all event types. What is the best approach for schema modeling? 1. Different datasources for each event type, with searches conducted in parallel across all datasources by 'UserId'. 2. One datasource containing all fields, with searches conducted by 'UserId' within that single datasource. The primary goal is to achieve the fastest search results for 'UserId' across all event types, with the number of events reaching billions per day.
k
Have one data source with range partitioning on user id.
Maybe event type as well if you have a lot of q's which need event type specifically
👍 1
a
Thank you. Your idea with different partitions looks really good I will check it in practice
b
Although Druid is very good at handling sparse datasets, if the number of distinct fields between events is very high, or if it is something that regularly changes, you may want to consider using a JSON field to hold the event-specific dimensions.