This message was deleted.
# general
s
This message was deleted.
s
You can query directly from Deep Storage using MSQ starting with release 28. You can see an example of this in the learn-druid project. It has a notebook called full-timeline-queries that walks through how to set it up and how to query it. In the current version of the learn-druid notebooks it is called
03-query/14-full-timeline-queries.ipynb
. Give it a shot and let me know if you have any questions. You'll need docker desktop on your computer.
j
i see, so the MSQ is the only way i can do this. Right?
s
Correct and you use retention rules to define which time frame of segments are not cached in the historicals.
j
After i specify the retention rules, there will be some segments in deep-storage because out of the retention rules and some in local of historical because still within the retention. In this case, can i get combination result from both data within one query, or it has to be separate query?
s
You need to specify the retention rule for the segments that are meant to stay in deep storage, but you do not include any replicants. This tells the coordinator to not mark these segments for deletion until the period completes and the segments fall to a DropForever rule that marks them for deletion. the example in the notebook uses this:
Copy code
[
    {"type":"loadByPeriod", "period":"P1M", "tieredReplicants": { "_default_tier": 1} },
    {"type":"loadByPeriod", "period":"P3M", "tieredReplicants": { }, "useDefaultTierForNull": False }, 
    {"type":"dropForever" }
]
in this example, the segments for the latest month are kept in "_default_tier" and segments for data between 1 and 3 months old will stay in Deep Storage. After they age out of the 3M limit they are marked for deletion, with the dropForever rules.
j
thanks, in this case will a single query with the most recent 3 month interval return all the data from both deep storage and historical local disk?
s
There are different APIs for cached vs deep storage queries. This notebook explains the use of each one and which portion of the data they will see. To answer your question. You can: • use /druid/v2/sql endpoint for cached results, 1 month in the above example • use /druid/v2/sql/statements for querying from deep storage, 3 months in the above example.
Today using the /druid/v2/sql/statements API means you read all the segments from deep storage and hence have access to the 3 months of data. There is work being done that will combine the historicals cached based results with the non-cached portion from deep storage, so stay tuned for that.
j
Got it, thanks every much!!!