<@U030C8H59T8> <@U0344FW86DD> wanted to get your t...
# dev
j
@Clint Wylie @Gian Merlino wanted to get your thoughts on a few PRs: • https://github.com/apache/druid/pull/19431 • https://github.com/apache/druid/pull/19413
For the latter in particular, this is trying to cover some edge cases that clone historicals currently misses: 1. Handoff gap – I'm not sure there is a way to correctly "hand-off" segments to ensure no temporary gap in timeline with cloned historicals. Even if, say, you atomically "swapped" the clone list (e.g. A clones B to B clones A), segment handoffs would still potentially have some lag. You would also need to update the default query clone behavior to include/exclude clones. 2. We cannot force all traffic to route to new router/broker/historicals AND ensure brokers query new historicals atomically. In other words, we cannot guarantee queries strictly v1/v2 router/broker/historical (it might be a mix). I'm trying to add support for ensuring this atomic cutover.
g
With 19413 are you intending to use that alongside the clone feature? Or by itself without cloning?
About the issues, with (1) what is the issue exactly? I don't quite follow
j
Right now, without cloning
g
Are you talking about the realtime -> historical handoff or something else?
j
About the issues, with (1) what is the issue exactly? I don't quite follow
realtime <--> historical handoff. Since query routing is not atomically tied to the set of servers being cloned, I'm pretty sure we cannot guarantee zero gaps in timeline for cutover queries (e.g. if you want to have new version of routers/brokers querying new historicals).
Basically, if I want to: • Ensure "new" Druid versions only query other "new" Druid versions • Cutover between historical replicas (e.g. hot tier v1 to hot tier v2) without exposing a temporary gap in timelines (queries querying the v1 historicals switching to v2 missing realtime handed-off segments due to lag between segment load duty and clone duty on coordinator, etc.)
kind of like the worker.version concept with overlord + middlemanager/indexer
g
I think it should be possible to avoid issue (1) with cloning. If you have a full tier B cloning a full tier A, and then you remove the clone settings once tier B is populated, then you should be safe to cut over queries to tier B when you've (a) removed clone settings; (b) observed loadstatus showing all segments available in that tier. I think?
For issue (2), isolating a full stack of the old version vs. new version across router-broker-historical wasn't a goal of the cloning feature
The cloning feature was mainly meant as a way of doing a blue/green update at the level of a tier of Historicals
j
For 2) yes this is a net-new feature
g
Potentially a very fast one if used in concert with turbo-loading
j
For 1), you'd need to ensure that no new segments are loaded strictly on the old tier. I think you also could potentially run into cases where the data is loaded onto an old historical but the clone() configuration was removed before the segment could be loaded onto the clone?
Since the clone is effectively a "dumb" copy, coordinator doesn't check (for example) if segment cloning has been done for handoff purposes. I suspect you'd still need to do the disable/drain/delete on old version historicals due to this potential diff
> For 1), you'd need to ensure that no new segments are loaded strictly on the old tier. This can be ensured by swapping the clone/clonee clone configs via one PUT to coordinator dynamic config, but I still think it could potentially leave a gap in the case of realtime segment handoff.
g
The way we do it involves breaking off the clone, removing load rules from the old tier, then confirming that all segments are available on the new tier, before terminating the old one
anyway I think I get what you're getting at
j
anyway I think I get what you're getting at
What do you mean? The handoff part?
g
I mean I think I get what you're looking for with the feature: seems like you're looking for a way to do a full stack blue/green deployment If you were only trying to deal with the handoff issue then I think there's probably a way you can use cloning today, or maybe some changes you can make to cloning to make it more convenient. In this case I might have tried to convince you to use cloning instead But since you additionally want the router + broker to participate in the blue/green deployment, the cloning feature isn't really the right thing to use, since it's not meant to operate at that level
j
> But since you additionally want the router + broker to participate in the blue/green deployment, the cloning feature isn't really the right thing to use, I think we can work around this limitation via service names, unless you think this would be acceptable to put in OSS?
g
what do you mean by "this" in "this would be acceptable to put in OSS"?
j
We could get around the router/broker stuff by using specific versioned service names
but to me that seems like so many moving parts (updating configs, etc.) just to ensure some isolation
in the above PR,
deploymentGroup
would be analogous to a worker.version for MMs/overlords
@Gian Merlino one other question: using the clone method, how are you folks ensuring that between the time when a new historical comes up/announces itself and the coordinator dynamic config being updated, a segment doesn't get loaded onto the new historical? Are you preemptively decommissioning the node?
@Clint Wylie any thoughts here^
@Abhishek Balaji Radhakrishnan do you folks use the clone feature?
g
the
deploymentGroup
could be a useful idea. It seems different enough from cloning that the features can both exist > are you folks ensuring that between the time when a new historical comes up/announces itself and the coordinator dynamic config being updated, a segment doesn't get loaded onto the new historical In general we're doing deployments at the level of the entire tier. Cloning is mainly being used as a way to ensure the replacement tier comes up balanced the same way as the old tier. As to segment availability, I believe we're only terminating the old tier once we're sure that the new tier is in steady state and has everything available. At this point the clone would have been deconfigured for some time
j
As to segment availability, I believe we're only terminating the old tier once we're sure that the new tier is in steady state and has everything available.
Ah – I guess I was referring to boot-time, not terminate-time. For example, booting a new historical (before it is placed as a clone in coordinator dynamic config), coordinator will see new historical and will try to load segments/queries execute on it no?
e.g. you'd want a node to effectively "boot" into cloning status to avoid queries being prematurely planned onto that node for segments
> the
deploymentGroup
could be a useful idea. It seems different enough from cloning that the features can both exist
deploymentGroup
can also be replicated with a combination of: • Historical tier aliasing • These configs (for controlling broker -> historical): https://druid.apache.org/docs/latest/operations/mixed-workloads/#restrict-broker-visibility-to-specific-tiers • And custom routing logic for brokers based on service name
g
We populate the
cloneServers
list before launching the new Historicals, using deterministic hostnames
j
makes sense, yeah we don't have that yet 🙃
Basically the long way around what I was mentioning in https://github.com/apache/druid/pull/19413 (e.g. deploymentGroup) is something like:
Copy code
1. spin up NEW ASGs
2. NEW brokers only watch/listen to NEW tiers (via druid.broker.segment.watchedTiers)
3. routers only query NEW brokers (via custom router logic)
4. update tier alias coordinator dynamic config with the new historical ASG versioned tiers
5. wait for segments to load
6. run query tee, etc. through the NEW routers/brokers, hitting NEW historicals, using the realtimeSegmentsMode=exclude to avoid hitting realtime tasks
7. enable NEW broker/router
8. disable OLD router/broker
9. update tier alias coordinator dynamic config to remove the OLD historical ASG versioned tiers
10. scale down OLD coordinator (force leadership change to NEW coordinator)
11. wait for things to stabilize
12. scale down OLD overlord (force leadership change to NEW coordinator)
13. drain tasks from OLD MMs
14. delete OLD historicals (alternatively this can be done much earlier, but putting last for completeness).
It's just using configs from a few places (IMO, it'd be nicer to abstract this into a single config, which was what that PR was trying to do).
My main goal with https://github.com/apache/druid/pull/19413 was just to avoid "leaking" deployment version (e.g. OLD/NEW build) into data/query service isolation (hot tier NEW/OLD or hot broker NEW/OLD)
LMK what you think; I suspect exposing a
druid.version
config for all Druid node types would be generally helpful on its own, not necessarily adding any extra logic initially.
a
I'm not fully caught up on this thread but some drive-by-comments:
@Abhishek Balaji Radhakrishnan do you folks use the clone feature?
We don't use this feature yet, but we plan to
LMK what you think; I suspect exposing a
druid.version
config for all Druid node types would be generally helpful on its own, not necessarily adding any extra logic initially.
The Druid servers expose a
version
and
buildRevision
extracted from build artifacts automatically w/o a config: https://druid.apache.org/docs/latest/querying/sql-metadata-tables/#servers-table (and metric dimension) - would that suffice? a new server in the "green" group would effectively have a different version and/or build revision that's different from the servers in the "blue" group for major/minor Druid upgrades. We can pause deployments to validate deployments in the new fleet of servers before resuming an upgrade to reminder of the servers etc. I was thinking it'd be helpful to also have a hash of the server properties (configSha) or so as well, but haven't thought about how we'd do this
j
@Abhishek Balaji Radhakrishnan you folks are using the operator right?
a
yes
j
What does your deployment look like there? I'm assuming it's rolling? The new repo doesn't have many docs, I was planning on exploring it to see what it offers
a
Currently only rolling deployments. Our fork of the druid-operator has diverged quite a bit from upstream...but I believe the upstream druid-operator should support B/G upgrades too
j
> but I believe the upstream druid-operator should support B/G upgrades too Hmmm – does this cover ingestion downtime?
Part of what I'm trying to solve for here is maintain B/G semantics RE query/data isolation while still getting benefits of zero-downtime ingestion (e.g. not shutting down tasks) that you get with rolling deployments
The Druid servers expose a
version
and
buildRevision
extracted from build artifacts automatically w/o a config: https://druid.apache.org/docs/latest/querying/sql-metadata-tables/#servers-table (and metric dimension) - would that suffice? a new server in the "green" group would effectively have a different version and/or build revision that's different from the servers in the "blue" group for major/minor Druid upgrades.
Yeah – I think we can make a revision into
version
(buildRevision seems a bit scary to use since you could be using same code commit but different configs, etc.)
👍 1
@Gian Merlino off the top of your head, do you think anything might break if we were to make a config that asserted a segment's load status as: "a segment is only considered loaded if at least 1 replica exists in all the tiers it has retention rules configured for." This would be different compared to the traditional definition as: "at least 1 replica exists in the timeline." While this might work in the traditional loading case, in the case of a retention-based move (hot -> cold, etc.) this might get complicated, where you could have a segment that, according to retention rules, should be in cold tier (but is currently in hot) and therefore is not loaded (according to the former definition). But according to the latter (current) definition, it is loaded (because at least 1 replica exists in the timeline).
I guess we could scope this new definition to handoff endpoint-only (e.g. the coordinator endpoint tasks hit to check loadstatus).
this would also help in cases where you have different broker/historical tiers querying the same data (but isolated based on query pattern – for example, one tier pair services lighter queries over the same data while another services heavier queries), and you want to make sure the ingestion tasks only complete once handoff indicates data is loaded on both tiers.
cc @Clint Wylie