This message was deleted.
# dev
s
This message was deleted.
c
assuming this is similar to kafka consumer groups, this thread https://apachedruidworkspace.slack.com/archives/C0303FDCZEZ/p1663601546431449 has some details of why kafka ingestion uses low level stream/partition positions
other abstractions are totally possible to define, i suggest updating https://github.com/apache/druid/issues/13724 to elaborate on what the abstraction is if you’ve got something in mind to take the discussion there if the low level model of the ‘seekable stream’ abstraction isn’t what you plan to use
or it could probably live entirely in the extension if the model doesn’t really apply to other types of streams, pretty much everything is extendable with some varying amount of effort
j
Thank you for the information @Clint Wylie. We took this into consideration and we are looking to see how we can still use the ReaderGroup abstraction level for the plugin and have it live entirely within the extension. Currently we were thinking about creating two ReaderGroups (two partitions) so that Druid is able to use its own parallelism login within the Seekable stream supervisor. Initially we were just going to use one readergroup to manage all readers so we are still investigating
g
If your ReaderGroups are something like Kafka consumer groups, you may find that the SeekableStream abstraction is not a great fit. It really wants to think of things in terms of partitions and offsets, and it really wants to handle its own checkpointing. I bet you can come up with a simpler design by creating a new Supervisor and Task type
j
@Gian Merlino Sorry for the late reply, I am just seeing this now. Yes the readergroups abstraction are similar to consumer groups abstraction. We are currently trying our best to ingest the seekablestreamsupervisor code and see how we can make it work with readergroups but it is a little tough. Would creating a new supervisor and task type compatible with a readergroup/consumergroup abstraction be more straightforward?
Because the Pravega client API isn't the best for accessing partitions and offsets at the lower level, readergroups are used to read streams and checkpoint at a higher level
d
@Gian Merlino -- Yes, Pravega reader groups are similar to Kafka consumer groups, but slightly more powerful, supporting checkpoints across the whole reader group (like commitSync or commitAsync in Kafka, but committing offsets to a checkpoint across all readers in the group). We are seeing that rewriting a new Supervisor would greatly simplify Druid's stream ingestion and autoscaling code, given the more complete abstractions provided by Pravega. We are new to Druid, so studying the SeekableStreamSupervisor ecosystem is how we're making our way through understanding the Supervisor/Task framework. We are attempting to write an initial integration using a single reader in a reader group, thus it comes out as a single-task single-virtual-partition ingestion. Being able to discuss issues with someone who knows the Supervisor/Task framework really well would likely accelerate our velocity.