Hi, wanted to open a thread on what folks think on...
# dev
j
Hi, wanted to open a thread on what folks think on storing commit metadata for multiple supervisors in the same datasource (e.g the operation here). Currently, Druid stores commit metadata 1:1 with datasource. This update is done either in a shared tx with segment publishing, or in an isolated commit. From what I can see, implementors of
DataSourceMetadata
are solely supervisor-based (either materialized view or seekable stream).
ObjectMetadata
seems to only be used in tests. The way I see it there are ≥ 2 options: • Commit a datasource metadata row per supervisor (likely the easiest, but will take some re-workings on the
SegmentTransactionalInsertAction
API and others, who assume these rows are keyed by
datasource
) – I'm currently doing this and it seems to work fine. • Commit a single row per datasource, storing partitions per supervisor ID and doing merges in the
plus
minus
methods ◦ Something like the payload being: ▪︎ map[supervisor_id] =
SeekableStreamSequenceNumbers
◦ This might suffer from write contention since N supervisors * M tasks per supervisor will be attempting to write new updates in the commit payload to this row in the DB.
g
what's your use case for having multiple supervisors per datasource?
j
Querying seems to work fine (results returned from both sets of peons) and with solution 1 from above it works
I just don't know if the metadata updates mentioned would pass review (or I'm missing some other use-case)
g
👀 1
👍 1
IMO, for metadata, solution 1 is best. It may require a database migration(?) but we might need one anyway for identifying supervisors in a different way
j
Yeah. I'm not sure what pains will come with a DB migration. I was thinking we could avoid the schema change by just populating
dataSource
column with
supervisorId
instead (which would be the datasource name for backwards compatibility anyways). I'm not sure if there's every a need to query all supervisor
DatasourceMetadata
for a given datasource (all metadata operations are currently local to a single supervisor). That being said, I could totally see someone wanting to say, reset all supervisors' metadata offsets for a specific datasource (in which case having the single column keyed by supervisor ID would make that tough).
i.e the current schema:
Copy code
|dataSource|created_date|commit_metadata_payload|commit_metadata_sha1
vs
Copy code
|dataSource|supervisorId|created_date|commit_metadata_payload|commit_metadata_sha1
the latter would give a bit more redundancy in storing the mapping of datasource -> list of supervisor IDs, since that mapping would really only be persisted in the supervisors table itself otherwise.
But either would work in the current state of things
Cc @kfaraz @Maytas Monsereenusorn
• PR ready for a review: https://github.com/apache/druid/pull/18082 • Left some notes in the TODO section with some qs
👍 1