Hi druids! My team is coming across a data integri...
# dev
m
Hi druids! My team is coming across a data integrity issue using Druid. We've got Kafka as input and the same topics serving Druid and Snowflake. In Snowflake we get full granularity of the data and in Druid we get it rolled up. The problem is that Druid is creating duplicates and rolling them up, leading to metrics doubling their expected values. This happens to a small percentage of the total ingested data but it's having a negative effect. Did anyone else come across this issue before? How can this be mitigated? Is this a "best-effort rollup" limitation? Essentially the column publisher revenue gbp in Druid should be 0.0045 instead of 0.009 based on the rows ingested. You can see that the column count shows the number of rolled-up rows. Thanks
g
i am not aware of anything in druid that can generate duplicates out of thin air. our kafka ingestion is exactly-once by design, so it would definitely be a bug if a duplicate ever got created by druid. is it possible the upstream kafka topic has its own duplicates, which are being de-duped by the snowflake ingest, but not by the druid ingest? (if duplicates exist upstream, druid will not de-duplicate them)
m
did you check if service side / any component retrying something and generates duplicate on kafka ?