We do, at Verizon we ingest data on close to 140M devices in the US, north of 280 Bil+ messages per day in real time. We use druid to model customer behavioral trends in real time (data usage patterns, RF signaling performance etc...)
Challenges:
We ended up building a native connector into Apache Pulsar for streaming persistent and extended the druid streaming api to include non-persistent topics (Indefinite streams) as well for stateless topics.
We also ended up building out our own concurrent intermediate persist implementation for real time streaming (similar to what Rivian did) to speed up ingestion throughput for large multi-dimension topics (6K+ columns we handle today). We noticed as you scale up the number of dimensions in druid, segment generation becomes exponentially slow, implementing a concurrent implementation for intermediary persists largely solved our issue and allowed us to better utilize the H/W for our middlemanagers.