This message was deleted.
# general
s
This message was deleted.
m
Hey @Maytas Monsereenusorn if you search this Slack room you'll see this discussion, that might help: https://apachedruidworkspace.slack.com/archives/C0303FDCZEZ/p1705051436852789
g
starrocks has done a comparison, although i haven't reviewed their work to see how "fair" it is my general thought is that when it comes to these modern columnar DBs, they all operate on similar principles, and so the one that will be fastest / lowest cost for a given workload would generally vary by workload based on what specific optimizations and features have been implemented in each particular DB unfortunately i don't know starrocks well enough to say which workloads it would excel at and which ones druid would excel at. i am just saying i expect you'd find that it's a mixed bag if you looked at a wide variety of workloads. (this is generally what i've seen when comparing druid with other columnar DBs)
also, the one that will be fastest / lowest cost for a given workload can potentially change over time, as each of these modern columnar DBs is constantly being improved and is therefore a moving target personally i like druid's general design / architecture best so i spend my effort making druid better 🙂
👍 3
🙏 1
a
One particular thing I have noted is some of the databases have a "_take everything I have_" design that works very well for a single-user, one-query-at-a-time workload but often falls flat when there are multiple concurrent users in the system and you do not want one single large query to hog all the resources. This former approach makes sense for databases like DuckDB that are meant to be run on your laptop. It's just one user running queries on their computer unless the kids in the family know how to run SQL queries 🙂 , But that's not a typical workload in production. To get a sense of real-world performance, I would also recommend a concurrent workload that runs for a while. In addition, in your production deployment, you will want to reserve ingestion capacity isolated from query load, something that Druid assumes even on a single node. That nuance is often missed by the person running the benchmark. All that said, do let us know what you find and if there are gaps. We are constantly looking to improve. In the long run, the databases that win are the ones that are receptive to community feedback. Everything else is secondary (not that we are not good at those secondary things ). We have a lot of cool stuff coming in 2024, and any feedback from the community will help us point that roadmap in the right direction.
🙌 1