:wave: Hi everyone! I am in the process of design...
# general
c
👋 Hi everyone! I am in the process of designing and implementing a stock management system for a government(you can imagine millions of items to keep track of) . Came across Druid and it looks like exactly what I am looking for. Looking through the examples and docs it seems that I will be able to do what I want. Just to confirm, with Druid I can build dynamically a stock item dictionary as the items come in, we can dynamically get aggregates like current count of item type in stock, item events like, delivered, tagged, ordered, checked out etc. I see that Druid has this concept of Lookups which I think would serve for my “Item Dictionary”. However, the documentation states that it is not recommended. Any other suggestion on how to implement this item dictionary? One other requirement I have: some of the events in the system, like receiving goods will produce a Goods Received Note (GRN) which is a pdf document that is digitally signed (legally required). Is there an extension to Druid to handle these documents? Normally we would store these documents in mongodb with sharding. Any advice and pointers would be really appreciated, thanks
j
Hi Chris, Afaik Druid does not support document or binary object storage by default ... and I haven't seen any extensions in that area. It is plausible that you could convert the binary object to base64 and store it in string format, maybe others here will have examples ... Druid was designed for high speed ingestion and analytics of fairly structured data ... so if you don't need to search the document content directly then one strategy to use the analytic query capabilities of Druid might be to store the documents in a file system somewhere (e.g. S3) and hold the file pointers in Druid to use at query time as needed. As far as lookups, they are useful to hold slowly changing dimensional data that can be updated periodically. You can update them in batch jobs or through streaming ingestion. Please take a look at this doc page and let us know if you have any additional questions. https://druid.apache.org/docs/latest/querying/lookups Thanks. John
c
Thanks, John, for your swift response. As far as the docs, I think will keep them in mongo for now. As far as the Lookup, let me give you a scenario; A Stock item is delivered by the supplier, it gets put into the stock, now during the lifecycle of this item not many properties will change, just state, who is using it etc. The basic information like name, barcode etc will not change. My thought was to make it a Lookup. Now all the events that change the state of the item will be a timeseries pointing to the lookup. That will give me the history of each item. For statistical purpose I thought of a table with current item counts, like laptops 100 at warehouse X. This will help me also fulfill orders. This simple model I have implemented in Mongo now, but I would like to use Druid for this. I am coming from the relational, nosql and graph world …. Need a bit help. Your help is greatly appreciated
j
Hi Chris, Check out Kafka-based Lookups ... they can update key/value pairs directly from a stream, so you will have the latest info at any given moment. Since you have multiple attributes that will be changing, there is a new enhanced version of Lookups under development, called "Dimension tables", which will operate like an upsert to an in-memory table ... same Kafka ingestion stream, but this time an entire record of data (PK + attributes) can be ingested and will either insert or update the existing data. I think this is coming to OS Druid (vs Imply only) but not sure ... if not, you can always implement multiple Lookups, one for each changing attribute. Alternately, if you just want to ingest the updated records sequentially, you can append them into a table with timestamp, and (since Druid is a timeseries database) use a Latest() function to always pull the values from the latest record. This gives you the ability to pull the latest info or the history all from one table.
c
Thanks for your feedback , will proceed and yell if I have some more questions
👍 1
Hi John, another question : Since my schemas in the app are always changing, how can I handle these changes on the Druid side ? I see that Druid has support for json columns. We just have one table with a json column ? I think thats not the right approach .... your thoughts
j
You can use Nested JSON columns if your data changes from row to row. Some people use this for multi-tenant tables, where based on the tenant ID the attribute set might be completely different ... so tuck them all under a nested column (e.g. call it "attributes" or "properties") and then schema looks much cleaner. If you are changing the schema once in a while and just want to add a new top level column to the datasource, you can just change the ingestion spec and ingest it. Querying old data will just show that column as empty/NULL, and the new records will show the values you ingested. Druid does not maintain a centralized datasource schema, that schema is determined dynamically by union'ing together the schemas of all of the segments in that datasource.
c
Tks