Hi, I am creating a custom druid metric column whi...
# dev
b
Hi, I am creating a custom druid metric column which will store a custom object and serialize it to a byte []. Size estimation is ~1KB. If within my ingestion spec i specify the metricCompression = lz4 (the default one), will the data in this particular custom metric column also get compressed? The documentation on this is a bit unclear because it says "Compression format for primitive type metric columns" Looking at some of the code it doesn't seem like there would be any compression https://github.com/apache/druid/blob/master/processing/src/main/java/org/apache/druid/segment/IndexMergerV9.java#L711-L716
c
you need to provide the serializer/deserializer yourself to get compression
there are a handful of things you could use to achieve this without writing the whole thing
json columns use a thing called
CompressedVariableSizedBlobColumnSerializer
and
CompressedVariableSizedBlobColumnSupplier
which allows you to write values into compressed blocks similar to how those metrics columns do, see https://github.com/apache/druid/blob/master/processing/src/main/java/org/apache/druid/segment/nested/NestedDataColumnSerializer.java and https://github.com/apache/druid/blob/master/processing/src/main/java/org/apache/druid/segment/nested/NestedDataColumnSupplier.java though they aren’t really driven through
ComplexMetricSerde
directly anymore, the older versions were https://github.com/apache/druid/blob/master/processing/src/main/java/org/apache/druid/segment/nested/NestedDataComplexTypeSerde.java
🙌 1
if your values are fixed width, you could also use the stuff that
CompressedVariableSizedBlobColumnSerializer
and
CompressedVariableSizedBlobColumnSupplier
are built on
there is also some different stuff used by things like the compressed bigdecimal extension, which i personally find a bit more complicated to follow, but is also used by some things, can find it in here https://github.com/apache/druid/tree/master/processing/src/main/java/org/apache/druid/segment/serde/cell
s
🙌 1
b
Thanks a lot for all the inputs!! Really appreciate the guidance here. Since we will need to store the timestamp column along with this byte[] and then do custom aggregations/post aggregations etc. based on that timestamp, we have decided to implement our own extension. We will serialize/deserialize the actual byte[] within our app layer itself (probably can be done at druid also but that's besides the point here) because we don't need to open/index it within druid store.