Hey, I’m thinking of making `ResultFormat` open fo...
# dev
e
Hey, I’m thinking of making
ResultFormat
open for extension, it will allow to create arbitrary
ResultFormat.Writer
for exporting formats, for example here instead of
CSV
a user could specify
PARQUET
,
XSL
(and so on) by loading a relative Druid extension.
Copy code
INSERT INTO
 EXTERN(
  S3(bucket => 'your_bucket', prefix => 'prefix/to/files')
 )
AS CSV
SELECT
 <column>
FROM <table>
The problem is
ResultFormat
is enum right now. I propose instead of using
values()
in
ResultFormat.fromString()
- there should be a
Guice MapBinder
that can collect
ResultFormat.Writer
implementations from different extensions (as well as from core lib). I wonder if anyone working on that right now or has some insight/plans?
🙌 2
a
I think that would be a pretty useful feature. However, I'm not sure if continuing to use ResultFormat would be the best approach. ResultFormat is used in a couple of other places, and just adding a new extension might not mean that we want to allow its usage everywhere else. For instance, we might not want regular queries results to be returned in parquet. Since this would be a slightly different usecase, it might make more sense to shift from ResultFormat to a new writer class to handle these formats with export queries.
e
As I understand, we can pass the writer here instead of ResultFormat https://github.com/gorillio/rill-druid/blob/93eeb05eaf5e9a90a475d6d82b6eeefef6b66a[…]y/src/main/java/org/apache/druid/msq/sql/MSQTaskQueryMaker.java So instead of refactoring ResultFormat we can search a writer implementation in ResultFormat.values() and in extensions
a
Would that not also allow those formats to be used anywhere, including the normal query API? ResultFormat.Writer is a bit old and has some quirks, like CSV writer adding a new line at the end of the file, which might be better as a configurable option when exporting, so in the future, it might be good to use something other than ResultFormat writer. I'm not sure about the answer here, but this might not be a serious problem.
e
hmm, the current problem - turned out that
storageConnector
opens a file for every incoming frame, while Parquet API doesn’t allow to append to a file. So possible options: • either rewriting all storage connector interface and implementations to give out a file handle instead • saving a new file for each frame (inefficient) I guess there can be some other approach I’m not yet aware
a
Ah, I was not aware Parquet had a restriction like that, that is the opposite of how runIncrementally would function. I am also not able to think of a better solution than the first one you mentioned currently. Would keeping the stream open in between runIncrementally calls be an option? It might not be the best approach though
e
I currently have a working implementation that keeps the stream open until the processor is cleaned up. That solves the problem. I’m going to test it on a larger cluster soon