Hi , new to using apache pinot... i want to know w...
# pinot-dev
a
Hi , new to using apache pinot... i want to know which config is actually used to flush segments to deep store , is it?
realtime.segment.flush.threshold.size
or
realtime.segment.flush.threshold.segment.size
?
why do i see
"segment.flush.threshold.size": "100000",
if i have set table json to 64MB.
Copy code
if i have set below conf in table json
"tableIndexConfig": {
    "loadMode": "MMAP",
    "streamConfigs": {
      "streamType": "kafka",
      "stream.kafka.consumer.type": "simple",
      "stream.kafka.topic.name": "test",
      "stream.kafka.decoder.class.name": "org.apache.pinot.plugin.stream.kafka.KafkaJSONMessageDecoder",
      "stream.kafka.consumer.factory.class.name": "org.apache.pinot.plugin.stream.kafka20.KafkaConsumerFactory",
      "stream.kafka.broker.list": "broker list",
      "realtime.segment.flush.threshold.rows": 0,
      "realtime.segment.flush.threshold.size": "64M",
      "realtime.segment.serverUploadToDeepStore": true,
      "stream.kafka.consumer.prop.auto.offset.reset": "smallest"
    }
f
Segment builder will learn step by step how much rows it will need to reach the 64M provided 🙂
a
By in object store i am seeing 34MB size files...
Also do you know who writes to object store, is it pinot server or controller?
I have this config in pinot-server.conf
Copy code
pinot.server.instance.segment.store.uri=s3://***/pinot/data/server
pinot.server.segment.store.uri=s3://***/pinot/data/server
I also have
Copy code
controller.data.dir=s3://***/pinot/data/controller
I am seeing
34MB
files present in
data/controller
not
data/server
.... What could be the reason?
f
First segment will be smaller following one will increase steb by step to reach you 64Mb 😉
a
why it is going under
data/controller
and not
data/server
m
Controller (deepstore) is where the golden back up copy is saved. Servers host a copy for serving
a
who does the job of writing to object store, is it controller or server?
m
It is coordinated by controller. For offline push with metadata push, the push job will directly write to deepstore. For realtime, servers will write. But in both cases, it is coordinated by the controller.
🍷 1