why do i see size as 0 , i do see data in query co...
# pinot-dev
a
why do i see size as 0 , i do see data in query console, i have set ""realtime.segment.flush.threshold.size": "16M"" but still i dont see segments if i query
Copy code
curl "<http://ip:9000/tables/table_name/size?detailed=true>" 2> /dev/null | jq -c
{"tableName":"table_name","reportedSizeInBytes":0,"estimatedSizeInBytes":0,"offlineSegments":null,"realtimeSegments":{"reportedSizeInBytes":0,"estimatedSizeInBytes":0,"missingSegments":0,"segments":{}}}
kafka topic has around GB's of data but no segments in pinot... When i run producer to add records in pinot , i do see increase in number of records via pinot query console (ran count(*))... what could be the issue here, why segments are not getting created?
Screenshot 2023-07-20 at 7.45.12 PM.png
If i check in UI , all segments are in consuming status with following metadata
Copy code
{
  "segment.realtime.numReplicas": "2",
  "segment.creation.time": "1689846621631",
  "segment.flush.threshold.size": "100000",
  "segment.realtime.startOffset": "0",
  "segment.realtime.status": "IN_PROGRESS"
}
j
That api doesn’t report the consuming segments size I believe. So once those have sealed, you should see the size
a
ok, how much time does it takes? Given that i have explicitly mentioned
Copy code
realtime.segment.flush.threshold.size": "16M"
in table json... Its been 5-6 hours and still in consuming state....
j
i think it’s some nuance around the first segment and only setting the
size
threshold
i think you want to set
realtime.segment.flush.threshold.rows=0
, otherwise i think pinot waits for 100k rows
a
ohh , currently it has 30k rows in it... I have around 32 segments...
will it work if i change the table properties via ui and add
realtime.segment.flush.threshold.rows
to some value?
j
it will, but it won’t apply until these segments seal. you can add it then run the forceCommit API to make it go faster
a
wait i have set realtime.segment.flush.threshold.rows=0
Copy code
{
  "REALTIME": {
    "tableName": "table_name",
    "tableType": "REALTIME",
    "segmentsConfig": {
      "timeType": "MILLISECONDS",
      "schemaName": "table_name",
      "replicasPerPartition": "2",
      "timeColumnName": "timestamp",
      "peerSegmentDownloadScheme": "http",
      "minimizeDataMovement": false
    },
    "tenants": {
      "broker": "DefaultTenant",
      "server": "DefaultTenant"
    },
    "tableIndexConfig": {
      "streamConfigs": {
        "streamType": "kafka",
        "stream.kafka.consumer.type": "simple",
        "stream.kafka.topic.name": "kafka-topic",
        "stream.kafka.decoder.class.name": "org.apache.pinot.plugin.stream.kafka.KafkaJSONMessageDecoder",
        "stream.kafka.consumer.factory.class.name": "org.apache.pinot.plugin.stream.kafka20.KafkaConsumerFactory",
        "stream.kafka.broker.list": "broker_list",
        "realtime.segment.flush.threshold.rows": "0",
        "realtime.segment.flush.threshold.segment.size": "16M",
        "realtime.segment.serverUploadToDeepStore": "true",
        "stream.kafka.consumer.prop.auto.offset.reset": "smallest"
      },
      "rangeIndexVersion": 2,
      "autoGeneratedInvertedIndex": false,
      "createInvertedIndexDuringSegmentGeneration": false,
      "loadMode": "MMAP",
      "enableDefaultStarTree": false,
      "enableDynamicStarTreeCreation": false,
      "aggregateMetrics": false,
      "nullHandlingEnabled": false,
      "optimizeDictionary": false,
      "optimizeDictionaryForMetrics": false,
      "noDictionarySizeRatioThreshold": 0
    },
    "metadata": {
      "customConfigs": {}
    },
    "isDimTable": false
  }
}
j
tbh i’m not sure then. there’s a lot of code using these values doing slightly different things. I still recommend force committing these first segments and seeing what it does after
a
I did on 1 table , size shows around 8 MB, does it mean it has less data
Is it possible that pinot is not consuming all data from the topic of kafka? How can i check whether kafka topic has how much data and how much pinot has consumed?
j
that’s also the size across all replicas of each partition, so there was even less data in each individual segment
you’d have to write a separate application to read from kafka so you can compare. there’s not a great other way
a
there is no command to check data in kafka topic?
@Johan Adami do you think uniqueness in record can play a factor in creation of less data in segment? I see records in kafka topic are duplicate...
j
pinot won’t do anything special with the data for dedup (unless you go down the deduping in Pinot route which is a separate topic)
you should be able to tail the kafka topic yourself and compare to what you see in pinot. it just depends how robust your testing needs to be
a
So there is a command using which we can get data in kafka topic
Copy code
kafka-log-dirs.sh --describe --bootstrap-server servers --topic-list kafka-topic-1
It has around 2.5 GB in each partition, there are 32 partitions so 80 GB of data. You are saying duplicate records does not have impact on pinot and then i am not sure why pinot is not showing the data in segments....
I also verified that number of records in kafka topic and my table in exactly same.
Copy code
bin/kafka-run-class.sh kafka.tools.GetOffsetShell   --broker-list localhost:9092   
--topic kafka-topic-1   | awk -F  ":" '{sum += $3} END {print "Result: "sum}'
and
Copy code
select count(*) from table_name;
gave same result.