Hi Team, We were trying out to move the completed...
# troubleshooting
a
Hi Team, We were trying out to move the completed segments to different hosts, but the completed segment is still on REALTIME Server and not moved to OFFLINE Server. Reference: https://docs.pinot.apache.org/operators/operating-pinot/tuning/realtime#moving-completed-segments-to-different-hosts Servers are tagged like below, server1: "listFields": { "TAG_LIST": [ "DefaultTenant_OFFLINE" ] } Serve2: "listFields": { "TAG_LIST": [ "DefaultTenant_REALTIME" ] } Table Config: "tenants": { "broker": "DefaultTenant", "server": "DefaultTenant", "tagOverrideConfig": { "realtimeConsuming": "DefaultTenant_REALTIME", "realtimeCompleted": "DefaultTenant_OFFLINE" } }, can someone help?
l
Hey Anish. The documentation is out of date, you can accomplish this using the instanceAssignmentConfigMap
I’ll send you an example
Copy code
"instanceAssignmentConfigMap": {
    "CONSUMING": {
      "tagPoolConfig": {
        "tag": "DefaultTenant_REALTIME"
      },
      "replicaGroupPartitionConfig": {
        "replicaGroupBased": true,
        "numReplicaGroups": 2,
        "numInstancesPerReplicaGroup": 2
      }
    },
    "COMPLETED": {
      "tagPoolConfig": {
        "tag": "DefaultTenant_OFFLINE"
      },
      "replicaGroupPartitionConfig": {
        "replicaGroupBased": true,
        "numReplicaGroups": 2,
        "numInstancesPerReplicaGroup": 4
      }
    }
  },
And get rid of the tagOverrideConfig, the above is the recommended approach
You will also have to execute a rebalance with reassign instances on your table
a
So, i have two offline servers and 2 realtime servers. what should be the value for numReplicaGroups and numInstancesPerReplicaGroup , if i want 2 replicas
l
Do you want two replica groups?
Or do you not want replica groups?
There are many different ways to assign instances, you should read this section https://docs.pinot.apache.org/operators/operating-pinot/instance-assignment#replica-group-instance-assignment
a
so this is for realtime table
l
Right. For replication, you can simply set
Copy code
"replicasPerPartition": 2
How many partitions do you have?
a
currently 1, for testing purpose
l
Ok no problem
I think you could do your replicaGroupPartitionConfig like this
Copy code
"replicaGroupPartitionConfig": {
           "numInstances": 2,
           "numPartitions": 1
      }
Let me know how that goes
a
same for COMPLETED?
l
Yeah if you have 2 offline and 2 realtime instances
When you change the table config, you can execute a dry run rebalance and see if that’s what you expect
a
Copy code
{
  "tableName": "reporting_aggregations_REALTIME",
  "tableType": "REALTIME",
  "segmentsConfig": {
    "timeType": "HOURS",
    "schemaName": "reporting_aggregations_schema",
    "retentionTimeUnit": "DAYS",
    "retentionTimeValue": "14",
    "timeColumnName": "stats_date_hour",
    "allowNullTimeValue": false,
    "replicasPerPartition": "2",
    "completionConfig": {
      "completionMode": "DOWNLOAD"
    }
  },
  "instanceAssignmentConfigMap": {
    "CONSUMING": {
      "tagPoolConfig": {
        "tag": "DefaultTenant_REALTIME"
      },
      "replicaGroupPartitionConfig": {
        "replicaGroupBased": true,
        "numReplicaGroups": 2,
        "numInstancesPerReplicaGroup": 2
      }
    },
    "COMPLETED": {
      "tagPoolConfig": {
        "tag": "DefaultTenant_OFFLINE"
      },
      "replicaGroupPartitionConfig": {
        "replicaGroupBased": true,
        "numReplicaGroups": 2,
        "numInstancesPerReplicaGroup": 4
      }
    }
  },
  "tableIndexConfig": {
    "streamConfigs": {
      "streamType": "kafka",
      "stream.kafka.consumer.type": "lowlevel",
      "stream.kafka.topic.name": "c8.max_reporting_aggregations_olap",
      "stream.kafka.decoder.class.name": "org.apache.pinot.plugin.stream.kafka.KafkaJSONMessageDecoder",
      "stream.kafka.consumer.factory.class.name": "org.apache.pinot.plugin.stream.kafka20.KafkaConsumerFactory",
      "stream.kafka.broker.list": "brokers",
      "realtime.segment.flush.threshold.rows": "0",
      "realtime.segment.flush.threshold.size": "0",
      "realtime.segment.flush.threshold.time": "24h",
      "realtime.segment.flush.threshold.segment.size": "500M",
      "realtime.segment.flush.autotune.initialRows": "10000000",
      "stream.kafka.consumer.prop.auto.offset.reset": "largest"
    },
    "rangeIndexVersion": 1,
    "autoGeneratedInvertedIndex": false,
    "createInvertedIndexDuringSegmentGeneration": false,
    "loadMode": "MMAP",
    "enableDefaultStarTree": false,
    "enableDynamicStarTreeCreation": false,
    "aggregateMetrics": false,
    "nullHandlingEnabled": true
  },
  "metadata": {
    "customConfigs": {}
  },
  "routing": {
    "instanceSelectorType": "strictReplicaGroup"
  },
  "upsertConfig": {
    "mode": "FULL",
    "comparisonColumn": "process_datetime"
  },
  "isDimTable": false
}
this looks fine ? @Lars-Kristian Svenøy
l
Hmm
So based on that, you need 8 offline tenants and 4 realtime ones
you’ll end up with 2 replica groups for each
a
my bad, changed to "instanceAssignmentConfigMap": { "CONSUMING": { "tagPoolConfig": { "tag": "DefaultTenant_REALTIME" }, "replicaGroupPartitionConfig": { "numInstances": 2, "numPartitions": 1 } }, "COMPLETED": { "tagPoolConfig": { "tag": "DefaultTenant_OFFLINE" }, "replicaGroupPartitionConfig": { "numInstances": 2, "numPartitions": 1 } } }
l
I think that’s fine
But only way to know is to try apply it and run a dry run rebalance with reassign instances
I haven’t used that specific way of setting instances, I always use replica groups
a
Hey @Lars-Kristian Svenøy, seems like it got re balanced , with previous setting only. I just saw before making the config change.
n
Just fyi, your previous config would've also worked fine. But the movement happens only periodically (hourly by default). You have to run a rebalance with dry run false, downtime true, bootstrap true if you want to see it immediately. If you're not seeing it move periodically, check controller logs for lines printed by Relocator
a
thanks @Neha Pawar for the info.