This message was deleted.
# general
s
This message was deleted.
d
In a production failure scenario, a reinstall should be a last resort option IMHO. This issue with a fresh reinstall and geo-replication is timing and cursor positioning. From a timing perspective, the data to transmitted between regions when it is first published. So the data cannot be sent to a cluster that is created day, weeks, months after the data was first published. Secondly, there replication subscriptions inside Pulsar that account for lagging data replication. The problem is that these are implemented as subscriptions, just like any other subscription in Pulsar. So once the data has been replicated, the cursor position on the replicated subscriptions is updated (there is the potential for the data to fall outside the retention period at this point if the replication subscription was the last/only active subscription, etc). Once the cursor is moved, it can’t go back.
One thing you might want to try is using a different cluster name when you re-create A. This will use a different subscription name as least, so you can avoid the cursor issue since I am 99% sure that the subscription name is derived from the cluster name.
e
Thanks for the reply @David K, I'll go into more detail about what I've done. Our goal is really to test the extreme case where production A has a system problem and must be reinstalled with the same name as earlier. We enabled retention on the georeplicated namespace so the data is still present on cluster B when A is reinstalled. We tested deactivating replication so that the cursor is no longer present on the topics and reactivating it by creating a subscription from the earliest messageId.
Copy code
pulsar-admin topics create-subscription -s pulsar.repl.new-cluster -m earliest <topic>
Even with the cursor removed, old messages are not duplicated, but new messages sent are duplicated if I send messages in topics to cluster B.
d
How do you know that the above subscription name is correct? I think you need to find the name of the subscription created by the geo-replication mechanism and then reset the cursor on it.
e
I've retrieved the name to use from the documentation that talks about geo replication. (the note in the end) We tried to reset the cursor on the subscription created by the geo-replication but it's not implemented : here is a issue https://github.com/apache/pulsar/pull/14887 where there talking about this
👀 1
in the ticket they seem to say it works but in our case it doesn't, maybe we've missed something.
d
I would agree. Do you get any sort of error when attempting to reset the cursor, etc?
e
I don't get any error when I run the reset-cursor on my subscription and my subcription look like this with the stats
Copy code
"subscriptions" : {
    "pulsar.repl.pulsar-preprod" : {
      "msgRateOut" : 0.0,
      "msgThroughputOut" : 0.0,
      "bytesOutCounter" : 0,
      "msgOutCounter" : 0,
      "msgRateRedeliver" : 0.0,
      "messageAckRate" : 0.0,
      "chunkedMessageRate" : 0,
      "msgBacklog" : 0,
      "backlogSize" : 0,
      "earliestMsgPublishTimeInBacklog" : 0,
      "msgBacklogNoDelayed" : 0,
      "blockedSubscriptionOnUnackedMsgs" : false,
      "msgDelayed" : 0,
      "unackedMessages" : 0,
      "type" : "None",
      "msgRateExpired" : 0.0,
      "totalMsgExpired" : 0,
      "lastExpireTimestamp" : 0,
      "lastConsumedFlowTimestamp" : 0,
      "lastConsumedTimestamp" : 0,
      "lastAckedTimestamp" : 0,
      "lastMarkDeleteAdvancedTimestamp" : 0,
      "consumers" : [ ],
      "isDurable" : true,
      "isReplicated" : false,
      "allowOutOfOrderDelivery" : false,
      "consumersAfterMarkDeletePosition" : { },
      "nonContiguousDeletedMessagesRanges" : 0,
      "nonContiguousDeletedMessagesRangesSerializedSize" : 0,
      "delayedMessageIndexSizeInBytes" : 0,
      "subscriptionProperties" : { },
      "filterProcessedMsgCount" : 0,
      "filterAcceptedMsgCount" : 0,
      "filterRejectedMsgCount" : 0,
      "filterRescheduledMsgCount" : 0,
      "durable" : true,
      "replicated" : false
    }
  },
i'm running out of ideas ...
d
can you share the output of
stats-internal
?
e
This one is from my topic on cluster B
Copy code
{
  "entriesAddedCounter": 0,
  "numberOfEntries": 2,
  "totalSize": 395,
  "currentLedgerEntries": 0,
  "currentLedgerSize": 0,
  "lastLedgerCreatedTimestamp": "2023-08-10T19:08:44.837Z",
  "waitingCursorsCount": 1,
  "pendingAddEntriesCount": 0,
  "lastConfirmedEntry": "23:1",
  "state": "LedgerOpened",
  "ledgers": [
    {
      "ledgerId": 23,
      "entries": 2,
      "size": 395,
      "offloaded": false,
      "underReplicated": false
    },
    {
      "ledgerId": 37,
      "entries": 0,
      "size": 0,
      "offloaded": false,
      "underReplicated": false
    }
  ],
  "cursors": {
    "pulsar.repl.pulsar-preprod": {
      "markDeletePosition": "23:1",
      "readPosition": "23:2",
      "waitingReadOp": true,
      "pendingReadOps": 0,
      "messagesConsumedCounter": 0,
      "cursorLedger": 38,
      "cursorLedgerLastEntry": 2,
      "individuallyDeletedMessages": "[]",
      "lastLedgerSwitchTimestamp": "2023-08-16T13:33:25.216Z",
      "state": "Open",
      "active": false,
      "numberOfEntriesSinceFirstNotAckedMessage": 1,
      "totalNonContiguousDeletedMessagesRange": 0,
      "subscriptionHavePendingRead": false,
      "subscriptionHavePendingReplayRead": false,
      "properties": {}
    }
  },
  "schemaLedgers": [
    {
      "ledgerId": 24,
      "entries": 1,
      "size": 1133,
      "offloaded": false,
      "underReplicated": false
    }
  ],
  "compactedLedger": {
    "ledgerId": -1,
    "entries": -1,
    "size": -1,
    "offloaded": false,
    "underReplicated": false
  }
}