Slackbot
10/24/2022, 11:50 AMSergio Ferragut
10/24/2022, 3:42 PMoverlord and the middle manager/indexers pods.Julian Reyes
10/24/2022, 3:43 PMJulian Reyes
10/24/2022, 3:45 PMSergio Ferragut
10/24/2022, 3:58 PMJulian Reyes
10/24/2022, 4:03 PMJulian Reyes
10/24/2022, 4:06 PMSergio Ferragut
10/24/2022, 4:09 PMSergio Ferragut
10/24/2022, 4:10 PMJulian Reyes
10/24/2022, 4:10 PMSergio Ferragut
10/24/2022, 4:11 PMJulian Reyes
10/24/2022, 4:13 PMJulian Reyes
10/24/2022, 4:14 PMJulian Reyes
10/24/2022, 4:26 PMSergio Ferragut
10/24/2022, 4:30 PMRemoteTaskRunner - Worker . the messages look like this:
2022-10-24T16:22:40,320 INFO [Curator-PathChildrenCache-3] org.apache.druid.indexing.overlord.RemoteTaskRunner - Worker[172.20.0.8:8091] wrote RUNNING status for task [query-acee308e-d6b7-4067-8d07-f328d72c67d4-worker0] on [TaskLocation{host='172.20.0.8', port=8101, tlsPort=-1}]
...
2022-10-24T16:24:10,528 INFO [Curator-PathChildrenCache-3] org.apache.druid.indexing.overlord.RemoteTaskRunner - Worker[172.20.0.8:8091] wrote SUCCESS status for task [query-acee308e-d6b7-4067-8d07-f328d72c67d4-worker0] on [TaskLocation{host='172.20.0.8', port=8101, tlsPort=-1}]Julian Reyes
10/24/2022, 4:31 PMJulian Reyes
10/24/2022, 4:32 PMJulian Reyes
10/24/2022, 4:36 PM/druid/indexer/v1/workers I get running tasks, for the offending pods, if I search for some of them, they show up as failedSergio Ferragut
10/24/2022, 5:17 PMJulian Reyes
10/24/2022, 5:19 PMJulian Reyes
10/24/2022, 5:20 PMSergio Ferragut
10/24/2022, 5:22 PMJulian Reyes
10/24/2022, 5:25 PMSergio Ferragut
10/24/2022, 5:27 PMJulian Reyes
10/24/2022, 9:36 PMaws sts assume-role-with-web-identity .. in the offending node itself and using the token existing in the serviceAccount works.
After issuing above command, running aws sts get-caller-identity will give us the role serviceAccount uses (once you set the corresponding output env vars). Then running aws kinesis get-shard-iterator --region us-east-1 --stream-name .. works no problem once assumed the role.
I have not recycled the nodes as I wanted to make sure the node itself had access to kinesis by assuming the role the serviceAccount uses, but that might be the next step however I am not sure if that is going to fix the problem really.Sergio Ferragut
10/25/2022, 2:57 PMJulian Reyes
10/25/2022, 2:59 PMDidip Kerabat
10/25/2022, 4:43 PMkubectl run, deploying debugging Docker image into the suspicious EC2 nodes.
I would first check of all EC2 nodes have the same labels, etc.Julian Reyes
10/25/2022, 5:36 PMJulian Reyes
10/25/2022, 5:36 PMSergio Ferragut
10/25/2022, 5:41 PMJulian Reyes
10/25/2022, 5:45 PMSergio Ferragut
10/25/2022, 5:47 PMJulian Reyes
10/25/2022, 5:49 PMJulian Reyes
10/25/2022, 5:51 PMDidip Kerabat
10/25/2022, 5:52 PMAWS_DEFAULT_REGION=us-east-1
AWS_REGION=us-east-1
AWS_ROLE_ARN=arn:aws:iam::1234567890:role/your-druid-role
AWS_WEB_IDENTITY_TOKEN_FILE=/var/run/secrets/eks.amazonaws.com/serviceaccount/token
AWS_STS_REGIONAL_ENDPOINTS=regional
and make sure role/your-druid-role has all the access to Kinesis.Julian Reyes
10/25/2022, 6:18 PMdruid_indexer_task_baseDir: /opt/druid/var/tmp
druid_indexer_task_gracefulShutdownTimeout: PT120S
druid_indexer_task_restoreTasksOnRestart: true
druid_node_type: middleManager
druid_service: druid/middleManager
druid_worker_capacity: 12
AWS_DEFAULT_REGION: us-east-1
AWS_REGION: us-east-1
AWS_ROLE_ARN: arn:aws:iam::12345:role/druid-cluster-test-1234b8c
AWS_WEB_IDENTITY_TOKEN_FILE: /var/run/secrets/eks.amazonaws.com/serviceaccount/token
Mounts:
/opt/druid/var/druid/ from data (rw)
/var/run/secrets/eks.amazonaws.com/serviceaccount from aws-iam-token (ro)
/var/run/secrets/kubernetes.io/serviceaccount from druid-cluster-test-token-12flz (ro)
Conditions:
Type Status
Initialized True
Ready True
ContainersReady True
PodScheduled True
Volumes:
aws-iam-token:
Type: Projected (a volume that contains injected data from multiple sources)
TokenExpirationSeconds: 86400
data:
Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
ClaimName: data-druid-cluster-test-middle-manager-4
ReadOnly: false
druid-cluster-test-token-12flz:
Type: Secret (a volume populated by a Secret)
SecretName: druid-cluster-test-token-12flz
Optional: false
QoS Class: BestEffort
Node-Selectors: <none>
Tolerations: <http://aws.amazon.com/eks-local-ssd=true:NoSchedule|aws.amazon.com/eks-local-ssd=true:NoSchedule>
<http://infura.io/subnet=public:NoSchedule|infura.io/subnet=public:NoSchedule>
<http://node.kubernetes.io/not-ready:NoExecute|node.kubernetes.io/not-ready:NoExecute> op=Exists for 300s
<http://node.kubernetes.io/unreachable:NoExecute|node.kubernetes.io/unreachable:NoExecute> op=Exists for 300s
Events: <none>Julian Reyes
10/25/2022, 6:18 PMdruid_indexer_task_baseDir: /opt/druid/var/tmp
druid_indexer_task_gracefulShutdownTimeout: PT120S
druid_indexer_task_restoreTasksOnRestart: true
druid_node_type: middleManager
druid_service: druid/middleManager
druid_worker_capacity: 12
AWS_DEFAULT_REGION: us-east-1
AWS_REGION: us-east-1
AWS_ROLE_ARN: arn:aws:iam::12345:role/druid-cluster-test-1234b8c
AWS_WEB_IDENTITY_TOKEN_FILE: /var/run/secrets/eks.amazonaws.com/serviceaccount/token
Mounts:
/opt/druid/var/druid/ from data (rw)
/var/run/secrets/eks.amazonaws.com/serviceaccount from aws-iam-token (ro)
/var/run/secrets/kubernetes.io/serviceaccount from druid-cluster-test-token-12flz (ro)
Conditions:
Type Status
Initialized True
Ready True
ContainersReady True
PodScheduled True
Volumes:
aws-iam-token:
Type: Projected (a volume that contains injected data from multiple sources)
TokenExpirationSeconds: 86400
data:
Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
ClaimName: data-druid-cluster-test-middle-manager-5
ReadOnly: false
druid-cluster-test-token-12flz:
Type: Secret (a volume populated by a Secret)
SecretName: druid-cluster-test-token-12flz
Optional: false
QoS Class: BestEffort
Node-Selectors: <none>
Tolerations: <http://aws.amazon.com/eks-local-ssd=true:NoSchedule|aws.amazon.com/eks-local-ssd=true:NoSchedule>
<http://infura.io/subnet=public:NoSchedule|infura.io/subnet=public:NoSchedule>
<http://node.kubernetes.io/not-ready:NoExecute|node.kubernetes.io/not-ready:NoExecute> op=Exists for 300s
<http://node.kubernetes.io/unreachable:NoExecute|node.kubernetes.io/unreachable:NoExecute> op=Exists for 300s
Events: <none>Julian Reyes
10/25/2022, 6:20 PMarn:aws:iam::12345:role/druid-cluster-test-1234b8c has the required permissions and it’s the role defined in the service accountDidip Kerabat
10/25/2022, 7:42 PMJulian Reyes
10/25/2022, 8:01 PMSergio Ferragut
10/25/2022, 8:15 PMCory Johannsen
10/25/2022, 8:21 PMJulian Reyes
10/25/2022, 8:25 PMSergio Ferragut
10/26/2022, 1:23 AMUser: arn:aws:sts::123456789012:assumed-role/cluster-test-instanceRole-role-1abcd121/i-101a23761bcd34aa56
From your configs:
AWS_ROLE_ARN: arn:aws:iam::12345:role/druid-cluster-test-1234b8cJulian Reyes
10/26/2022, 10:28 AMJulian Reyes
10/26/2022, 10:34 AMUser: arn:aws:sts::123456789012:assumed-role/cluster-test-instanceRole-role-1abcd121/i-101a23761bcd34aa56 is different that the serviceAccount has AWS_ROLE_ARN: arn:aws:iam::12345:role/druid-cluster-test-1234b8c. The goal for serviceAccount roles in EKS is that the latter gets assumed by the former. https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.htmlSergio Ferragut
10/26/2022, 4:45 PMSergio Ferragut
10/26/2022, 4:47 PM,"druid-lookups-cached-global"]
success log
,"lookups-cached-global"]Julian Reyes
10/26/2022, 8:47 PMlookups-cached-global could not be found.Didip Kerabat
10/27/2022, 6:45 AMJulian Reyes
10/27/2022, 10:19 AMSergio Ferragut
10/27/2022, 2:58 PMJulian Reyes
10/27/2022, 5:34 PMJulian Reyes
10/27/2022, 8:23 PMGetShardIterator ! Most of the events were ListShards and the username is something like aws-sdk-java-1666253500488Julian Reyes
10/27/2022, 8:26 PMSergio Ferragut
10/27/2022, 8:27 PMJulian Reyes
10/27/2022, 8:28 PMSergio Ferragut
10/27/2022, 8:28 PMJulian Reyes
10/27/2022, 8:32 PM{
"eventVersion": "1.08",
"userIdentity": {
"type": "AssumedRole",
"principalId": "ABCD1234:aws-sdk-java-1666253500490",
"arn": "arn:aws:sts::1234567890:assumed-role/druid-cluster-us2-abcd1234/aws-sdk-java-1666253500490",
"accountId": "1234567890",
"accessKeyId": "12341bcd123",
"sessionContext": {
"sessionIssuer": {
"type": "Role",
"principalId": "ABCD1234",
"arn": "arn:aws:iam::1234567890:role/druid-cluster-us2-abcd1234",
"accountId": "1234567890",
"userName": "druid-cluster-us2-abcd1234"
},
"webIdFederationData": {
"federatedProvider": "arn:aws:iam::1234567890:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/F12345",
"attributes": {}
},
"attributes": {
"creationDate": "2022-10-27T19:56:46Z",
"mfaAuthenticated": "false"
}
}
},
"eventTime": "2022-10-27T20:28:53Z",
"eventSource": "<http://kinesis.amazonaws.com|kinesis.amazonaws.com>",
"eventName": "ListShards",
"awsRegion": "us-east-1",
"sourceIPAddress": "10.3.103.119",
"userAgent": "aws-sdk-java/1.12.264 Linux/5.4.209-116.367.amzn2.x86_64 OpenJDK_64-Bit_Server_VM/25.275-b01 java/1.8.0_275 vendor/Oracle_Corporation cfg/retry-mode/legacy",
"requestParameters": {
"streamName": "geotrips"
},
"responseElements": null,
"requestID": "c9f57a44-1b3f-c110-9221-9601afd22fa4",
"eventID": "4289e9fb-111f-475a-8e12-fb56d8122e13",
"readOnly": true,
"eventType": "AwsApiCall",
"managementEvent": true,
"recipientAccountId": "1234567890",
"vpcEndpointId": "vpce-081234",
"eventCategory": "Management",
"tlsDetails": {
"tlsVersion": "TLSv1.2",
"cipherSuite": "ECDHE-RSA-AES128-GCM-SHA256",
"clientProvidedHostHeader": "<http://kinesis.us-east-1.amazonaws.com|kinesis.us-east-1.amazonaws.com>"
}
}Sergio Ferragut
10/27/2022, 8:43 PMsourceIPAddress shown from a good MM or a bad one? Can you compare one from good MM to one from bad MM and see if the assumed role is different?Julian Reyes
10/27/2022, 8:46 PMJulian Reyes
10/27/2022, 8:50 PMJulian Reyes
10/27/2022, 8:54 PMSergio Ferragut
10/27/2022, 11:16 PMThe AWS access key ID and secret access key are used for Kinesis API requests. If this is not provided, the service will look for credentials set in environment variables, via Web Identity Token, in the default profile configuration file, and from the EC2 instance profile provider (in this order).
I'm walking through the code, to get a sense of what Druid is doing, for Kinesis ingestion tasks, this is when the AWS client is generated for individual task JVMs (peons) to obtain credentials.
As far as I understand this, It uses defaultAWSCredentialsProviderChain created here:
You can see that the provider chain used is:
• ConfigDrivenAwsCredentialConfigProvider - accessKey, secretKey specified.
• LazyFileSessionCredentialsProvider - uses a specified properties file to get sessionToken, accessKey, secretKey
• EnvironmentVariableCredentialsProvider - uses env AWS_ACCESS_KEY_ID, AWS_SECRET_KEY_ID
• SystemPropertiesCredentialProvider - aws.accessKeyId and aws.secretKey Java system properties.
• WebIdentityTokenCredentialsProvider - this is the one you are expecting to use with AWS_ROLE_ARN and AWS_WEB_IDENTITY_TOKEN_FILE
• ProfileCredentialsProvider - use default profile
• EC2ContainerCredentialsProviderWrapper -
• InstanceProfileCredentialsProvider
I still don't understand what's going on in your scenario, but hopefully seeing the actual provider chain will help you figure it out.
One other thought is that the druid.indexer.runner.javaOptsArray in the middle managers can change setting for the task's JVM, but that would be across the board on all MMs, so it still doesn't seem relevant.Didip Kerabat
10/27/2022, 11:17 PMJulian Reyes
10/28/2022, 10:55 AM