This message was deleted.
# troubleshooting
s
This message was deleted.
s
Hi @Julian Reyes, The pods that require access for streaming ingestion are the
overlord
and the
middle manager/indexers
pods.
j
I think they have it. I set the permissions through a service account that was applied to Deployment and Statefulset k8s resources
There is ingestion happening as I have data in the datasources however it seems like something is missing
s
the error is intermittent? Or just some of the tasks fail? Are they consistently from a particular indexer/MM?
j
I have 12 middleManagers pods and errors seem to come from 3 of them always
they seem to have the correct service account applied
s
That is an interesting clue. Thinking out loud, Were those pods started at an earlier time? Perhaps before you applied the service account? Have you tried restarting them?
The other thought.. are all kinesis tasks assigned to those MMs failing?
j
I have not restarted them, wanted to see what your opinion was. They are just 4days old which means they should have the settings properly applied since they were applied long time ago
s
That’s odd.
j
how can I find out that? what I see in services is that two of the problematic MM are blacklisted
there are 9 slots used of 12 across pretty much all MM except the ones blacklisted
I have restarted them, lets see what happens, would’ve liked to know what happened before
s
You can look at the overlord log to see where each task is allocated by searching for
RemoteTaskRunner - Worker
. the messages look like this:
Copy code
2022-10-24T16:22:40,320 INFO [Curator-PathChildrenCache-3] org.apache.druid.indexing.overlord.RemoteTaskRunner - Worker[172.20.0.8:8091] wrote RUNNING status for task [query-acee308e-d6b7-4067-8d07-f328d72c67d4-worker0] on [TaskLocation{host='172.20.0.8', port=8101, tlsPort=-1}]
...
2022-10-24T16:24:10,528 INFO [Curator-PathChildrenCache-3] org.apache.druid.indexing.overlord.RemoteTaskRunner - Worker[172.20.0.8:8091] wrote SUCCESS status for task [query-acee308e-d6b7-4067-8d07-f328d72c67d4-worker0] on [TaskLocation{host='172.20.0.8', port=8101, tlsPort=-1}]
j
all right, I restarted the offending pods and the new ones are getting same issue
is there an API where same info gets displayed?
if I do a GET at
/druid/indexer/v1/workers
I get running tasks, for the offending pods, if I search for some of them, they show up as failed
s
are there any succeeding stream tasks on the offending MMs? How is the service account/assumed role being provided? environment variables? EC2 instance profile ?
j
No streams tasks succeeded in offending MMs. It's being provided by EC2 instances assuming a role. https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html
Only 3 MMs are failing the other ones are working fine and same service account is applied to all of them
s
I’m no expert on this🤔 but can you redeploy those EC2 instances? Could it be that they need to refresh their role after some previous permission change?
j
I think they are spot EC2 instances, I could try that but I think they should be able to just assume the role. I could also try to deploy aws cli to some of those nodes and manually check
s
Those sounds like good troubleshooting tests. Let us know how it goes.
j
well, issuing
aws sts assume-role-with-web-identity ..
in the offending node itself and using the token existing in the serviceAccount works. After issuing above command, running
aws sts get-caller-identity
will give us the role serviceAccount uses (once you set the corresponding output env vars). Then running
aws kinesis get-shard-iterator --region us-east-1 --stream-name ..
works no problem once assumed the role. I have not recycled the nodes as I wanted to make sure the node itself had access to kinesis by assuming the role the serviceAccount uses, but that might be the next step however I am not sure if that is going to fix the problem really.
s
So I guess the question still remaining is what is the difference between the MMs that work and those that don't?
j
I do not see difference really by describing their pods
d
You can do
kubectl run
, deploying debugging Docker image into the suspicious EC2 nodes. I would first check of all EC2 nodes have the same labels, etc.
j
Yeah a strategy I saw was deploying aws cli image in to the node and test the service account but I think that will work
I would need to see the labels for the nodes itself
s
@Julian Reyes one thing you mentioned about your test is that you had to define the environment variables for it to work. Could that be a difference in the pods that work and don't? Perhaps the ones that work have the environment variables already defined?
j
What env variables? The ones I defined in the nodes was the ones related to AWS, key, secret, token, etc so the assumerole works
s
I’m not really familiar with how the authentication stack works… so it is just a thought about what might be diff in the good va bad pods.
j
I think all of them have the same settings as they were all deployed using the helm chart and defining the replcias
The node assumes a role (service account role) and the Aws SDK in druid Auths by making use of Web identify token. It's a pretty common scenario in k8s
d
if you are using AWS’s IMDS v2 auth (auth using EKS’s service account), there will be AWS specific environment variables. Make sure they all the same on every pod and every node, e.g.
Copy code
AWS_DEFAULT_REGION=us-east-1
AWS_REGION=us-east-1
AWS_ROLE_ARN=arn:aws:iam::1234567890:role/your-druid-role
AWS_WEB_IDENTITY_TOKEN_FILE=/var/run/secrets/eks.amazonaws.com/serviceaccount/token
AWS_STS_REGIONAL_ENDPOINTS=regional
and make sure
role/your-druid-role
has all the access to Kinesis.
j
some of the env vars for one of the offending pods:
Copy code
druid_indexer_task_baseDir:                                     /opt/druid/var/tmp                                                                                                                                                          
       druid_indexer_task_gracefulShutdownTimeout:                     PT120S                                                                                                                                                                      
       druid_indexer_task_restoreTasksOnRestart:                       true                                                                                                                                                                        
       druid_node_type:                                                middleManager                                                                                                                                                               
       druid_service:                                                  druid/middleManager                                                                                                                                                         
       druid_worker_capacity:                                          12                                                                                                                                                                          
       AWS_DEFAULT_REGION:                                             us-east-1                                                                                                                                                                   
       AWS_REGION:                                                     us-east-1                                                                                                                                                                   
       AWS_ROLE_ARN:                                                   arn:aws:iam::12345:role/druid-cluster-test-1234b8c                                                                                                                    
       AWS_WEB_IDENTITY_TOKEN_FILE:                                    /var/run/secrets/eks.amazonaws.com/serviceaccount/token                                                                                                                     
     Mounts:                                                                                                                                                                                                                                       
       /opt/druid/var/druid/ from data (rw)                                                                                                                                                                                                        
       /var/run/secrets/eks.amazonaws.com/serviceaccount from aws-iam-token (ro)                                                                                                                                                                   
       /var/run/secrets/kubernetes.io/serviceaccount from druid-cluster-test-token-12flz (ro)                                                                                                                                                       
 Conditions:                                                                                                                                                                                                                                       
   Type              Status                                                                                                                                                                                                                        
   Initialized       True                                                                                                                                                                                                                          
   Ready             True                                                                                                                                                                                                                          
   ContainersReady   True                                                                                                                                                                                                                          
   PodScheduled      True                                                                                                                                                                                                                          
 Volumes:                                                                                                                                                                                                                                          
   aws-iam-token:                                                                                                                                                                                                                                  
     Type:                    Projected (a volume that contains injected data from multiple sources)                                                                                                                                               
     TokenExpirationSeconds:  86400                                                                                                                                                                                                                
   data:                                                                                                                                                                                                                                           
     Type:       PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)                                                                                                                                              
     ClaimName:  data-druid-cluster-test-middle-manager-4                                                                                                                                                                                           
     ReadOnly:   false                                                                                                                                                                                                                             
   druid-cluster-test-token-12flz:                                                                                                                                                                                                                  
     Type:        Secret (a volume populated by a Secret)                                                                                                                                                                                          
     SecretName:  druid-cluster-test-token-12flz                                                                                                                                                                                                    
     Optional:    false                                                                                                                                                                                                                            
 QoS Class:       BestEffort                                                                                                                                                                                                                       
 Node-Selectors:  <none>                                                                                                                                                                                                                           
 Tolerations:     <http://aws.amazon.com/eks-local-ssd=true:NoSchedule|aws.amazon.com/eks-local-ssd=true:NoSchedule>                                                                                                                                                                                     
                  <http://infura.io/subnet=public:NoSchedule|infura.io/subnet=public:NoSchedule>                                                                                                                                                                                               
                  <http://node.kubernetes.io/not-ready:NoExecute|node.kubernetes.io/not-ready:NoExecute> op=Exists for 300s                                                                                                                                                                        
                  <http://node.kubernetes.io/unreachable:NoExecute|node.kubernetes.io/unreachable:NoExecute> op=Exists for 300s                                                                                                                                                                      
 Events:          <none>
and env vars from one of the working pods:
Copy code
druid_indexer_task_baseDir:                                     /opt/druid/var/tmp                                                                                                                                                          
       druid_indexer_task_gracefulShutdownTimeout:                     PT120S                                                                                                                                                                      
       druid_indexer_task_restoreTasksOnRestart:                       true                                                                                                                                                                        
       druid_node_type:                                                middleManager                                                                                                                                                               
       druid_service:                                                  druid/middleManager                                                                                                                                                         
       druid_worker_capacity:                                          12                                                                                                                                                                          
       AWS_DEFAULT_REGION:                                             us-east-1                                                                                                                                                                   
       AWS_REGION:                                                     us-east-1                                                                                                                                                                   
       AWS_ROLE_ARN:                                                   arn:aws:iam::12345:role/druid-cluster-test-1234b8c                                                                                                                    
       AWS_WEB_IDENTITY_TOKEN_FILE:                                    /var/run/secrets/eks.amazonaws.com/serviceaccount/token                                                                                                                     
     Mounts:                                                                                                                                                                                                                                       
       /opt/druid/var/druid/ from data (rw)                                                                                                                                                                                                        
       /var/run/secrets/eks.amazonaws.com/serviceaccount from aws-iam-token (ro)                                                                                                                                                                   
       /var/run/secrets/kubernetes.io/serviceaccount from druid-cluster-test-token-12flz (ro)                                                                                                                                                       
 Conditions:                                                                                                                                                                                                                                       
   Type              Status                                                                                                                                                                                                                        
   Initialized       True                                                                                                                                                                                                                          
   Ready             True                                                                                                                                                                                                                          
   ContainersReady   True                                                                                                                                                                                                                          
   PodScheduled      True                                                                                                                                                                                                                          
 Volumes:                                                                                                                                                                                                                                          
   aws-iam-token:                                                                                                                                                                                                                                  
     Type:                    Projected (a volume that contains injected data from multiple sources)                                                                                                                                               
     TokenExpirationSeconds:  86400                                                                                                                                                                                                                
   data:                                                                                                                                                                                                                                           
     Type:       PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)                                                                                                                                              
     ClaimName:  data-druid-cluster-test-middle-manager-5                                                                                                                                                                                           
     ReadOnly:   false                                                                                                                                                                                                                             
   druid-cluster-test-token-12flz:                                                                                                                                                                                                                  
     Type:        Secret (a volume populated by a Secret)                                                                                                                                                                                          
     SecretName:  druid-cluster-test-token-12flz                                                                                                                                                                                                    
     Optional:    false                                                                                                                                                                                                                            
 QoS Class:       BestEffort                                                                                                                                                                                                                       
 Node-Selectors:  <none>                                                                                                                                                                                                                           
 Tolerations:     <http://aws.amazon.com/eks-local-ssd=true:NoSchedule|aws.amazon.com/eks-local-ssd=true:NoSchedule>                                                                                                                                                                                     
                  <http://infura.io/subnet=public:NoSchedule|infura.io/subnet=public:NoSchedule>                                                                                                                                                                                               
                  <http://node.kubernetes.io/not-ready:NoExecute|node.kubernetes.io/not-ready:NoExecute> op=Exists for 300s                                                                                                                                                                        
                  <http://node.kubernetes.io/unreachable:NoExecute|node.kubernetes.io/unreachable:NoExecute> op=Exists for 300s                                                                                                                                                                      
 Events:          <none>
role
arn:aws:iam::12345:role/druid-cluster-test-1234b8c
has the required permissions and it’s the role defined in the service account
d
weird, maybe this is a complication due to using spot instances? I never used spot instances before.
j
not sure… also both instances (offending and working ones) have same age (brought up online 27 days ago)
s
I know @Cory Johannsen has experience with spot instances. Cory, any ideas on why some MMs may work and others not in terms of authentication with AWS Identity Token?
c
We run on spot instances but aren't using the token approach for auth in the same manner. I haven't see this occur.
j
it’s hard to debug actually
s
What Druid version are you on? This doesn't explain the difference across pods, but there is an STS related PR included in Druid 24.0. @Julian Reyes just a few more thoughts, are there any warnings in the failed task logs prior to the " role does not have permission ERROR"? Is there a chance that a different role is being assumed in the chain? Perhaps another way to address this is to compare the beginning of the successful task log to beginning of the unsuccessful one and see what material differences they have at startup. Actually in the failed log, isn't it reporting a different role, from task log at beginning of the thread:
Copy code
User: arn:aws:sts::123456789012:assumed-role/cluster-test-instanceRole-role-1abcd121/i-101a23761bcd34aa56
From your configs:
Copy code
AWS_ROLE_ARN: arn:aws:iam::12345:role/druid-cluster-test-1234b8c
j
@Sergio Ferragut the difference between the role arns is because when I pasted them here I modified them so not to expose the real name, and in the process probably I did not keep the same name, but in reality the names are the same. I am using 0.24.0 version currently. I am going to upload a failed task log and a success one. I could not see any difference other than the error. In the success one there are warns about lookups but the lookups work properly, not sure why warns show up in the indexing tasks.
But obviously, the role the node has
User: arn:aws:sts::123456789012:assumed-role/cluster-test-instanceRole-role-1abcd121/i-101a23761bcd34aa56
is different that the serviceAccount has
AWS_ROLE_ARN: arn:aws:iam::12345:role/druid-cluster-test-1234b8c
. The goal for serviceAccount roles in EKS is that the latter gets assumed by the former. https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html
s
Thanks for sharing the logs and for your patience while I learn about identities and service accounts in AWS 🙂 I found a difference in the logs that probably isn't the culprit, but it does indicate two different MM configurations:
Line 79 in both has a difference in the extensions.LoadList: failed log
Copy code
,"druid-lookups-cached-global"]
success log
Copy code
,"lookups-cached-global"]
j
Hi Sergio, I checked again live logs and I think the difference you saw was because I modified the logs to change some keywords. I double check tasks logs again for success and failed ones and both have the same load list. I think otherwise an exception would show up since
lookups-cached-global
could not be found.
d
OK, silly question, can you check your ingestion JSON and see if you manually defined an AWS role in there?
j
Nop, there are no AWS role in the spec itself, with the given cluster config, all tasks would fail if I attempt doing so
s
I got a suggestion from someone who know AWS well. He suggested "look at AWS cloudtrail for all these API calls and slice and dice them by different fields (role arn, etc, source ip, response status code etc..) Maybe from there you can see a pattern. AWS support can also help with these permission issues, they're pretty good especially to rule out an issue on EKS side."
j
Thanks, going to have a look at cloudtrail, although might be difficult to find out given the amount of pods, etc we have
All right this is interesting, I went to CloudTrail and I did not see any event
GetShardIterator
! Most of the events were
ListShards
and the username is something like
aws-sdk-java-1666253500488
but perhaps it makes sense that it does not show up in cloudtrail as the api call was not successful? returning the error of permission denied?
s
Not sure, but in that case you should see the successful ones from the successful tasks, right?
j
thats right, however there are no events for that action at all
s
Do the events for assuming a role show up? Is that an event?
j
yeah (note I modify some keywords)
Copy code
{
  "eventVersion": "1.08",
  "userIdentity": {
    "type": "AssumedRole",
    "principalId": "ABCD1234:aws-sdk-java-1666253500490",
    "arn": "arn:aws:sts::1234567890:assumed-role/druid-cluster-us2-abcd1234/aws-sdk-java-1666253500490",
    "accountId": "1234567890",
    "accessKeyId": "12341bcd123",
    "sessionContext": {
      "sessionIssuer": {
        "type": "Role",
        "principalId": "ABCD1234",
        "arn": "arn:aws:iam::1234567890:role/druid-cluster-us2-abcd1234",
        "accountId": "1234567890",
        "userName": "druid-cluster-us2-abcd1234"
      },
      "webIdFederationData": {
        "federatedProvider": "arn:aws:iam::1234567890:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/F12345",
        "attributes": {}
      },
      "attributes": {
        "creationDate": "2022-10-27T19:56:46Z",
        "mfaAuthenticated": "false"
      }
    }
  },
  "eventTime": "2022-10-27T20:28:53Z",
  "eventSource": "<http://kinesis.amazonaws.com|kinesis.amazonaws.com>",
  "eventName": "ListShards",
  "awsRegion": "us-east-1",
  "sourceIPAddress": "10.3.103.119",
  "userAgent": "aws-sdk-java/1.12.264 Linux/5.4.209-116.367.amzn2.x86_64 OpenJDK_64-Bit_Server_VM/25.275-b01 java/1.8.0_275 vendor/Oracle_Corporation cfg/retry-mode/legacy",
  "requestParameters": {
    "streamName": "geotrips"
  },
  "responseElements": null,
  "requestID": "c9f57a44-1b3f-c110-9221-9601afd22fa4",
  "eventID": "4289e9fb-111f-475a-8e12-fb56d8122e13",
  "readOnly": true,
  "eventType": "AwsApiCall",
  "managementEvent": true,
  "recipientAccountId": "1234567890",
  "vpcEndpointId": "vpce-081234",
  "eventCategory": "Management",
  "tlsDetails": {
    "tlsVersion": "TLSv1.2",
    "cipherSuite": "ECDHE-RSA-AES128-GCM-SHA256",
    "clientProvidedHostHeader": "<http://kinesis.us-east-1.amazonaws.com|kinesis.us-east-1.amazonaws.com>"
  }
}
s
Is the
sourceIPAddress
shown from a good MM or a bad one? Can you compare one from good MM to one from bad MM and see if the assumed role is different?
j
that source IP is from a good one, I tried to filter out by a bad source IP but it seems I cant filter by IP
it’s weird, most of them come from same source IP
somehow I am leaning towards thinking that the real issue is that seemingly randomly druid ISN’T assuming the role as expected and making the kinesis call with the instance role instead of the service account role
s
Yeah...I think the strange thing about your scenario is that it seems to be specific to individual MM pods/nodes, so it doesn't seem that random. From the docs the order of the provider chain is:
Copy code
The AWS access key ID and secret access key are used for Kinesis API requests. If this is not provided, the service will look for credentials set in environment variables, via Web Identity Token, in the default profile configuration file, and from the EC2 instance profile provider (in this order).
I'm walking through the code, to get a sense of what Druid is doing, for Kinesis ingestion tasks, this is when the AWS client is generated for individual task JVMs (peons) to obtain credentials. As far as I understand this, It uses defaultAWSCredentialsProviderChain created here: You can see that the provider chain used is: • ConfigDrivenAwsCredentialConfigProvider - accessKey, secretKey specified. • LazyFileSessionCredentialsProvider - uses a specified properties file to get sessionToken, accessKey, secretKey • EnvironmentVariableCredentialsProvider - uses env AWS_ACCESS_KEY_ID, AWS_SECRET_KEY_ID • SystemPropertiesCredentialProvider -
aws.accessKeyId
and
aws.secretKey
Java system properties. • WebIdentityTokenCredentialsProvider - this is the one you are expecting to use with AWS_ROLE_ARN and AWS_WEB_IDENTITY_TOKEN_FILE • ProfileCredentialsProvider - use default profile • EC2ContainerCredentialsProviderWrapper - • InstanceProfileCredentialsProvider I still don't understand what's going on in your scenario, but hopefully seeing the actual provider chain will help you figure it out. One other thought is that the
druid.indexer.runner.javaOptsArray
in the middle managers can change setting for the task's JVM, but that would be across the board on all MMs, so it still doesn't seem relevant.
d
Another idea, if there’s enough budget, build a brand new EKS, deploy Druid there, rerun the test, and see if the same thing happened there.
j
Didip, we are in the middle of doing what you suggested, we have severeal EKS cluster and we are working towards getting an EKS cluster just for my team and for the workload I need for Druid and other services, so hopefully I will soon find out if the error repeats by itself or it was some hiccup in AWS/EKS/EC2. Also, it looks like CloudTrail does not log every Kinesis API Call, so thats the reason we did not see the call erroring out https://docs.aws.amazon.com/streams/latest/dev/logging-using-cloudtrail.html