Hi I have ingested 116554 tables to datahub, the w...
# ingestion
m
Hi I have ingested 116554 tables to datahub, the web UI began to crash(waiting forever) when 10k tables were ingested. I am not sure what was going on.
datahub docker check
shows no issue The context is I am running
docker-compose.quickstart.yml
in my dev machine
l
We will get back to you soon on this
thank you 1
b
is there enough resources for elasticsearch? just wondering
m
@better-orange-49102 How to check that?
b
does
Copy code
curl -X GET <http://elasticsearch:9200/_cat/indices>
show any red indices?
m
(venv) zxu@dev-zxu:~/code$ curl -X GET http://elasticsearch:9200/_cat/indices curl: (6) Could not resolve host: elasticsearch
b
ah, maybe localhost:9200 from terminal then
m
(venv) zxu@dev-zxu:~/code$ curl -X GET http://localhost:9200/_cat/indices yellow open datajobindex_v2 uHp8QHEoS9Cjootv-rS43A 1 1 0 0 208b 208b yellow open dataset_datasetprofileaspect_v1 NNIxnkZTSaKleIN2CB15jg 1 1 0 0 208b 208b yellow open mlmodelgroupindex_v2 cYJABGi8Ttis0HJsAqKT8g 1 1 0 0 208b 208b yellow open mlmodelindex_v2 HPHXTT-4R8GtNsbSKWfJSw 1 1 0 0 208b 208b yellow open dataflowindex_v2 T_uXSqkuTrePNKZux9TZ6A 1 1 0 0 208b 208b yellow open datahubpolicyindex_v2 I8TSi1qtRsaNtWPQNb7gUw 1 1 0 0 208b 208b yellow open corpuserindex_v2 RYqafQmxRJiZJe4keLzUXw 1 1 1 0 5.3kb 5.3kb yellow open dataprocessindex_v2 KeCTP2qMRGGCoOaUVm010A 1 1 0 0 208b 208b yellow open chartindex_v2 e6_HRnmkTASvNZMgilAlGQ 1 1 0 0 208b 208b yellow open tagindex_v2 kdIhpNw4S2yUwG1XkXKNSw 1 1 1 0 17.8kb 17.8kb yellow open mlmodeldeploymentindex_v2 sWeST4wVQ8C9SWgaU_v3_A 1 1 0 0 208b 208b yellow open datajob_datahubingestioncheckpointaspect_v1 BNex0yZYRk6_5_A4vAjqQw 1 1 0 0 208b 208b yellow open dashboardindex_v2 zWLRfXjpST2hO00QIP9GIw 1 1 0 0 208b 208b yellow open .ds-datahub_usage_event-000001 Jh9tpT_sRDySCql77atN1g 1 1 74 0 168.1kb 168.1kb yellow open datasetindex_v2 Ic0mV5TDQi2MBOTU17GRgg 1 1 116554 332330 106.8mb 106.8mb yellow open mlfeatureindex_v2 7QXf2uZWSd2r8ow9E-rjiQ 1 1 0 0 208b 208b yellow open dataplatformindex_v2 E-t_J0FfSkG-J4aW1jS83Q 1 1 0 0 208b 208b yellow open datajob_datahubingestionrunsummaryaspect_v1 djrn2TCOSOiOVluEu0X9iQ 1 1 0 0 208b 208b yellow open glossarynodeindex_v2 jRQ7beghR9SkpacYFDpzpw 1 1 0 0 208b 208b yellow open datahubretentionindex_v2 lBNlx7VdSf-QwaojQSmfUw 1 1 0 0 208b 208b yellow open system_metadata_service_v1 C3XdbBEWSDyIYwnAjpQL5Q 1 1 699382 0 65.9mb 65.9mb yellow open schemafieldindex_v2 bV4OS18bQGOYa9bJA4uxJA 1 1 0 0 208b 208b yellow open mlfeaturetableindex_v2 EeBpkhwKR_uLpoLnNxBu3A 1 1 0 0 208b 208b yellow open glossarytermindex_v2 s1L2226kSbiUYeJ5Dsg0PA 1 1 0 0 208b 208b yellow open mlprimarykeyindex_v2 Ti9MD4PPR1K4W2-h4dGxZA 1 1 0 0 208b 208b yellow open corpgroupindex_v2 jywAD6mDQGe37wq53RweJw 1 1 0 0 208b 208b yellow open dataset_datasetusagestatisticsaspect_v1 45ISXlA0QuKJ2jq7W1g0cQ 1 1 0 0 208b 208b (venv) zxu@dev-zxu:~/code$
b
i guess it shd be fine then
thank you 1
s
How many tables do you have? You mentioned 116554 tables as well as 10k tables. Can you clarify exactly how many tables do you have? I can try to reproduce the issue. I have tried the web UI homepage with around 40k tables and there was some slowness but did not crash. Is this some different page or the homepage or all pages? Can you check the logs of the frontend pod and share them? Maybe there are some errors happening there
I ingested 100k datasets on my local and UI is working fine for me. Does your system have enough resources for running the quickstart?
Can you use this to find out how much memory do you have?
Copy code
docker info | grep Mem
g
Is it possible that you were trying to use the app while ingesting? I have found that running a large ingestion job on my local computer can cause the rest of my computer to slow down- if you were trying to use the app while also running a large ingestion job that could have contributed. Generally ingestion is run on a different machine than Datahub in production deployments.
👍 1
m
Thank you for your quick reply. This morning I tried the UI again. It works. So for now I am not able to repro it.
🙌 1
g
Great! Glad to hear things resolved themselves.
Let us know if the issue comes back & we can try to debug more
m
Sure. Thanks!
g
I know of Datahub deployments working with orders of magnitude larger scale than 10k datasets
so I would not expect it to be a problem!
m
docker info | grep Mem WARNING: No swap limit support Total Memory: 68.59GiB
g
ahhh
this is an older log ^ ?
or is that happening currently?
m
Currently
g
can you break that down by container?
docker stats
it would be helpful to know which container is being the memory hog
m
CONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM % NET I/O BLOCK I/O PIDS b7202ffcfc19 datahub-frontend-react 1.03% 443.2MiB / 68.59GiB 0.63% 2.7MB / 13.4MB 0B / 12.3kB 114 dc9b1a9ba0ad datahub-gms 0.89% 1.247GiB / 68.59GiB 1.82% 3.78GB / 4.1GB 0B / 91.6MB 186 6363c80910c9 schema-registry 0.19% 614.2MiB / 68.59GiB 0.87% 26.5MB / 25.4MB 0B / 205kB 107 5fbf1943ed04 broker 1.58% 835.3MiB / 68.59GiB 1.19% 768MB / 793MB 0B / 70.9MB 138 d763ba3030a0 neo4j 0.62% 1.468GiB / 68.59GiB 2.14% 26.9kB / 19.2kB 0B / 120MB 103 3d16bfaea5c7 mysql 0.10% 1.029GiB / 68.59GiB 1.50% 2.01GB / 2.06GB 0B / 6.74GB 38 39c81aa6878a zookeeper 0.15% 226.1MiB / 68.59GiB 0.32% 5.44MB / 3.69MB 0B / 1.02MB 183 0fd24b86c59d elasticsearch 0.37% 797.3MiB / 1GiB 77.86% 1.32GB / 640MB 81.6MB / 18.1GB 204 0080fa47211f beacon 0.11% 457.1MiB / 1GiB 44.64% 0B / 0B 66.9MB / 57.4GB 49 a2aab75971e9 envoy-A 11.52% 768.1MiB / 2GiB 37.50% 0B / 0B 69.4MB / 54.9MB 19 696c5133bccd envoy-endpoint-register 0.32% 93.11MiB / 128MiB 72.74% 0B / 0B 22MB / 4.1kB 6 6f239f3ffbac pastis_updater_envoy 0.00% 76.58MiB / 68.59GiB 0.11% 0B / 0B 21MB / 243MB 28 1428f9608f7d envoy-sds 0.00% 99.17MiB / 128MiB 77.48% 0B / 0B 24.4MB / 0B 38 769e137a5a21 singer 0.24% 1.024GiB / 8GiB 12.80% 0B / 0B 94.7MB / 125MB 98 082202b15cf1 ratelimit-A 10.65% 110.8MiB / 1GiB 10.82% 0B / 0B 39.4MB / 54.8MB 11 2d0cbfb10f92 mcrouter 0.03% 110.2MiB / 68.59GiB 0.16% 0B / 0B 45.2MB / 893MB 8 15f1e23ad485 tcollector 0.09% 206.7MiB / 68.59GiB 0.29% 0B / 0B 47MB / 56.6MB 8 3574d2038b39 pastis_proxy 0.02% 82.89MiB / 1GiB 8.10% 0B / 0B 25.5MB / 0B 44 48fcd88f9624 pastis_updater 0.01% 45.13MiB / 1GiB 4.41% 0B / 0B 11.6MB / 0B 26 429232c02d81 metrics-agent 0.72% 124.1MiB / 68.59GiB 0.18% 0B / 0B 45.5MB / 0B 179 f9c21bfb18f1 zk_update_monitor 0.00% 341.7MiB / 68.59GiB 0.49% 0B / 0B 131kB / 63.3GB 6 FYI this morning I checked the status
Copy code
datahub docker check
The following issues were detected:
- elasticsearch-setup container is not present
- kafka-setup container is not present
Then I ran
Copy code
datahub docker quickstart --quickstart-compose-file ./docker/quickstart/docker-compose.quickstart.yml
again.
Copy code
The following issues were detected:
- kafka-setup container is not present
- elasticsearch-setup container is not present
This happens again
Currently in pinterest-hive I have 116.6k tables
Where to check the logs?
g
have you tried running just w/
datahub docker quickstart
? without specifying a compose file?
but for errors- yes!
you can do
docker logs --follow datahub-gms
that will give you a live feed of the metadata service’s logs
run that command & refresh the page & let me know what prints out!
btw- it doesn’t look like your datahub containers are taking up too much memory
it might be that you’ve given too many resources to docker as a whole
if you go to docker desktop > settings > resources, how many GB of memory have you allocated to docker?
m
have you tried running just w/ 
datahub docker quickstart
? without specifying a compose file?
The reason I did not do that is I have a port conflict with 8081. I have to modify the compose file
you can do 
docker logs --follow datahub-gms
Looks like keep poping up a lot of message with old timestamp
Copy code
03:44:38.160 [I/O dispatcher 1] INFO  c.l.m.s.e.update.BulkListener:28 - Successfully fed bulk request. Number of events: 1 Took time ms: -1
03:44:38.219 [qtp544724190-58] INFO  c.l.metadata.entity.EntityService:549 - INGEST urn urn:li:dataset:(urn:li:dataPlatform:pinterest-hive,weiranliu.no_sa_active_users_base,PROD) with system metadata {lastObserved=1640231078216, runId=pinterest-hive-2021_12_22-23_45_58}
03:44:38.224 [mae-consumer-job-client-0-C-1] INFO  c.l.m.k.MetadataAuditEventsProcessor:101 - {urn=urn:li:dataset:(urn:li:dataPlatform:pinterest-hive,weiranliu.no_sa_active_users_base,PROD), aspects=[{com.linkedin.metadata.key.DatasetKey={name=weiranliu.no_sa_active_users_base, platform=urn:li:dataPlatform:pinterest-hive, origin=PROD}}, {com.linkedin.common.Status={removed=false}}]}
03:44:38.229 [I/O dispatcher 1] INFO  c.l.m.s.e.update.BulkListener:28 - Successfully fed bulk request. Number of events: 1 Took time ms: -1
03:44:38.229 [mae-consumer-job-client-0-C-1] INFO  c.l.m.k.MetadataAuditEventsProcessor:101 - {urn=urn:li:dataset:(urn:li:dataPlatform:pinterest-hive,weiranl
if you go to docker desktop > settings > resources, how many GB of memory have you allocated to docker?
I am running datahub in my devserver in ubuntu
Running
docker logs --follow datahub-gms --since 2021-12-23T16:33:02Z
g
oh, what if you just did
docker logs --follow datahub-gms
without the since? and then refresh?
m
g
I am running datahub in my devserver in ubuntu
got it- what are the resources allocated to the machine overall?
I see the error message,
Copy code
Caused by: org.elasticsearch.ElasticsearchException: Elasticsearch exception [type=too_many_buckets_exception, reason=Trying to create too many buckets. Must be less than or equal to: [65535] but was [65536]. This limit can be set by changing the [search.max_buckets] cluster level setting.]
what version of datahub are you running?
this was an error we saw earlier in the case that a user had many many browse paths. however this should have been fixed in a later version
it also exists only for browse- can you verify search works?
m
what are the resources allocated to the machine overall?
Copy code
vmstat -s
     71926472 K total memory
     13810968 K used memory
     21901644 K active memory
     12815148 K inactive memory
     31164780 K free memory
      1190164 K buffer memory
     25760560 K swap cache
            0 K total swap
            0 K used swap
            0 K free swap
     22794417 non-nice user cpu ticks
      3286348 nice user cpu ticks
      9313976 system cpu ticks
   1471659686 idle cpu ticks
       169654 IO-wait cpu ticks
            0 IRQ cpu ticks
      1194434 softirq cpu ticks
        23468 stolen cpu ticks
      5761549 pages paged in
    386317985 pages paged out
            0 pages swapped in
            0 pages swapped out
   1108917083 interrupts
   2096488378 CPU context switches
   1639858236 boot time
      5937327 forks
what version of datahub are you running?
I git pull the repo yesterday. I am developing my ingestion source so the DataHub CLI version is unavailable
Yes. The search works
g
Ok--- maybe that is it!
so, there generally should only be a browse path for each dataset container
👍 1
i wonder if you are creating more browse paths w/ your ingestion source
than other ingestion sources 🤔
and thats causing scale issues for the browse endpoint
are you generating custom browse path aspects for your custom source?
or are you using default browse logic?
m
are you generating custom browse path aspects for your custom source?
No
g
ok- can you share some urns that are being generated for you custom source?
g
ok! and then another one?
g
hmm ok that seems fine
is it always
<namespace>.<table_name>
?
m
yes!
g
how many different namespaces do you have?
m
one moment
g
also, datahub offers a hive ingestion source out of the box.
are you on an incompatible hive version?
if you just are making a custom source to extend properties like tags, ownership, dataset name, etc you can also use the out of the box hive source w/ transformers to enrich the data
that way you don’t have to maintain a fork of the hive connector
(but that’s an unrelated concern)
m
are you on an incompatible hive version?
Yes. Hive 1.2 that is incompatible
✔️ 1
how many different namespaces do you have?
2133
g
oh
that’s not that many
hmm 😕
so in your video you get the error when visiting pinterest-hive path
m
Right
g
what happens if you go to pinterest-hive/default instead?
does that work?
m
That works
g
which is the largest namespace?
w/ the most datasets?
m
one moment
The lagest namespace has 15254, and web UI is fine
g
gotcha.. hmm
that also seems fine
m
Maybe we can follow up this later, browse namespaces is not in our major user workflow. I can also debug it.
g
ok sounds good!
I’ll look into this separately. I’m still surprised you’re seeing this issue given the numbers of entities
m
Thank you so much for your help!
teamwork 1
b
so, there generally should only be a browse path for each dataset container
hey @green-football-43791, can I check if this is the case in latest datahub version? i still see that the browse path is still 1:1 bound to dataset name
For context, we also encountered too many buckets error in es as we have >89K datasets in one Hive source 😄
l
Thanks for the context @boundless-student-48844 - we will take a look at this
b
Thanks @loud-island-88694! For now, we just increased the max_buckets size in ES (took some time to justify that to our managed ES team 😛). Do let us know if there’s any potential fix on browser path. We can then contribute a PR for that 🙇 cc @victorious-truck-11787
b
Browse paths strikes again!
l
We are revisiting how browse works. @big-carpet-38439 will be in touch soon.
👍 1
😱 1
m
I am having the same problem, datahub-gms log says
Copy code
Caused by: org.elasticsearch.ElasticsearchException: Elasticsearch exception [type=too_many_buckets_exception, reason=Trying to create too many buckets. Must be less than or equal to: [65535] but was [65536]. This limit can be set by changing the [search.max_buckets] cluster level setting.]
b
Yes this is because of a ton of assets inside a browse path is my guess. We are currently working on migrating existing browse experience to one based on search filters
This should help resolve most of these issues
I'd say eta is next 2-3 weeks if all goes well
m
Thank you so much
cc @curved-librarian-24314
h
Sorry to resurrect an old thread but we are seeing the issue discussed in this thread on a datahub instance that is loaded with about 40 million metadata items. When browsing the datasets, we get this error:
Copy code
13:11:56.109 [Thread-143] ERROR c.l.m.s.e.query.ESBrowseDAO:126 - Browse query failed: Elasticsearch exception [type=search_phase_execution_exception, reason=]
13:11:56.109 [Thread-143] ERROR c.l.d.g.r.browse.BrowseResolver:60 - Failed to execute browse: entity type: DATASET, path: [prod, postgres, kafka_alias101], filters: null, start: 0, count: 10 Browse query failed:
13:11:56.113 [Thread-143] ERROR c.l.d.g.e.DataHubDataFetcherExceptionHandler:21 - Failed to execute DataFetcher
java.util.concurrent.CompletionException: java.lang.RuntimeException: Failed to execute browse: entity type: DATASET, path: [prod, postgres, kafka_alias101], filters: null, start: 0, count: 10
        at java.util.concurrent.CompletableFuture.encodeThrowable(CompletableFuture.java:273)
        at java.util.concurrent.CompletableFuture.completeThrowable(CompletableFuture.java:280)
        at java.util.concurrent.CompletableFuture$AsyncSupply.run(CompletableFuture.java:1606)
        at java.lang.Thread.run(Thread.java:748)
Caused by: java.lang.RuntimeException: Failed to execute browse: entity type: DATASET, path: [prod, postgres, kafka_alias101], filters: null, start: 0, count: 10
        at com.linkedin.datahub.graphql.resolvers.browse.BrowseResolver.lambda$get$1(BrowseResolver.java:68)
        at java.util.concurrent.CompletableFuture$AsyncSupply.run(CompletableFuture.java:1604)
        ... 1 common frames omitted
Caused by: com.datahub.util.exception.ESQueryException: Browse query failed:
        at com.linkedin.metadata.search.elasticsearch.query.ESBrowseDAO.browse(ESBrowseDAO.java:127)
        at com.linkedin.metadata.search.elasticsearch.ElasticSearchService.browse(ElasticSearchService.java:105)
        at com.linkedin.entity.client.JavaEntityClient.browse(JavaEntityClient.java:175)
        at com.linkedin.datahub.graphql.types.dataset.DatasetType.browse(DatasetType.java:170)
        at com.linkedin.datahub.graphql.resolvers.browse.BrowseResolver.lambda$get$1(BrowseResolver.java:52)
        ... 2 common frames omitted
Caused by: org.elasticsearch.ElasticsearchStatusException: Elasticsearch exception [type=search_phase_execution_exception, reason=]
        at org.elasticsearch.rest.BytesRestResponse.errorFromXContent(BytesRestResponse.java:187)
        at org.elasticsearch.client.RestHighLevelClient.parseEntity(RestHighLevelClient.java:1892)
        at org.elasticsearch.client.RestHighLevelClient.parseResponseException(RestHighLevelClient.java:1869)
        at org.elasticsearch.client.RestHighLevelClient.internalPerformRequest(RestHighLevelClient.java:1626)
        at org.elasticsearch.client.RestHighLevelClient.performRequest(RestHighLevelClient.java:1583)
        at org.elasticsearch.client.RestHighLevelClient.performRequestAndParseEntity(RestHighLevelClient.java:1553)
        at org.elasticsearch.client.RestHighLevelClient.search(RestHighLevelClient.java:1069)
        at com.linkedin.metadata.search.elasticsearch.query.ESBrowseDAO.browse(ESBrowseDAO.java:96)
        ... 6 common frames omitted
        Suppressed: org.elasticsearch.client.ResponseException: method [POST], host [<https://vpc-datahubpoc2-3y2kiko4mdy2r6avw2syw67fby.us-east-1.es.amazonaws.com:443>], URI [/datasetindex_v2/_search?typed_keys=true&max_concurrent_shard_requests=5&ignore_unavailable=false&expand_wildcards=open&allow_no_indices=true&ignore_throttled=true&search_type=query_then_fetch&batched_reduce_size=512&ccs_minimize_roundtrips=true], status line [HTTP/1.1 503 Service Unavailable]
{"error":{"root_cause":[],"type":"search_phase_execution_exception","reason":"","phase":"fetch","grouped":true,"failed_shards":[],"caused_by":{"type":"too_many_buckets_exception","reason":"Trying to create too many buckets. Must be less than or equal to: [65535] but was [65536]. This limit can be set by changing the [search.max_buckets] cluster level setting.","max_buckets":65535}},"status":503}
                at org.elasticsearch.client.RestClient.convertResponse(RestClient.java:302)
                at org.elasticsearch.client.RestClient.performRequest(RestClient.java:272)
                at org.elasticsearch.client.RestClient.performRequest(RestClient.java:246)
                at org.elasticsearch.client.RestHighLevelClient.internalPerformRequest(RestHighLevelClient.java:1613)
                ... 10 common frames omitted
Caused by: org.elasticsearch.ElasticsearchException: Elasticsearch exception [type=too_many_buckets_exception, reason=Trying to create too many buckets. Must be less than or equal to: [65535] but was [65536]. This limit can be set by changing the [search.max_buckets] cluster level setting.]
        at org.elasticsearch.ElasticsearchException.innerFromXContent(ElasticsearchException.java:496)
        at org.elasticsearch.ElasticsearchException.fromXContent(ElasticsearchException.java:407)
        at org.elasticsearch.ElasticsearchException.innerFromXContent(ElasticsearchException.java:437)
        at org.elasticsearch.ElasticsearchException.failureFromXContent(ElasticsearchException.java:603)
        at org.elasticsearch.rest.BytesRestResponse.errorFromXContent(BytesRestResponse.java:179)
        ... 13 common frames omitted
Full stack trace is in the attached file. As the error message says, we could increase the number of buckets in elastic search to higher than 65535 but 65535 already sounds like a lot of buckets. @big-carpet-38439 Is this related to the issue you were fixing above? I assume that by now the changes you mentioned above are in the product.
s
Request you to please add the version numbers that you are using for everything. That should help once someone can look at this
h
We are on Datahub version 0.8.39 and AWS opensearch 1.2
b
Hey @big-carpet-38439, do you have any update on this? We are also having the same problem and our default bucket limit (10K) is even lower than the default. We face it on the following flow: DataHub UI -> Clicked Datasets -> Click DataPlatform more than 10K dataset -> 500 error ❌ I’ve attached an image shows exactly where we get an error and the related logs:
Copy code
06:23:57.918 [pool-10-thread-1] INFO  c.l.m.filter.RestliLoggingFilter:55 - GET /entitiesV2?ids=List(urn%3Ali%3Acorpuser%3Asalih.can@udemy.com) - batchGet - 200 - 4ms
06:23:57.944 [I/O dispatcher 1] INFO  c.l.m.k.e.ElasticsearchConnector:41 - Successfully feeded bulk request. Number of events: 1 Took time ms: -1
	... 27 common frames omitted
06:23:57.856 [Thread-25115] ERROR c.datahub.graphql.GraphQLController:98 - Errors while executing graphQL query: "query getMe {\n  me {\n    corpUser {\n      urn\n      username\n      info {\n        active\n        displayName\n        title\n        firstName\n        lastName\n        fullName\n        email\n        __typename\n      }\n      editableProperties {\n        displayName\n        title\n        pictureLink\n        teams\n        skills\n        __typename\n      }\n      __typename\n    }\n    platformPrivileges {\n      viewAnalytics\n      managePolicies\n      manageIdentities\n      generatePersonalAccessTokens\n      manageIngestion\n      manageSecrets\n      manageDomains\n      manageTests\n      manageGlossaries\n      manageUserCredentials\n      __typename\n    }\n    __typename\n  }\n}\n", result: {errors=[{message=An unknown error occurred., locations=[{line=2, column=3}], path=[me], extensions={code=500, type=SERVER_ERROR, classification=DataFetchingException}}], data={me=null}, extensions={tracing={version=1, startTime=2022-07-18T06:23:57.826Z, endTime=2022-07-18T06:23:57.856Z, duration=29691126, parsing={startOffset=236621, duration=217511}, validation={startOffset=384372, duration=138868}, execution={resolvers=[{path=[me], parentType=Query, returnType=AuthenticatedUser, fieldName=me, startOffset=418447, duration=28545139}]}}}}, errors: [DataHubGraphQLError{path=[me], code=SERVER_ERROR, locations=[SourceLocation{line=2, column=3}]}]
06:23:58.437 [Thread-29132] ERROR c.l.m.s.e.query.ESBrowseDAO:126 - Browse query failed: Elasticsearch exception [type=search_phase_execution_exception, reason=all shards failed]
06:23:58.437 [Thread-29132] ERROR c.l.d.g.r.browse.BrowseResolver:60 - Failed to execute browse: entity type: DATASET, path: [prod, hive], filters: null, start: 0, count: 10 Browse query failed:
06:23:58.437 [Thread-29132] ERROR c.l.d.g.e.DataHubDataFetcherExceptionHandler:21 - Failed to execute DataFetcher
java.util.concurrent.CompletionException: java.lang.RuntimeException: Failed to execute browse: entity type: DATASET, path: [prod, hive], filters: null, start: 0, count: 10
	at java.util.concurrent.CompletableFuture.encodeThrowable(CompletableFuture.java:273)
	at java.util.concurrent.CompletableFuture.completeThrowable(CompletableFuture.java:280)
	at java.util.concurrent.CompletableFuture$AsyncSupply.run(CompletableFuture.java:1606)
	at java.lang.Thread.run(Thread.java:748)
Caused by: java.lang.RuntimeException: Failed to execute browse: entity type: DATASET, path: [prod, hive], filters: null, start: 0, count: 10
	at com.linkedin.datahub.graphql.resolvers.browse.BrowseResolver.lambda$get$1(BrowseResolver.java:68)
	at java.util.concurrent.CompletableFuture$AsyncSupply.run(CompletableFuture.java:1604)
	... 1 common frames omitted
Caused by: com.datahub.util.exception.ESQueryException: Browse query failed:
	at com.linkedin.metadata.search.elasticsearch.query.ESBrowseDAO.browse(ESBrowseDAO.java:127)
	at com.linkedin.metadata.search.elasticsearch.ElasticSearchService.browse(ElasticSearchService.java:105)
	at com.linkedin.entity.client.JavaEntityClient.browse(JavaEntityClient.java:175)
	at com.linkedin.datahub.graphql.types.dataset.DatasetType.browse(DatasetType.java:170)
	at com.linkedin.datahub.graphql.resolvers.browse.BrowseResolver.lambda$get$1(BrowseResolver.java:52)
	... 2 common frames omitted
Caused by: org.elasticsearch.ElasticsearchStatusException: Elasticsearch exception [type=search_phase_execution_exception, reason=all shards failed]
	at org.elasticsearch.rest.BytesRestResponse.errorFromXContent(BytesRestResponse.java:187)
	at org.elasticsearch.client.RestHighLevelClient.parseEntity(RestHighLevelClient.java:1892)
	at org.elasticsearch.client.RestHighLevelClient.parseResponseException(RestHighLevelClient.java:1869)
	at org.elasticsearch.client.RestHighLevelClient.internalPerformRequest(RestHighLevelClient.java:1626)
	at org.elasticsearch.client.RestHighLevelClient.performRequest(RestHighLevelClient.java:1583)
	at org.elasticsearch.client.RestHighLevelClient.performRequestAndParseEntity(RestHighLevelClient.java:1553)
	at org.elasticsearch.client.RestHighLevelClient.search(RestHighLevelClient.java:1069)
	at com.linkedin.metadata.search.elasticsearch.query.ESBrowseDAO.browse(ESBrowseDAO.java:96)
	... 6 common frames omitted
	Suppressed: org.elasticsearch.client.ResponseException: method [POST], host [<https://es.udemy.com:9999>], URI [/datahub_datasetindex_v2/_search?typed_keys=true&max_concurrent_shard_requests=5&ignore_unavailable=false&expand_wildcards=open&allow_no_indices=true&ignore_throttled=true&search_type=query_then_fetch&batched_reduce_size=512&ccs_minimize_roundtrips=true], status line [HTTP/1.1 503 Service Unavailable]
{"error":{"root_cause":[{"type":"too_many_buckets_exception","reason":"Trying to create too many buckets. Must be less than or equal to: [10000] but was [10001]. This limit can be set by changing the [search.max_buckets] cluster level setting.","max_buckets":10000}],"type":"search_phase_execution_exception","reason":"all shards failed","phase":"query","grouped":true,"failed_shards":[{"shard":0,"index":"datahub_datasetindex_v2_1656325420723","node":"lvOhAjVGQh-HO97e1HpyDg","reason":{"type":"too_many_buckets_exception","reason":"Trying to create too many buckets. Must be less than or equal to: [10000] but was [10001]. This limit can be set by changing the [search.max_buckets] cluster level setting.","max_buckets":10000}}]},"status":503}
		at org.elasticsearch.client.RestClient.convertResponse(RestClient.java:302)
		at org.elasticsearch.client.RestClient.performRequest(RestClient.java:272)
		at org.elasticsearch.client.RestClient.performRequest(RestClient.java:246)
		at org.elasticsearch.client.RestHighLevelClient.internalPerformRequest(RestHighLevelClient.java:1613)
		... 10 common frames omitted
I believe this is not related to specific DataHub version but we use
0.8.38
and ES version
7.8.1
. Maybe we need to change the way we query ES through GraphQL, seems like the queries are not so efficient.
h
If anyone is still waiting on a fix for this, we are not seeing this issue after upgrading to DataHub version 0.9.0.
b
Yes - we've recently made our browse queries more efficient which should help resolve this for many cases
However, you should indeed be able to bump the
max_buckets
for your elastic instance, we've also seen this work!