Hey folks, I have executed `Redshift-usage` inges...
# ingestion
p
Hey folks, I have executed
Redshift-usage
ingestion and after that operation
Queries
tab is enabled. However, I’m not able to see any query in it. When I check the documents of
dataset_datasetusagestatisticsaspect_v1
index at ElasticSearch, I’ve validated that there are SQL queries in
topSqlQueries
fields.
Copy code
"topSqlQueries": [ - 
              "/* ' Query generated by Chartio {\"reason\":\"dashboard_refresh_data\"} */ <REDACTED_SQL_COMMAND>",
               "-- Looker Query Context '{\"user_id\":111,\"history_id\":22222,\"instance_slug\":\"asdqwezxc\"}' WITH company_wide_cdf AS (/* Primary owner: Seref */ <REDACTED_SQL_COMMAND>"
               ]
Could you please help me about this issue?
b
@miniature-tiger-96062 Are you able to take a look?
m
Hi @polite-flower-25924, I will try to reproduce and take a look. The issue we faced earlier was because there was a default limit on the length of the field that stored top queries, which was solved as part of the fix.
@polite-flower-25924, just confirming if you were running the updated datahub version ?
m
@polite-flower-25924: We just released 0.8.16 which should have fixes for this issue (as well as a host of other features!)
p
@miniature-tiger-96062 it is 0.8.14 🙈
let me check with the latest version.
m
👍
p
Copy code
Caused by:
org.elasticsearch.client.ResponseException: method [POST], host [<http://elasticsearch-master:9200>], URI [/_reindex?scroll=5m&slices=1&requests_per_second=-1&wait_for_completion=true&timeout=1m], status line [HTTP/1.1 400 Bad Request]|{"took":1107,"timed_out":false,"total":3410,"updated":0,"created":994,"deleted":0,"batches":1,"version_conflicts":0,"noops":0,"retries":{"bulk":0,"search":0},"throttled_millis":0,"requests_per_second":-1.0,"throttled_until_millis":0,"failures":[{"index":"dataset_datasetusagestatisticsaspect_v1_1634882071332","type":"_doc","id":"0d79b2909ccc3fc97a98d68f63ac6831","cause":{"type":"illegal_argument_exception","reason":"Document contains at least one immense term in field=\"topSqlQueries\" (whose UTF8 encoding is longer than the max length 32766), all of which were skipped.  Please correct the analyzer to not produce such terms.  The prefix of the first immense term is: '[91, 34, 47, 42, 32, 39, 32, 81, 117, 101, 114, 121, 32, 103, 101, 110, 101, 114, 97, 116, 101, 100, 32, 98, 121, 32, 67, 104, 97, 114]...', original message: bytes can be at most 32766 in length; got 32900","caused_by":{"type":"max_bytes_length_exceeded_exception","reason":"max_bytes_length_exceeded_exception: bytes can be at most 32766 in length; got 32900"}},"status":400},{"index":"dataset_datasetusagestatisticsaspect_v1_1634882071332","type":"_doc","id":"a50804d8beaac5b9e743fba4fb3353d5","cause":{"type":"illegal_argument_exception","reason":"Document contains at least one immense term in field=\"topSqlQueries\" (whose UTF8 encoding is longer than the max length 32766), all of which were skipped.  Please correct the analyzer to not produce such terms.  The prefix of the first immense term is: '[91, 34, 115, 101, 108, 101, 99, 116, 32, 100, 105, 115, 116, 105, 110, 99, 116, 32, 112, 114, 95, 49, 46, 99, 111, 117, 114, 115, 101, 105]...', original message: bytes can be at most 32766 in length; got 34226","caused_by":{"type":"max_bytes_length_exceeded_exception","reason":"max_bytes_length_exceeded_exception: bytes can be at most 32766 in length; got 34226"}},"status":400},{"index":"dataset_datasetusagestatisticsaspect_v1_1634882071332","type":"_doc","id":"2dd4b33e0871a1ca67760624525758e7","cause":{"type":"illegal_argument_exception","reason":"Document contains at least one immense term in field=\"topSqlQueries\" (whose UTF8 encoding is longer than the max length 32766), all of which were skipped.  Please correct the analyzer to not produce such terms.  The prefix of the first immense term is: '[91, 34, 112, 97, 100, 98, 95, 102, 101, 116, 99, 104, 95, 115, 97, 109, 112, 108, 101, 58, 32, 115, 101, 108, 101, 99, 116, 32, 42, 32]...', original message: bytes can be at most 32766 in length; got 35563","caused_by":{"type":"max_bytes_length_exceeded_exception","reason":"max_bytes_length_exceeded_exception: bytes can be at most 32766 in length; got 35563"}},"status":400},{"index":"dataset_datasetusagestatisticsaspect_v1_1634882071332","type":"_doc","id":"fe4f1ce1fe093bd9897db999501b1989","cause":{"type":"illegal_argument_exception","reason":"Document contains at least one immense term in field=\"topSqlQueries\" (whose UTF8 encoding is longer than the max length 32766), all of which were skipped.  Please correct the analyzer to not produce such terms.  The prefix of the first immense term is: '[91, 34, 119, 105, 116, 104, 32, 115, 117, 109, 109, 97, 114, 121, 32, 97, 115, 32, 40, 32, 115, 101, 108, 101, 99, 116, 32, 108, 100, 97]...', original message: bytes can be at most 32766 in length; got 36921","caused_by":{"type":"max_bytes_length_exceeded_exception","reason":"max_bytes_length_exceeded_exception: bytes can be at most 32766 in length; got 36921"}},"status":400},{"index":"dataset_datasetusagestatisticsaspect_v1_1634882071332","type":"_doc","id":"29beaeee8c233a99a885ea7ad411d1a0","cause":{"type":"illegal_argument_exception","reason":"Document contains at least one immense term in field=\"topSqlQueries\" (whose UTF8 encoding is longer than the max length 32766), all of which were skipped.  Please correct the analyzer to not produce such terms.  The prefix of the first immense term is: '[91, 34, 47, 42, 32, 39, 32, 81, 117, 101, 114, 121, 32, 103, 101, 110, 101, 114, 97, 116, 101, 100, 32, 98, 121, 32, 67, 104, 97, 114]...', original message: bytes can be at most 32766 in length; got 38281","caused_by":{"type":"max_bytes_length_exceeded_exception","reason":"max_bytes_length_exceeded_exception: bytes can be at most 32766 in length; got 38281"}},"status":400},{"index":"dataset_datasetusagestatisticsaspect_v1_1634882071332","type":"_doc","id":"b0879f92dfadeb4166011b2013d993f3","cause":{"type":"illegal_argument_exception","reason":"Document contains at least one immense term in field=\"topSqlQueries\" (whose UTF8 encoding is longer than the max length 32766), all of which were skipped.  Please correct the analyzer to not produce such terms.  The prefix of the first immense term is: '[91, 34, 45, 45, 32, 76, 111, 111, 107, 101, 114, 32, 81, 117, 101, 114, 121, 32, 67, 111, 110, 116, 101, 120, 116, 32, 39, 123, 92, 34]...', original message: bytes can be at most 32766 in length; got 34431","caused_by":{"type":"max_bytes_length_exceeded_exception","reason":"max_bytes_length_exceeded_exception: bytes can be at most 32766 in length; got 34431"}},"status":400}]}
^ when I upgraded the datahub version to 0.8.16, I faced the message above 😕
It seems to be an ElasticSearch Issue. Even though I’m not an expert of ElasticSearch, I need to disable index and change the type to
text
according to this post: https://discuss.elastic.co/t/bytes-can-be-at-most-32766-in-length/213650/7 ?
Hmm, I think I need to set “index” false
Copy code
"topSqlQueries": { - 
          "type": "text",
          "fields": { - 
            "keyword": { - 
              "type": "keyword",
              "ignore_above": 256
            }
          }
        },
b
Thanks for the info @polite-flower-25924 - we need to look into this one a bit more
p
Thank you @big-carpet-38439. I will also look at this. According to my last investigation, reindex operation is executed from the source index
dataset_datasetusagestatisticsaspect_v1
to a temporary index. This temporary index mapping is not aligned with the source index.
The related stack trace:
Copy code
org.elasticsearch.client.RestHighLevelClient.parseResponseException(RestHighLevelClient.java:1872)
    at org.elasticsearch.client.RestHighLevelClient.internalPerformRequest(RestHighLevelClient.java:1626)
    at org.elasticsearch.client.RestHighLevelClient.performRequest(RestHighLevelClient.java:1583)
    at org.elasticsearch.client.RestHighLevelClient.performRequestAndParseEntity(RestHighLevelClient.java:1553)
    at org.elasticsearch.client.RestHighLevelClient.reindex(RestHighLevelClient.java:557)
    at com.linkedin.metadata.search.elasticsearch.indexbuilder.IndexBuilder.buildIndex(IndexBuilder.java:82)
    at com.linkedin.metadata.timeseries.elastic.indexbuilder.TimeseriesAspectIndexBuilders.buildAll(TimeseriesAspectIndexBuilders.java:29)
    at com.linkedin.metadata.timeseries.elastic.ElasticSearchTimeseriesAspectService.configure(ElasticSearchTimeseriesAspectService.java:114)
    at com.linkedin.metadata.kafka.MetadataChangeLogProcessor.<init>(MetadataChangeLogProcessor.java:81)
I have to delete the
dataset_datasetusagestatisticsaspect_v1
index. Then, I trigger a Redshift usage ingestion job again and everything seems fine. I’m not sure that if it’s a breaking change or not 🙂
👍 1
r
Hey there! 👋 Make sure your message includes the following information if relevant, so we can help more effectively! 1. Are you using UI or CLI for ingestion? 2. Which DataHub version are you using? (e.g. 0.12.0) 3. What data source(s) are you integrating with DataHub? (e.g. BigQuery)
Oh man. Glad you were able to get this resolved...