Hello, I have been tasked to debug errors message...
# all-things-deployment
f
Hello, I have been tasked to debug errors messages appearing on our datahub deployment on kubernetes using the helm chart provided by acryldata bumped chart version from
0.2.181
to
0.2.182
We are ingesting metadata from Databricks, Sagemaker, dbt. When displaying the logs of these 3 jobs, I get recurring similar errors being : • datahub-gms return status 500 (see picture) • Error registering Avro Schema Even after trying the solution propsed in this issue the problem keeps recurring. Thank you for the help!
a
@dazzling-yak-93039 might be able to speak to this!
d
Hi, what schema registry are you using?
f
I think it is the provided by the helm charts version
0.2.182
a
There is a known issue with internal and glue schema registries in the last two versions of datahub, which could cause this issue, can you make sure you are using the confluent schema registry?
f
so I need to change the schema registry image in the values.yaml file? https://github.com/acryldata/datahub-helm/blob/master/charts/datahub/values.yaml
lines 487 to 498
d
Yes! That's right. Sorry, I think we had a different chart version that changed the default back to what works, but it looks like it's not this one.
f
so there need to be an update of the prerquisites helm charts?
d
Yeah, you may need to install the schema registry which will be used.
f
image.png
this is what I currently have in my values.yaml
d
This looks good to me
f
It is with this config that I get the error describes 😕
d
Did you make the change in values.yaml as well?
f
My bad, I meant this is the current content of my values.yaml file
So it is currently failing with these values, so what do you suggest to replace them with?
a
I see I see. Actually, if GMS is coming up, then I think you have that part configured correctly. @famous-waitress-64616 can you weigh in on the ingestion failures?
f
🙏
f
Are you running ingestion from the UI? And are you sure datahub-datahub-gms:8080 is the correct url from inside your cluster? You can try
curl datahub-datahub-gms:8080/health
to check
f
Thank you @mammoth-zebra-81887 While waiting to be merged and released, is there a way to change the value of this environment variable through the values.yaml? I want to test this on our dev environement
Here is what I ended up doing in my helm releases, it is inspired by the PR. • pre-reqs values yaml file :
Copy code
cp-helm-charts:
  enabled: true
  cp-schema-registry:
    enabled: true
    kafka:
      bootstrapServers: "datahub-prerequisites-kafka:9092"
• values.yaml :
Copy code
global:
  kafka:
    bootstrap:
      server: "datahub-prerequisites-kafka:9092"
    schemaregistry:
      type: KAFKA
      url: "<http://datahub-prerequisites-cp-schema-registry:8081/>"

# ADDED
datahubSystemUpdate:
  extraEnvs:
    - name: "KAFKA_PROPERTIES_AUTO_REGISTER_SCHEMAS"
      value: "true"
d
Sweet, thanks for sharing!
m
While waiting to be merged and released, is there a way to change the value of this environment variable through the values.yaml? I want to test this on our dev environement
I think there is no way through the values.yaml 🤦‍♂️ You can patch the environment variable, though. (If the
INTERNAL
schema registry is unstable, it might be better to use the
cp-schema-registry
)
f
I see, Got help from a colleague experienced with helm, he came to the same conclusion. After investigating the chart, we discovered the
extraEnvs
block as well as the environment use when schema Registry type is KAFKA. that solution has stabilized our dev deployment of datahub, will trying PROD next week
m