And now ... some more challenges! I have the helm...
# replication-ask-ai
s
And now ... some more challenges! I have the helm chart running with one major exception. Airbyte-temporal will not start, complaining that the database already exists. Something seems to happen that leaves the database in a semi-useable stage as the `airbyte/temporal-auto-setup:1.13.0`pod fails to initialize correctly or something. Anyone had to deal with this?
k
A message from kapa.ai
New message text here
j
This is a repeat issue many people seem to be having: https://airbytehq.slack.com/archives/C01AHCD885S/p1689780217867399 Would love if the team could address this.
🙏 1
m
@Conor Barber (Airbyte) do you have any idea about this problem?
🙏 1
s
Help would be appreciated here. I'm on day 3 of trying to make the Helm chart produce a usable installation 😉
@Conor Barber (Airbyte) 🙂
@Marcos Marx (Airbyte) - he is not a part of this channel. Anyone else that could possibly help?
c
Hi Stefán! Taking a look now
🙏 1
Hi Stefán, we are trying to reproduce and see if it’s correlated with the issue that Jason linked. Could you provide the exact error message you get so we can be sure?
s
Copy code
+ temporal-sql-tool --plugin postgres --ep postgresql-postgresql-ha-pgpool.default -u postgres -p 5432 create --db temporal
Fri, Jul 21 2023 5:39:10 pm
2023-07-21T17:39:10.882Z	ERROR	Unable to create SQL database.	{"error": "pq: database \"temporal\" already exists", "logging-call-at": "handler.go:97"}
Fri, Jul 21 2023 5:39:10 pm
2023/07/21 17:39:10 Loading config; env=docker,zone=,configDir=config
@Conor Barber (Airbyte) - The database a) did not exist and b) should be synced rather than failing
BTW. this is on a fresh install with everything purged/cleaned
c
Thanks Stefán. Trying to reproduce now
🙏 1
s
Rancher, Kubernetes, PostgresHA (pool connection)
c
My current guess is, looking at the script that handles this, you ended up in a weird state from a previous install and our shell script does not gracefully handle an existing db anymore
If this is just a db created only for temporal as part of this fresh install and you aren’t using an existing db for it, one workaround is simply to delete the existing temporal db in postgres that was created from a previous install
s
@Conor Barber (Airbyte) - this is what I do before every try. (See message from before)
btw. uninstalling the helm script leaves a lot of artefacts: • airbyte-mino (StatefulSet) [Understandable?] • a completed airbyte-airbyte-bootloader (Job) [Strange] • airbyte-airbyte-secrets (Opaque Secret) [Understandable?] • airbyte-airbyte-env (ConfigMap) [Understandable?] @Conor Barber (Airbyte), I clean these up as well (before each attempt)
👀 1
And, it's easy to argue that the installation should survive an existing database. A proper database migration is a better path for install then failing. What would happen if someone was trying to do a recovery install on top of an older/failed installation where postgres held all the key information?
c
The thing is, it should work as you describe. Can you compare your chart configuration/env to this and see if anything (that you wouldn’t normally expect) is out of kilter? Ensure appropriate files are present that the chart wants, etc.
Copy code
---
# Source: airbyte/charts/temporal/templates/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: airbyte-temporal
  labels:
    <http://helm.sh/chart|helm.sh/chart>: temporal-0.47.9
    <http://app.kubernetes.io/name|app.kubernetes.io/name>: temporal
    <http://app.kubernetes.io/instance|app.kubernetes.io/instance>: airbyte
    <http://app.kubernetes.io/version|app.kubernetes.io/version>: "0.50.9"
    <http://app.kubernetes.io/managed-by|app.kubernetes.io/managed-by>: Helm
spec:
  replicas: 1
  selector:
    matchLabels:
      <http://app.kubernetes.io/name|app.kubernetes.io/name>: temporal
      <http://app.kubernetes.io/instance|app.kubernetes.io/instance>: airbyte
  template:
    metadata:
      labels:
        <http://app.kubernetes.io/name|app.kubernetes.io/name>: temporal
        <http://app.kubernetes.io/instance|app.kubernetes.io/instance>: airbyte
    spec:
      serviceAccountName: airbyte-admin
      containers:
      - name: airbyte-temporal
        image: airbyte/temporal-auto-setup:1.13.0
        imagePullPolicy: IfNotPresent
        env:
          - name: AUTO_SETUP
            value: "true"
          - name: DB # The DB engine to use
            value: "postgresql"
          - name: DB_PORT
            value: "5432"
          - name: POSTGRES_USER
            valueFrom:
              secretKeyRef:
                name: airbyte-airbyte-secrets
                key: DATABASE_USER
          - name: POSTGRES_PWD
            valueFrom:
              secretKeyRef:
                name: airbyte-airbyte-secrets
                key: DATABASE_PASSWORD
          - name: POSTGRES_SEEDS
            valueFrom:
              configMapKeyRef:
                name: airbyte-airbyte-env
                key: DATABASE_HOST
          - name: DYNAMIC_CONFIG_FILE_PATH
            value: "config/dynamicconfig/development.yaml"
        # Values from secret

        # Values from env
        ports:
        - containerPort: 7233
        volumeMounts:
        - name: airbyte-temporal-dynamicconfig
          mountPath: "/etc/temporal/config/dynamicconfig/"
        resources:
          limits: ***
          requests: ***
      volumes:
      - name: airbyte-temporal-dynamicconfig
        configMap:
          name: airbyte-temporal-dynamicconfig
          items:
          - key: development.yaml
            path: development.yaml
---
s
@Conor Barber (Airbyte) - I will compare. I'm also looking into the Postgres cluster being out of sync.
That was not it. I will look into the values.yaml file now
@Conor Barber (Airbyte) I'm only using the root yaml file from the /airtable directory. The rest is standard stuff. What should I be looking at? (with that file in mind and the
temporal:
section.)
it has this image reference (different from the one you reference):
Copy code
image:
  repository: airbyte/temporal-auto-setup
@Conor Barber (Airbyte) - the image used to create temporal `airbyte/temporal-auto-setup`is two years old. (Both good and bad news)
c
we also noticed that. I don’t think it’s actively used, but we still pull it as part of the charts. Perhaps it is getting accidentally used in lieu of the actual image here in your config?
s
@Conor Barber (Airbyte) - it is in your values.yaml file: https://github.com/airbytehq/airbyte-platform/blob/main/charts/airbyte/values.yaml Line 1106 As per instructions from your website: https://docs.airbyte.com/deploying-airbyte/on-kubernetes-via-helm#deploy-airbyte (Custom deployment)
c
Precisely, yes. We actually upgraded the temporal version and stopped using that image just last week and it doesn’t look like the docs got updated: https://github.com/airbytehq/airbyte-platform/commit/db4fe060e82d7c7ad7c9f04c9be7ddb47ac77d19
When did you pull the images/version of airbyte initially? I wonder if you got caught up in the transition and it put your install in a weird state. Other end users have reported the new install is working for them
from that change, you can see the new image is supposed to be temporalio/auto-setup:1.20.1
s
@Conor Barber (Airbyte) - I have clean/pulled/upgraded before each run.
@Conor Barber (Airbyte) - I can see that the Dockerfile changed. Your helm chart is incorrect. Please let me know when that has been corrected. (I'm sorry, but I'm a bit frustrated having spent 13+ hours on trying to fix "my end" of this)
c
We will get you sorted, don’t worry! I think we might have narrowed down the issue, my suspicion is that this script (which is used by both helm and docker compose) was updated but the charts were not, and the mismatch is causing your issue. So I am looking into updating the version in the chart to rectify this
🙏 1
Stefán, very sorry for the trouble this caused for you. I have an internal PR up to bump this image version (we merge to our internal repo and then it bubbles out to airbyte-platform) so we should get a new chart out once the helm test suite finishes in about 15 minutes. Let’s verify with that new chart whether your issue is resolved
🙏 1
Rest assured, this is not the experience that we expect end users to have when installing helm charts. We will look into why this happened and ensure it doesn’t happen again.
s
Thank you. I will let you know how it goes
c
Stefán, I just merged the update. It will take a few more minutes to bubble out and publish to the chart repo. I’ll let you know when that is done. I can’t guarantee this will fix your issue (the tests for that script on our end still worked, but it looks like our tests for this only run from a clean room approach instead of an existing install, which we should also capture), but it will at least eliminate that possibility if it doesn’t
0.47.11 just published. Give that a whirl and see if it helps: https://github.com/airbytehq/helm-charts/blob/main/worker-0.47.11.tgz Hopefully if not it will give us a better error message from the tool
🙏 1
s
@Conor Barber (Airbyte) It is looking a lot better, thank you! We are almost there! I have a pod `airbyte-connector-builder-server`that is not up and running. I see no logs for that
c
Hi Stefán, what does it say when you
kubectl describe
the pod?
btw, I created an issue to update our test suite to try and catch the specific scenario you ran into with your install: https://github.com/airbytehq/airbyte/issues/28610#issue-1818379117