Hi all - I'm running into 502 errors on Airbyte ho...
# replication-ask-ai
z
Hi all - I'm running into 502 errors on Airbyte hosted on our EC2. I want to confirm that Airbyte is throwing the 502 Bad Gateway errors, but am unsure on where to look for airbyte logs on our EC2 - any ideas on where they'd be located? ChatGPT mentions a airbyte-server.log but I'm unable to find it
k
You can find the logs for your Airbyte instance running on EC2 in a couple of ways. Firstly, you can get your logs from your Airbyte Docker containers in CloudWatch. You can follow this [guide](https://aws.amazon.com/pt/premiumsupport/knowledge-center/cloudwatch-docker-container-logs-proxy/) to set this up. If you want to access the logs directly from your EC2 instance, you can use the
docker logs
command. For example, to get the logs for the Airbyte server, you can use the following command:
Copy code
bash docker logs airbyte-server > server.log
This command will output the logs from the
airbyte-server
container to a file named
server.log
. You can replace
airbyte-server
with the name of any other Airbyte container to get its logs. Please note that you need to run these commands on your EC2 instance where Airbyte is running. You can do this by SSHing into your EC2 instance. Sources: - [Deploy Airbyte on AWS](https://docs.airbyte.com/deploying-airbyte/on-aws-ec2) - [Airbyte Support Conversation in Slack](https://airbytehq.slack.com/archives/C021JANJ6TY/p1664834193615179)
j
Usually a 502 is your load balancer not being able to talk to the backend service.
Are you using an ALB to handle ingress to the service? Is your health check endpoint set correctly?
z
Thanks for the quick response, @Jason Reslock! Hmm I believe so - it happens seemingly at random - for example, the 502 only occurred this past Thursday and Sunday (we run jobs daily) from 3:15am EST -> 3:35am EST
j
If it is intermittent like that and you can correlate it with big sync jobs or other activity it is probably overloaded and not responding. If you have any sort of cloudwatch metrics you should be able to look at CPU/memory utilization, network I/O, etc. over time to see if there are spikes that correlate with your outages.
z
I did notice that on one of them, the very small tasks prior (like 2 rows of data) took ~25 minutes to sync, then the failures occurred. It could be an overload, but there's just not that much going on. It appears to hang on this message for 25 minutes (I have run into this before):
Copy code
2023-09-15 07:00:20 INFO i.a.c.t.s.DefaultTaskQueueMapper(getTaskQueue):31 - Called DefaultTaskQueueMapper getTaskQueue for geography auto
2023-09-15 07:24:55 INFO i.a.w.t.TemporalAttemptExecution(get):136 - Docker volume job log path: /tmp/workspace/2218/0/logs.log
Do you have any ideas on what could be causing that hang? We boot the Ec2 up 30 minutes before starting any jobs - does it sometimes take longer for airbyte to start up/be ready to execute connections? I'll have to take a look at those cloudwatch metrics - thanks for that pointer!
j
My airbyte systems all take a long time to start up due to pulling the images for every component (look a the compose file and you will see how many images it needs just to start up) and also all images for source/destination.
During startup I always see 502 errors until the system is truly up and running/responding to health checks from the ALB
z
This is so good to know! We currently give it 30 minutes which is usually enough, but perhaps I should lengthen this out to a full hour: When does Airbyte start-up occur - on EC2 start-up? Or when the first connection runs?
j
I think that depends on where you are pulling images from, how many custom source/destination images you have configured, etc.
z
We just have SQL Server, Salesforce, and Redshift sources. 1 destination - Snowflake. Airbyte is baked into the EC2 image. Are all of these images pulled when the EC2 is started?
j
I believe all of the images are pulled unless they already exist locally on the EC2 VM. The service startup is basically just
docker compose up
I have that wrapped in a systemd unit on my Airbyte VM but there are a number of ways to handle it.
z
Ah I see! Looks like we have ours wrapped in a systemctl process as well I believe I've nailed down the issue - the airbyte-temporal service is timing out During the ~30 minutes that we were getting 502 errors, we received this error message ~1000 times (compared to normal jobs, we get that error message 1 time every 30 minutes)
Sep 15 07:59:58 ip-10-0-2-140.ec2.internal docker-compose[5150]: airbyte-temporal                  | {"level":"info","ts":"2023-09-15T07:59:58.992Z","msg":"query directly through matching on sticky timed out, attempting to query on non-sticky","service":"history","shard-id":1,"address":"172.18.0.2:7234","shard-item":"0xc000844480","component":"history-engine","wf-namespace":"default","wf-id":"connection_manager_28fc97a3-af61-4992-b00b-4fbbd5801f64","wf-run-id":"566e16d6-30f5-4c1b-83aa-f71b0db72459","wf-query-type":"getState","wf-task-queue-name":"","wf-next-event-id":37,"logging-call-at":"historyEngine.go:923"}
It appears to be timing-out when trying to connect to some IP address (presumably where the Temporal service is located). Is Temporal a separate service that a locally hosted airbyte API would reach out to? Does it experience latency from time to time? Thinking about a smart way to handle for this...
j
Temporal is one of the services that airbyte starts up as part of the core system.