Imagine that I would like to make sure my airbyte ...
# replication-ask-ai
g
Imagine that I would like to make sure my airbyte instance on GCP is only up when my pipelines are running. I thought about using Prefect to get my instance up and then wait for airbyte to load completely and then triggering the next steps of the flow (airbyte syncs, dbt, census, etc...) including getting the instance down after syncs are processed. • Would that be the recommended way ? • How would you make sure Airbyte is fully operational before triggering the syncs through prefect ? Thanks
k
New message text here
c
i guess a couple of questions 1. Where would you persist state? Would this be a helm chart or a docker compose installation?
2. Airbyte exposes healthchecks that prefect can probe to see if it’s online.
in the end, airbyte wants a database to store its state (postgres), so this seems theoretically possible, it’s just not a typical use case to bring the non-stateful aspect up and down like that (though could be a good cost savings measure if you don’t use airbyte constantly)
In other words, the kind of stateless workflow you are describing seems theoretically possible but it’s not something that I would describe as battle tested on our end or well supported - you would for sure have to ensure that only one airbyte instance is running at any given time, since definitely having multiple airbytes touching the same db is a no-go. If you didn’t care about sync history and had a way to dynamically create connections like using the api or terraform you could probably spin up a fresh db every time. We do this in some of our test suites but having something e2e that uses the API to do that is also not super well tested.
g
thanks The gcp instance would be up only while pipelines are up however the storage would remain
So I wouldnt call it stateless
c
ok, so you do want to persist state in between spin up and spin down. Then yeah I would say the primary concern is making sure there will only ever be one airbyte hitting that db So I would have prefect spin it up (docker compose, helm, whatever), then probe the health endpoints exposed in the api yaml, then kick off the syncs using airbyte-api
then probably poll the sync status via the api for completion, since I don’t believe we expose event hooks, long polling, or callback url
and again I’ll note that this kind of workflow is not a super tested thing, so you may find some bugs involved. I would especially make sure that there are no jobs in flight when airbyte goes down
if prefect is the only thing that can touch the airbyte instance it seems pretty safe though. You would want to put some concurrency limitation in prefect so that two prefect’s aren’t trying to run at once, and that the entire airbyte lifecycle (spin up, job monitoring, spin down) is managed by prefect and that no second prefect flow could do anything while the first flow is still running
g
oh right
thanks for the tips
c
good luck!
g
Honestly i have the feeling that it's a lot of hassle to save 100 dollars per month of instance cost
c
probably so. It depends on your scale and when you want to run syncs, obviously
if you do a ton of work once a week for 3 hours it’s probably worth it. If you do a decent amount of work once a day maybe not
g
true