I'm looking for some advice around airbyte transformations and DBT - in particular running a lot of pipelines at scale...
We are currently:
• running DBT + Snowflake + Airbyte
• have 1xCRM, 1xJIRA, 1xCustom Product A, 30x Custom Product B, 100x Custom Product C all about to be ingested into the data warehouse
Syncing data for the 1x options is easy; but there are 30x and 100x of Product A and Product B respectively (and these numbers will continue to grow). Both products have data separated at the source via data schemas (they are not multi-tenant).
I'm after some advice re setup so that we don't get into a huge mess with this many pipelines. We want to be able to have snowflake shares & reporting per product (for our clients) so need this to be data striped and have role based security, but also require cross-client aggregated reporting across all. This drives the need to have per-pipeline transformations & other transformations
I'm curious to know how people have done this at scale with airbyte, specifically re:
1. Many pipelines: There doesn't appear to be an easy way in Airbyte to copy source/destination/connections with a few edits (particularly programatically) which means a huge volume of clicking. Am I missing something obvious here?
2. Transformations with DBT: I've just started reading the Airbyte transformations documentation, and it could be possible to run the per-pipeline transformations from Airbyte (all connected to the same dbt source). However, it looks like I would need to have a dbt project per pipeline (for the separate config). Am I missing something obvious here?
So really my questions is about running a lot of pipelines at scale within Airbyte/DBT (and Snowflake) and I don't want to get into a huge mess
Thanks in advance 🙂
FYI @Varun Khanna