Hi! Today I opened this topic about some problems...
# replication-troubleshooting
o
Hi! Today I opened this topic about some problems I've been facing lately with syncs for high volume tables from Oracle: https://discuss.airbyte.io/t/poor-sync-performance-for-high-volume-loads-100-gb/3284 Has anyone faced similar performance issues as me? 🤔
🐙 1
🙏 2
💯 1
🙌 2
u
Hello Oriol, it's been a while without an update from us. Are you still having problems or did you find a solution?
d
Hello. I’ve been working with @Oriol on the same project for a while using Airbyte. We still have the same issues with performance. Could you please help us? We would like to know if it is a common limit with this tool or we are doing something wrong 🥲
Hey there @Marcos Marx (Airbyte) . Could you please help us with this performance issue? Any news?
s
Hey @Daniel and @Oriol, sorry for the delay in response. Airbyte can be fragile with high volume syncs. How many tables are you trying to sync and have you tried creating multiple syncs for different tables (to achieve a sense of parallelism?) Increasing the number of workers may also improve throughput: https://discuss.airbyte.io/t/understanding-workers-jobs-and-concurrency/1164 Finally, it might be worth creating an issue on our github relating to this so that it can be triaged and viewed by our engineering team. The Oracle source is still in alpha but we’re working on trying to make it better so it would be good to have more information about how our users are using certain connectors and what their needs are. We appreciate you giving us a try. I’ll try reading over your thread again on discourse and circulate it with the rest of the team for thoughts. Thanks!
d
Hey @Saj Dider (Airbyte), thank you for the answer! Currently we have created one sync for every table in order to achive a good parallelism and schedule control. We used all the resources of our Airbyte instance in the tests we described in the Discuss issue (we didn't have multiple syncs running in parallel), so I'm not sure if increasing the number of workers will improve the performance, but we will give it a try 🙂 We're also going to open an issue in the Github repo
s
Hey @Daniel thanks for clarifying. I’m not 100% convinced increasing the number of workers will lead to such a dramatic improvement in performance (especially the improvement you’re looking for) but it’s worth a shot. Another thing i’ve seen folks do is deploy on kubernetes but that seems to be more flaky and requires more tinkering and might not be suitable for your use case. Either way, please open an issue on our Github repo!
👍 1
d
Thank you @Sajarin Dider . Kubernetes is a very atractive option for us in the future. Do you have any solid performance test results proving that the Kubernetes alternative is a better option in terms of performance in a single connection (obviously it will scale better)? It would be nice to justify the effort before such a big change.
s
Hey Daniel, sorry for the delay again. I don’t have any concrete numbers to prove that Kubernetes would lead to better performance, my earlier claim is based on anecdotal evidence from other users who use Airbyte on kubernetes. Sorry I can’t be of more help here.
👍 1
d
Thank you @Sajarin Dider ! We’ll have this in mind 🙂 Meanwhile we are looking for some alternatives in order to extract these big tables. We are testing AWS DMS (Database Migration Service) with incredible results. We just expect that Airbyte could improve this performance issues in the near future