Hi! đź‘‹
Question surrounding incremental sync and substreams.
I have an API that I want to incrementally request data from. The parent stream outputs the ID to request, and I would like to incrementally request the data for each child. There is around 1000 known parent ids.
Basically can I apply a state to each substream and have it store its state (so 1000 internal states).
Other side is do I need to accept that there is going to be duplication and request smaller slices (hourly for example) the dedup early in the transform layers.
Output data is ID + date time as the state.
If it fails I need it to get the data on the next run, so worst case I duplicate.
Ideal case is each stream can run on its own little world happily keeping its own state, but I also don’t want 1000+ tables created if possible…
Thanks!