*Is this your first time deploying Airbyte*: No *O...
# replication-troubleshooting
f
Is this your first time deploying Airbyte: No OS Version / Instance: AWS EC2 Linux Deployment: Docker Airbyte Version: 0.40.18 Step: AWS EC2 Linux Connection details: MySQL 1.0.12 -> Redshift 0.3.51 Description: I’ve been running into an issue several times over the past few weeks (I posted about it once on this channel but got no response). In short, I am syncing a table that results in duplicates when COPY’d into Redshift. This has occurred when I have run Full Refresh and when Ive run Incremental refreshes; and seems to occur randomly for different tables. Here is an example where I run an Incremental Append refresh for the 1st time on a table, but when it is written, the data is duplicated. This can be seen in the logs as well, where the sync summary claims that the correct amount of rows are written to the table,
Copy code
{
  "streamName": "table_name",
  "stats": {
    "recordsEmitted": 428544,
    "bytesEmitted": 260633069,
    "recordsCommitted": 428544
  }
}
but if I look at the normalization step it is showing the true amount of rows inserted - 857086 (which is duplicate of the data)
Copy code
2022-11-26 20:50:46 [42mnormalization[0m > 17 of 117 OK created incremental model schema.table_name..................................................... [[32mINSERT 0 857086[0m in 85.81s]
Would be great to get some help trouble shooting what is going wrong
Looking at the _airbyte_raw_table_name table, which I assume is used by airbyte for staging purposes, I can see that the duplicates exists there too. Each _airbyte_ab_id has 2 rows.
Here is where I see the duplication potentially taking place in the logs :
Copy code
2022-11-26 20:38:07 [32mINFO[m i.a.w.i.DefaultAirbyteStreamFactory(internalLog):120 - Copying stream table_name of schema schema into tmp table _airbyte_tmp_ueh_table_name to final table _airbyte_raw_table_name from stage path data_sync/prod/schema_table_name/2022_11_26_19_b68a2550-ba59-4650-9bd7-5e2878f8a04f/ with 2 file(s) [2022_11_26,2022_11_26]
2022-11-26 20:38:07 [32mINFO[m i.a.w.i.DefaultAirbyteStreamFactory(internalLog):120 - Starting copy to tmp table from stage: _airbyte_tmp_ueh_table_name in destination from stage: data_sync/prod/schema_table_name/2022_11_26_19_b68a2550-ba59-4650-9bd7-5e2878f8a04f/, schema: schema, .
2022-11-26 20:38:30 [32mINFO[m i.a.w.i.DefaultAirbyteStreamFactory(internalLog):120 - Copy to tmp table schema._airbyte_tmp_ueh_table_name in destination complete.
All non duplicated tables (regardless of size) say copying stream ….. with 1 file(s). However the duplicated tables all show copying stream…. with 2 file(s)
I see earlier in the logs that the tables with 2 File(s) have have a situation where Airbyte Closes the buffer stream for that table, then right away Starts a new buffer stream for that table. Thus, uploading the data to S3 again
Looking at the actual S3 bucket, (I turned Prune files option off to debug), I can see that the File only exists 1 time in the Bucket (so the file is overwritten), but the manifest file has duplicate entries for the URL property. So it is eventually copying the file into redshift twice. Why isnt Airbyte able to see that the file already exists? Does it have to do with the latest Redshift PR that was merged which has to do with S3 bucket paths?
n
Hi Frank, please review our code of conduct: double posting a question multiple times goes against it. Please make sure to only post a question once per platform/thread. https://docs.airbyte.com/project-overview/slack-code-of-conduct/#rule-3-dont-double-post Let's continue in the discussion in the forum.
✅ 1
u
Hello Frank Kody, it's been a while without an update from us. Are you still having problems or did you find a solution?