Gustavo Maia
11/14/2022, 9:30 PMMarcos Marx (Airbyte)
11/14/2022, 9:35 PM"data": [{"user_id": 1, "name": "marcos"}, {"user_id": 2, "name": "gustavo"}]Gustavo Maia
11/14/2022, 10:49 PMMarcelo Pio de Castro
11/15/2022, 12:40 AMGustavo Maia
11/16/2022, 12:07 PMMarcelo Pio de Castro
11/16/2022, 10:57 PMMarcelo Pio de Castro
11/16/2022, 11:03 PMMarcelo Pio de Castro
11/16/2022, 11:04 PMGustavo Maia
11/17/2022, 2:16 PMEli Sigal
11/17/2022, 4:56 PMimport requests
s = requests.Session()
list_of_dfs=[]
with s.get(url, headers=None, stream=True) as resp:
for line in resp.iter_lines():
if line:
list_of_dfs.append(pd.DataFrame(line))
Just make sure to map the data acordingly since streaming data is not row by row as you would expect.
or use iter_content (iter content also receives param chunk size) and with data you have a buffer that you can put into pd.read_csv and then use chunks to split even more the data
you can also use concurrent futures multithreading to create faster stream.
This way you can then use pd.concat([list_of_dfs])
after all data is ready ( if needed ofcourse)
or if this is big data you can use dask read csv and create partitions or pyspark for the same use cases.
if the separator or delimiters are not consistent you will need to work around the data for each stream and create try except block to create a dataframe with the correct data.
hope this helpsuser
12/06/2022, 8:25 PMGustavo Maia
12/06/2022, 8:27 PMSaj Dider (Airbyte)
12/30/2022, 6:32 PM