Hi all. i'm trying to evaluate datahub for my use ...
# getting-started
w
Hi all. i'm trying to evaluate datahub for my use case and i'm not certain if it's a fit.... I have a data lake with bunch of immutable parquet files in a (theoretical)
<hdfs://my-orders>
folder. a new file is uploaded each day, but i want my logical data set to be “orders.” i have an airflow job that runs every week on the last week's files. for example, airflow job
build-weekly-summary/__scheduled_2023-09-09T01:00:00
reads [
<hdfs://my-orders/2023-09-03.parquet>
,
<hdfs://my-orders/2023-09-04.parquet>
, etc] and writes
<hdfs://order-summary/2023-09-09.parquet>
. I want to track lineage that shows which files were accessed and written by a specific run of a job. but i can’t find a way to register a file like s3://my-orders/2023-09-06.parquet to the orders data set in data hub. effectively i want to: 1. Go to data hub and click on the "Orders" logical dataset 2. See that this data set is composed of 24 files in data lake, including
<hdfs://my-orders/2023-09-03.parquet>
3. Click on
<hdfs://my-orders/2023-09-03.parquet>
and see (via lineage) that it was read by the
build-weekly-summary/__scheduled_2023-09-09T01:00:00
job. 4. See that this job passed all the validation checks. 5. See that this job also wrote out a file to
<hdfs://order-summary/2023-09-09.parquet>
Is that possible with Data Hub? it seems like the s3 data lake tooling supports something similar but hdfs is tied to hive? if it's not natively supported, would it be realistic to implement a custom source. edited: replaced s3 path examples with hdfs as that's the store we're actually using in our environment.
m
Hey @wonderful-library-51057, while DataHub doesn't have the concept of a logical dataset, there is the concept of a Container and the concept of a Data Product. You could use either of these to model what you're expecting here.
I think Container would map better
w
ah, that's helpful. thanks. is there an example of something like this in the online demo?
m
doesn't look like it... but we can spin something up for you. cc @delightful-ram-75848
w
oh sweet
that'd be amazing
m
you could also look at the golden files that are checked into the repo to get a sense of the kind of metadata events that the s3 source produces
which should give you an idea for how to emit the metadata to get the effect you need
essentially the folder becomes a container
and all the individual files become datasets with the container aspect set to the folder's container urn
w
is there something similar for hdfs? i see it referenced in the data_lakes category here: https://github.com/datahub-project/datahub/blob/4ffad4d9b91c25d9f8380fba7d81f65fedfd188c/metadata-ingestion/src/datahub/lite/duckdb_lite.py#L548 but couldn't find much documentation on it outside of hive.
m
unfortunately, we don't have an out of box crawler for hdfs ... but it should be easy to extend the datalake
connector
w
that's good to know. we might be able to demo the s3 integration for a stakeholders but then talk through how it'd map to hdfs. would a custom crawler implementation look something like this? https://github.com/acryldata/meta-world/blob/master/custom_sources/src/my-source/custom_ingestion_source.py
or is there a separate set of data lake interfaces we'd have to implement
this is also prob a basic question, but is it safe to assume that it's not possible to swap mysql for postgres as a backing store for datahub?
i saw
acryldata/datahub-postgres-setup
but wasn't certain if that was a native alternative to mysql or something different
m
This is probably better as a separate thread for discoverability purposes. Mind starting a new one?
w
sure