Hi! I'm trying to set up lineage for a system invo...
# ingestion
q
Hi! I'm trying to set up lineage for a system involving AWS Glue, AWS Athena and metabase. A series of spark jobs produce data in the glue catalog, metabase then reads it using athena as a backend. When ingesting this into datahub, the metabase source creates a new Athena datasource for all source tables, which already exist in the glue platform. Is there a good way to solve this?
m
Hi @quaint-branch-37931, I need to get back to you. Glue doesn't really know where the source table actually came from (athena or redshift, etc.) So there is a possiblity that a particular dataset gets registered twice, once as glue and then as the specific platform. Do I understand your issue correctly?
q
Hey @miniature-tiger-96062, I think what you said is correct. My issue however is slightly different: we have the data in glue primarily (it was put there by EMR), so athena is not the underlying data source. Athena can run queries against the glue metadata catalog, we use this functionality to query our data lake in a serverless fashion. So while from metabases point of view (which the ingestion script sees), this data is in athena. While actually athena just 'reads through' to the actual data, which is in glue and registered in our datahub instance as such.
I opened a pull request to allow configuring the mapping 🙂
👀 1
m
Hi @quaint-branch-37931, Thanks a lot for the pr. if you use the engine mapping to map athena -> glue, what happens if metabase is actually referring to an athena dataset?
q
Then it will be incorrectly mapped to the glue platform, I guess. I'm not 100% sure on the details but I think Glue serves as a data catalog for athena, so all tables in athena exist in Glue. The issue is that athena also has views, which don't exist in Glue. So those views would unfortunately result in an error. I don't think we can know whether a dataset is a view or a table without querying athena, which would require a specific integration with the metabase plugin, as it stands.
It is a bit unfortunate that these kinds of situations will always have to be dealt with in the specific ingestion source that encounters the issue - in this case metabase. In an ideal world, I guess there would be an instance-wide config that would know how to handle athena sources. No idea how something like that would look, though.
m
Right! So I think passing the mapping is not the right way to handle this. Currently, in the metabase ingestion we rely on metabase to give us the underlying platform name, which in your case becomes athena. Happy to brainstorm with though, on first defining the ideal behavior and then thinking of how to get it done.
q
While it works for my specific usecase, I agree that this isn't the nicest way of solving the more general issue. I think the ideal solution would be a change to how Athena is handled as a platform in DataHub. What I personally would expect: all athena tables live in the glue platform (which backs them, https://docs.aws.amazon.com/athena/latest/ug/glue-athena.html) Athena views could be their own dataset in the athena platform, showing lineage back to the source glue tables. I don't think something like that could be implemented as it stands, because the dataset urn depends on the platform (and thus, in this scenario, on whether the resource is a table or a view). Maybe a solution could be something like a "redirect" dataset? Such that you'd create all glue tables in athena and mark them as a reference to the underlying glue source. Happy to discuss more!
I guess this is actually a bit similar to the situation with dbt (https://datahubspace.slack.com/archives/CUMUWQU66/p1642065169140100): both have some kind of "underlying" node. Are there more platforms that have this concept?