Hi :datahub: community! How are you modelling eg a...
# ingestion
w
Hi datahub community! How are you modelling eg a kafka streams application or a microservice reading from a kafka topic and writing to another one? For the topics, it is clear that I can use Dataset entity. But what about the process in-between? I find DataFlow+DataJob not fitting the purpose here. Related to this I raised this feature request https://feature-requests.datahubproject.io/b/Developer-Experience/p/entity-model-for-streaming-applications-or-microservices Anyway, advice and suggestions are welcome. Thanks!
l
I've been having the same questions myself 😛. For the most part I've settled on just linking the two topics enough since all I care about is the metadata (and where it came from). Any additional info like the microservice itself could be placed in the docs section but yeah it doesn't visualize well
w
The node in-between is important in our organisation so we can track the ownership of that process, since the owner of the process can be different from the owners of the datasets on the edges.
l
In my case the owners have always been the producers to the topic (since ultimately, they own the stream they produce to like they would own any rest api they expose) but yeah i can see the conflict
w
@little-megabyte-1074 Do you think a feature like this could be interesting for the open source project? If so, I could help with a contribution.
l
@witty-butcher-82399 YES!!! Absolutely - this would be such a fantastic contribution back to the OSS community
I’d be more than happy to pull in folks from the core DataHub team to provide consultation/modeling feedback as you’re going along
w
Sure, I will need some help. Actually, my fist question is regarding the existing DataProcess entity. That entity fits our requirements, however we are not using it because it was marked as deprecated and I remember someone mentioning some issues with it. Should we start with a new entity from scratch? or instead reborn the DataProcess one and fixing the issues if any? Thanks
I spent some time on this. Here is the update and findings: 1. My first attempt was to update existing DataProcess entity by adding missing aspects (I guess you get the idea from the attached path 😅 ) and still more aspects could be added (eg about tags). This results in DataProcess aggregating existing definitions from both DataFlow and DataJob. So, the resulting model is quite redundant. 2. Then I realized DataJob has everything I need
DataJobKey, DataJobInfo, DataJobInputOutput, EditableDataJobProperties, Ownership, Status, GlobalTags, BrowsePaths, GlossaryTerms, InstitutionalMemory, DataPlatformInstance
and even DataFlow reference in DataJob is optional
flowUrn: optional DataFlowUrn
. So using DataJob without having a parent DataFlow may fit my requirements! Unfortunately that’s not possible because DataJob URN does include the DataFlow: eg
urn:li:dataJob:(urn:li:dataFlow:(airflow,dag_abc,PROD),task_123)
So I would like some advice from the modelling team. Thanks!
i
Hey Sergio, So DataProcess is something that was discussed by the core team as being redundant with DataFlows + DataJobs. Your use-case seems to fit a gap that we did not consider. Would you consider creating a PR to un-deprecate DataProcess, create an integration for it so for analysis? That way we can move the discussion there with the right folks and iron out any concerns that might exist?
w
I agree, we can have a more focused discussion in the PR. I’ll ping you when I have one ready.
As agreed, here it is the PR so we can follow up on the discussion https://github.com/linkedin/datahub/pull/4114 thankyou1
thank you 1