Hi
@Je Sum Yip! First of all, i'm still learning Pulsar Python functions too, so this is just from my understanding and if you have different thoughts let me know.
Pulsar functions use a Serialization & De-Serialization routine (SerDe) when publishing to topics. So, in Python, types that are primitive (or more specifically, serializable by default) automatically choose that "output schema" type that the producer function returns. This is the identity SerDe.
For something like a custom class or some other object that's not serializable, the two other options are:
- Pickle (usually works well)
- A Custom SerDe class (when you need / want full control on the serialization)
Basically, between those three choices (identity, pickle, custom class) you don't really need to explicitly specify an output schema, since that will handle anything from primitive types to custom classes, and ultimately everything custom is serialized to bytes and then de-serialized back to the native python object.
Now with all that said, the Java docs for Pulsar functions do say:
If the input and output topics have schema, Pulsar Functions use schema for SerDe.
Which doesn't seem to be the case for Python. So i think you're right that it's not specifiable in the same way. But i'm thinking depending on the function you're developing you may not have to, based on the way the serialization operates