This message was deleted.
# general
s
This message was deleted.
j
I think you have to identify the handler class as part of the function's config. Are you asking if there is some adapter for Spring that will auto generate a function config from "auto detected/injected" handlers/implementations?
e
I don't know exactly what you mean when you say auto generating a function config from auto detected/injected handlers/implementations. Maybe? I'm just trying to understand the lifecycle of a function that is deployed within pulsar. In other words, based on the fact that you have to provide the classname within the deployed jar file that implements the Function interface, I understand that somehow, pulsar must be instantiating an instance of that class and then invoking the
process
method each time a message is consumed from an input topic. I guess what I'm looking for is some way to make the spring framework own/manage the lifecycle of that function. I would like to be able to provide a jar file to pulsar that, ideally, would be a spring boot application. And then pulsar could load a bean that is a
org.apache.pulsar.functions.api.Function
instance. I'm not sure exactly how this would work. Here is an example of something that I'm currently trying to experiment with that I haven't been able to test quite yet. This is not the ideal example because ultimately, if this does work, pulsar would still be responsible for invoking the function's constructor rather than just loading the function from the spring application context: Some service interface:
```public interface TransformerService {
String transform(String s);
}```
Some service implementation:
```import org.springframework.stereotype.Service;
@Service
public class UpperCaseTransformerService implements TransformerService {
@Override
public String transform(String s) {
return s.toUpperCase();
}
}```
Some Function:
```import org.apache.pulsar.functions.api.Context;
import org.apache.pulsar.functions.api.Function;
import org.springframework.boot.SpringApplication;
import org.springframework.boot.autoconfigure.SpringBootApplication;
import org.springframework.context.ConfigurableApplicationContext;
@SpringBootApplication
public class UppercaseTransformer implements Function<String, String> {
final ConfigurableApplicationContext ctx;
public UppercaseTransformer() {
ctx = SpringApplication.run(UppercaseTransformer.class);
}
public String process(String s, Context context) throws Exception {
final TransformerService springBean = ctx.getBean(TransformerService.class);
return springBean.transform(s);
}
}```
Ideally I would want use start.spring.io to create a simple uber executable jar file that configures and launches the spring application. The function would be a spring bean, and pulsar would have the ability to load that spring bean from the application context.
Maybe pulsar could provide an api that gives a little more control over the function's lifecycle (ie: initialization and teardown). Maybe something like: Function Supplier:
(
org.apache.pulsar.functions.api.Context
context) ->
org.apache.pulsar.functions.api.Function
This would allow the jar file to provide a function initializer that pulsar could use to get a handle on the function instead of creating it itself. Maybe there would also be a tear down method in this api that pulsar could invoke to gracefully shutoff the Function, or stop the app. One advantage of this approach would be to simplify pulsar's Function interface such that any java.util.Function could be used without loosing the ability to pass in a configuration object. I'm rather new to pulsar, but is it common for the Context object to have different values for each invocation of the function when a message is processed? Or would it make sense to just pass the context once during a function initialization, and then use the function for each message without passing the context?
j
But there are a lot of things that go into deploying a function beyond the implementation of it, including topic inputs/outputs that might not be known to the function itself. In your example, how would Pulsar know how to "deploy" the
UppercaseTransformer
and set all the other values available here: https://pulsar.apache.org/docs/pulsar-admin/#create-1 Spring could certainly tell Pulsar the input and output schemas/types and the function implementation class (and perhaps assume guess a name) but not much else besides that... unless you're saying all of that would come through Spring's configuration mechanisms?
e
I'll have to take a look at the link you referenced above when I get back to my computer. I guess I'm just thinking that whatever pulsar needs to be able to deploy a function could be encapsulated into an API or interface, an instance of which could be provided by some provider or factory function
j
I actually definitely want to move in that direction on my thread- my thought was annotations, your thought is Spring, but we're both getting at the same thing! šŸ™‚
šŸ‘ 1
e
@Devin G. Bost you may be interested in reviewing this thread. ā˜ļø
šŸ‘ 1
d
@John Dimeo What kind of annotations are you referring to? I was under the impression that lombok annotations can be used, but I'm not a Java expert, so maybe I'm misunderstanding.
e
I also think that updating the pulsar code to define an interface that clearly specifies everything Pulsar needs to create/manage a function would make Pulsars own codebase better. Imagine what it would be like if in order to deploy a function in pulsar, all pulsar needed was a jar file. But that jar file provided it's own provider/factory method for instantiating the function for Pulsar using a service provider interface (SPI) or something similar. The Pulsar codebase wouldn't need to use some funky java reflection (which I imagine is a whole different can of worms we could discuss when it comes Pulsars compatibility with JPMS in current java versions better than 9) and the instantiation of the actual function could be left to the implementor of the jar file. This would provide benefits regardless if the jar implementor wanted to use spring or not.
As an example of the SPI see here: https://www.baeldung.com/java-spi
In the above example, imagine that the
ExchangeRateProvider
was an interface called
PulsarFunctionProvider
where
PulsarFunction
is an interface that specifies everything Pulsar internals need to properly configure/use a
PulsarFunction
.
j
@Devin G. Bost not Lombok (in this case! I love it in general!) rather using annotations to track function config in the code itself. Then I reflect those annotations to build the CLI
pulsar-admin
command to deploy. For example:
Copy code
@Doc(name = "odf-finsummaryfn")
@Requires(@Doc(name = SupporterSummaryChange.TOPIC))
@Produces(@Doc(name = "odf-financial-summary", type = FinancialSummary.class))
@Parallelism(3)
public class FinancialSummaryFn implements Function<SupporterSummaryChange, FinancialSummary>, JSONUtils, IsDocumented {
    ....
@Doc
is just a little annotation I wrote to track some documentation metadata on a class. I'm using it here to specify the function name, the topics it requires/is an input, the topic it produces/is the output, and the default/starting parallelism I then do this kind of thing later:
Copy code
String[] args = new String[] {
	"functions",              mode,
	"--jar",                  "/tmp/" + jar + ".jar",
	"--inputs",               first(fnConfig.getInputs()),
	"--classname",            fnConfig.getClassName(),
	"--name",                 fnConfig.getName(),
	"--subs-position",        pos.name(),
	"--function-config-file", "/tmp/function-config.yaml",
	"--parallelism",          fnConfig.getParallelism().toString()
};
and use the annotations to create a
FunctionConfig
and use
ProcessBuilder
to call the function deploy command for me
e
@John Dimeo how compatible is all that reflection in Java 17 with the JPMS?
j
It's all just annotation processing so nothing concerning, right? Here's one example for
@Parallelism
Copy code
val instances = Optional.ofNullable(fn.getClass().getAnnotation(Parallelism.class)).map(Parallelism::value).orElse(1);
e
What are your thoughts pertaining to the ServiceLoader John?
j
I like it, but I don't think we can completely specify a function from within code. Function config falls into two big categories: • Ontological: how the topics are wired up, what schema they expect/produce, etc. This is stuff that would work really nicely inside code and what I have tried to capture in annotations • Environmental: which tenant/namespace to deploy to, might even argue subscription position (since that can change depending on needs). This is stuff that would be painful to have inside the .jar and does work better on the CLI or in the function YAML I guess my ideal world would be that in official Pulsar you can do any of the following: • Implement
Function
, manage all other config externally (today) • Implement
Function
, annotate your class with some things, manage the rest through CLI or YAML (my set up) • Implement a more complete
Function
interface that specifies the same things you can annotate, manage the rest through CLI or YAML (your suggestion)
e
Maybe we’re conflating different things: object instantiation, and function configuration (ie: how the function is wired to everything else in pulsar). From my perspective, in order for me to implement a pulsar function today, I just need to create a class that implements
Function<I, O>
and package that in a jar file. Then there is the question, how does that function get deployed into Pulsar? We need to use the Pulsar Admin API. Based on the examples in the deploying functions documentation it looks like this:
Copy code
$ bin/pulsar-admin functions create \
  --jar <jarFile> \
  --classname <classname> \
  --inputs <inputs> \
  --output <output>
The parameters
--jar
and
--classname
are the only parameters the admin api uses to instantiate the function. And Pulsar internals currently uses reflection to invoke a no arg constructor to instantiate an object of the type specified by the classname located within the jar. The other parameters:
--inputs
,
--outputs
, etc... are what I think you referred to above as ā€œenvironmentalā€ configuration. These have nothing to do with instantiating a function, but they are used for wiring up (configuring) a function. Can you provide me with the names of the source code files within the Pulsar codebase that handles all this initialization and configuration? I’ll have a look at it to see if I can better understand.
j
I actually think that
--inputs
is a mix of both: the topic name (if you really leverage tenants and namespaces and don't inject environmental information into your topic names) and schema are ontological and should be about instantiating the function. The function will break if the data coming in isn't the entity/data contract it expects. However, the tenant and namespace are environmental. In my example above:
Copy code
@Doc(name = "odf-finsummaryfn")
@Requires(@Doc(name = SupporterSummaryChange.TOPIC))
@Produces(@Doc(name = "odf-financial-summary", type = FinancialSummary.class))
It deliberately does not include things like CPU and RAM requirements or if this is the DEVINT or PROD cluster. But still things that are very helpful to define with the function implementation itself instead of always having to "remember" that function X consumes data contract Y during deployment. So yes, we are conflating, because we must :-)