This message was deleted.
# general
s
This message was deleted.
e
Source? You need CDC
j
No. I need a few dozen billion objects I've saved to be moved into pulsar
I'm not interested in any sort of change data, and it's more of the batch source rather than a streaming source, in that after the source load is done we won't trigger it again, but I am looking at a highly scalable way to bring in a huge data set.
We have looked at a few other etl type options, and the performance has not been all that impressive, though I'm open to suggestions. Ive done some other work writing some functions that broke the load job up into batches (published to Pulsar topics) and then processed the batches in a highly parallel way with Pulsar functions and got much closer to the performance targets I am looking for.
a
I wrote an Elasticsearch source for Flink some time ago. You want to use
search_after
and create a
PIT
to ensure good performance. I think the biggest hurdle you will face is that the Pulsar IO framework does not provide abstractions that help with coordination of the connector instances. I鈥檇 be happy to provide some guidance if you start working on this
j
I was looking at the BatchSource interface and PIP: https://github.com/apache/pulsar/wiki/PIP-65:-Adapting-Pulsar-IO-Sources-to-support-Batch-Sources Have you used that?
I was thinking of a sliced scroll (since I'm not interested in realtime data)
I've also achieved very high performance with sliced scrolls previously. I've not used PIT.
a
BatchSource
will work if you are going to assign all work packages to the instances right at the beginning, yes
j
my initial idea on reading that PIP was to assign each slice of the scroll to one instance
馃憤 1
a
The docs explicitly advise against sliced scrolls for larger datasets (see here https://www.elastic.co/guide/en/elasticsearch/reference/current/paginate-search-results.html#scroll-search-results)
j
yeah, I read it
specifically, I'm not interested in preserving index state, I'm not going to sort in elasticsearch, and I have not seen how you can do a PIT across multiple threads.
The design there would be to have a query that allowed it self to be partitioned to generate the parallel PITs
Same idea, however we slice it. The idea would be to generate n number of tasks, where each task is a slice, then I believe BatchSource can help with managing the different instances, etc.
@Enrico Olivelli in creating the config for an ElasticSearchBatchSource, what is the preferred model for encapsulating common properties between the Sink and the Source? An interface with the full list of fields in each implementation? A common parent class?
e
Probably a common parent class may work. it is not common to have both a source and a sink for the same system all together (they are usually quite different)
maybe you can take a look to KafkaSource and KafkaSink
j
Thank you. I've been working on creating a common parent so far today, with a few minor issues. I'll definitely look at how KafkaSink and Source are doing it. Its always nice to have a model.
In this case I want to be able to reuse all of the common ElasticSearch client code, because that is the same across both implementations. There are differences, certainly, but clearly a lot of work was done to support ElasticSearch, OpenSearch, and the different versions.
KafkaSink and Source Config classes are completely independent and just have the same fields. If its ok with you I'd like to continue to try to go down the common parent path.
e
sure
馃憤 1
j
I implemented the sliced PIT search for Elasticsearch, but when working on OpenSearch, I noticed that it only supports PIT from 2.4 (and not at all in the version of the REST client used in Pulsar). @Enrico Olivelli @Alexander Preu脽 do you think I should also implement the scroll based search since that will have the most backwards compatibility? Also, is there a reason Pulsar isn't using the OpenSearch java client? That library is compatible back to OpenSearch 1.3.9 (https://github.com/opensearch-project/opensearch-java/blob/main/COMPATIBILITY.md). What should the goal for backwards compatibility be?
e
@Nicol贸 Boschi may have a good answer here
Regarding backward compatibility, it is important that if an user upgrades Pulsar the system continues to work seamlessly. We should avoid breaking changes as much as possible.
j
Ok. Given that the current functionality (bulk indexing) is basically the same going back to ES 5 or maybe even ES 2, I think we may have to think about a minimum supported version for using ES as a batch source. I've implemented the core methods to do both PIT and Scroll searches for ElasticSearch and OpenSearch. Now I am at the point of working on what the common interface for all of this is in the
RestClient
to then be called by the
ElasticSearchBatchSource
. It is probably a good idea to implement some version checks if a user has configured PIT, but their ES / OS version doesn't support it.
a
Personal opinion: Even ES6 is ancient at this point, if there is no strong need in the community I think supporting versions going all the way back to ES5 is not necessary. I鈥檓 not familiar with how the OpenSearch versions relate to the Elastic ones so I can鈥檛 really comment on that. Implementing some checks on startup sounds like a good idea 馃憤
j
@Alexander Preu脽 completely agree. I thought I would be able to use PIT for both, but given that 2.4 is when OpenSearch introduced PIT (latest OpenSearch version is 2.8), it felt like that might be too much of a constraint. Additionally, I've implemented sliced scroll searches dozens of times, so it didn't take me long and would provide the farthest back compatibility if needed by a user.
a
Yeah, sounds reasonable
e
think we may have to think about a minimum supported version for using ES as a batch source.
Makes sense. The versions supported by the Sink can be different from the versions supported by the Source
j
Sounds good. I keep plugging away at this.
Ok, so I built out the tests for PIT and Scroll searches, and found an interesting edge case. For ES 7 the code currently uses the OpenSearchHighLevelRestClient. Fine. Except for the fact that PIT searches use a different endpoint in ES and OpenSearch. So using the OpenSearchHighLevelRestClient isn't going to work for ES 7 clusters. I looked and the ES Java client works back to 7.14, so if you all think its prudent, I can make the change to use the ES java client if a supported version of ES is detected. I believe the OpenSearch rest client should be compatible with 7.10 and previous versions of ES...so I'm not really sure what to do for the 7.10 - 7.14 versions of ES. Additionally, I believe scroll searches work just fine with the OpenSearch rest client against an ES cluster, as those have the same endpoints, so another option to just just not allow PIT searches on those and have the user configure a scroll search instead.
a
These client compatibility topics always seem to be a nightmare 馃榿 Let鈥檚 say the implementation would only be based on the scrolled search right now - does this guarantee data isn鈥檛 read twice? IIRC this was one of the main reasons for using the PIT (besides the performance on large searches)
j
the one thing I've learned in my last 7 years of using ElasticSearch at scale is there is really no such thing as a guarantee. Scroll search appears to be deprecated in ES 8, so we definitely need to leave the PIT search there. Both PIT and Scroll claim to effectively snapshot the index state at the start time. PIT has advantages that multiple queries can be run and you can move forwards and backwards in pages. Scroll only allows the one query to run, and you can only move forward in results. There is also some documentation saying that the slicing mechanism used for PIT is more efficient than for Scroll.
@Alexander Preu脽 @Enrico Olivelli it gets worse. I wanted to add OpenSearch 2.4+ into testing so that I could validate that the PIT search works as expected on OpenSearch. Well, the OpenSearch 1.2.4 rest client being used isn't able to communicate with OpenSearch 2.x at all. I'm going to try to see if the later versions of the OpenSearch high level rest client can work with OpenSearch 1.X. Otherwise, I guess I can look at implementing the OpenSearch java client (similar to the Elasticsearch Java client).
OpenSearch 2.7 is just the right version. The high level rest client works with 1.2.x as well as 2.x. The gotcha there is that support for the es type has been completely removed. So again, we have to pick our poison on ES compatibility. I would say this might be where I copy everything into a new module, but honestly, the es sink should support OpenSearch 2.x as well, and types have been known to be a bad idea since ES 5 if not earlier, and deprecated for a long time.
Please let me know how youd like me to proceed. I can also push everything to a branch so you can see what it looks like so far.
e
yes, creating a draft PR would help
cc @Nicol贸 Boschi
j
@Enrico Olivelli Thank you. Will do. The main things I have left to finish are key generation from the ES result and any schema handling
@Enrico Olivelli @Alexander Preu脽 @Nicol贸 Boschi Here is the draft PR in my repo : https://github.com/Q6Cyber/pulsar/pull/2 I greatly appreciate any feedback at this stage. I still have a fair amount of work to do, but I think I've made some solid progress and hopefully have things on a good path.
I'm headed out on vacation for a week. Thank you for the help and guidance with this. I'm looking forward to getting it wrapped up when I come back.
a
I鈥檒l have a look at the PR over the next few days. Enjoy your vacation!
馃憤 1
e
Thanks