Hi folks, I am writing `TransformFunction` impleme...
# troubleshooting
a
Hi folks, I am writing
TransformFunction
implementations for geospatial support functionality. Many of them are designed to compare a given static value against table column values. Their current implementations allow to pass the static parameter in any of the two arguments of the function signature. My question is: is there an easy way to identify a static value passed as argument to a
TransformFunction
, e.g. from a
Projectionblock
?
m
If you look at
StWithinFunction
this is how you detect it inside the function:
Copy code
transformFunction instanceof LiteralTransformFunction
a
Got it. Thanks!
m
which functions are you adding out of interest?
a
All of them 😉
m
haha cool
I think a lot of them are already supported by the GIS library that's being used
but needs to be wired up
a
Yes, for the 2D Cartesian case (
GEOMETRY
) we can simply wrap JTS, and we'll add a set of issues and respective PRs sometimes next week, I hope. But I need to build geodetic support libraries first (to support
GEOGRAPHY
), and hopefully be able to help implementing the common MBR R-Tree to actually make many of these functions performant.
m
not sure if relevant to you...but at the moment I think the geospatial indexing only works for the st_distance function
I think there is support in the H3 library to get that working for points inside a polygon
using the
polyfill
function
but I didn't get the chance to try that out yet
a
And that's fine, but H3 is a Point-first library, meaning that its application centers around Point geometry predicates and proximity - which is equally true for the index implementation (thus
ST_Distance
is currently the only function that can benefit of the H3 index structure, and that only for Point proximity). The polyfill is essentially an in-memory index over a hex-gridded Polygon, which is nice and fast for approximate PIP searches. But we are in need to being able to have seamless support for all geometry types.
r
could you create an issue to discuss before submitting PRs please?
👍 1
a
Of course.
r
fully fledged geospatial support is going to be great 👍
r
note: there were some limitation of the LiteralTransformFunction - it always converts itself to a STRING expression and reinterpret its type based on the context. this will probably limit how you work with the static value because technically it should be a byte type but it cannot convey that type info into the transform function. please also tag me when you create the github issue. I can comment more on that
a
Will do. I've noticed that it has no
bytes
parsing implemented; we can discuss then, but maybe just a short info on this: could we enhance it (or create an extending Class) to try/catch either the WKB or WKT serializer implementation you think?
NVM I see how it's done then. The important thing at init time for geospatial functions is only to know which argument is static (or if both are); this allows us to cache it (and, for set theoretic functionality like union, intersection, containment etc. their internal graph representation) into class state to improve consecutive calls, and only deserialize the running table values. Huge win.
👍 1
r
yes I think you probably get what i said previously. to summarize there are 2 things we need to optimize if there's literal shape passed to the geo spatial transform functions. 1. when literal is passed in, we do not need to run GEO spatial compilation on each projection block elements - instead we should only run it once (either during init as a private field or during actual projectblock transform) 2. However transform function is init on their own groups of segments to process, therefore it will still be overhead on each transform function, @Richard Startin created this to share across different transform functions for literals https://github.com/apache/pinot/pull/7720, but as I mentioned previously it only works with type STRING, so in order to pre-cache the geo spatial object we need to extend the LiteralTransformFunction to support geo spatial type
IMO first step will give the most bang on the bucks
O(N_of_rows)
, second step is less
O(N_of_segments)
r
it's not really true that it only works for strings - it uses the raw literal as the key, but that raw literal can be parsed to anything
FTR I don't really like LIteralTransformFunction at all, it's very inefficient, I would like to have a concept of a scalar expression, where all expressions are currently vectors
👍 1
but that would block progress on geometry literals
r
yeah share the same thoughts, but i guess with or w/o literal transform function sharing; we can at least reduce the compute from a per-block perspective.
https://github.com/apache/pinot/pull/8167 let me know what you think.