Join us for January’s events :snowman: : Office ho...
# events
y
Join us for January’s events ☃️ : Office hour: Performance Tuning and optimizations w @Gonzalo Ortiz | January 8th, Online. RSVP> Security Features in Apache Pinot w @Barkha Herman , January 15th, Online. RSVP>
r
@Gonzalo Ortiz Could we resume our conversation on optimizing joins using co-location and other techniques on this call?
IIRC, last time we took the conversation offline and I began looking into co-location via segment assignment. But could achieve co-location because segment assignment couldn't provide enough placement control.
By "enough placement control" I mean ensuring the same set of keys are on the same servers.
g
Yes, sure
r
In the case of a fact-fact joins, co-partitioning is required to achieve co-location. And to get the maximum benefit from partition parallelism, it appears that a higher CPU core to server ratio is better
Shuffling will occur if the data isn’t co-partitioned on your join key. What can be done in this case?
In the case of fact-dim table join, the new lookup join strategy in the upcoming 1.3 release can be used
g
Shuffling will occur if the data isn’t co-partitioned on your join key. What can be done in this case?
Not much. We want to work on these new types of join strategies in the future
r
Looks like the plan is to broadcast the smaller table to the larger table and to re-partition during the join? If so, this would perform as well as a pre-colocated join if re-broadcasting can be eliminated or reduced.
Assuming partition skew (hot partition) negatively impacts co-located join performance, what can we do to reduce the impact?
Are there any plans to infer the query hint for the new lookup join?
@Aditya Verma I've been dropping questions here for Gonzalo for this week's meetup
a
Hi @robert zych got it. Let me think of few and then post them.
r
The one we discussed about partitioning on time was a good one. I sent you a tip about that on another thread earlier today.
g
Hi there! It would be great if you could write your questions in the prepared form included in the meetup event instead of here. It would make it easier to keep the event organized
a
we can't writemuich questions on the form , possible I can paste them here @Gonzalo Ortiz
m
You can put them here as well, @Aditya Verma
r
Since the question line in the form has a 129 character limit, I will repost my questions here: 1. Is it true that to get the most out of partition parallelism a higher CPU core to server ratio is better? 2. Assuming a hot partition negatively impacts co-located join performance, what can we do to minimize the impact? 3. Are there any plans to infer the query hint for the new lookup join?
a
1. understanding of Node rebalancing and data distribution 2. Which ec2 node would be best for my use case(Aggragation,time range) 3. kafka partition and table partitioning (Lets say we have 8 partitions and 8 segments will be created for them as to avid the data duplication we will send the data from kafka by key. so when all those 8 segments wil be filled then how does that work it will crate new segemnts which, how will they know what data other segmenst will have even if its key based.
g
Ey @Aditya Verma! I can try, but my forté is in query optimization. Although I can provide some general guidelines, you would probably need to talk with someone with more knowledge on the ingestion side
a
Hi @Gonzalo Ortiz please refer me to someone who is more informative on this part . This is one of the major blockers and concern right now
r
@Yarden Rokach @Maximilian Lau Will the recording from last week's session with Gonzalo be posted?
m
Sadly, there was a mix-up with the Zoom links for this session, so we did not record that one. Sorry about that!
r
again 😞