Hi Team, We are observing very high latency when ...
# troubleshooting
a
Hi Team, We are observing very high latency when querying with lookup . Without lookup: 500ms With lookup: 10secs Can someone help please Query:
SELECT
lookUp('dim_tabl1', 'display_name', 'dim_joinkey', fact_joinkey),
sum(metric1) AS sum_1
FROM reporting_aggregations
WHERE stats_date_hour between '2022012000' and '2022012223'
GROUP BY lookUp('dim_tabl1', 'display_name', 'dim_joinkey', fact_joinkey)
ORDER BY sum(metric1)
DESC LIMIT 10000
k
do you have the response stats from the query?
a
Copy code
timeUsedMs:12581
numDocsScanned:42834397
totalDocs:4017929622
numServersQueried:2
numServersResponded:2
numSegmentsQueried:444  
numSegmentsMatched:3
numConsumingSegmentsQueried:3
numEntriesScannedInFilter:1
numEntriesScannedPostFilter:85668794   
numGroupsLimitReached:false
partialResponse:-
minConsumingFreshnessTimeMs:1643726394648   
offlineThreadCpuTimeNs:0
realtimeThreadCpuTimeNs:0
k
that seems slow to lookup 4Million entries.. @Yupeng Fu, @Jackie ^^
s
Lookup use case is mostly for Fact table lookups with DIM tables. In this case DIM table is of size 700KB(2k records) and FACT table is 15M/day.
m
would be running the lookup 85,668,794 times right? (Not sure if it does any caching in case of doing the same lookup)
k
my bad its 42M and not 4 Million
10 seconds is still quite high
I looked at the code and looks like there is some locking going on for every key lookup and few other things which is definitely time consuming. Is it possible to create an issue and attach the profile output? @Richard Startin can guide you on how to generate the jfr profile dump
a
We have an existing olap setup on druid, with similar data set and volume. We are actually trying to replace this druid setup only with Pinot for upsert functionality. Lookups on druid works pretty fast(in subseconds).
r
hi @Anish Nair can you log on to the pod and run
jcmd <pid> JFR.start duration=60s settings=profile filename=slow_lookups.jfr
while querying the server with lookup queries please?
a
@Sumit Lakra can you do the above please
r
this should be fairly conclusive what's going on, especially if it's contention
a
yes. @Sumit Lakra will provide the dump in sometime.
s
Hi, @Richard Startin here is the jfr file created while running the query using the shared command.
r
it looks like there's some actionable low hanging fruit in lookup, so we should be able to improve this, are you able to deploy from nightly OSS releases to try it out?
thanks @Sumit Lakra
s
Yes, we can try out nightly releases
r
very conclusive profile, thanks. A couple of issues: 1. a lot of time waiting on locks in
DimensionTableDataManager
2. it's not the bottleneck because point 1. is throttling progress, but the top allocations come from unnecessary boxing in the lookup function implementation itself, if 1. were fixed this would become limiting
this is going to be a lot more straightforward than the issues you've had with upsert, I'll get back to you about when this will be addressed
👍 3
s
Thanks Richard
a
@Richard Startin possible to share the PR, so that we can track the progress
actually we were planning to push this Pinot setup in a production next week and this issue is a blocker. It would be really helpful if you could share information about approximate time it can be fixed so that we can plan our release accordingly. @Kishore G @Richard Startin
k
We will try our best. Thanks again for reporting the issue. Will keep you posted
r
this has been merged, please experiment with nightly builds tomorrow https://github.com/apache/pinot/pull/8102
a
thanks @Richard Startin, we are trying to do the deploy this build. will inform once its done.
s
Hi @Richard Startin @Kishore G I cloned the latest nightly release and tried to run the same with our existing pinot-server.conf file. The server starts up but stops on its own with no relevant error logs. This here shared is the output of the process. Can you help us debug it please
This is our server’s conf
r
there seems to be a mix of plugin versions (0.9.2 and 0.10.0-SNAPSHOT)
a
so we should be using the latest i.e 0.10 , is that correct ?
r
yes, it looks like you're loading 2 of everything
a
Hey @Richard Startin, we deployed the build. But now when we are querying on table with lookups. It is returning "null". Interestingly, when querying on OFFLINE table, values are returning, but timings are not improved, it is same as yesterday.
r
can you share another profile (it will confirm whether you are running the right build or not)
if you have frames like these, it's not the right build or the build doesn't have the fix
if you don't have these frames, there is a new bottleneck
s
@Anish Nair @Richard Startin here are the profiles from the 2 servers
r
you're not running with the fix
did you get the image from here? https://hub.docker.com/r/apachepinot/pinot/tags
s
No, we are not using docker. We cloned the repo from
Copy code
<https://github.com/apache/pinot.git>
r
ok
you need to pull changes from upstream then
that lock taking the time in your offline profile doesn't exist any more
s
So earlier we were using pinot binaries from their download site. I cloned the repo for the first time today so the master branch must be up to date
r
then you can't have deployed the right binaries
s
I am not sure I understand if the present build is wrong or our last one. Earlier we used 0.9.2 version downloaded from https://pinot.apache.org/download/ . Today we tried cloning the git repo and building from it
r
can you query the /info endpoint - that has git commit hashes
s
Screenshot 2022-02-02 at 17.15.03.png
this is git log --oneline output
r
you need to verify you've deployed the right binaries, because the stack frame in your profile doesn't exist on master
so query the cluster's /info endpoint, it has commit hashes for each component
s
okay let me check that
r
sorry i meant the /version endpoint, that's the one with commit hashes
s
Copy code
{
  "pinot-protobuf": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-kafka-2.0": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-avro": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-distribution": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-csv": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-s3": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-segment-uploader-default": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-yammer": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-thrift": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-confluent-avro": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-batch-ingestion-standalone": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-orc": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-azure": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-gcs": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-dropwizard": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-hdfs": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-adls": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-kinesis": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-json": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-minion-builtin-tasks": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-parquet": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438",
  "pinot-segment-writer-file-based": "0.10.0-SNAPSHOT-ea2f0aa641e17301293662c8e79dfd94d8568438"
}
looks different
r
let me take a look, you've definitely deployed something you've built
can you make sure the offline server is redeployed
ea2f0aa641e17301293662c8e79dfd94d8568438 is indeed the latest commit on master, but the offline server you took a profile from can't be running that code
let me see if the profile captured any version info for that offline server
here's its classpath:
Copy code
CLASSPATH_PREFIX	/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-client-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-httpfs-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-native-client-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-nfs-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/common/hadoop-common-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/guava-11.0.2.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/hadoop-annotations-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/hadoop-auth-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/gson-2.2.4.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/commons-configuration2-2.1.1.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/htrace-core4-4.1.0-incubating.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-databind-2.7.8.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-annotations-2.7.8.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-core-2.7.8.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-core-asl-1.9.13.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-jaxrs-1.9.13.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-mapper-asl-1.9.13.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-xc-1.9.13.jar
sorry that's the CLASSPATH_PREFIX you set
here's the classpath
Copy code
java.class.path	/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-client-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-httpfs-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-native-client-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/hdfs/hadoop-hdfs-nfs-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/common/hadoop-common-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/guava-11.0.2.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/hadoop-annotations-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/hadoop-auth-3.0.0.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/gson-2.2.4.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/commons-configuration2-2.1.1.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/htrace-core4-4.1.0-incubating.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-databind-2.7.8.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-annotations-2.7.8.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-core-2.7.8.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-core-asl-1.9.13.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-jaxrs-1.9.13.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-mapper-asl-1.9.13.jar:/pinot/hadoop-3.0.0/share/hadoop/common/lib/jackson-xc-1.9.13.jar:/pinot/lib/pinot-all-0.9.0-jar-with-dependencies.jar:/pinot/lib/jmx_prometheus_javaagent-0.12.0.jar:/pinot/lib/jmx_prometheus_javaagent-0.16.1.jar:/pinot/lib/pinot-all-0.10.0-SNAPSHOT-jar-with-dependencies.jar:/pinot/plugins/pinot-input-format/pinot-csv/pinot-csv-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-input-format/pinot-csv/pinot-csv-0.9.0-shaded.jar:/pinot/plugins/pinot-input-format/pinot-avro/pinot-avro-0.9.0-shaded.jar:/pinot/plugins/pinot-input-format/pinot-avro/pinot-avro-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-input-format/pinot-parquet/pinot-parquet-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-input-format/pinot-parquet/pinot-parquet-0.9.0-shaded.jar:/pinot/plugins/pinot-input-format/pinot-thrift/pinot-thrift-0.9.0-shaded.jar:/pinot/plugins/pinot-input-format/pinot-thrift/pinot-thrift-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-input-format/pinot-json/pinot-json-0.9.0-shaded.jar:/pinot/plugins/pinot-input-format/pinot-json/pinot-json-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-input-format/pinot-protobuf/pinot-protobuf-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-input-format/pinot-protobuf/pinot-protobuf-0.9.0-shaded.jar:/pinot/plugins/pinot-input-format/pinot-orc/pinot-orc-0.9.0-shaded.jar:/pinot/plugins/pinot-input-format/pinot-orc/pinot-orc-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-input-format/pinot-confluent-avro/pinot-confluent-avro-0.9.0-shaded.jar:/pinot/plugins/pinot-input-format/pinot-confluent-avro/pinot-confluent-avro-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-batch-ingestion/pinot-batch-ingestion-standalone/pinot-batch-ingestion-standalone-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-batch-ingestion/pinot-batch-ingestion-standalone/pinot-batch-ingestion-standalone-0.9.0-shaded.jar:/pinot/plugins/pinot-batch-ingestion/pinot-batch-ingestion-spark/pinot-batch-ingestion-spark-0.9.0-shaded.jar:/pinot/plugins/pinot-batch-ingestion/pinot-batch-ingestion-hadoop/pinot-batch-ingestion-hadoop-0.9.0-shaded.jar:/pinot/plugins/pinot-segment-uploader/pinot-segment-uploader-default/pinot-segment-uploader-default-0.10.0-SNAPSHOT.jar:/pinot/plugins/pinot-segment-uploader/pinot-segment-uploader-default/pinot-segment-uploader-default-0.9.0.jar:/pinot/plugins/pinot-stream-ingestion/pinot-kafka-2.0/pinot-kafka-2.0-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-stream-ingestion/pinot-kafka-2.0/pinot-kafka-2.0-0.9.0-shaded.jar:/pinot/plugins/pinot-stream-ingestion/pinot-kinesis/pinot-kinesis-0.9.0-shaded.jar:/pinot/plugins/pinot-stream-ingestion/pinot-kinesis/pinot-kinesis-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-segment-writer/pinot-segment-writer-file-based/pinot-segment-writer-file-based-0.10.0-SNAPSHOT.jar:/pinot/plugins/pinot-segment-writer/pinot-segment-writer-file-based/pinot-segment-writer-file-based-0.9.0.jar:/pinot/plugins/pinot-minion-tasks/pinot-minion-builtin-tasks/pinot-minion-builtin-tasks-0.9.0.jar:/pinot/plugins/pinot-minion-tasks/pinot-minion-builtin-tasks/pinot-minion-builtin-tasks-0.10.0-SNAPSHOT.jar:/pinot/plugins/pinot-environment/pinot-azure/pinot-azure-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-environment/pinot-azure/pinot-azure-0.9.0-shaded.jar:/pinot/plugins/pinot-metrics/pinot-dropwizard/pinot-dropwizard-0.9.0-shaded.jar:/pinot/plugins/pinot-metrics/pinot-dropwizard/pinot-dropwizard-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-metrics/pinot-yammer/pinot-yammer-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-metrics/pinot-yammer/pinot-yammer-0.9.0-shaded.jar:/pinot/plugins/pinot-file-system/pinot-hdfs/pinot-hdfs-0.9.0-shaded.jar:/pinot/plugins/pinot-file-system/pinot-hdfs/pinot-hdfs-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-file-system/pinot-adls/pinot-adls-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-file-system/pinot-adls/pinot-adls-0.9.0-shaded.jar:/pinot/plugins/pinot-file-system/pinot-gcs/pinot-gcs-0.9.0-shaded.jar:/pinot/plugins/pinot-file-system/pinot-gcs/pinot-gcs-0.10.0-SNAPSHOT-shaded.jar:/pinot/plugins/pinot-file-system/pinot-s3/pinot-s3-0.9.0-shaded.jar:/pinot/plugins/pinot-file-system/pinot-s3/pinot-s3-0.10.0-SNAPSHOT-shaded.jar
it's running pinot 0.9.0, not 0.10.0-SNAPSHOT
because the jars are mixed up, and 0.9.0 always comes first
if you fix that and redeploy, you should see how big the improvement is, and then we can look at next steps, although we need to be more measured about more changes to the lookup function
a
@Richard Startin we redeployed. Now the query is returning null for both the queries
image.png
r
how frequently are you creating new segments?
a
3-4 per day is the current freq for realtime.
r
do you have metrics to confirm that?
a
i am checking the segments list of Realtime Table.
r
the problem is, the only way this could return null is if the reference data map is empty, which would happen if it's being updated a lot, which would also explain the high contention
the assumption behind lookup is that the reference data changes very rarely
a
you mean, the dimension table?
r
yes
a
dimension tables data is not changed.
r
then there wouldn't have been any contention...
a
i can see data , if i query directly on dimension table.
r
I'm looking at your realtime profile, this was definitely running the new code, and you said it was returning null?
a
yes. it was returning null.
r
let me try to reproduce this, it's very unlikely to have been caused by the change
by very unlikely i mean inconceivable, there may be something else on master which triggered this
but first: before, when the query was slow, was it producing correct results or timing out?
a
results were getting produced, when timeout was increased to 15-30seconds
r
ok, if that's changed it will be something else on master
I'll add some observability surrounding this map so we can know what's in it and how big it is without taking a heap dump
could you take a heap dump and inspect the map?
if you search the heap dump for
DimensionDataTableManager
you can see what's in
_lookupTable
a
@Sumit Lakra check this once.
r
Copy code
jcmd <realtime_pid> GC.heap_dump filename=realtime
and
Copy code
jcmd <offline_pid> GC.heap_dump filename=offline
a
also, i am seeing error while running batch ingestion now.
Copy code
Caught temporary exception while pushing table: dim_partners segment: dim_partners_OFFLINE_0 to <http://d9-max-insert-2.srv:9000>, will retry
org.apache.pinot.common.exception.HttpErrorStatusException: Got error status code: 500 (Internal Server Error) with reason: "Exception while uploading segment: null" while sending request: <http://d9-max-insert-2.srv:9000/v2/segments?tableName=dim_partners&tableName=dim_partners&tableType=OFFLINE> to controller: d9-max-insert-2.srv, version: Unknown
r
so is the dimension table there or not?
can you look at your realtime and server logs for stack traces?
a
its there currently. this is different dimension table, i was trying to ingest. The one which returned null is different.
s
@Anish Nair @Richard Startin heap dump reports for offline. For realtime, analysis is underway
r
these aren't going to be very useful if I can't follow the links again
s
I understand. For that I will have to share all the index files, and the dump too I think. Offline dump was 180 MB but realtime is 20 GB
r
can you just find
DimensionDataTableManager
 you can see what's in 
_lookupTable
instead?
s
let me check
r
things that would explain the wrong query result: • the map is empty • if it's not empty, pick a key from the main table and check if it's in the map
another thing, I'm not exactly sure how this query will get planned:
Copy code
SELECT 
lookUp('dim_tabl1', 'display_name', 'dim_joinkey', fact_joinkey), 
sum(metric1) AS sum_1
FROM reporting_aggregations
WHERE stats_date_hour between '2022012000' and '2022012223'
GROUP BY lookUp('dim_tabl1', 'display_name', 'dim_joinkey', fact_joinkey)
ORDER BY sum(metric1) 
DESC LIMIT 10000
can you try to rewrite it as
Copy code
SELECT 
lookUp('dim_tabl1', 'display_name', 'dim_joinkey', fact_joinkey) as dim, 
sum(metric1) AS sum_1
FROM reporting_aggregations
WHERE stats_date_hour between '2022012000' and '2022012223'
GROUP BY dim
ORDER BY sum(metric1) 
DESC LIMIT 10000
a
it is resulting into error. PQLParsingError:
r
did you replace the template names properly? aliasing should work fine
a
Hey, syntax issue resolved. but still seeing null.
r
sure, but keep the query that way for once this is resolved
a
sure.
r
also change to
ORDER by sum_1
👍 1
I haven't been able to reproduce what you're seeing on another system except when the fact table and dim table have no keys in common
so I will wait for you to introspect the heap dump to figure out if that's the case
a
can i try rebuilding dimension table? you see anything related ? select on dim table is giving result. so not sure
r
have you checked they actually join?
a
yes. on 0.9.2 version it did.
same data
r
have you checked?
a
Values are present. It should join.
r
there's not a great deal I can do to help you without getting the info I've requested because I can't reproduce the problem without constructing a data set which doesn't join.
a
Hey @Richard Startin, we are not able to find the class in heap dump. will it be possible to get on quick call ?
r
evidence of an empty lookup table via the heap dump would suggest that building the lookup table failed somehow, but I doubt that's happened
a
Hey @Richard Startin @Kishore G We are trying to create heap Dump by doing the following activity. 1. Starting the heap dump 2. hitting the queries for a 1-2mins. 3. Stop the heap dump Is this the right way? Please guide us.
r
we'd need to know whether the map called
_lookupTable
in the class
DimensionDataTableManager
is empty or not, you can do this by introspecting a heap dump, you don't need to be running queries at the same time
s
This is what I am able to see
r
can you scroll down a bit, want to see what's in the field
table
s
I don’t see anything with just name ‘table’. There’s _tableDataDir, _tableSchema alongside _lookupTable. Can you help with the hierarchy a bit please
r
table should be within _lookupTable - it's a hash map
you've expanded the class, which isn't super interesting
s
This is the top branch
r
i see, this is MAT, in JVisualVM you can basically browse the objects
s
I see. Let me try with that
r
let me put together an example quickly
s
Is this useful ?
Similarly table = null for DimensionTableDataManager#2 and DimensionTableDataManager#3 as well.
r
so here you can see you can browse the contents of the lookup table
yours is empty, which explains why the join is failing
or not resolving
I think this is a symptom of another problem. Are there any logs indicating the dimension table data failed to load?
s
Let me check for them
a
@Richard Startin, the dimension table was loaded when we were on 0.9.2. and when i am doing select * from dim_table. data is showing up.
r
right, but you had seriously high contention on 0.9.2 suggesting the write lock was often held
a
currently, when i am trying to upload again, it is failing. Batch ingestion is not working. But Segment is still present over servers.
r
now it uses a lock-free approach with a CAS loop, and the update rate might be so high that the update never succeeds
a
So , what do you suggest to move forward? any troubleshoot
r
it's difficult for a couple of reasons 1) this is not my primary focus or enterprise support, I have other things to do 🙂 2) it feels like there is an underlying issue with your cluster that would need to be diagnosed, the contention should not have been so high, and the CAS loop should have succeeded, unless there is a very high update rate
I would suggest to go back to 0.9.2, keep the query changes, and see if they help
in the meantime, I will change the lock free update logic to not ensure update ordering, which will guarantee progress, but we need to understand how to reproduce high update rates
a
can you help me understand update rate here?
r
I'm sorry, I've already spent a lot of unplanned time on this and am going to spend more to refine the fix to remove contention
as mentioned before, 0.9.2 is safe, and you may be able to improve performance by streamlining the query by using aliases
s
We understand. Thank you for your time and support nonetheless :)
a
We tried with 0.9.2 with aliases. Without Lookup:
image.png
With Lookup:
No improvement in timings.
r
I'm fairly certain that the map being empty is caused by this; https://apache-pinot.slack.com/archives/C011C9JHN7R/p1643876335533179 and not by removing the locks
so we may have a controller bug on master
a
@Xiang Fu
x
I see, we should fix this, can you create a github issue and describe how to reproduce this ? @Anish Nair
a
sure.
@Xiang Fu issue created
r
it looks like this will fix your issue https://github.com/apache/pinot/pull/8132
a
Hey @Richard Startin, this fixes the batch ingestion ?
r
yes, and as mentioned before, joins not working on 0.10.0-SNAPSHOT is likely a symptom of that problem
so what happened is you picked up another bug with the fix, this PR fixes the new bug
a
Okay. Thanks for the update. We will try the build in coming days.
This build worked. Now the latency is minimal and batch ingestion is also working. thanks @Richard Startin @Xiang Fu @Kishore G
👍 1
r
yes, it was unfortunate you picked up a regression along with the perf fix, thanks @Mark Needham for the bug fix!
a
Hey guys, faced another issue regarding this. The Lookups were working fine until now, currently when did a lookup query it was returning null again. No config change or anything from our side. Reran the ingestion batch again, after ingestion values were visible. Any know reason ? @Richard Startin @Xiang Fu
r
I'll have a think later
we've discovered other issues with things like schema updates (it's a community feature) so there may be other issues with updates