we have only one server, where that table has 35B ...
# troubleshooting
s
we have only one server, where that table has 35B rows
m
Need more info then
s
as of now we are using single node cluster
m
Is it happening for all queries or just one query? Was there an expensive query that was fired? What does the cpu/mem usage of the server look like
s
most of the queries its happening
cpu - going beyond 30
Screenshot from 2022-01-31 10-42-10.png
btw, we have 40 cpu core m/c
i noticed, there is suddenly surge in cpu%
m
My guess is expensive query(s) are eating up resources for that one node.
s
yeah
it seems so
m
Have you overridden the default time out of 10s?
s
yes
to 30sec
m
See if num docs/entries scanned metrics went up? That will confirm the theory
s
Result: {"exceptions":[{"errorCode":427,"message":"1 servers [ipaddr_O] not responded"}],"numServersQueried":1,"numServersResponded":0,"numSegmentsQueried":0,"numSegmentsProcessed":0,"numSegmentsMatched":0,"numConsumingSegmentsQueried":0,"numDocsScanned":0,"numEntriesScannedInFilter":0,"numEntriesScannedPostFilter":0,"numGroupsLimitReached":false,"totalDocs":0,"timeUsedMs":25000,"offlineThreadCpuTimeNs":0,"realtimeThreadCpuTimeNs":0,"segmentStatistics":[],"traceInfo":{},"numRowsResultSet":0,"minConsumingFreshnessTimeMs":0}
this is the result
n
Am I reading the error correctly? Is it actually saying <host.ip_O>? As in the letter O?
s
yes you are reading correct
🆗 1
any way to solve this?
either adding one more server?
or any other thing
m
_O in the name implies it was offline server.
👍 1
I meant to look at the metrics for numEntriesScanned (if you have monitoring set up with something like Grafana), If that went up, then it was due to expensive queries.
Another thing to check is if read qps went up.
s
after increasing timeout , query was success
Result: {"resultTable":{"dataSchema":{"columnNames":["avg(request_length)","ssl_protocol"],"columnDataTypes":["DOUBLE","STRING"]},"rows":[[376.10587859996525,"-"],[279.46690930735764,"TLSv1.2"],[113.68766410250731,"TLSv1.3"]]},"exceptions":[],"numServersQueried":1,"numServersResponded":1,"numSegmentsQueried":3561,"numSegmentsProcessed":3561,"numSegmentsMatched":3561,"numConsumingSegmentsQueried":0,"numDocsScanned":35249695297,"numEntriesScannedInFilter":0,"numEntriesScannedPostFilter":70499390594,"numGroupsLimitReached":false,"totalDocs":35249695297,"timeUsedMs":33517,"offlineThreadCpuTimeNs":0,"realtimeThreadCpuTimeNs":0,"segmentStatistics":[],"traceInfo":{},"numRowsResultSet":3,"minConsumingFreshnessTimeMs":0}
query: select AVG(a), b from "table_name" group by b order by b limit 1000
m
Yeah you are finding avg on 35B rows and grouping + sorting the data using single node.
Using star tree index will make this query much faster
s
yeah
i missed to add this in startree index conf
already using startree index
if i add this col in index config under startree, it will reindex almost 3500 segments
which is costly one, right?
m
Yes