exception_server_logs.txt
# troubleshooting
p
exception_server_logs.txt
this was the exception before i restarted server instance
exceptions in controller
m
What's your data dir on controller and server? Are you using
/tmp
? If so, perhaps that was getting cleaned up?
p
data dir on controller is /tmp on r4.4xlarge (ebs volume of size 1 TB), and on server is /data (mounted nvme ssd volume of size 7.5TB on i3en.3xlarge instance). i am not running anything else on the instance that is cleaning it up. i have configured s3 as deep store. server disc is at 2.5TB so far, and controller disc is at 4.8 GB so far.
i have two clusters set up. the one which is healthy has a much lesser disc usage on the zookeeper node in comparison to this unhealthy one. do you think that is the root cause or is that a symptom? both zookeeper nodes are running on same machine configuration and heap configuration.
3 GB v/s 87 GB
m
If you have s3 as deep store, then controller data dir should point to that and not /tmp?
p
you are correct. it is pointing to s3
Copy code
controller.data.dir=<s3://roku-dea-dev/pinot-segment-store/pinot-poc/controller-data>
controller.local.temp.dir=/tmp/pinot/
running into this problem again. i also noticed that the segments for the pinot cluster with star-tree index were not moved to Deleted_Segments in s3. segments were also not deleted locally from data servers. data from the newer segments is not available for query, and only data from older segments show up in query results.
here is my table config
is only controller leader backing up segments to s3? or other controllers in the group also be involved? i am noticing high memory usage on the pinot cluster with star-tree index in comparison to one without any index.
m
Retention is a periodic job (6hr iirc). So it may take time before the next retention kicks in. It should not have anything to do with whether there is star tree index or not.