Hi guys! Hey, I was publishing a lot of old logs I...
# troubleshooting
d
Hi guys! Hey, I was publishing a lot of old logs I have to Pinot, and it stopped being able to commit segments because the Controller got out of space, even though I have 20G in it - I reserved far more space in the Server instances though. How much space am I supposed to have for the Controller instance(s)? Why does it use so much space?
m
The deep store is expected to store the entire copy of segments across all tables in tar.gz compressed format. I think you are storing data on controller local disk. If you replace with deepstore then you won’t have this issue
d
Oh... I didn't know that... so the default deep store, if I don't explicitly specify one, is the Controller filesystem then? (Just to know how it works by default)
m
yeh if you don't configure anything it'll go under
/tmp/data/PinotController
d
Ouch... thanks for the heads up, guys! I'll make sure we have that properly configured, then.
Also, is there any docs for Controller replicas, like how to properly configure it to avoid conflicts in the committing of segments for example?
@Mayank, @Mark Needham, would it be OK to use EFS mounted in the Controller(s) as a deep store? Or, if S3 is the better solution, is there any documentation on how to integrate to it while still using the Pinot Helm chart?
m
EFS is independent of instances right?
if so it should be ok, but will let Mayank confirm
but maybe we need to explain how to integrate S3 with Helm also
m
Yes EFS works too, if you can mount controller data dir. for controller replication the only constraint is all replicas point to same data dir (deep store or efs). S3 is the most popularly used so may have all the configs available
d
Got it. EFS is a rather simple approach for us, but more expensive than S3, so ideally we'd like to use S3, but we need to learn how to integrate to it while still using the Helm chart. Is there any example of explanation of how to configure the chart for that?
From the Helm chart statefulset template (
<https://github.com/apache/pinot/blob/master/kubernetes/helm/pinot/templates/controller/statefulset.yaml>
), I saw that the config file is located at
/var/pinot/controller/config/pinot-controller.conf
, and that the volume is mounted from the ConfigMap from name
config
, which gets the Helm values from
pinot.controller.config
. So I just need to add those S3 deep store configurations to the
pinot.controller.config
value in my Statefulset deployment values, right?
I found out that the configs should go in the
controller.extras.config
value, actually, and
server.extras.config
as well. I have one more question, though, guys: I noticed that the configs for the S3 deep store are pointing to directories that differ from the ones I'm currently running at; For example, in one of my Servers I have the indexes at
/var/pinot/server/data/index
, which seem to be the default configuration, but then the deep store configs for the Server point to these directories instead:
Copy code
pinot.server.instance.dataDir=/tmp/pinot-tmp/server/index
pinot.server.instance.segmentTarDir=/tmp/pinot-tmp/server/segmentTars
Is this difference expected? Shouldn't the
pinot.server.instance.dataDir
config point to
/tmp/pinot-tmp/server/index
instead? Or is that because of the difference between normal ("hot") storage of segments and deep store?
Looks like I got it all wrong. I didn't realize that "deep store" is actually synonymous to "segment store", and I though the segments were actually stored locally to the Server instances and the "deep store" was something meant to be only a backup if the data was ever needed (like if segments corrupted for example). So how relevant is it the disk space for the Server instances? Currently I have them with a gp2 of 1TB each, but this is now looking like much more than they need; How much will we really need, how can I assess that?
@Mayank if you have a moment, please let me know if my latest assumption here is correct (sorry to keep bugging you)
m
Sure np. No docs for EFS (not aware of anyone using it). For S3: https://docs.pinot.apache.org/users/tutorials/use-s3-as-deep-store-for-pinot
Server tmp dir is used for tmp files only. Segments are on dataDir
d
Got it, I read that part and added some configs according to those docs, for my sysops to review and then redeploy our cluster; However, I'd like to understand how much space I'll need in the Servers' local disks - since the segment store will actually be in S3
m
Deep Store is indeed for the backup copy of segments. Servers have a local copy on EBS attached for serving (so if serer loses data, it downloads from deepstore)
d
Ah, then it does need to have a big storage space too, right? Like, a big EBS block?
m
Yes
d
Ah, ok then. I think I need to change that directory however; I was using
/var/pinot/server/data/index
before, if I change the dataDir to
/tmp/pinot-tmp/server/index
it means I'll end up storing the local segments out of the EBS blocks, right? (BTW I'm not even sure where
/var/pinot/server/data/index
came from, doesn't seem to be a configuration default)