Hi guys, I tried setting up S3 deep store for my P...
# troubleshooting
d
Hi guys, I tried setting up S3 deep store for my Pinot cluster, and used this part of config in the
pinot.yaml
file, in the
controller
config, for deploying via the official Helm chart:
Copy code
extra:
    # Note: Extra configs will be appended to pinot-controller.conf file
    configs: |-
      pinot.set.instance.id.to.hostname=true
      controller.task.scheduler.enabled=true
      # Note: change this to the real bucket, after creating it in S3
      controller.data.dir=s3://<redacted>
      controller.local.temp.dir=/tmp/pinot-tmp-data/
      pinot.controller.storage.factory.class.s3=org.apache.pinot.plugin.filesystem.S3PinotFS
      pinot.controller.storage.factory.s3.region=<redacted>
      pinot.controller.segment.fetcher.protocols=file,http,s3
      pinot.controller.segment.fetcher.s3.class=org.apache.pinot.common.utils.fetcher.PinotFSSegmentFetcher
Is this the correct configuration to do - to append these configs to the existing ones? Or should the approach have been different? I'm asking this because I found a weird directory in one of the Controller local filesystems:
Copy code
root@pinot-controller-1:/opt/pinot# du -sh /var/pinot/controller/data\,s3\:/
0	/var/pinot/controller/data,s3:/
By the looks of it, it seems like a configuration mistake somewhere [SOLVED! - SOLUTION:] Because the official Helm chart already defines that configuration, it ends up concatenating the one from the extras and the default one when loading from the final file, because that setting ends up defined twice in
pinot-controller.conf
inside the Controller container. The solution for this is to, instead of putting
controller.data.dir
as a one-line string in the
extra
configs
, just define that setting starting from the
controller
options in that YAML file, then
data
instead of
extra
, then
dir
, so that the option replaces the default value.
🌟 1
j
According to the following document, the plugin
pinot-s3
needs to be added to the JVM args of controller (and minion)
-Dplugins.dir=/opt/pinot/plugins -Dplugins.include=pinot-s3
d
From our logs, though, that plugin seems to have been loaded already
j
Ok, then I guess this JVM arg is no longer needed. That was the only difference I could see between my configuration. Otherwise, my config map is the same as yours. I also appended the S3 details to the existing one.
d
Ah, cool! Thanks man! 🙂
m
re: adding the plugin explicitly - I think you only need to do that if you don't want ALL plugins to be loaded. re: The
controller.data.dir
- you are right that it seems to be appending that value to another value. I'm not sure how we can get it to override the initial value. I guess @Mayank might know
d
Weird... well, I'm keeping an eye on things...
m
the problem is that it may be writing the segment backup files to that weird concatenated path! Keep an eye if it's doing that. If you see the segment files in S3 then it's probably fine.
d
Correct, and it's not fine, it's putting the files there indeed:
Copy code
root@pinot-controller-0:/opt/pinot# du -sh /var/pinot/controller/data,s3:/
166M	/var/pinot/controller/data,s3:/
and I see no objects in my S3 bucket
m
although I can't tell how exactly Ahmed removed one of those
data.dir
values
d
Ah, thanks man! Lemme try to find that out...
Thanks for the thread, guys! I made the same mistake, added the data dir to the extra configs, and they ended up duplicated in the final config file used by the Controller, just concatenating the data dirs. It's fixed now on my side too!
👍 1
m
I suppose we should improve user experience by avoiding/flagging such side effects.
d
If that's possible to do, I think it would be a good idea... or perhaps improving the docs, maybe, I dunno...
Just an update: now it's working fine, I'm getting the segments backed up to S3! Yay! 🙂
m
winning! Could you share the working config and where in the docs needs updating and then I'll update the docs tomorrow?
m
Thanks @Mark Needham @Diogo Baeder
d
Sure! By the way, it would be nice to have some docs about configuring all this together, perhaps - deploying Pinot in Amazon EKS with the Helm chart and S3 as deep store. I think it might be a common thing for other projects, maybe.
m
yeh that's a good idea
we can do some variations showing different deep stores
or whatever else
d
Yeah, that would be great! Maybe some recipes, I dunno...
These are the relevant parts of our configuration, I removed particular configs like node tolerance etc.
m
perfect. Thanks very much!
d
❤️
m
d
That's correct, although it's not clear to me exactly how they're merged with the other values files (I'm not even sure if we use the official values files, perhaps we don't, but even so these configs I sent I believe are relevant - but might require some reviewing from your guys, of course)
(Before using it for documentation, I mean)