Alexandr Yankov
12/07/2025, 4:55 PMKevin Cai
12/08/2025, 2:35 AMAlexandr Yankov
12/08/2025, 6:33 AMKevin Cai
12/08/2025, 6:36 AMdevice or resource busy from pod fe-1, is suspecious.Alexandr Yankov
12/08/2025, 6:36 AMAlexandr Yankov
12/08/2025, 6:36 AMAlexandr Yankov
12/08/2025, 6:37 AMKevin Cai
12/08/2025, 6:38 AMAlexandr Yankov
12/08/2025, 6:40 AMlooked at the latest logsAlexandr Yankov
12/08/2025, 6:41 AMAlexandr Yankov
12/08/2025, 6:42 AMAlexandr Yankov
12/08/2025, 6:46 AMAlexandr Yankov
12/08/2025, 6:46 AMAlexandr Yankov
12/08/2025, 6:48 AMAlexandr Yankov
12/08/2025, 6:48 AMAlexandr Yankov
12/08/2025, 6:48 AMAlexandr Yankov
12/08/2025, 6:50 AMHe tries to get a token, but since the service has not yet started, he gets a connection refused message. I checked the ports, 8030 and the others are available on the leader, but not on this follower.Alexandr Yankov
12/08/2025, 6:51 AMKevin Cai
12/08/2025, 7:05 AMAlexandr Yankov
12/08/2025, 7:09 AMAlexandr Yankov
12/08/2025, 7:09 AMIt doesn't start, which means the service isn't sending traffic to it. There's no endpoint.Kevin Cai
12/08/2025, 7:11 AMfe-search service, not this fe-service service.Alexandr Yankov
12/08/2025, 7:11 AMKevin Cai
12/08/2025, 7:11 AMfe-searchAlexandr Yankov
12/08/2025, 7:13 AMAlexandr Yankov
12/08/2025, 7:14 AMUntil the main service is up, it won't route traffic or open ports. I even tried deleting the probes, but the result is the same.Kevin Cai
12/08/2025, 7:15 AMfe-search, it sets to ready immediately without checking the pod's probing.Alexandr Yankov
12/08/2025, 7:16 AMAs I understand it, he doesn't consider it ready for himself until he exchanges with the leader. I'm trying to solve this problem.Alexandr Yankov
12/08/2025, 7:18 AMI ran nginx on the same node, and when it started, traffic started flowing. It turns out the k8s node itself doesn't open the port until the pod starts.Alexandr Yankov
12/08/2025, 11:34 AMAlexandr Yankov
12/08/2025, 12:05 PMThe probes are running on the same port as the web server.Kevin Cai
12/08/2025, 12:14 PMfe-search in your env doesn't have fe-1 endpoint registered when the pod is up (not ready yet, just up), must be something wrong this service setup.Alexandr Yankov
12/08/2025, 12:15 PMKevin Cai
12/08/2025, 12:20 PMAlexandr Yankov
12/08/2025, 12:21 PMAlexandr Yankov
12/08/2025, 12:23 PMKevin Cai
12/08/2025, 12:24 PMAlexandr Yankov
12/08/2025, 12:28 PMAlexandr Yankov
12/08/2025, 12:32 PMAlexandr Yankov
12/08/2025, 12:33 PMAlexandr Yankov
12/08/2025, 12:46 PMAlexandr Yankov
12/08/2025, 12:50 PMAlexandr Yankov
12/08/2025, 12:51 PMKevin Cai
12/08/2025, 1:06 PMfe-search configuration, as you can see, you can't even reach the fe-1 8030 by ip address, it doesn't go through the fe-search serviceKevin Cai
12/08/2025, 1:06 PMAlexandr Yankov
12/08/2025, 1:20 PMAlexandr Yankov
12/08/2025, 1:23 PMAlexandr Yankov
12/08/2025, 1:33 PMKevin Cai
12/08/2025, 1:34 PM2025-12-08 13:31:58.629Z WARN (nioEventLoopGroup-4-1|147) [MetaBaseAction.isFromValidFe():114] request is not from valid FE. client: 172.21.0.101
fe-1
2025-12-08 13:32:38.687Z WARN (main|1) [NodeMgr.getFeNodeTypeAndNameFromHelpers():523] failed to get fe node type from helper node: 172.21.0.100:9010. response code: 400
2025-12-08 13:32:38.687Z WARN (main|1) [NodeMgr.getClusterIdAndRoleOnStartup():372] current node is not added to the group. please add it first. sleep 5 seconds and retry, current helper nodes: [172.21.0.100:9010]Kevin Cai
12/08/2025, 1:34 PMalter system add follower ...
fe-1
2025-12-08 13:32:43.696Z INFO (main|1) [NodeMgr.getFeNodeTypeAndNameFromHelpers():553] get fe node type FOLLOWER, name 172.21.0.101_9010_1765200762950 from 172.21.0.100:8030
2025-12-08 13:32:43.705Z INFO (main|1) [NodeMgr.getVersionFileFromHelper():701] Downloading version file from <http://172.21.0.100:8030/version>
2025-12-08 13:32:43.763Z INFO (main|1) [NodeMgr.getNewImageOnStartup():743] skip download image for /opt/starrocks/fe/meta/image, current version 0 >= version 0 from 172.21.0.100:9010
2025-12-08 13:32:43.763Z INFO (main|1) [NodeMgr.getClusterIdAndRoleOnStartup():499] Current run_mode is shared_nothing
2025-12-08 13:32:43.764Z INFO (main|1) [NodeMgr.getClusterIdAndRoleOnStartup():504] Got role: FOLLOWER, node name: 172.21.0.101_9010_1765200762950 and run_mode: shared_nothing
2025-12-08 13:32:43.765Z INFO (main|1) [BDBEnvironment.ensureHelperInLocal():351] start to check if local replica environment from /opt/starrocks/fe/meta/bdb contains 172.21.0.100:9010Alexandr Yankov
12/08/2025, 1:37 PMKevin Cai
12/08/2025, 1:38 PM10.230.x.x, and it shows the request from 10.128.0.16 in fe-0 log, which the fe-1's http request get rejected.Kevin Cai
12/08/2025, 1:40 PMAlexandr Yankov
12/08/2025, 1:48 PMAlexandr Yankov
12/08/2025, 1:49 PMKevin Cai
12/08/2025, 1:50 PMKevin Cai
12/08/2025, 1:50 PMAlexandr Yankov
12/08/2025, 1:52 PMKevin Cai
12/08/2025, 1:56 PMAlexandr Yankov
12/08/2025, 1:56 PMAlexandr Yankov
12/08/2025, 1:59 PMAlexandr Yankov
12/08/2025, 2:00 PMKevin Cai
12/08/2025, 2:01 PMkubectl get nodes -o wideAlexandr Yankov
12/08/2025, 2:03 PMKevin Cai
12/08/2025, 2:05 PMAlexandr Yankov
12/08/2025, 2:06 PMAlexandr Yankov
12/08/2025, 2:07 PMAlexandr Yankov
12/08/2025, 2:08 PMAlexandr Yankov
12/08/2025, 2:18 PMKevin Cai
12/08/2025, 2:27 PMKevin Cai
12/08/2025, 2:29 PMKevin Cai
12/08/2025, 2:31 PMsr-port-service, not sure if it impacts somehow. this is just my wild guess, may not make sense at all.Alexandr Yankov
12/08/2025, 2:35 PMAlexandr Yankov
12/08/2025, 2:47 PMexperimented for a couple of hours.
What assortment would you recommend?Alexandr Yankov
12/08/2025, 6:05 PMAlexandr Yankov
12/08/2025, 7:06 PMAlexandr Yankov
12/08/2025, 8:27 PMAlexandr Yankov
12/08/2025, 9:41 PMKevin Cai
12/09/2025, 12:02 AMAlexandr Yankov
12/09/2025, 6:40 AMAlexandr Yankov
12/09/2025, 7:00 AMAlexandr Yankov
12/09/2025, 7:02 AMA client connected via a load balancer makes a request from the Kubernetes node, not through the service....Kevin Cai
12/09/2025, 7:09 AMAlexandr Yankov
12/09/2025, 7:22 AMKevin Cai
12/09/2025, 9:33 AMAlexandr Yankov
12/10/2025, 5:45 AMAlexandr Yankov
12/10/2025, 5:49 AMThe K8S load balancer has a static public IP address, and the airflow request goes over the internet. Using virtual cables, you get two computers from different networks connected via a single cable. This makes them accessible via OSI Layer 2 over MAC. I'll test this theory today.Alexandr Yankov
12/10/2025, 11:51 AMFor some reason, a request is coming from the node's IP address?Kevin Cai
12/10/2025, 12:16 PMAlexandr Yankov
12/10/2025, 12:59 PMAlexandr Yankov
12/10/2025, 1:00 PM2025-12-10 16:00:00.286+03:00 WARN (starrocks-mysql-nio-pool-3|190) [StmtExecutor.execute():800] execute Exception, sql SHOW COMPUTE NODES
`org.apache.thrift.TApplicationException: Internal error processing forward`Alexandr Yankov
12/10/2025, 1:00 PMHow is this pool assembled? Can I look at the code?Kevin Cai
12/10/2025, 1:11 PMAlexandr Yankov
12/10/2025, 1:15 PMAlexandr Yankov
12/10/2025, 1:15 PMAlexandr Yankov
12/10/2025, 1:45 PMWe did this, and it handles single requests fine. But when we launch Airflow, we get this error. As far as I understand, some requests are coming from the Kubernetes node's IP and are being rejected.
I've contacted Cloud support, and maybe they can explain.Alexandr Yankov
12/10/2025, 1:50 PMrequests once per minute, request with Internal error processing forward, once every three minutes.Alexandr Yankov
12/10/2025, 2:36 PMAlexandr Yankov
12/12/2025, 8:37 AMCan you tell me what problem you're seeing? Again?
Cloud support said everything's fine, check your Starrocks settings.
I've read a lot of issues and seen that detailed problems are resolved by initing the container that parses DNS names, and then configuring fe.conf based on that.
Something like this.
- bash
- '-c'
- |
set -ex
hostname=`hostname`
ip=`ifconfig eth0|grep inet|awk '{print $2}'`
echo "$(sed 's/^'"$ip"'.*/'"$ip"' '"$hostname"'.starrocks-shared-data-fe-service.***.svc.dev-blz1.hefei '"$hostname"'/g' /etc/hosts)" > /etc/hosts
[[ $hostname =~ -([0-9]+)$ ]] || exit 1
ordinal=${BASH_REMATCH[1]}
if [[ $ordinal -eq 0 ]]; then
#first pod as leader default, FQDN host type
/opt/starrocks/fe/bin/start_fe.sh --host_type FQDN
else
#other pods as followers,helper is the first pod's service, FQDN host type
/opt/starrocks/fe/bin/start_fe.sh --helper starrocks-shared-data-fe-0.starrocks-shared-data-fe-service.***.svc.dev-blz1.***i:9010 --host_type FQDN
fiKevin Cai
12/12/2025, 9:06 AMKevin Cai
12/12/2025, 9:06 AMAlexandr Yankov
12/12/2025, 12:12 PMAlexandr Yankov
12/12/2025, 12:13 PMOf course, you can try and then look at the service manifests; there might be some option that will solve everything.Alexandr Yankov
12/12/2025, 1:16 PMAlexandr Yankov
12/12/2025, 1:21 PMIt would be interesting to understand and solve the problem)Kevin Cai
12/12/2025, 2:40 PMKevin Cai
12/12/2025, 2:42 PMAlexandr Yankov
12/12/2025, 3:41 PMIs there a way to specify allowed networks? By narrowing the range?Alexandr Yankov
12/12/2025, 3:46 PMAlexandr Yankov
12/12/2025, 3:47 PMYou can say, "I need manifests for this and that," I'll generate them and send them to your email. Take a look.Alexandr Yankov
12/12/2025, 4:44 PMAlexandr Yankov
12/14/2025, 7:41 AMAlexandr Yankov
12/14/2025, 7:42 AM