<@U0A71G31CDV> I am trying to select the data pres...
# questions-and-troubleshooting
r
@Rocky I am trying to select the data present in the s3 bucket using parquet file but getting this error:
Copy code
select * from files  ( "path"="<s3://my-bukcet/test_load.parquet>", "format"="parquet", "aws.s3.region"="us-east-1",  "aws.s3.use_instance_profile"="false", "aws.s3.use_aws_sdk_default_behaviour"="true" );
ERROR 1064 (HY000): Access storage error. Error message: Forbidden
I added the access to this s3 bucket in the policy attached to the IAM role linked with the ServiceAccount of the starrocks pods.
r
Hi! I'm Rocky. I see you're encountering a
Forbidden
error when using the
FILES
table function to query S3 with an IAM Role for Service Accounts (IRSA). The
Forbidden
error typically indicates that the AWS credentials identified by StarRocks do not have the necessary permissions to access the bucket or the specific file, or that the credential provider chain is not correctly picking up your IRSA settings. Here are a few things to check and try: 1. Verify Property Settings When using IRSA in Kubernetes, StarRocks relies on the AWS C++ SDK to find the credentials via environment variables (
AWS_ROLE_ARN
and
AWS_WEB_IDENTITY_TOKEN_FILE
). Try changing your query to set
"aws.s3.use_instance_profile"="true"
. In many StarRocks versions, this flag acts as a toggle to enable the default AWS credential provider chain, which includes IRSA.
Copy code
sql
SELECT * FROM FILES (
    "path" = "<s3://my-bukcet/test_load.parquet>",
    "format" = "parquet",
    "aws.s3.region" = "us-east-1",
    "aws.s3.use_instance_profile" = "true"
);
(Note: Ensure there are no angle brackets
< >
around your S3 URI unless they were just placeholders in your example.)
2. Check FE Pod Permissions The
FILES
table function requires both the FE and the BE to have access to S3. * FE needs access to fetch the file schema and metadata. * BE needs access to scan the actual data. Make sure your
ServiceAccount
(with the IAM role annotation) is assigned to both the FE and BE deployments/statefulsets in your Kubernetes cluster. 3. IAM Policy Requirements Ensure the IAM role attached to your ServiceAccount has a policy that includes at least the following permissions: *
s3:GetBucketLocation
on
arn:aws:s3:::my-bukcet
*
s3:ListBucket
on
arn:aws:s3:::my-bukcet
_
s3:GetObject
on
arn:aws:s3:::my-bukcet/_
Example policy snippet:
Copy code
json
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Action": [
                "s3:GetBucketLocation",
                "s3:ListBucket"
            ],
            "Resource": "arn:aws:s3:::my-bukcet"
        },
        {
            "Effect": "Allow",
            "Action": "s3:GetObject",
            "Resource": "arn:aws:s3:::my-bukcet/*"
        }
    ]
}
4. Trust Relationship Confirm that the IAM Role's Trust Relationship correctly allows the OIDC provider of your EKS cluster to assume the role for the specific ServiceAccount and namespace you are using. 5. Debugging If it still fails, check the
<http://be.INFO|be.INFO>
or
fe.log
on the StarRocks pods. They often contain more detailed error messages from the AWS SDK that can help pinpoint if it's a "Token not found" or an "Access Denied" from S3. Let me know if changing
aws.s3.use_instance_profile
to
true
resolves the issue!
r
@Rocky This error while executing the query:
Copy code
2026-03-03 00:50:08.067+08:00 WARN (starrocks-mysql-nio-pool-94|44445) [StmtExecutor.execute():901] execute Exception, sql: {}, select * from FILES ( "path"="<s3://my-bukcet/test_load.parquet>", "format"="parquet", "aws.s3.region"="us-east-1",  "aws.s3.use_instance_profile"="false", "aws.s3.use_aws_sdk_default_behaviour"="true" ) limit 5
 com.starrocks.sql.analyzer.StorageAccessException: Access storage error. Error message: Forbidden
        at com.starrocks.sql.analyzer.QueryAnalyzer.resolveTableFunctionTable(QueryAnalyzer.java:1691)
r
Hi, I'm Rocky! The "Forbidden" error usually indicates that StarRocks is attempting to access S3 with credentials that do not have the required permissions for the specified bucket or path, or the bucket name itself might be incorrect. Here are a few things you should check to resolve this: 1. Check for Typos in the Path In your query, the path is
<s3://my-bukcet/test_load.parquet>
. Please verify if the bucket name is actually
my-bukcet
or if it should be
my-bucket
. S3 returns a "Forbidden" (403) error sometimes if the bucket name is wrong but similar to one you don't have access to, or due to DNS/routing issues. 2. Verify Authentication Parameters You are using
"aws.s3.use_aws_sdk_default_behaviour"="true"
. This tells StarRocks to use the default AWS credential provider chain (Environment variables, Java system properties,
~/.aws/credentials
file, or Instance Profile). If you want to rule out issues with the environment-level credentials, try providing the credentials explicitly in the query to see if it works:
Copy code
sql
SELECT * FROM FILES (
    "path" = "<s3://your-bucket-name/test_load.parquet>",
    "format" = "parquet",
    "aws.s3.region" = "us-east-1",
    "aws.s3.access_key" = "YOUR_ACCESS_KEY",
    "aws.s3.secret_key" = "YOUR_SECRET_KEY"
) LIMIT 5;
3. Check IAM Permissions The identity being used (whether via explicit keys, instance profile, or default SDK behavior) must have the following permissions in AWS: *
s3:GetObject
on the file path (
arn:aws:s3:::my-bucket/test_load.parquet
). *
s3:ListBucket
on the bucket (
arn:aws:s3:::my-bucket
). If you are using an IAM Role (Instance Profile), ensure that the StarRocks FE and BE nodes both have the role attached and that the trust relationship is configured correctly. 4. Region Mismatch Ensure that
"aws.s3.region"="us-east-1"
matches the actual region where your S3 bucket resides. While S3 is a global namespace, requesting the wrong region can sometimes result in access errors depending on the client configuration. 5. Troubleshooting Steps * Test with AWS CLI: From one of the StarRocks BE nodes, try running
aws s3 cp <s3://my-bukcet/test_load.parquet> -
to see if the machine itself has access using the same credentials. * Check FE/BE Logs: Look at
fe.log
and
<http://be.INFO|be.INFO>
for more detailed stack traces if the error persists after correcting the bucket name and credentials. If you fix the bucket name and still see the error, I recommend trying the explicit
access_key
and
secret_key
method first to confirm connectivity.
r
@Rocky If I give access to S3 bucket to the role attached with Node under whiich starrocks fe pod is running and update the query like this:
Copy code
select * from files  ( "path"="<s3://my-bucket/test.parquet>", "format"="parquet", "aws.s3.region"="us-east-1",  "aws.s3.use_instance_profile"="true", "aws.s3.use_aws_sdk_default_behaviour"="false" );
Then I am able to fetch the data. Why di I need to provide access to S3 bucket to node role when I am using serviceaccount which has the role attached which has pemirsison to s3 bucket
r
Hi there! I'm Rocky. The reason you are seeing this behavior is due to the specific parameters you've set in your
files()
table function. By setting
"aws.s3.use_instance_profile"="true"
, you are explicitly telling StarRocks to bypass other authentication methods and use the EC2 Instance Metadata Service. This service retrieves credentials from the IAM role attached to the Kubernetes Node (the physical/virtual machine) rather than the IAM role attached to your ServiceAccount. Why your current configuration requires the Node Role: 1.
"aws.s3.use_instance_profile"="true"
: This forces the StarRocks components (FE and BE) to look at the host machine's identity. 2.
"aws.s3.use_aws_sdk_default_behaviour"="false"
: This prevents the AWS SDK from using its standard "Credential Provider Chain," which is the mechanism that typically discovers the IRSA (IAM Roles for Service Accounts) tokens injected by Kubernetes. How to use the ServiceAccount (IRSA) instead: To use the IAM role attached to your ServiceAccount, you need to allow StarRocks to use the standard AWS SDK search path, which looks for the
AWS_WEB_IDENTITY_TOKEN_FILE
environment variable injected by the IRSA mechanism. Try updating your query to:
Copy code
sql
SELECT * FROM files (
    "path" = "<s3://my-bucket/test.parquet>",
    "format" = "parquet",
    "aws.s3.region" = "us-east-1",
    "aws.s3.use_instance_profile" = "false", -- Disable instance profile
    "aws.s3.use_aws_sdk_default_behaviour" = "true" -- Enable default SDK chain to find IRSA
);
Important considerations: * FE and BE Configuration: When using
files()
or External Catalogs, the FE handles metadata/schema inference and the BEs handle the actual data scanning. Both the FE and BE pods must be configured with the same
serviceAccountName
in their Kubernetes deployment manifests to ensure they both have permission to access the bucket. * Trust Relationship: Ensure your IAM Role's Trust Relationship allows the OIDC provider of your EKS cluster to assume the role for the specific namespace and service account name used by StarRocks. If you use the settings above, you should be able to remove the S3 permissions from the Node role and rely solely on the ServiceAccount role.
r
@Rocky this gives me this error:
Copy code
SELECT * FROM files (     "path" = "<s3://my-bucket/test.parquet>",     "format" = "parquet",     "aws.s3.region" = "us-east-1",     "aws.s3.use_instance_profile" = "false",      "aws.s3.use_aws_sdk_default_behaviour" = "true"  );
ERROR 1064 (HY000): Access storage error. Error message: failed to get file schema, path: <s3://my-bucket/test.parquet>, error: [Init parquet reader fail. IOError: BE access S3 file failed, SdkResponseCode=403, SdkErrorType=15, SdkErrorMessage=No response body., filename: <s3://my-bucket/test.parquet>]
r
Hello! I'm Rocky, and I'd be happy to help you with this error. The error
SdkResponseCode=403
specifically means "Forbidden." This indicates that AWS S3 is rejecting the request from the StarRocks Backend (BE) because it cannot verify the credentials or the credentials provided do not have permission to access the bucket/file. Based on your query, here are the most likely causes and how to fix them: 1. Missing or Incorrect Credentials In your query, you set
"aws.s3.use_aws_sdk_default_behaviour" = "true"
. This tells StarRocks to look for credentials in the environment variables (like
AWS_ACCESS_KEY_ID
) or the
~/.aws/credentials
file on every BE node. If those are not configured on the servers running StarRocks, the request will fail. Solution: Try providing the credentials directly to see if it resolves the issue:
Copy code
sql
SELECT * FROM files (
    "path" = "<s3://my-bucket/test.parquet>",
    "format" = "parquet",
    "aws.s3.region" = "us-east-1",
    "aws.s3.access_key" = "YOUR_ACCESS_KEY",
    "aws.s3.secret_key" = "YOUR_SECRET_KEY"
);
2. Using Instance Profiles (IAM Roles) If your StarRocks BE nodes are running on EC2 and have an IAM role attached, you should set
"aws.s3.use_instance_profile" = "true"
. Corrected Query for IAM Roles:
Copy code
sql
SELECT * FROM files (
    "path" = "<s3://my-bucket/test.parquet>",
    "format" = "parquet",
    "aws.s3.region" = "us-east-1",
    "aws.s3.use_instance_profile" = "true"
);
3. Check for Literal Angle Brackets In your error message, the path is shown as
<s3://my-bucket/test.parquet>
. If you literally included the
< >
brackets in your SQL string, S3 will treat that as part of the path and fail. Make sure the path is just the URL: * Incorrect:
"path" = "<s3://bucket/file.parquet>"
* Correct:
"path" = "<s3://bucket/file.parquet>"
4. IAM Permissions Ensure that the IAM user or role being used has the following permissions on the S3 bucket: *
s3:GetObject
(to read the file) *
s3:ListBucket
(required for StarRocks to locate the files) Summary of Parameters Parameter | Description
aws.s3.use_instance_profile
| Set to
"true"
if using an IAM role attached to the EC2 instance.
aws.s3.access_key
/
secret_key
| Use these for manual credential entry.
aws.s3.use_aws_sdk_default_behaviour
| Uses the standard AWS credential provider chain on the BE host. Try providing the
access_key
and
secret_key
first—if that works, you'll know it's a configuration issue with how the BE nodes are retrieving credentials!
r
@Rocky
Copy code
select * from FILES ( "path"="<s3://my-bucket/test.parquet>", "format"="parquet", "aws.s3.region"="us-east-1",  "aws.s3.use_instance_profile"="false", "aws.s3.use_aws_sdk_default_behaviour"="true" ) limit 5;
ERROR 1064 (HY000): Access storage error. Error message: Forbidden
[10:44 PM] But if i provide access to Node role, and update the query like this:
Copy code
select * from FILES ( "path"="<s3://my-bucket/test.parquet>", "format"="parquet", "aws.s3.region"="us-east-1",  "aws.s3.use_instance_profile"="true",) limit 5;
it worked without any error. I am not sure what I am missing here. I cross checked the ServciceAccount attached to the starrocks pods is having the role attached to it which has having the access to this bucket.
r
It looks like you are trying to use IAM Roles for Service Accounts (IRSA) with the
FILES()
table function in StarRocks, but you're hitting a
Forbidden
error. The issue is likely due to a small typo in your parameter name and how StarRocks interacts with the AWS SDK. 1. Correct the spelling:
behavior
vs
behaviour
StarRocks (and the underlying AWS SDK integration) uses the US English spelling for parameters. In your query, you used
behaviour
(with a 'u'), which is not recognized. Because it's unrecognized and
use_instance_profile
is set to
false
, the SDK fails to find valid credentials and defaults to an anonymous request, resulting in
Forbidden
. Try updating your query to:
Copy code
sql
SELECT * FROM FILES (
    "path" = "<s3://my-bucket/test.parquet>",
    "format" = "parquet",
    "aws.s3.region" = "us-east-1",
    "aws.s3.use_instance_profile" = "false",
    "aws.s3.use_aws_sdk_default_behavior" = "true" -- Corrected spelling
) LIMIT 5;
2. Why IRSA needs this setting When you set
"aws.s3.use_instance_profile"="true"
, StarRocks explicitly asks the SDK to look at the EC2 Metadata Service (the Node role). When you want to use IRSA (ServiceAccount): *
"aws.s3.use_instance_profile"
must be
"false"
. *
"aws.s3.use_aws_sdk_default_behavior"
must be
"true"
. This tells the AWS SDK to follow its default credential provider chain, which includes checking for the environment variables injected by IRSA (
AWS_WEB_IDENTITY_TOKEN_FILE
and
AWS_ROLE_ARN
). 3. Additional Troubleshooting If the spelling fix doesn't resolve it, please check the following: * Pod Environment Variables: Verify that the StarRocks BE (Backend) pods actually have the IRSA environment variables. Run:
Copy code
bash
    kubectl exec <be-pod-name> -n <namespace> -- env | grep AWS
You should see
AWS_ROLE_ARN
and
AWS_WEB_IDENTITY_TOKEN_FILE
. * Trust Relationship: Ensure the IAM Role's trust policy allows the OIDC provider of your EKS cluster to assume the role for the specific namespace and ServiceAccount name used by the StarRocks pods. * Explicit Role ARN: Sometimes, it helps to explicitly provide the IAM role ARN in the query to guide the SDK:
Copy code
sql
    SELECT * FROM FILES (
        "path" = "s3://...",
        "format" = "parquet",
        "aws.s3.iam_role_arn" = "arn:aws:iam::123456789012:role/your-irsa-role",
        "aws.s3.use_instance_profile" = "false",
        "aws.s3.use_aws_sdk_default_behavior" = "true"
    ) LIMIT 5;
Note: Ensure you are using a relatively recent version of StarRocks (2.5+), as IRSA support and the
use_aws_sdk_default_behavior
property were introduced/refined in later versions.
r
This fixed the issue: Correct the spelling:
behavior
vs
behaviour
😀 1
@Rocky Why this configuration not attaching PVC with CN pods?
Copy code
starrocksCnSpec:
  replicas: 1
  config: |
   datacache_disk_size = 80%
r
In the StarRocks Operator, simply adding a configuration like
datacache_disk_size
in the
config
section will not automatically trigger the creation or attachment of a Persistent Volume Claim (PVC). The
config
block only manages settings within the
cn.conf
file, but it does not manage Kubernetes infrastructure like storage. To attach a PVC to your CN pods for data caching, you must explicitly define the storage requirements using the
storageSpec
or
storageVolumes
field within the
starrocksCnSpec
. Recommended Configuration You need to add a
storageSpec
section to your YAML. This tells the Operator to create a PVC and mount it to the default storage directory (usually
/opt/starrocks/cn/storage
).
Copy code
yaml
starrocksCnSpec:
  replicas: 1
  # Add this section to provision and attach the PVC
  storageSpec:
    name: cn-cache-storage
    storageSize: 200Gi  # Specify the size of your cache disk
    storageClassName: "standard" # Use your K8s storage class (e.g., gp3, local-path, etc.)
  config: |
    # This setting now refers to 80% of the disk provisioned above
    datacache_disk_size = 80%
Why this happens: 1. Infrastructure vs. Application:
datacache_disk_size
is an application-level setting that tells the StarRocks process: "Use 80% of the disk space available in my storage directory." It cannot "ask" Kubernetes for a new disk. 2. Default Behavior: Without a
storageSpec
, the Operator deploys CN pods with ephemeral storage (the pod's root filesystem). If the pod restarts, any cached data is lost, and you risk filling up the node's root partition. 3. Storage Path: By default, the Operator mounts the
storageSpec
volume to
/opt/starrocks/cn/storage
. Since StarRocks looks for cache data in
storage_root_path
(which defaults to that same path), the PVC will be used for your data cache automatically once defined. Note: If you are using a shared-data (Compute-Storage Separation) architecture, ensuring your CN has persistent/large local storage for the Data Cache is critical for performance. Referencesquick_start/helm.mdsql-reference/sql-statements/cluster-management/storage_volume/CREATE_STORAGE_VOLUME.mdknowledge/usage/storage/hdfs_storage_volume_extended.md
r
@Rocky What is the recommended configuration for CN data caching. Should I go with PVC storage or default ephermal? And what happpens if I go with the default one. WIll it not cache the data if I am using S3 for data storage?
r
Hi there! I'm Rocky. For CN (Compute Node) data caching in a shared-data architecture (like when using S3), here is the recommended approach and the trade-offs between PVC and ephemeral storage: Recommended Configuration: PVC (SSD) For production environments, it is highly recommended to use PVC storage with SSDs for your CN data cache. * Persistence: A PVC ensures that the cached data survives pod restarts or rescheduling. If a CN pod restarts, it can immediately resume serving "hot" data from its local disk without having to fetch it from S3 again. * Performance: SSDs provide the low-latency I/O required to make the cache effective. The performance of a shared-data cluster matches a shared-nothing cluster only when the local disk cache is hit. * Stability: Using a dedicated PVC prevents the container's root filesystem from filling up, which could lead to node pressure or pod eviction. What happens if you use the "Default" Ephemeral Storage? If you don't configure a PVC, the CN will use ephemeral storage (typically the container's writable layer or an
emptyDir
). 1. Will it still cache data from S3? Yes. Even with ephemeral storage, StarRocks will still cache data. When you query data from S3, the CN downloads the blocks and stores them in the directory specified by
storage_root_path
(defaulting to
${STARROCKS_HOME}/storage/starlet_cache
). 2. The "Cold Start" Problem: The biggest issue with ephemeral storage is that the cache is wiped whenever the pod restarts. After a restart, every query will be a "cold start," meaning the CN must fetch data from S3 again. This leads to significantly higher query latency until the cache is "warmed up" again. 3. Risk of Disk Full: By default, the cache can occupy up to 80% of the available disk space. If this is running on the container's root disk, it can cause the entire node or pod to become unstable if not managed carefully via Kubernetes resource limits. Summary Table Feature | PVC (Recommended) | Ephemeral (Default) Cache Persistence | Yes (survives restarts) | No (lost on restart) Performance | High (especially with SSD) | Variable (depends on host disk) S3 Caching | Enabled by default | Enabled by default Best For | Production, stable performance | Testing, dev, or very high-frequency scaling Key Configurations In your
cn.conf
(or via Helm/Operator): * `storage_root_path`: Defines where the cache is stored (e.g.,
/opt/starrocks/storage
). * `datacache.enable`: Set to
true
(default) in your table properties to ensure data is cached. * `datacache_disk_size`: Can be used to limit the total cache size on disk. If you are using the StarRocks Operator, you can easily configure the
storageSpec
for CNs to use a
PersistentVolumeClaim
. Referencesdeployment/sr_operator.mdquick_start/helm.mddeployment/shared_data/hdfs.mdquick_start/shared-data.md
r
@Rocky I am using autoscaling in CN pods. WIll it create same number of PVCs' as number of CN pods or CN will share the PVC for data caching?
r
When using autoscaling for Compute Node (CN) pods in a Kubernetes environment (typically via the StarRocks Operator), each CN pod will have its own dedicated Persistent Volume Claim (PVC). They do not share a single PVC for data caching. Here is a breakdown of why and how this works: 1. One PVC per Pod CN pods are typically deployed as a
StatefulSet
(or managed by the StarRocks Operator in a similar fashion). Because the Data Cache depends on local disk I/O performance, StarRocks is designed to use storage that is "local" to the compute instance. In Kubernetes, this is achieved by using
volumeClaimTemplates
. * When the horizontal pod autoscaler (HPA) triggers the creation of a new CN pod, Kubernetes creates a new, unique PVC for that specific pod. * This ensures that each node has its own isolated cache space, avoiding the I/O bottlenecks and file-locking conflicts that would occur if multiple pods tried to write to the same shared volume. 2. Why they don't share a PVC * Performance: Sharing a single PVC (using
ReadWriteMany
mode) across multiple pods usually involves network-attached storage (like NFS), which is significantly slower than local disk or block storage (like EBS or local SSDs). This would negate the performance benefits of having a Data Cache. * Cache Management: The StarRocks cache engine manages files at the local path level. If multiple instances wrote to the same directory, they would overwrite or corrupt each other's cache metadata. 3. How "Cache Sharing" works instead While they don't share the same physical storage, StarRocks supports a feature called Cache Sharing (available in shared-data mode). * Instead of sharing a disk, CN nodes can fetch cache data from each other over the network. * If Pod A needs data that it doesn't have in its local PVC, it can check if Pod B has it in its PVC and pull it directly. * This feature minimizes the "cold start" performance hit when new CN pods are added during autoscaling. Summary for your setup: * Autoscaling: Will create $N$ PVCs for $N$ CN pods. * Storage Path: Configured via
storage_root_path
in
cn.conf
. * Recommendation: Use high-performance storage (like SSDs) for these PVCs to get the most out of the Data Cache. If you are worried about "cold" caches after scaling up, enable the
enable_datacache_sharing
system variable. Referencesdata_source/data_cache.mddeployment/shared_data/feature-support-shared-data.mddeployment/shared_data/hdfs.mddeployment/shared_data/s3.md
r
@Rocky How to enable
enable_datacache_sharing
in CN?
r
To enable the
enable_datacache_sharing
feature (also known as Cache Sharing) in StarRocks (available starting from v3.5.1), you need to configure both an FE configuration and a system variable. Cache Sharing allows nodes (CN or BE) to fetch cache data from other nodes over the network if a local cache miss occurs, which is particularly useful for maintaining performance during elastic scaling. Steps to enable Cache Sharing: 1. Enable FE Historical Node Tracing You must first enable the FE configuration item
enable_trace_historical_node
. This is a mutable parameter and can be set dynamically:
Copy code
sql
    ADMIN SET FRONTEND CONFIG ("enable_trace_historical_node" = "true");
2. Enable the System Variable Set the
enable_datacache_sharing
system variable to
true
. While this variable is
true
by default in v3.5.1 and later, it only takes effect if the FE configuration above is also enabled.
Copy code
sql
    SET GLOBAL enable_datacache_sharing = true;
Related Configurations *
datacache_sharing_work_period
: This system variable (default
600
seconds) defines the duration for which Cache Sharing remains active after a cluster scaling event. * CN Requirements: Ensure that the Data Cache itself is enabled on your CN nodes (which is the default). In
cn.conf
,
datacache_enable
should be
true
and a
storage_root_path
must be configured for the cache storage. How to Monitor You can verify if Cache Sharing is working by checking the following metrics in the Query Profile: * `DataCacheReadPeerCounter`: Number of reads from other nodes. * `DataCacheReadPeerBytes`: Total bytes read from other nodes. * `DataCacheReadPeerTimer`: Total time spent fetching cache data from peer nodes. Referencesdata_source/data_cache.mdsql-reference/System_variable.mddeployment/shared_data/feature-support-shared-data.mdusing_starrocks/caching/block_cache.mddeployment/shared_data/s3.mdadministration/management/FE_configuration.md
r
@Rocky What will happen if data cache volume storage gets filled up?
r
When the data cache (also known as Block Cache) storage reaches its configured limit, StarRocks handles it through an automatic eviction and overwrite mechanism. Here is what happens in detail: 1. Automatic Eviction Data Cache does not stop working when it reaches its limit. Instead, it triggers an eviction process based on the configured replacement policy (LRU or SLRU). * Default Limit: By default, StarRocks starts evicting data when disk usage reaches 80% of the disk capacity (or the value set in
storage_root_path
). * LRU (Least Recently Used): The blocks that haven't been accessed for the longest time are marked for eviction first. * SLRU (Segmented LRU): This policy splits the cache into "protected" and "evicted" segments to ensure that frequently accessed "hot" data isn't kicked out by one-time large scans. 2. "Overwrite" vs. Deletion A key behavior of the StarRocks Data Cache is that it does not delete files from the disk to free up space in the traditional sense. * When a block is "evicted," it is simply marked as overwritable. * New data coming into the cache will overwrite these old blocks directly. * Resulting behavior: If you use commands like
du
or
ls
on the cache directory, you will notice the disk usage stays at its historical peak and does not decrease even after eviction. This is normal and does not mean the cache is "stuck." 3. Impact on Performance * Cache Misses: Once the cache is full and starts evicting, you may experience more "cache misses" if your working set of data (the "hot" data) is larger than the allocated cache size. * Remote I/O: When a cache miss occurs, StarRocks must fetch the data from the remote storage (S3, HDFS, etc.), which may increase query latency compared to a cache hit. 4. What if the Physical Disk fills up? While the Data Cache manages its own quota (e.g., 80%), the physical disk can still reach 100% due to other factors like: * Runtime log files. * Core dump files from system crashes. * Persistent indexes for Primary Key tables. * Mixed deployments where BE/CN share disks with other services. If the physical disk becomes completely full, the backend (BE) or Compute Node (CN) may become unstable or crash. Important Configurations You can tune these behaviors using the following parameters: *
starlet_star_cache_disk_size_percent
: Sets the percentage of disk space the Data Cache is allowed to use (default is
80
). *
datacache_disk_size
: You can set a specific hard limit (e.g.,
2T
) instead of a percentage. *
block_cache_evict_policy
: Allows you to choose between
LRU
and
SLRU
. Summary: When the cache fills up, StarRocks gracefully cycles out the oldest data to make room for the newest, ensuring the system continues to function without manual intervention. Referencesusing_starrocks/caching/block_cache.mddata_source/data_cache_troubleshooting.mddata_source/data_cache.md
r
@Rocky Can we set this
enable_trace_historical_node
in the fe.conf during the server startup?
r
Yes, you can set
enable_trace_historical_node
in the
fe.conf
file before starting your Frontend (FE) server. According to the StarRocks documentation, this is an FE configuration parameter introduced in v3.5.1. Configuration Details: * Parameter name:
enable_trace_historical_node
* Default value:
false
* Type: Boolean * Is mutable: Yes (This means it can be modified both in the config file and at runtime using
ADMIN SET CONFIG
). * Description: This parameter allows the system to trace historical nodes. When enabled (
true
), it facilitates the Cache Sharing feature, allowing the system to select the appropriate cache nodes during elastic scaling. How to set it in `fe.conf`: Simply add the following line to your
fe.conf
file:
Copy code
properties
enable_trace_historical_node = true
After adding this and restarting the FE, the setting will take effect. Alternatively, because it is mutable, you can also enable it at runtime without a restart by executing the following SQL command:
Copy code
sql
ADMIN SET FRONTEND CONFIG ("enable_trace_historical_node" = "true");
Referencesadministration/management/FE_configuration.md
r
@Rocky What do you mean by
nable_trace_historical_node
is This parameter allows the system to trace historical nodes. When enabled (
true
), it facilitates the Cache Sharing feature, allowing the system to select the appropriate cache nodes during elastic scaling.
r
Hi! I'm Rocky, and I'm happy to help you understand the
enable_trace_historical_node
parameter. In StarRocks,
enable_trace_historical_node
is an FE (Frontend) configuration parameter introduced in v3.5.1. It plays a critical role in maintaining query performance during cluster scaling (adding or removing nodes) by enabling the Cache Sharing feature. Here is a breakdown of what it does and why it matters: 1. The Problem: Cache Invalidation during Scaling In a shared-data (compute-storage separation) architecture, BE nodes cache data from remote storage (like S3 or HDFS) on their local disks to speed up queries. StarRocks uses consistent hashing to determine which BE handles which data. * When you scale the cluster, the data-to-node mapping changes. * Data that was previously cached on "Node A" might now be assigned to "Node B". * Because "Node B" doesn't have that data in its local cache yet, it suffers a cache miss and must fetch the data from slow remote storage, leading to significant performance "jitter" or slowdowns during the scaling process. 2. The Solution: Cache Sharing The
enable_trace_historical_node
parameter allows the FE to "remember" or trace which nodes previously held the cache for specific data ranges before the scaling event occurred. When this parameter is enabled (
true
), it facilitates Cache Sharing: * Peer-to-Peer Retrieval: If a node (Node B) experiences a local cache miss for data that was recently moved to it, it doesn't immediately go to remote storage. * Historical Awareness: Instead, using the historical node information provided by the FE, Node B identifies that "Node A" previously cached this data. * Network Fetch: Node B fetches the cached data directly from Node A over the internal network. This is typically much faster than downloading it from remote storage. 3. How to Enable It To use this feature, you need to configure both the FE and the system variables: 1. FE Configuration (
fe.conf
):
Copy code
properties
    enable_trace_historical_node = true
2. System Variable (SQL):
Copy code
sql
    SET GLOBAL enable_datacache_sharing = true;
4. Related Parameters *
datacache_sharing_work_period
: (Default 600s) Defines how long the Cache Sharing feature remains active after a scaling event. Once the new nodes have "warmed up" their local caches, the system stops fetching from peers to save network bandwidth. *
historical_nodes_min_update_interval
: (Default 600s) Controls how often the system records snapshots of the cluster nodes to ensure the historical trace is accurate. Summary: This parameter is essential for Elastic Scaling scenarios where you want to add or remove compute power without sacrificing query latency during the transition period. Is there anything else I can help you with regarding cluster scaling or data caching? Referencesadministration/management/FE_configuration.mddata_source/data_cache.mdsql-reference/System_variable.mdadministration/management/BE_configuration.mddeployment/shared_data/feature-support-shared-data.md
r
@Rocky I attached PVC with CN but what does this disk Cache metrics means in :
show compute nodes\G;
Copy code
Starrocks > show compute nodes\G;
*************************** 1. row ***************************
        ComputeNodeId: 10023
                   IP: kube-starrocks-cn-0.kube-starrocks-cn-search.asd.svc.cluster.local
        HeartbeatPort: 9050
               BePort: 9060
             HttpPort: 8040
             BrpcPort: 8060
        LastStartTime: 2026-03-12 08:10:33
        LastHeartbeat: 2026-03-12 08:25:28
                Alive: true
 SystemDecommissioned: false
ClusterDecommissioned: false
               ErrMsg: 
              Version: 4.0.1-cd9df36
             CpuCores: 8
             MemLimit: 12.960GB
    NumRunningQueries: 0
           MemUsedPct: 2.08 %
           CpuUsedPct: 0.0 %
     DataCacheMetrics: Status: Normal, DiskUsage: 0B/0B, MemUsage: 36.7MB/2.5GB
       HasStoragePath: true
           StatusCode: OK
          StarletPort: 9070
             WorkerId: 1
        WarehouseName: default_warehouse
            TabletNum: 116
1 row in set (0.00 sec)
Why it is 0B/0B?
r
The
DiskUsage: 0B/0B
metric in
DataCacheMetrics
indicates that the StarRocks Compute Node (CN) has no disk quota allocated for the Data Cache (also known as Block Cache). Even if you have attached a Persistent Volume Claim (PVC) to the pod, the StarRocks process needs to be explicitly configured to use that storage for caching. Here is why it shows
0B/0B
and how to fix it: 1. The Quota is set to 0 by default In many configurations, the parameter
datacache_disk_size
defaults to
0
. When this is 0, StarRocks only uses memory for caching (as seen in your
MemUsage: 36.7MB/2.5GB
), and the disk cache is effectively disabled. To fix this: You need to set a non-zero value for the disk cache size in your
cn.conf
(or via the StarRocks Operator/Helm chart). You can specify a physical limit or a percentage:
Copy code
properties
# Example: Allocate 100GB for disk cache
datacache_disk_size = 100G

# OR allow it to use a percentage of the disk (e.g., 80%)
starlet_star_cache_disk_size_percent = 80
_Note: In shared-data mode, StarRocks takes the greater value between
datacache_disk_size
and
starlet_star_cache_disk_size_percent
._ 2.
storage_root_path
is not configured
StarRocks needs to know which directory to use for storing cache files. In a Kubernetes environment with a PVC, you must ensure the
storage_root_path
in
cn.conf
points to the exact mount point of your PVC. To fix this: 1. Identify where your PVC is mounted inside the CN pod (e.g.,
/opt/starrocks/storage
). 2. Update your `cn.conf`:
Copy code
properties
   storage_root_path = /opt/starrocks/storage
If
storage_root_path
is left empty, the disk cache will be disabled even if a quota is set. 3. Verify Configuration at Runtime You can check if the CN has recognized your settings by running the following SQL:
Copy code
sql
-- Check the disk paths being used
SELECT * FROM information_schema.be_configs WHERE NAME LIKE "%storage_root_path%";

-- Check the configured disk size
SELECT * FROM information_schema.be_configs WHERE NAME LIKE "%datacache_disk_size%" OR NAME LIKE "%starlet_star_cache_disk_size_percent%";
Summary of Metrics * DiskUsage: 0B / 0B: The first
0B
is the current used space; the second
0B
is the total allocated quota. Since the quota is
0B
, no data is being cached to disk. * MemUsage: 36.7MB / 2.5GB: Your memory-based cache is active and has a 2.5GB limit, which is why you see some usage there. Recommendation for Kubernetes If you are using the StarRocks Operator, make sure your
StarRocksCluster
YAML includes the path in the
config
section:
Copy code
yaml
spec:
  starRocksCnSpec:
    config: |
      storage_root_path = /opt/starrocks/storage
      datacache_disk_size = 100G
    storageVolumes:
      - name: storage
        mountPath: /opt/starrocks/storage
        storageClassName: standard
        storageSize: 200Gi
Once you apply these changes and restart the CN, the
0B/0B
should update to reflect your PVC's allocated cache size. Referencesadministration/management/BE_configuration.mdusing_starrocks/caching/block_cache.mddeployment/shared_data/feature-support-shared-data.md
r
@Rocky This is my cnspec conf :
Copy code
storageSpec:
    # Specifies the name prefix of the volumes to mount. If left unspecified,
    # `emptyDir` volumes will be used, which are ephemeral, and only for log.
    # The logs will be lost on pod restart when using emptyDir volumes.
    #
    # For persistent storage, specify a volume name prefix.
    # For example, using `cn` as the name prefix would BE appropriate.
    # Note: If the values of the following mountPaths are not default, you must set the name to cn.
    name: "cn-cache-storage"
    # The storageClassName represent the used storageclass name. If not set will use k8s cluster default storageclass.
    # You must set a name when you set storageClassName
    # Note: Because hostPath field is not supported here, hostPath is not allowed to be set in storageClassName.
    storageClassName: ""
    # the storage size of per persistent volume for data.
    storageSize: 100Gi
    # The number of persistent volumes for data.
    # If storageCount == 1,
    #   the storageMountPath field is used to specify the mount path of the persistent volume. If storageMountPath is empty,
    #   the storageMountPath will be set to /opt/starrocks/cn/storage.
    #   If storageMountPath is not /opt/starrocks/cn/storage, you must add in config the following configuration: storage_root_path = xxx.
    # If storageCount > 1,
    #   the storageMountPath field is used to specify the prefix of the mount path for the persistent volume. For example, if the
    #   storageMountPath is /opt/starrocks/cn/storage, the real mount path will be /opt/starrocks/cn/storage0, /opt/starrocks/cn/storage1, ...
    #   You must add in config the following configuration: storage_root_path = /opt/starrocks/cn/storage0;/opt/starrocks/cn/storage1;...
    storageCount: 1
    # see the comment of storageCount for the usage of storageMountPath.
    storageMountPath: ""
    # If not set will use the value of the storageClassName field.
    logStorageClassName: ""
    # The storage size of persistent volume for log, and the mount path is /opt/starrocks/cn/log.
    # If you set it to 0Gi, the related PVC will not be created, and the log will not be persisted.
    logStorageSize: 20Gi
    # If logMountPath is empty, the logMountPath will be set to /opt/starrocks/cn/log.
    # If logMountPath is not /opt/starrocks/cn/log, you must add in config the following configuration: sys_log_dir = xxx.
    logMountPath: ""
    # If not set will use the value of the storageClassName field.
    spillStorageClassName: ""
    # Setting this parameter can persist spill storage, and the mount path is /opt/starrocks/cn/spill.
    # If you set it to 0Gi, the related PVC will not be created, and the spill will not be persisted.
    # You need to add spill_local_storage_dir=/opt/starrocks/cn/spill in cn.conf.
    spillStorageSize: 0Gi
    # If spillMountPath is empty, the spillMountPath will be set to /opt/starrocks/cn/spill.
    # If spillMountPath is not /opt/starrocks/cn/spill, you must add in config the following configuration: spill_local_storage_dir = xxx.
    spillMountPath: ""
And also setting this in its helm values:
Copy code
starrocksCnSpec:
  config: |
   datacache_disk_size = 80%
But still when I run
Copy code
Starrocks > show compute nodes\G;
*************************** 1. row ***************************
        ComputeNodeId: 10005
                   IP: kube-starrocks-cn-0.kube-starrocks-cn-search.qwer.svc.cluster.local
        HeartbeatPort: 9050
               BePort: 9060
             HttpPort: 8040
             BrpcPort: 8060
        LastStartTime: 2026-03-12 14:47:31
        LastHeartbeat: 2026-03-12 14:53:51
                Alive: true
 SystemDecommissioned: false
ClusterDecommissioned: false
               ErrMsg: 
              Version: 4.0.1-cd9df36
             CpuCores: 1
             MemLimit: 1.620GB
    NumRunningQueries: 0
           MemUsedPct: 14.92 %
           CpuUsedPct: 0.4 %
     DataCacheMetrics: Status: Normal, DiskUsage: 0B/0B, MemUsage: 34.6MB/331.7MB
       HasStoragePath: true
           StatusCode: OK
          StarletPort: 9070
             WorkerId: 1
The DiskaUsage is 0?
r
Hello! I'm Rocky. It looks like you're having trouble with the Data Cache on your StarRocks Compute Nodes (CN). Based on the
show compute nodes
output and your configuration, the reason
DiskUsage
is showing
0B/0B
is likely because the disk cache quota was not correctly initialized. Here are the most likely causes and how to fix them: 1. Change the Disk Size Parameter The parameter
datacache_disk_size
typically expects an absolute value (like
80Gi
or
100G
). Using a percentage string like
80%
for this specific parameter might result in an invalid configuration, causing the quota to default to 0. For shared-data clusters (using CNs), you should use the dedicated percentage parameter: * Recommendation: Use
starlet_star_cache_disk_size_percent = 80
instead of
datacache_disk_size = 80%
. 2. Explicitly Set
storage_root_path
Although the default
storage_root_path
is usually
${STARROCKS_HOME}/storage
, in Kubernetes environments with the StarRocks Operator, the mount path can sometimes differ from what the CN process expects as its default. If your
storageSpec
is mounting the volume to
/opt/starrocks/cn/storage
, you should explicitly tell the CN process to use this path for the cache. 3. Ensure
datacache_enable
is True
While Data Cache is enabled by default in newer versions (v3.2.3+), it is best practice to set it explicitly in your configuration to avoid any ambiguity. ────────── Suggested Updated Configuration Try updating your
starrocksCnSpec
config as follows:
Copy code
yaml
starrocksCnSpec:
  config: |
    # Enable the Data Cache
    datacache_enable = true
    # Specify the path to the persistent volume mount
    storage_root_path = /opt/starrocks/cn/storage
    # Use the percentage parameter for shared-data cache
    starlet_star_cache_disk_size_percent = 80
Verification Steps 1. Restart the CN pods: Apply the changes to your Helm values or Custom Resource and wait for the pods to restart. 2. Check CN Logs: If it still shows
0B/0B
, check the CN logs (usually located at
/opt/starrocks/cn/log/cn.INFO
) for any errors related to "StarCache" or "DataCache". Common errors include permission issues or invalid paths. 3. Check Configuration via SQL: You can run this query to see what values the CN is actually using:
Copy code
sql
    SELECT * FROM information_schema.be_configs
    WHERE NAME LIKE "%datacache%" OR NAME LIKE "%storage_root_path%" OR NAME LIKE "%starlet_star%";
Let me know if this helps or if you see any specific error messages in the logs! Referencesdeployment/shared_data/feature-support-shared-data.mdusing_starrocks/caching/block_cache.mdadministration/management/BE_configuration.mddata_source/feature-support-data-lake-analytics.mddata_source/data_cache.md
r
@Rocky This is the out put:
Copy code
SELECT * FROM information_schema.be_configs
    ->     WHERE NAME LIKE "%datacache%" OR NAME LIKE "%storage_root_path%" OR NAME LIKE "%starlet_star%";
+-------+-------------------------------------------+---------------------------+--------+---------------------------+---------+
| BE_ID | NAME                                      | VALUE                     | TYPE   | DEFAULT                   | MUTABLE |
+-------+-------------------------------------------+---------------------------+--------+---------------------------+---------+
| 10005 | datacache_block_buffer_enable             | true                      | bool   | true                      |       0 |
| 10005 | datacache_block_size                      | 262144                    | int64  | 262144                    |       0 |
| 10005 | datacache_checksum_enable                 | false                     | bool   | false                     |       0 |
| 10005 | datacache_direct_io_enable                | false                     | bool   | false                     |       0 |
| 10005 | datacache_disk_adjust_interval_seconds    | 10                        | int64  | 10                        |       1 |
| 10005 | datacache_disk_idle_seconds_for_expansion | 7200                      | int64  | 7200                      |       1 |
| 10005 | datacache_disk_size                       | 80%                       | string | 100%                      |       1 |
| 10005 | datacache_enable                          | true                      | bool   | true                      |       0 |
| 10005 | datacache_eviction_policy                 | slru                      | string | slru                      |       0 |
| 10005 | datacache_inline_item_count_limit         | 130172                    | int32  | 130172                    |       1 |
| 10005 | datacache_max_concurrent_inserts          | 1500000                   | int64  | 1500000                   |       0 |
| 10005 | datacache_max_flying_memory_mb            | 2                         | int64  | 2                         |       0 |
| 10005 | datacache_mem_adjust_interval_seconds     | 10                        | int64  | 10                        |       1 |
| 10005 | datacache_mem_adjust_period               | 20                        | int64  | 20                        |       1 |
| 10005 | datacache_mem_size                        | 20%                       | string | 20%                       |       1 |
| 10005 | datacache_min_disk_quota_for_adjustment   | 107374182400              | int64  | 107374182400              |       1 |
| 10005 | datacache_persistence_enable              | true                      | bool   | true                      |       0 |
| 10005 | datacache_scheduler_threads_per_cpu       | 0.125                     | double | 0.125                     |       0 |
| 10005 | datacache_skip_read_factor                | 1                         | double | 1.0                       |       0 |
| 10005 | datacache_unified_instance_enable         | true                      | bool   | true                      |       0 |
| 10005 | enable_datacache_disk_auto_adjust         | true                      | bool   | true                      |       1 |
| 10005 | enable_datacache_mem_auto_adjust          | true                      | bool   | true                      |       1 |
| 10005 | internal_service_datacache_rpc_thread_num | -1                        | int32  | -1                        |       0 |
| 10005 | report_datacache_metrics_interval_ms      | 60000                     | int32  | 60000                     |       1 |
| 10005 | starlet_star_cache_async_init             | true                      | bool   | true                      |       0 |
| 10005 | starlet_star_cache_block_size_bytes       | 1048576                   | int32  | 1048576                   |       0 |
| 10005 | starlet_star_cache_disk_size_bytes        | 0                         | int64  | 0                         |       0 |
| 10005 | starlet_star_cache_disk_size_percent      | 80                        | int32  | 80                        |       0 |
| 10005 | starlet_star_cache_mem_size_bytes         | 134217728                 | int64  | 134217728                 |       1 |
| 10005 | starlet_star_cache_mem_size_percent       | 0                         | int32  | 0                         |       1 |
| 10005 | storage_root_path                         | /opt/starrocks/cn/storage | string | ${STARROCKS_HOME}/storage |       0 |
+-------+-------------------------------------------+---------------------------+--------+---------------------------+---------+
31 rows in set (0.17 sec)
What is wrong in this due to which
Copy code
DataCacheMetrics: Status: Normal, DiskUsage: 0B/0B, MemUsage: 34.6MB/331.7MB
DiskUsage is 0? this is : starrocks@kube-starrocks-cn-0:/opt/starrocks/cn$ df -h Filesystem Size Used Avail Use% Mounted on overlay 50G 13G 38G 25% / tmpfs 64M 0 64M 0% /dev /dev/root 315M 315M 0 100% /usr/local/sbin/modprobe /dev/nvme1n1p1 50G 13G 38G 25% /etc/hosts shm 64M 0 64M 0% /dev/shm /dev/nvme5n1 100G 791M 100G 1% /opt/starrocks/be/storage /dev/nvme4n1 20G 176M 20G 1% /opt/starrocks/be/log tmpfs 2.0G 4.0K 2.0G 1% /run/secrets/eks.amazonaws.com/serviceaccount tmpfs 7.8G 0 7.8G 0% /proc/acpi tmpfs 7.8G 0 7.8G 0% /sys/firmware
r
Hi! I'm Rocky, your StarRocks helper. Based on the configuration output you provided, your Data Cache (specifically the disk-based Block Cache) shows
DiskUsage: 0B/0B
(meaning the total allocated quota is 0), which usually happens for one of the following reasons: 1. Disk Auto-Adjustment (Common in v3.3+) You have
enable_datacache_disk_auto_adjust = true
. This feature dynamically scales the cache size based on the available disk space. * The Problem: If the disk where your
storage_root_path
(
/opt/starrocks/cn/storage
) resides is already quite full (e.g., above 70-80%), the system might automatically reduce the Data Cache quota to 0 to prevent the disk from running out of space. * Check: Run
df -h /opt/starrocks/cn/storage
on the node. If the usage is high, the auto-adjuster likely shrunk the cache. * Fix: You can try disabling the auto-adjust feature or setting a lower
disk_safe_level
(though increasing disk space is better).
Copy code
sql
    UPDATE be_configs SET VALUE = "false" WHERE NAME = "enable_datacache_disk_auto_adjust";
2. Percentage Calculation Issues You have
datacache_disk_size
set to
80%
. In some environments (like specific Docker or Kubernetes configurations), StarRocks might fail to correctly detect the total disk capacity, resulting in a calculated quota of 0. * Fix: Try setting an absolute value instead of a percentage to see if the cache initializes.
Copy code
sql
    UPDATE be_configs SET VALUE = "100G" WHERE NAME = "datacache_disk_size";
(Adjust
100G
to a value that fits your disk capacity).
3. Starlet Cache Parameter Override In some versions, the legacy parameter
starlet_star_cache_disk_size_bytes
might conflict if set to 0. Although you have the percent set to 80, try explicitly setting a size. * Note: Your output shows
starlet_star_cache_disk_size_bytes
is
0
and it is not mutable at runtime. If the absolute size remains 0 despite the percentage, you may need to set an absolute value in your
be.conf
and restart. 4. Permissions or Path Issues If the process cannot create the
datacache
subdirectory under
/opt/starrocks/cn/storage
, the initialization will fail. * Check: Look for a directory named
datacache
inside your storage path. * Logs: Check your
<http://be.INFO|be.INFO>
or
be.WARNING
logs for messages related to "Block Cache" or "Data Cache initialization". Search for errors like: *
failed to init block cache
*
no space left on device
*
permission denied
Recommendation 1. Set an absolute disk size to bypass detection issues:
UPDATE be_configs SET VALUE = "50GB" WHERE NAME = "datacache_disk_size";
2. Verify the disk usage on the host machine using
df -h
. 3. Check the logs (
storage/log/be.INFO
) to see if there were any errors during the Data Cache startup. Let me know if setting an absolute value changes the
DiskUsage
metrics! Referencesadministration/management/BE_configuration.mddeployment/shared_data/feature-support-shared-data.mdusing_starrocks/caching/query_cache.mdusing_starrocks/caching/block_cache.md
It appears there is a configuration mismatch between your StarRocks configuration and your actual disk mounts, which is the most likely reason why
DiskUsage
is showing as
0B/0B
. Based on the information you provided: 1. Path Mismatch: * In your
be_configs
, the
storage_root_path
is set to
/opt/starrocks/cn/storage
. * In your
df -h
output, your 100GB NVMe disk (
/dev/nvme5n1
) is mounted at
/opt/starrocks/be/storage
. * Result: StarRocks is looking for a storage directory at
/opt/starrocks/cn/storage
(which is likely on your 50G root partition
/
), while your intended cache disk is mounted elsewhere. 2. Why
0B/0B
?
* If the directory defined in
storage_root_path
does not exist, is not writable by the StarRocks user, or is not correctly initialized, the Data Cache (Block Cache) will fail to initialize its disk components. When initialization fails, the total quota defaults to
0B
. * Even if
/opt/starrocks/cn/storage
exists on the root partition, StarRocks might be failing to create the
datacache/
sub-directory required for the cache. Recommended Steps to Fix: 1. Correct the
storage_root_path
You need to ensure the StarRocks configuration matches your mount point. Since your disk is at
/opt/starrocks/be/storage
, update your
cn.conf
(or
be.conf
if applicable):
Copy code
properties
storage_root_path = /opt/starrocks/be/storage
2. Check Directory Permissions Ensure that the StarRocks user has full read/write permissions to the mount point:
Copy code
bash
sudo chown -R starrocks:starrocks /opt/starrocks/be/storage
sudo chmod -R 755 /opt/starrocks/be/storage
3. Restart the CN/BE Node The
storage_root_path
parameter has
MUTABLE = 0
, meaning it requires a restart of the process to take effect. 4. Verify in Logs After restarting, check the
<http://be.INFO|be.INFO>
log (typically in
/opt/starrocks/be/log/be.INFO
or your configured log path) for any errors related to "DataCache" or "BlockCache". Look for lines like: *
[warmup] failed to init block cache
*
block cache path not exists
5. Absolute Size Configuration (Optional) If the percentage (
80%
) is still not reflecting correctly after fixing the path, try setting an absolute byte value for
starlet_star_cache_disk_size_bytes
in your configuration to force a specific quota: * For 80GB (80% of 100GB):
starlet_star_cache_disk_size_bytes = 85899345920
Summary of findings: The metric
0B/0B
specifically means the "Total Quota" is zero. This almost always points to the Data Cache failing to initialize because it cannot find or access the directory specified in
storage_root_path
. Correcting the path to
/opt/starrocks/be/storage
and restarting should resolve this. Referencesusing_starrocks/caching/block_cache.mddata_source/data_cache_observe.mddata_source/data_cache_troubleshooting.mddata_source/data_cache.md
r
@Rocky Seeing these logs in CN pod /opt/starrocks/cn/log/WARN* files:
Copy code
W20260312 16:26:07.199614 140410819720768 cache_file.h:130] write star cache file data/0000000000132031_43bb6026-f5f8-4cdb-88b5-b5f294b715cf.dat failed, errmsg: INTERNAL: fail to write block because no valid cache space
r
The warning log
INTERNAL: fail to write block because no valid cache space
indicates that the Data Cache (StarCache) on your Compute Node (CN) is unable to write new data blocks to the local disk. This typically happens when the cache system has reached its configured storage limit or the physical disk is exhausted. In StarRocks shared-data clusters or when using Data Cache for external catalogs, the CN uses local disk space to cache remote data for better performance. Potential Causes 1. Disk Quota Reached: The cache has hit the limit defined by
starlet_star_cache_disk_size_percent
(default is 80%) or
datacache_disk_size
. 2. Physical Disk Full: The underlying disk is physically full due to other files (logs, core dumps, or persistent indexes), leaving no room even if the cache quota hasn't been met. 3. Metadata Corruption: In some older versions (v3.1.x/v3.2.x), metadata bloat or corruption could prevent the cache from allocating space correctly. Recommended Troubleshooting Steps 1. Check Physical Disk Usage Run
df -h
on the CN pod/node to check the available space on the partition where
storage_root_path
is located. * If the disk is at 100% or near it, check for large log files in
log/
or core dumps. * Check the size of the
datacache
or
starlet_cache
directory under your
storage_root_path
. 2. Verify Cache Configuration Check the current limits in your
cn.conf
or via SQL:
Copy code
sql
-- Check disk size percentage (default 80)
SELECT * FROM information_schema.be_configs WHERE NAME = 'starlet_star_cache_disk_size_percent';

-- Check absolute disk size limit
SELECT * FROM information_schema.be_configs WHERE NAME = 'datacache_disk_size';
If your disk also stores other data (like many logs), consider reducing
starlet_star_cache_disk_size_percent
to
70
or
60
to ensure the OS and other processes have enough headroom. 3. Clear the Cache (Emergency Fix) If the error persists and queries are failing, you can manually clear the cache. This will not cause data loss as the master data resides in remote storage (S3/HDFS), but it will temporarily slow down queries until the cache is refilled. 1. Stop the CN node. 2. Delete the cache data and metadata:
Copy code
bash
    # Path is usually under storage_root_path
    rm -rf <storage_root_path>/datacache/*
    # Or for older versions/specific configs
    rm -rf <storage_root_path>/starlet_cache/*
3. Restart the CN node. 4. Version-Specific Issue (v3.1.x / v3.2.x) If you are on an older version (e.g., v3.1.8 or v3.2.3), there is a known issue where cache metadata can consume excessive space. * Fix: Upgrade to v3.1.9+, v3.2.4+, or v3.3.x. * Workaround: Clear the
<storage_root_path>/starlet_cache/star_cache/meta
directory and restart the CN. Summary of configurations to check in `cn.conf`: * `storage_root_path`: Where the cache is stored. * `datacache_enable`: Should be
true
. * `starlet_star_cache_disk_size_percent`: The percentage of disk space dedicated to cache. Referencesusing_starrocks/caching/block_cache.mddata_source/data_cache_troubleshooting.md
r
@Rocky WHy these errors in CN pod logs?
Copy code
W20260312 19:32:10.201527 139816318520704 memory_lock.cpp:46] mlock failed for  143073280-349753344 (206680064 bytes): Cannot allocate memory
W20260312 19:32:10.327111 139816318520704 timezone_utils.cpp:93] not found timezone:Africa/Casablanca
W20260312 19:32:10.327775 139816318520704 timezone_utils.cpp:93] not found timezone:Africa/El_Aaiun
W20260312 19:32:10.451163 139816318520704 disk_space_monitor.cpp:192] The current available disk space is too small, so disable the disk cache directly. If you still need it, you could reduce the value of `datacache_min_disk_quota_for_adjustment`, path: /opt/starrocks/cn/storage/datacache
W20260312 19:32:12.062283 139816318520704 jit_engine.cpp:476] System or Process memory limit is less than 16GB, disable JIT. You can set jit_lru_cache_size or jit_lru_object_cache_size a properly positive value in BE's config to force enabling JIT
W20260312 19:32:12.077826 139816318520704 starrocks_be.cpp:148] 'thrift_port' is deprecated, please update be.conf to use 'be_port' instead!
W20260312 19:32:15.136991 139813040158272 heartbeat_server.cpp:222] kube-starrocks-cn-0.kube-starrocks-cn-search.olymsre01.svc.cluster.local not equal to to backend localhost 10.214.70.64
W20260312 19:32:27.171322 139814173079104 cache_file.h:130] write star cache file <<FC><A7>^s<E3>/   <C0>q)<99><8C>c<BC>h/SCHEMA_0000000000002724 failed, errmsg: INTERNAL: fail to write block because no valid cache space
W20260312 19:44:09.024158 139814223435328 cache_file.h:130] write star cache file data/000000000000004c_87f572a5-0d93-4b82-9175-6b78e3420e44.dat failed, errmsg: INTERNAL: fail to write block because no valid cache space
r
Hello! I'm Rocky. Looking at your CN (Compute Node) pod logs, these messages are a mix of environmental warnings, configuration deprecations, and resource limits. Here is a breakdown of what they mean and how to address them: 1. Memory and JIT Warnings *
mlock failed ... Cannot allocate memory
: * Reason: The CN process tried to lock memory to prevent it from being swapped out, but the system (or the Kubernetes pod's memory limit) denied the request. * Solution: Check your Pod's memory
limits
and
requests
. Also, check the host's
ulimit -l
(max locked-in-memory address space). In Kubernetes, you may need to increase the memory limit or ensure the node has enough free memory. *
System or Process memory limit is less than 16GB, disable JIT
: * Reason: StarRocks enables Just-In-Time (JIT) compilation by default only if the node has at least 16GB of memory. This is to prevent JIT's own memory overhead from causing OOM (Out Of Memory) issues on smaller nodes. * Solution: If you want JIT, increase the Pod's memory limit to 16GB or more. Otherwise, this warning is harmless and JIT will simply be disabled. 2. Disk and Data Cache Errors *
The current available disk space is too small, so disable the disk cache directly
: * Reason: The CN's
storage/datacache
directory is on a disk that is nearly full or doesn't meet the minimum space requirements for the Data Cache to function safely. * Solution: Increase the volume size assigned to your CN pods. You can also adjust the threshold using the BE/CN config
datacache_min_disk_quota_for_adjustment
, but increasing disk space is the recommended path. *
fail to write block because no valid cache space
: * Reason: This is a direct consequence of the disk space issue above. The CN is trying to cache data from remote storage (like S3/HDFS) to the local disk for performance, but it has run out of allocated or physical space. * Solution: Check your
storage_root_path
disk usage and the
datacache_disk_size
configuration in
cn.conf
. 3. Networking and Heartbeats *
heartbeat_server.cpp:222] ... not equal to to backend localhost ...
: * Reason: The FE (Frontend) is sending a heartbeat to the CN, but there is a mismatch between the address the CN is using to identify itself and the address the FE is using to reach it (e.g., Pod IP vs. FQDN). * Solution: Ensure you have configured
priority_networks
in your
cn.conf
so the CN picks the correct network interface. For example:
priority_networks = 10.214.70.0/24
. 4. Configuration and Environment *
'thrift_port' is deprecated, please update be.conf to use 'be_port' instead!
: * Reason: You are using an older configuration key. * Solution: In your
cn.conf
(or via the Helm chart
config
section), rename
thrift_port
to
be_port
. *
not found timezone: Africa/Casablanca
: * Reason: The container image's timezone database is missing specific timezone files. * Solution: This is generally harmless unless you are performing queries that specifically require these timezones. To fix it, you would need to ensure
/usr/share/zoneinfo
is fully populated in the image. Summary of Recommendations: 1. Increase Resources: Your CN appears to be running with less than 16GB of memory and very little disk space. Try increasing memory to at least 16GB and providing more storage for the Data Cache. 2. Fix Configs: Update
thrift_port
to
be_port
and set
priority_networks
to match your Pod network. Referencesusing_starrocks/caching/block_cache.mddata_source/data_cache_troubleshooting.md