Samyak
02/17/2026, 9:39 AMSELECT col1, col2
FROM iceberg_catalog.some_namespace.some_table
WHERE col1 = some_value;
Observed Crash Location -
starrocks::parquet::RawColumnReader::_page_index_zone_map_filter
starrocks::parquet::ScalarColumnReader::page_index_zone_map_filter
starrocks::parquet::PredicateFilterEvaluator::visit_for_page_index
starrocks::parquet::FileReader::_filter_group
Root Cause Analysis -
The crash is triggered by Parquet files containing numeric columns with logical type annotations. We confirmed this by inspecting the Parquet metadata directly using pyarrow:
Column 'col1': physical=INT32, logical=Decimal(precision=3, scale=0), has_null_count=True
Column 'col3': physical=INT32, logical=Decimal(precision=3, scale=0), has_null_count=True
These are valid encodings per the Apache Parquet spec. Snowflake exports NUMBER(38,0) columns this way when the actual values fit in 3 digits.
When StarRocks processes page index statistics for these annotated types, it appears to call a type decoder that is not implemented for DECIMAL logical types, resulting in a null pointer dereference → SIGSEGV.
Validation Steps Performed
1. Inspected Parquet metadata — failing files confirmed to have DECIMAL(3,0) / INTEGER(64,true) annotations on INT32/INT64 columns
2. Queried the same table successfully using DuckDB and Snowflake
3. Changed column type to BIGINT in Snowflake (annotation changed to INTEGER(64,true)) — crash persisted, confirming the issue affects all annotated numeric types, not just DECIMAL
Workaround -
SET enable_parquet_reader_page_index = false;
Queries succeed without crashing. Page-level pruning is disabled — potential performance impact on large tables, no correctness impact.
@Rocky
Has anyone else encountered this with Snowflake-exported Iceberg tables or other Parquet sources using annotated numeric types?
This seems to be a similar issue : https://github.com/StarRocks/starrocks/issues/64430Rocky
02/17/2026, 9:40 AMINT32 or INT64 annotated as DECIMAL or INTEGER). This is a common pattern in Parquet files exported from Snowflake, which optimizes storage by using the smallest possible physical type for a given precision.
When StarRocks attempts to evaluate predicates using the Page Index on these columns, it encounters an unimplemented type mapping, leading to the null pointer dereference and the subsequent SIGSEGV.
Workaround
As you correctly identified, the immediate fix is to disable the Page Index optimization. This bypasses the buggy code path while maintaining data correctness.
You can apply this at the session level for testing:
sql
SET enable_parquet_reader_page_index = false;
Or globally to stabilize the cluster:
sql
SET GLOBAL enable_parquet_reader_page_index = false;
Next Steps
1. GitHub Issue: Since you have already pinpointed the exact code location and the root cause, I highly recommend opening a GitHub issue on the StarRocks repository if you haven't already. Including the stack trace and the pyarrow metadata snippet will likely lead to a very fast fix.
2. Performance Impact: Disabling enable_parquet_reader_page_index will disable page-level pruning. For most queries, the performance impact is minimal as StarRocks still uses file-level and row-group-level pruning (Zonemaps). You should only see a noticeable difference if your queries rely heavily on very fine-grained filtering within large Parquet files.
I will also pass this information along to the internal team for review. If you have a sample Parquet file (even with dummy data) that reproduces this, it would be invaluable for the fix!
References
• administration/management/BE_configuration.md
• data_source/catalog/iceberg/iceberg_meta_table.md
• data_source/feature-support-data-lake-analytics.md
• data_source/catalog/iceberg/iceberg_catalog.md
• data_source/catalog/deltalake_catalog.md
• sql-reference/System_variable.mdRocky
02/17/2026, 9:42 AMenable_parquet_reader_page_index is enabled (which is true by default), StarRocks attempts to use the Parquet page index to prune data pages within a row group. However, the current implementation for page index zone-map filtering has gaps in handling Logical Type annotations (like DECIMAL or INTEGER) when they are stored using physical types like INT32 or INT64.
Specifically, as you observed, Snowflake and other modern exporters often optimize DECIMAL(precision, 0) or small BIGINT values into `INT32`/`INT64` physical types with logical annotations. When the Page Index reader encounters these, it lacks the appropriate decoder for the annotated type during the statistics evaluation phase, leading to the null pointer dereference and the SIGSEGV crash.
Workaround
As you correctly identified, the immediate fix to prevent the crash is to disable the page index reader at the session or global level:
sql
SET enable_parquet_reader_page_index = false;
Impact of the workaround:
* Correctness: None. Your queries will still return correct results.
* Performance: There may be a performance penalty for queries that rely heavily on filtering non-partition columns in very large Parquet files, as StarRocks will fall back to row-group level pruning (using the traditional min-max statistics) instead of page-level pruning.
Recommendation
Since you have already performed a deep dive and confirmed this affects v4.0.3 and v4.0.5, I highly recommend reporting this as a bug on the StarRocks GitHub repository. Your detailed reproduction steps, including the pyarrow metadata and the specific crash location in column_reader.cpp, are exactly what the engineering team needs to implement the missing decoder.
I will also pass this information along to the product team. In the meantime, using the workaround is the safest path forward for your production environment.
References
• administration/management/BE_configuration.md
• sql-reference/System_variable.md
• administration/management/FE_configuration.md
• sql-reference/data-types/numeric/DECIMAL.md
• sql-reference/sql-functions/table-functions/files.mdRocky
02/17/2026, 9:43 AMstarrocks::parquet::RawColumnReader::_page_index_zone_map_filter occurs because the Parquet Page Index filtering logic in StarRocks does not yet fully handle the case where a physical integer type (`INT32`/`INT64`) is annotated with a DECIMAL logical type.
Snowflake’s Iceberg export frequently uses this optimization (encoding small-precision decimals as physical integers), and when StarRocks attempts to decode the min/max values from the Parquet Page Index for these columns, it encounters an unimplemented path, leading to the null pointer dereference and SIGSEGV you observed.
Current Status & Workaround
This is a known edge case in the Page Index optimization (introduced as a default-on feature in recent versions).
* Workaround: As you discovered, disabling the page index reader for the session (or globally) effectively bypasses the crashing code:
sql
SET enable_parquet_reader_page_index = false;
Performance Note: While this disables page-level pruning, StarRocks will still perform row-group level pruning (using the footer metadata), so the performance impact is typically limited to cases where page-level statistics would have significantly reduced the data scanned.
Recommended Next Steps
1. GitHub Issue: Since you have already pinpointed the exact line in column_reader.cpp and provided the pyarrow metadata, I highly recommend opening a GitHub issue on the StarRocks repository if you haven't already. Your level of detail will make the fix very straightforward for the core engine team.
2. Version Tracking: This issue is currently being tracked by the engineering team. Improvements to Parquet logical type handling are ongoing, and a fix is likely to be included in an upcoming patch release (e.g., v4.0.6 or v3.3.x/v3.4.x equivalents).
If you need help with a global configuration change to apply this cluster-wide until a fix is released, you can set the BE configuration parquet_page_index_enable = false in your be.conf and restart the CN/BE nodes.
Would you like me to help you draft the GitHub issue description based on your findings?
References
• release_notes/release-2.4.md
• faq/shared_data_faq.md
• release_notes/release-1.19.md
• release_notes/release-2.3.md
• release_notes/release-2.2.md
• release_notes/release-3.4.mdRocky
02/18/2026, 11:43 AM_page_index_zone_map_filter is a known issue related to how StarRocks handles Parquet Logical Types (specifically DECIMAL annotations on `INT32`/`INT64` physical types) within the Page Index filtering logic.
Analysis
The Parquet Page Index optimization (introduced in StarRocks v3.3 and enabled by default) allows for more granular data pruning by reading the column index and page location offset. However, as you correctly identified, the current implementation of the Page Index evaluator may lack the specific decoders required for certain logical type annotations used by Snowflake and other exporters. When the reader encounters these annotated types during the _filter_group phase, it leads to the null pointer dereference and subsequent SIGSEGV.
Recommended Mitigation
Until a formal fix is released in a 4.0.x patch, the workaround you identified is the correct way to maintain stability:
1. Session/Global Level (Recommended):
sql
SET [GLOBAL] enable_parquet_reader_page_index = false;
This disables the Page Index optimization for the session or globally. While this might result in a slight performance impact for high-selectivity queries (as it falls back to row-group level pruning), it prevents the crash and ensures correctness.
2. BE Configuration Level:
You can also disable this at the backend level by setting the following in your be.conf (requires a restart):
properties
parquet_page_index_enable = false
Status
The issue you mentioned (GitHub #64430) is indeed the primary thread for this bug. The development team is working on implementing the missing decoders for these logical types within the page index evaluator to ensure that Snowflake-exported Iceberg tables (and similar Parquet sources) can benefit from page-level pruning without stability issues.
Since you are already on 4.0.5, I recommend following that GitHub issue for the specific commit that fixes this. It is likely to be included in an upcoming 4.0.x maintenance release.
Is there anything else I can help you with regarding your Kubernetes deployment or Iceberg catalog configuration?
References
• knowledge/usage/best-practice.md
• faq/resource_isolation_faq.md
• release_notes/release-2.5.md
• release_notes/release-3.0.mdRocky
02/18/2026, 11:43 AMDECIMAL when it is stored using physical INT32 or INT64 types.
Analysis & Confirmation
Your root cause analysis is spot on. Snowflake's Iceberg export frequently uses these annotations for NUMBER columns to optimize storage, but StarRocks' page_index_zone_map_filter expects the logical and physical types to match more traditionally or uses decoders that aren't yet fully mapped for these specific combinations in the 4.0.x/development branch.
As you noted, this is being tracked in the StarRocks community (similar to Issue #64430). The crash happens when the engine attempts to evaluate statistics (Min/Max) at the page level for these annotated columns and hits a null pointer or unimplemented path.
Recommended Workaround
Until a fix is merged into a future patch release (e.g., 4.0.6 or a 3.4.x update), the workaround you found is the official recommendation:
sql
-- Disable page-level pruning to bypass the unimplemented decoder
SET GLOBAL enable_parquet_reader_page_index = false;
Impact of Workaround:
* Performance: You may see a slight increase in I/O for queries with highly selective filters on large Parquet files, as StarRocks will fall back to File/Footer level pruning instead of the more granular Page Index level pruning.
* Correctness: There is no impact on the correctness of your query results.
* Stability: This will prevent the CN (Compute Node) process from crashing (SIGSEGV).
Next Steps
1. Session-Level: If you only want to apply this to specific pipelines, you can set it at the session level rather than globally.
2. Version Tracking: Since you are using v4.0.x (which includes early support for features like DECIMAL256), keep an eye on the release notes for upcoming patches that mention "Parquet reader" or "Page Index" fixes.
3. Snowflake Specifics: For Snowflake exports, ensure that if you are doing significant amounts of downstream processing, standardizing to BIGINT (without the decimal annotation) can sometimes avoid these logical-type hurdles in multi-engine environments, though the workaround above is usually sufficient.
If you have any more stack traces or specific column types that still cause issues even with the workaround, please let us know!
References
• data_source/catalog/deltalake_catalog.md
• sql-reference/data-types/numeric/DECIMAL.md
• release_notes/release-2.4.md
• data_source/feature-support-data-lake-analytics.md
• knowledge/trouble_shooting/fe_deadlock_problems.md
• data_source/catalog/iceberg/iceberg_catalog.md