## pola-rs/polars — rs-0.54.4…rs-0.55.1

_321+ commits._

### Features
- **feat: Add `struct.drop()` (#28666)** (bdd7f8e)
- **feat: Improve error message when CSV name de-duplication fails (#28658)** (73bc476)
- **feat: Always keep first metadata per source for Parquet (#28661)** (a203389)
- **feat: Improve plan-time row estimates for multi-file parquet scans (#28380)** (7394e54)
- **feat(python): Expose array and plugin function views in Python visitor (#28635)** (49a1931)
- **feat(python): Issue FutureWarning on using `from_arrow` of ArrowStreamExportable (#28442)** (a1cfca8)
- **feat(rust): Serialize and deserialize `SinkTypeIR::Partitioned` (#28112)** (3b507a7)
- **feat: Add `infer_schema_files` parameter to `scan_csv` (#28440)** (f85168b)
- **feat(rust): Support reading IEEE 754 total order Parquet column order (#27896)** (b78761c)
- **feat(rust): Expose more Expr nodes for cudf_polars (pt 2) (#28404)** (d757b6e)
- **feat: Allow callback sinks on cloud (#28458)** (bba7ac6)
- **feat: Pre-partition group-by on hive keys  (#28444)** (adb194d)
- **feat: Add ewm_sum and ewm_sum_by #28151 (#28215)** (997ce5c)
- **feat: Serve stale records from object_store DNS cache (#28256)** (bea2971)
- **feat(rust): Expose more expr nodes for cudf-polars (#28117)** (4065955)
- **Add partition_hive optimization flag** — Introduced a new optimization flag `partition_hive` that enables pre-partitioning of Hive-partitioned joins and group-by operations. This flag works in conjunction with predicate pushdown and is now configurable through the lazy frame optimization settings. (5252537)
- **Add Arrow C Stream scanning** — Introduced `scan_arrow_c_stream`, a new function that lets you read data from any object using the Arrow C Stream interface. This enables efficient zero-copy data exchange with other libraries that support the Arrow standard. (e69ee81)
- **Smarter parentheses in expression display** — Improved how Polars displays intermediate expressions by only adding parentheses when mathematically necessary for clarity, rather than wrapping every operation. This makes query explanations and debug output cleaner and easier to read. (17d6ebb)
- **Add Iceberg initial-default field support** — Enhanced native Iceberg scanning to support the `initial-default` field attribute alongside existing identity-transformed partition fields. This allows Polars to handle default values defined in Iceberg table schemas when reading data files. (3e0efd9)
- **Filter dtype mismatch errors** — When comparing DataFrames with mismatched data types, error messages now show only the columns that differ instead of all columns. This makes it easier to spot the actual problem when comparing large DataFrames with mostly matching schemas. (dc91f6a)

### Fixes
- **redo patch** (5958ff9)
- **fix: Preserve nulls when importing Arrow maps (#28680)** (103c8cf)
- **fix: Propagate null `by` column values in `rolling_*_by` (#27367)** (9336423)
- **fix: Fix self-referencing `field` in `struct.with_fields` with `over` (#28678)** (b507b75)
- **fix: Fix Arrow buffer offset for `Utf8` and `Binary` (#28662)** (2ecc3d8)
- **fix: Clamp group-by slice offset (#28579)** (a9d1c0e)
- **Fix categorical fill_null min/max** — Fixed the fill_null operation for categorical columns to properly use lexical (alphabetical) ordering when filling null values with "min" or "max" strategy. This ensures categorical data is filled correctly based on the actual category values rather than internal codes. (268a863)
- **Fix GIL deadlock in arrow stream** — Fixed a deadlock that occurred when resolving schemas in the `__arrow_c_stream__` method by releasing Python's Global Interpreter Lock (GIL) during schema resolution. This resolves issues when using Python-based dataset scans like `scan_delta` or `scan_iceberg` that need to call back into Python from another thread. (75d9820)
- **Fix Series data corruption from Arrow** — Fixed a bug where converting Arrow LargeList data into Polars Series would corrupt nested data. The fix properly handles nested types during conversion while preserving the fast_explode optimization flag when appropriate. (397f032)
- **Fix enum metadata propagation** — Fixed an issue where metadata wasn't being properly preserved when converting dictionary-type columns to Iceberg column mappings. The fix ensures that when processing enum columns, any associated metadata is now carried through the conversion process. (ae588a9)
- **Fix parquet field IDs for enums** — Parquet files with enum and categorical columns now correctly preserve field ID metadata when writing. Previously, field IDs were lost during the conversion of these special data types. (f6c766a)
- **Fix binary view Arrow export** — Fixed an issue where Arrow C interoperability was incorrectly exporting binary view buffer pointers, causing data corruption when inline values appeared before non-inlined ones. The fix ensures buffers are exported at the correct offset-adjusted pointer location rather than the allocation start. (1ec0668)
- **Optimize len() on concatenated data** — When calculating the length of concatenated or unioned DataFrames, the query optimizer now pushes the len() operation down to each input before combining them, then sums the results. This improves performance by computing lengths more efficiently rather than calculating length on the final combined data. The change includes tests to verify the optimization works correctly for both concat and union operations. (af24b92)
- **Fix duplicate hive partition values** — Fixed an issue where hive partition rewrites could produce duplicate values when partitions repeat across files. The fix deduplicates partition values and properly handles null partitions (represented as `__HIVE_DEFAULT_PARTITION__`) in both group-by and join operations. (e79bdb0)
- **Fix Arrow export offset calculation** — Corrected a bug where offset values were being double-counted when exporting sliced Series with Array types to Arrow format. The fix ensures that offset calculations properly handle alignment and avoid duplication when converting nested array structures. (7a6cdbc)
- **Fix equality checks in sorts and joins** — Fixed how Struct, List, and Array types are compared when sorting and joining data. The sort operation now properly respects ordering preferences, and joins now correctly handle null equality rules for nested data types. (a2b0b2a)
- **fix(rust): Use try_new in StructArray construction in polars-json (#27489)** (7d6483b)
- **fix(rust): Raise ComputeError instead of panicking in repeat_by when output exceeds IdxSize::MAX (#27892)** (ee2469b)
- **fix(rust): Run type coercion pass on pivot's internally generated group_by (#27897)** (c576781)
- **fix: More careful slice pushdown into joins (#28578)** (f16e181)
- **fix: Preserve ordering in sliced unions (#28576)** (0c886ca)
- **fix: Drop input sortedness when casting to a string (#28574)** (5a9e03a)
- **fix: Clear sortedness flags in `StringChunked` substring kernels (#28573)** (36e414b)
- **fix: Bad mask handling when reading optional parquet column (#28547)** (72c95bf)
- **fix: Any operation on `Unknown(Int)` and `Unknown(Float)` should result in `Unknown(Float)` (#28545)** (16554b0)
- **fix: Flip `nulls_last` after `Expr.reverse()` (#28572)** (138e72b)
- **fix: Fix high blocking thread use in sink_parquet with async local path (#28543)** (a33dc3d)
- **fix: Serialize LazyFrames backed by bytes (#28568)** (56092e4)
- **fix: Release GIL in `SQLContext.execute()` (#28549)** (345ee88)
- **fix: Propagate `nulls_last` in `function_expr_sortedness` (#28544)** (cbabf96)
- **fix: Incorrect slicing when a join requires sorting (#28541)** (8122e93)
- **Fixed crash on self-join queries** — Resolved a bug that caused the system to panic when performing self-joins on Iceberg scan operations. A test case was added to prevent this issue from recurring. (b1268f1)
- **Fix worker count tracking** — Fixed a bug where the ParkGroup worker count wasn't being decremented when workers exited, which could cause incorrect worker tracking in the async executor. (0d4a49a)
- **Fix async connector Send bounds** — Added missing Send trait bounds to the Connector's SenderExt and ReceiverExt types to ensure thread-safe behavior in async operations. This prevents potential data races when these types are used across thread boundaries. (a3982dd)
- **Fix undefined behavior in null checks** — Fixed unsafe memory access in the first_non_null and last_non_null functions that could crash when processing empty data chunks. The code now safely finds the first non-empty chunk before checking for null values instead of directly accessing memory. (49ceb73)
- **Fix scalar expression broadcasting in optimizer** — Fixed a bug in the common subexpression optimizer (streaming engine) that was incorrectly broadcasting non-column height expressions, which caused scalar literals to be treated as if they should expand across rows. A new check ensures that only true column-producing expressions are cached and reused in element-wise operations. (1f63626)
- **Fix sort optimization in gather operations** — Corrected how Polars tracks sorted data through Gather operations, ensuring that sort information is properly maintained when combining gather and sort operations. This fixes a bug where sort optimization wasn't working correctly in certain query patterns. (f75402e)
- **Fix NULL handling in SQL NOT IN** — Corrected SQL query behavior when using NOT IN with NULL values in sets or subqueries. The fix implements proper three-valued logic (3VL) where comparisons against NULL should return UNKNOWN rather than FALSE, ensuring correct filtering behavior in joins and subqueries. (b72bea9)
- **Fix literal value comparison handling** — Updated how literal values are compared to properly handle floating-point numbers using total comparison semantics, which correctly treats NaN values as equal to themselves. This ensures that expressions with NaN literals are compared consistently. (1a384cf)
- **Fix SQL aggregate null handling** — Fixed SUM and CORR SQL aggregate functions to properly return NULL when all inputs are null, and added support for the TOTAL aggregate function. Also refactored HAVING clause processing to use Polars' native having() method instead of post-filter logic. (af10a52)
- **fix: Share `null_count_dtype` helper between Delta and Iceberg, fixing `SchemaError` (#28479)** (7b612c4)
- **fix: Remove non-output columns from the equi-join and semi/anti-join operators (#28446)** (b2c3491)
- **fix: Fix dropped slice on multiple unions (#28477)** (3707b24)
- **perf: Optimize not(bool_f) to not_bool_f (#28474)** (74d3d2d)
- **fix: Fix in-memory engine incorrect slice on maintain order join (#28478)** (f3cd4c3)
- **fix: Check join schema by position (#28455)** (2f71d86)
- **fix: Block predicate pushdown past overwritten window keys (#28429)** (39fd098)
- **fix: Avoid panic when union slice skips all rows (#28420)** (0bf71d9)
- **fix: Solve panic in `dt.replace` when there were multiple chunks (#28437)** (8752751)
- **fix: Propagate `is_scalar` from the input to the output of `.sort()` and `.sort_by()` (#28438)** (0189b46)
- **fix(rust): Insert missing coercions from `Unknown(_)` in list/array arithmetic (#28411)** (ce6e4de)
- **fix: Resolve CSV column names overwrite in DSL->IR conversion (#28383)** (c6e05d1)
- **fix: Avoid IEJoin rewrite for Categorical comparisons (#28427)** (0cdafb4)
- **fix: Do not rewrite `sort().reverse()` when `maintain_order=True` (#28403)** (7b67ac4)
- **fix: Ensure BinaryView offset+len does not exceed i32::MAX where possible (#28048)** (6478498)
- **fix: Invalid offset in strptime (#28388)** (00e0078)
- **fix: Incorrect schema type for decimal <-> primitive division (#28373)** (5826d2f)
- **fix: Panic in in-memory CSEE handling (#28371)** (c3e9aa9)
- **fix(rust): Apply same type coercion to IsBetween as binary comparisons (#28300)** (7a99142)
- **fix: Fix offset in arrow ffi export of sliced struct arrays (#28369)** (b26546a)
- **fix: Fix cross filter not applied with sink and CSE (#28297)** (d62a534)
- **perf: Split multiplexers that directly scan from in-memory DataFrame (#28376)** (10b1a12)
- **perf: Make DNS cache global (#28352)** (77f19be)
- **perf: Do not remove cache if predicates not pushed to all inputs (#28341)** (9c22354)
- **perf: Environment variable for logging slow DNS lookup (#28211)** (fee57e2)
- **perf: Pre-partition on left, right and semi joins on hive partitioned data (#28374)** (ec98255)
- **fix: Fix panic on projection pushdown with caches (#28280)** (cd22493)
- **fix: Fix `write_json()` null values in `Array` columns being written incorrectly as `null` (#28330)** (fd6549f)
- **fix: Fix EntityTooSmall on sink_ipc to S3 (#28255)** (f640848)
- **fix: Raise error instead of silent wrapping for `select(len())` (#28355)** (aa25a0e)
- **fix: Float16 groupby aggregates (#28361)** (7a21b5b)
- **fix: Respect lexical ordering of Categorical in `top_k`/`bottom_k` (#28359)** (61a08f1)
- **Add dtype overrides by column position** — Added support for overriding data types in CSV files by column position (in addition to the existing column-name-based override). Users can now pass a list of data types that will be applied to columns in order, making it easier to fix data type inference without needing to know column names. (7301041)
- **Fix type coercion in fused math** — Fixed a bug where fused multiply-add operations would fail when operands had unknown or mismatched types. The optimizer now properly determines a common type for all operands and casts them accordingly before performing the fused operation. (20dc03f)
- **Respect custom AWS checksum settings** — Fixed AWS S3 configuration to honor user-provided checksum algorithms instead of always overriding with CRC64NVME. When no checksum algorithm is specified, it defaults to the efficient CRC64NVME option. (dccffc2)
- **Optimize joins on partitioned data** — Improved performance for inner joins on Hive-partitioned datasets by rewriting them as unions of filtered joins on individual partitions. This allows the query optimizer to prune unnecessary partitions and reduce data scanned during join operations. (0fdc9a5)
- **Update PyArrow IPC API usage** — Replaced deprecated PyArrow Feather API calls with the current IPC module API in Polars' file reading functionality. This ensures compatibility with newer PyArrow versions and adds support for column selection by name when using PyArrow as the backend. (57f20ca)
- **Restore HF_TOKEN environment variable support** — Fixed a regression that prevented Polars from reading the HuggingFace token from the HF_TOKEN environment variable when accessing Hugging Face datasets. The fix removes an overly restrictive early-exit condition and adds a test to verify the token is properly sourced and used. (7bfbe0f)
- **Block math between date/time and numbers** — Added validation to prevent adding or subtracting temporal data types (dates, times, durations) with non-temporal types like integers. Operations that mix these incompatible types now raise an error instead of attempting a coercion. (9235715)
- **Fix Series construction from unaligned arrays** — Fixed a crash that occurred when creating a Series or DataFrame from NumPy arrays with unaligned memory layouts. The fix now properly handles and aligns such arrays before processing, and includes tests to prevent regression. (cf8323f)

### Backend
- **lockfile** (fb25e49)
- **pyo3-polars** (a88efdc)
- **Rust Polars 0.55.x release** (f5cab50)
- **depr(python): Deprecate `Expr.rechunk()` (#28692)** (2e64f66)
- **depr: Deprecate `struct.rename_fields()` with an incorrect number of fields (#28672)** (2893048)
- **Release Python Polars 1.43.2** — Updated Python Polars to version 1.43.2, bumping version numbers across all runtime packages and configuration files. (12f83f5)
- **depr: Deprecate casts from `Categorical` to integer dtypes (#28525)** (a122259)
- **depr: Deprecate not setting the `plan_stage` argument in `show_graph()` (#28391)** (df13a11)
- **Python Polars 1.43.1 (#28523)** (8d438be)
- **Python Polars 1.43.0 (#28394)** (434fa10)
- **depr: Deprecate casting numeric types to categoricals (#28349)** (fd4574f)
- **depr: Deprecate `cat.get_categories()` and `cat.to_local()` (#28299)** (d6ef91a)

### Tests
- **test: Fix flaky test (#28454)** (a184230)
- **test: Fix flaky ordering expectation in top_k_by (#28435)** (3cb01f4)
- **test(python): Ignore all deprecation warnings in Python doctest (#28398)** (8588201)
- **test(python): Ignore deprecation warnings on `to_struct()` in doctest (#28370)** (262c8fb)

### Docs
- **docs(python): Fix Polars Cloud API reference link (#28693)** (9ce74db)
- **docs: Update and restructure README (#28490)** (f6d2bec)
- **Fix Spark migration guide examples** — Corrected output shape value, added missing import, fixed typos, and improved code formatting in the migration guide documentation to ensure examples are accurate and consistent. (16ae441)
- **docs: Relocate Polars Cloud & On-Prem User Guide (#28462)** (c0b036a)
- **docs: Add notes on k8s operator (#28445)** (640ac29)
- **docs: Update comparison page (#28418)** (d02100a)
- **docs: Update GPU support documentation with the cudf-polars 26.06 release (#27830)** (7f326b9)
- **Fix dev docs SEO indexing** — Updated search engine robot directives to prevent development documentation versions from being indexed as canonical by search engines. Development builds now include a noindex meta tag, while allowing search engines to crawl the stable documentation. (409a87f)
- **HDFS support documentation added** — Added comprehensive documentation for experimental HDFS (Hadoop Distributed Filesystem) support in Polars on-premises, including installation steps, configuration instructions, and usage examples for both direct data access and Iceberg metadata operations. The documentation also updates the configuration reference to document the new `extras.hdfs.enabled` and `extras.pyiceberg.enabled` settings. (f8bcc3d)

### Chore
- **build: Bump quinn-proto from 0.11.14 to 0.11.16 (#28524)** (2eccb46)
- **refactor(rust): Add (with_)context to PolarsResult for easier annotation (#28677)** (446de71)
- **chore(rust): Clean up polars-utils import structure (#28668)** (9e291f4)
- **chore: Re-enable `test_extension()` for streaming engine (#28611)** (712abc7)
- **chore(rust): Clean up vec utils (#28667)** (13cb50a)
- **chore(rust): Remove Rust compiler intrinsics, use nightly functions (#28665)** (8b9536c)
- **chore: Update `test_hive_join_rewrite_semi_join` test to work with streaming engine (#28610)** (a1711c2)
- **ci: Fix duckdb delta extension install collision (#28607)** (9aca544)
- **refactor: Make hive_part extraction a function and public (#28507)** (a3e282a)
- **Consolidate expression comparison logic** — Refactored expression equality checks to eliminate duplicate code by centralizing comparison logic into a single `is_expr_equal_shallow` method. This improves consistency in how expressions are compared across the codebase and makes the logic easier to maintain. (b6b709f)
- **refactor(rust): Assert parameter in `filter_scan_ir` (#28472)** (37964ac)
- **chore(rust): Make the fmt macro exportable (#28459)** (bdc5be5)
- **chore(python): Bump `ruff` and `mypy` package versions (#28456)** (82c7d29)
- **chore: Update AI policy for comments (#28436)** (e5b36aa)
- **build(rust): Propagate `bigidx` to `polars-plan` indirectly included via `pyo3-polars` (#28406)** (1dbd4bc)
- **build(rust): Propagate bigidx to polars-plan (#28396)** (af17ec6)
- **refactor: Re-work `shuffle` parameter for `sample()` (#27460)** (3bba9f1)
- **refactor(rust): SpillFrame instead of DataFrame in Morsel (#28270)** (ea9ce90)
- **build: Set the default `dev` profile to `line-tables-only` (#28358)** (2aa533a)
- **ci: Downgrade the `release-drafter` ci version back to version 6 (#28362)** (a265f6a)
- **ci: Bump the ci group across 1 directory with 7 updates (#28169)** (bc6a7a7)
- **Update spin dependency to v0.10.1** — Updated the spin Rust dependency from version 0.10.0 to 0.10.1, which includes a new checksum reflecting the updated package. (22dfc3a)
- **Reduced internal dispatch overhead** — Optimized numeric type handling in several core operations (is_first_distinct, is_last_distinct, is_unique, rle, and unique_counts) by replacing broad macro-based dispatch with targeted float-specific dispatch and explicit bit representation matching for integers, reducing unnecessary code generation. (d8a460c)
- **Shrink binary by optimizing SQL parser** — Updated build configuration to use size-optimized compilation for SQL parsing dependencies, reducing the overall binary size without impacting performance-critical code paths. (685c6a2)
- **Optimize function accepts intermediate represen…** — Refactored the optimize function to accept an already-converted intermediate representation (IR) instead of a raw logical plan, moving the conversion step earlier in the pipeline. This change simplifies the optimizer's responsibility and allows better control over optimization flags during the conversion process. (3a8c176)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/pola-rs/polars?utm_source=github-action)._