## apache/spark — v4.2.1-rc1…v4.3.0-rc1

_1513+ commits._

### Features
- **[SPARK-59060][PYTHON] Add missing cpuAmount to the benchmark mock task context** (2cee139)
- **[SPARK-58947][CORE] Handle Char in createArray and add tests for SparkCollectionUtils** (4efb478)
- **[SPARK-58940][SDP][DOC] Add migration notes for Declarative Pipelines case sensitivity** (06ce035)
- **[SPARK-58825][SQL] Support TimestampNTZType as a JDBC partition column** (952042f)
- **[SPARK-58915][CORE][K8S][YARN] Add `supportsExecutorHold` to the scheduler backends** (5ddd052)
- **[SPARK-58913][CORE] Add `TaskSchedulerImpl.hasPipelinedTaskSets`** (e87de79)
- **[SPARK-58914][CORE][UI] Add a reusable `confirm-link` for Web UI links** (1aafc54)
- **[SPARK-58903][SQL][TEST] Add BIN BY tests for resolving re-output DISTRIBUTE and metadata columns** (2ba0aac)
- **[SPARK-58871][INFRA] Add timeout to disk cleanup steps** (5644715)
- **[SPARK-58842][ML] Add descriptive messages to bare require checks in ALS** (09aa4ed)
- **[SPARK-58741][SQL][TEST] Add nanosecond-timestamp coverage for collect_list** (3755010)
- **[SPARK-58823][SQL][TEST] Add nanos-timestamp tests for scalar subquery / EXISTS / NOT IN (SELECT)** (cba8f05)
- **[SPARK-58689][CORE][TEST] Add tests for SparkStringUtils padding and abbreviation helpers** (ea3a34a)
- **[SPARK-58761][SQL][TEST] Add dedicated e2e test coverage for GROUP BY over nanosecond-precision timestamp keys** (3683d14)
- **[SPARK-58742][SQL][TEST] Add nanosecond-timestamp coverage for collect_set** (307c12f)
- **[SPARK-58062][SQL] Add `variant_strip_nulls` expression** (a5dbd23)
- **[SPARK-57228][SS] Support transformWithState in Real-Time Mode** (ee905e3)
- **[SPARK-58761][SQL][TEST] Add dedicated e2e test coverage for DISTINCT over nanosecond-precision timestamp columns** (2417114)
- **[SPARK-58761][SQL][TEST] Add DESCRIBE / SHOW CREATE TABLE tests for nanosecond-timestamp columns** (f45bbfa)
- **[SPARK-58743][SQL][TEST] Add nanosecond-timestamp coverage for mode** (fa49c0b)
- **[SPARK-58740][SQL][TEST] Add to_xml test coverage for nanosecond-precision timestamp types** (f31412b)
- **[SPARK-58569][SDP][DOC] Add documentation for Auto CDC (SCD Type 1)** (256e58d)
- **[SPARK-58698][SS][TEST] Add additional tests for RTM** (8ccbf4e)
- **[SPARK-58737][SQL] Support nanosecond-precision timestamps in ColumnarRow/ColumnarBatchRow copy() and get()** (ba17d5c)
- **[SPARK-57235][DOC][SS] Add documentation for stateful queries in RTM** (a1f2be7)
- **[SPARK-58409][SDP] Add SCD2 AutoCDC end-to-end test suites** (bdbbfab)
- **[SPARK-57460][SQL] Support nanosecond-precision timestamp types in the JDBC datasource** (f223b52)
- **[SPARK-57823][SQL] Support ORC predicate pushdown for nanosecond-precision timestamps** (55600bf)
- **[SPARK-57822][SQL] Support Parquet predicate pushdown for nanosecond-precision timestamps** (96e7763)
- **[SPARK-57832][SQL] Support timestamp - timestamp subtraction for nanosecond-precision timestamps** (482d461)
- **[SPARK-57837][SQL] Support the precision argument in current_timestamp(p) and localtimestamp(p)** (6e8c5ef)
- **[SPARK-58523][SQL] Add Catalyst runtime filtering interface for DSv2 scans** (87064f6)
- **[SPARK-58609][SQL] Add archivePathFilter option to select inner archive entries** (fc8f161)
- **[SPARK-58635][SS] Support for streaming aggregation in Real-Time Mode (RTM)** (644f242)
- **[SPARK-58678][CORE][TEST] Add `OpenHashMapBenchmark`** (62e79ed)

### Fixes
- **[SPARK-59084][CORE] Fix anchors styled as buttons to keep the button hover style** (834ca7e)
- **[SPARK-59066][CORE] Fix History Server Event Log `Download` button to use a Bootstrap 5 class** (2852b96)
- **[SPARK-59067][CORE] Fix `Download` to be a button aligned with the other buttons** (cdcc179)
- **[SPARK-59054][CORE][SQL] Fix `KeyGroupedPartitioner` to compare partition keys by value in storage-partitioned join shuffles** (b46334b)
- **[SPARK-55271][SS] Fix NullPointerException in Kafka micro-batch streaming metrics reporting** (2514a7d)
- **[SPARK-59008][K8S] Fix NPE in ExecutorPodsLifecycleManager when pod is deleted between get() calls** (d1dc748)
- **[SPARK-59025][SQL] Fix subset join key grouping when the keyed side reports a PartitioningCollection** (7cc9194)
- **[SPARK-58988][SQL] Fix PartitioningCollection invariant violation exception when allowKeysSubsetOfPartitionKeys** (bfeaa49)
- **[SPARK-58994][ML] Fix prediction column metadata in GBTRegressionModel** (360428a)
- **[SPARK-58816][SQL] Fix INSERT with column list to resolve structs inside arrays and maps positionally** (cfe1f16)
- **[SPARK-58428][SQL] Fix optimizer hang in DSv2 expression pushdown** (33ccc39)
- **[SPARK-58627][SQL] Mark raise error Throwable & fix sequence throwable** (aac2bd1)
- **[SPARK-58920][DOC] Fix typos in the Quick Start and Security docs** (b80495e)
- **[SPARK-58921][CONNECT][PYTHON] Fix misspelled messageParameters argument in Spark Connect tag validation** (7367a27)
- **[SPARK-58887][CORE] Fix barrier job to be cancellable during slot check retries** (dd9403e)
- **[SPARK-58886][CORE] Fix Int overflow in `CoarseGrainedSchedulerBackend.requestExecutors`** (c88db74)
- **[SPARK-58844][UDF] Fix a Cancel-after-Finish flake in the UDF worker protocol test client** (1f9fe79)
- **[SPARK-58782][SQL] Fix DSv2 pushdown null literal serialization bug** (1308d4c)
- **[SPARK-58690][CORE] Fix typos in storage package comments** (cb7db40)
- **[SPARK-58843][DOC] Fix grammar errors in the SQL tuning, GraphX and MLlib clustering guides** (bdfd9e0)
- **[MINOR][DOC] Fix broken submitting-applications link in Kafka guide** (33a6b89)
- **[SPARK-57956][SQL] Fix AQE incorrectly eliminating a global LIMIT over a non-materialized query stage** (3c6ecb2)
- **[SPARK-58746][INFRA] Remove a dead lsof workaround and fix stale references in release tooling** (7681d2e)
- **[SPARK-58759][BUILD] Fix grammar in build file comments** (726e4df)
- **[SPARK-58714][SS] Fix AdminClient leak in KafkaTokenUtil.obtainToken** (00bd78f)
- **[MINOR][DOC][SQL] Fix minor gaps and problems in Thrift server docs** (527a26d)
- **[SPARK-58712][EXAMPLES] Fix incorrect usage strings in four examples** (b515835)
- **[SPARK-57665][SQL] Fix slice() returning empty for large length in the interpreted path** (3a9db9a)
- **[SPARK-57787][4.3][CONNECT][PYTHON][FOLLOWUP] Fix syntax error in `local_server.py`** (21021bb)
- **[SPARK-57353][SQL] Fix single-pass resolver crash on GROUPING SETS with HAVING/ORDER BY** (32274cd)
- **[SPARK-58544][SQL] Fix vector distance and norm functions returning wrong results from intermediate float overflow** (da54921)
- **[SPARK-58663][BUILD] Fix typos and grammar in root pom.xml comments** (78b925c)
- **[SPARK-58679][DOC] Fix typos and grammar in the SQL name resolution reference doc** (b013c37)
- **[SPARK-48973][SQL] Fix mask to handle characters outside the BMP** (f447372)

### Backend
- **Preparing Spark release v4.3.0-rc1** (3db8134)
- **[SPARK-58750][4.3][CORE] Tolerate FileAlreadyExistsException from checkpoint part file rename** (c14f72f)
- **[SPARK-58392][SQL] Use loadTable for RelationCatalog reads with table-state options** (675fc98)
- **[SPARK-59009][SQL] Re-map outputOrdering in InMemoryRelation.newInstance()** (bf536dc)
- **[SPARK-58933][4.3][SQL] Resolve expressions in INSERT target IDENTIFIER clauses** (4e111b4)
- **[SPARK-58974][SQL] Apply the narrowed-partitioning skew guard regardless of requireAllClusterKeysForDistribution** (ed82c61)
- **[SPARK-58572][SDP] AutoCDC Cross SCD Convergence Suite** (f68e4e8)
- **[SPARK-59027][SQL] Share grouped partition key ordering between `createShuffleSpec` and `GroupPartitionsExec`** (fd87d09)
- **[SPARK-59022][SQL] Keyed shuffle must follow the declared partition key order** (568b2c2)
- **[SPARK-58932][SS] Close the accepted socket in TransformWithStateInPySparkStateServer** (818c2c8)
- **[SPARK-58601][PYTHON] Tighten mapInPandas return-value contract to require a strict Iterator** (decb734)
- **[SPARK-58977][SS][PYTHON] Handle state server shutdown before Python connects** (8fb82fe)
- **[SPARK-58935][DOC] Document `executorManagement` and `shared` queue metrics of `LiveListenerBus`** (743d71d)
- **[SPARK-58931][SQL] Reject negative randstr length during analysis** (cfa026c)
- **[SPARK-58976][UI] Remove unused generateOutputOperationStatusForUI in BatchPage** (7202bf7)
- **[SPARK-58650][PYTHON][SQL] Improve source extraction during transpilation** (daf88e4)
- **[SPARK-58848][CORE][SS] Prevent unbounded shuffle reader connection waits** (21ba279)
- **[SPARK-58948][SQL] Deduplicate nulls-buffer creation in columnar decompressors** (0d76ca0)
- **[SPARK-58941][SDP] Sort schema inference flows by identifier parts to avoid dotted-name collisions** (f15b18b)
- **[SPARK-58939][PYTHON] Move Arrow collect batch reordering into ArrowCollectSerializer** (0fd18e8)
- **[SPARK-58166][SQL] Replace TypeTag on PhysicalDataType with ClassTag** (33a9221)
- **[SPARK-58168][SQL] Do no use Array[Nothing] in Flatten** (b037b8c)
- **[SPARK-58169][SQL] Use java reflection for finding the SparkSession companions** (9029a69)
- **[SPARK-57548][SQL] Avoid eager scala.reflect.runtime.universe initialization in ScalaReflection** (6e8f581)
- **[SPARK-58937][SDP] Consider unified timeline when determining affected rows for SCD2** (8a1cd77)
- **[SPARK-58872][K8S] Warn when driver credentials drop the driver service account** (263b7c5)
- **[SPARK-57980][4.3][SQL] Extract the HashAggregateExec hash-map spill machinery into a shared helper** (5644a6a)
- **[SPARK-58894][SQL] Detect cyclic view references in nested subqueries** (ef4fd13)
- **[SPARK-58922][CORE] Extract a null/empty MDC-array check in SparkLogger** (afd1458)
- **[SPARK-58849][SQL] Make DESCRIBE TABLE resilient to corrupt partition metadata** (aa6e654)
- **[SPARK-58925][SS][DOC] Document maxOpenFiles troubleshooting guidance** (34c5360)
- **[SPARK-58699][PYTHON] Short-circuit identity conversion for string and binary in LocalDataToArrowConversion** (c9cb188)
- **[SPARK-58779][SQL] Make InlineCTE tolerate a CTE reference without its definition during analysis** (75fcd00)
- **[SPARK-58792][SQL] Clone Hive GenericUDF per copy** (1d2b179)
- **[SPARK-58850][SQL] Update stale plan node names in AQE test comments** (e5d8cc5)
- **[SPARK-58213][SQL] corr should return NULL when a column has zero variance** (3a07a3d)
- **[SPARK-49442][SS][FOLLOWUP] Use monotonic clock for Kafka partition cache TTL** (3e25b4c)
- **[SPARK-58857][K8S] Bind the result of Utils.randomize in LocalDirsFeatureStep** (143090e)
- **[SPARK-58389][SQL][FOLLOWUP] Pass table-state options while loading write targets** (0dfb24d)
- **[SPARK-58658][CONNECT] Do not return redacted configurations from the Config RPC** (df4af9b)
- **[SPARK-58827][DOC] Document SPARK_LOCAL_IP for local RemoteClassLoaderError test failures** (756c6e3)
- **[SPARK-58826][SQL] Hoist repeated session.implicits imports in AutoCDC tests** (2688d47)
- **[SPARK-34254][SQL][DOC] Document CREATE EXTERNAL TABLE for data source tables** (687159a)
- **[SPARK-58711][SQL] Use StructType.getFieldIndex in Parquet and ORC aggregate push down** (72df4e9)
- **[SPARK-58822][SQL][TEST] Extract a nonEmptyLines helper in LogicalPlanDifferenceSuite** (52bb9b4)
- **[SPARK-58725][K8S] Resolve the pod before deciding whether --kill found it** (ce42012)
- **[SPARK-58841][K8S][DOC] Update `YuniKorn` docs with `1.9.0`** (9752bad)
- **[SPARK-58834][INFRA] Bump docker CI actions to latest ASF-allowlisted versions** (50ca093)
- **[SPARK-58724][SS] Incremental state cleanup for streaming dropDuplicates** (dfad9c7)
- **[SPARK-38794][K8S] Create executor ConfigMap before requesting executors** (8253c33)
- **[SPARK-58784][PYTHON] Merge the batched-UDF mapper into an explicit SQL_BATCHED_UDF branch in read_udfs** (f78dcff)
- **[SPARK-58709][CORE][TEST] Use assert(!x) instead of assert(x === false) in SecurityManagerSuite** (223b56c)
- **[SPARK-58619][CONNECT][FOLLOWUP] Map JDBC connection errors by SQLSTATE class 08 instead of hard-coded condition names** (d50b1bb)
- **[SPARK-58756][4.3][BUILD][CORE] Revert OIDC credential propagation from `branch-4.3` (moved to `4.4.0`)** (8afeb5b)
- **[SPARK-58684][PYTHON] Make PySpark error sub-conditions inherit the parent condition error state** (e441259)
- **[SPARK-58619][CORE][CONNECT] Assign SQLSTATE 08003 to the INVALID_HANDLE session sub-conditions** (91c0768)
- **[SPARK-58751][SS][PYTHON] Stop leaking Python workers on TransformWithState init failure** (bb5237b)
- **[SPARK-58389][SQL][FOLLOWUP] Pin DSv2 table instances by state options during analysis** (3a22f0d)
- **[SPARK-58760][SQL] Use consistent colName interpolation in CatalogColumnStat** (047c709)
- **[SPARK-58758][DOC] Remove the stale Highlights in 3.0 section from the ML guide** (e8969b4)
- **[SPARK-57523][SQL] Declarative Pipelines should not retry flows that fail due to streaming source changes** (b45a8c7)
- **[SPARK-58745][DOC] Document missing JSON and XML data source options** (5326cfd)
- **[SPARK-58707][SQL] Do not prune a JSON schema down to the corrupt record column** (d6209be)
- **[SPARK-58739][SQL][TEST] Test UnsafeRow nested struct/map round-trip for nanos timestamps** (185f135)
- **[SPARK-58731][SQL] Exempt empty sketches from the hll_union lgConfigK check** (db427b2)
- **[SPARK-58710][MLLIB] Remove unused members from the decision tree Node classes** (fdaaaae)
- **[SPARK-58659][SS] Set Real-Time Mode config defaults** (33638c9)
- **[SPARK-57787][CONNECT][FOLLOWUP] Harden persistent local Connect server management** (4b0de36)
- **[SPARK-58693][K8S] Skip unlistable directories in K8s shuffle data recovery** (022e5e7)
- **[MINOR][DOC] Document CSV read behaviour when multiLine is disabled** (a92d7af)
- **[SPARK-58457][SQL][DOC] Correct the PERMISSIVE mode claim that corrupt records are dropped** (1353482)
- **[SPARK-58210][SQL][FOLLOWUP] Extend CombineAdjacentAggregation to partial merge** (0d8a1e6)
- **[SPARK-57370][SQL][FOLLOWUP] Route codegen referencing an unnameable class to Janino when narrowing is unsound** (3e39b4a)
- **[SPARK-58513][CORE] UnifiedMemoryManager wrongly reports INVALID_DRIVER_MEMORY on the executor side** (c252290)
- **[SPARK-58660][INFRA] Self-heal the PR Build check when the notify workflow fails to create it** (93e5bf4)
- **[SPARK-38954][CORE][FOLLOWUP] Isolate credential provider requirement failures** (5ad868e)
- **[SPARK-58639][DOC] Replace stale listCachedTables reference in the SQL performance tuning guide** (7b7727a)
- **[SPARK-58662][SS] Remove redundant toString in RocksDBStateStoreProvider state transition messages** (323669d)
- **[SPARK-58680][CORE] Use isEmpty/nonEmpty for emptiness checks in ExternalAppendOnlyMap** (aa495df)
- **[SPARK-58682][DOC] Update outdated `OpenHashMap`/`OpenHashSet` performance claims** (c529597)
- **[SPARK-58588][SQL] Return the broadcast hash join build side from JoinSelectionHelper instead of a Boolean** (5ab2862)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/apache/spark?utm_source=github-action)._