## bodo-ai/Bodo — 2026.2…2026.4

_73 commits._

### Features
- **[GPU] Support Series.isna()/isnull() on GPU (#1112)** (1d091c1)
- **Add multi-node GPU benchmarking scripts (#1110)** (64c6999)
- **[GPU] Support Series.mean() on GPU (#1111)** (5df4caf)
- **[GPU] Support Series.round() on GPU (#1109)** (2ee5b1c)
- **[GPU] Support Series.dt.year on GPUs (#1105)** (8687aa1)
- **[GPU] support union all on GPU (#1101)** (f9f14d5)
- **[GPU] Support Reduce operations on GPU (#1094)** (1235307)
- **[GPU] BSE-5333: Support non-equi join conditions in cuda join (#1085)** (835b490)
- **[GPU] BSE-5333: Support left, right, and outer joins in cuda join (#1073)** (e981d40)
- **[GPU][BSE-5342] Add assert GPU tests run without CPU fallback (#1069)** (5bf05ab)
- **[GPU] Add two more GPU tests. (#1049)** (73beaee)
- **[GPU] Add GPU CTE support. (#1048)** (574eb03)
- **[GPU][BSE-5315] Add tests for CPU-GPU transfer (#1047)** (1fbf64e)
- **Support S3 in GPU Parquet write (#1041)** (3efb63a)
- **Support S3 in GPU parquet read (#1038)** (5bb53cc)

### Fixes
- **[GPU] fix batch nullptr bug. (#1114)** (eb50a4e)
- **[GPU] Fix node/gpu detection for multi-node non-block/non-default rank assignments (#1107)** (f78d4a9)
- **[GPU] Add check around MPI_Allreduce to fix mpich gpu build (#1104)** (5043b9b)
- **Fix multi-node spawn mode with OpenMPI (#1100)** (1ce63ad)
- **Fix C++ ternary operator bug (#1097)** (8d9553f)
- **[GPU] Fix Parquet output schema (#1087)** (36b9589)
- **[GPU] Fix broadcast join hang for TPC-H Q2 (#1084)** (a93fbc4)
- **[GPU] Fix memory leak in Parquet reader (#1080)** (01bf3ed)
- **[GPU] Fix GPU shuffle empty condition (#1074)** (1030683)
- **[GPU] Enable address sanitizer for GPU builds and fix issues (#1071)** (aea59fd)
- **Fix Snowflake nested data tests (#1064)** (33faaae)
- **BSE-5332: Fix column pruning under aggregate (#1058)** (1233973)
- **Fix windows build (#1056)** (62f4b8b)
- **Fix missing symbol error in CPU build (#1050)** (36dd5a8)
- **Fix MacOS Arm delocate command in pip wheels (#1044)** (d97e98f)
- **Fix nightly tests after Pandas 3 upgrade (#1034)** (3a1f28c)

### Backend
- **[BSE-5360] Build GPU conda package (#1115)** (5154fa8)
- **[GPU] Upgrade to rapids 26.04 (#1116)** (18908f1)
- **GPU limit. (#1108)** (f871844)
- **[GPU] BSE-5362: Cuda Sort (#1106)** (d16950c)
- **Split join with non-equi conditions. (#1102)** (851e4ca)
- **[GPU] Enable GPU direct RDMA path for MPICH/UCX (#1103)** (6d3a477)
- **[GPU][BSE-5357] Make sure GPU buffers are ready before calling MPI (#1075)** (ec80a51)
- **[GPU] BSE-5333: Gpu mark join (#1090)** (5812e0c)
- **Remove old comment (#1096)** (52e01eb)
- **[GPU] enable TPC-H Q3/Q4 tests on GPU (#1093)** (f350dde)
- **[GPU] Avoid UDF for Series.str.startswith/endswith on CPU and GPU (#1088)** (0879ada)
- **[BSE-5349] GPU Benchmark TPCH Q5 (#1052)** (d605231)
- **[GPU] Increase default batch size (#1083)** (a56c3dc)
- **[GPU] use hwloc for node topology detection (#1082)** (c854ce5)
- **[GPU] Use MPI large counts if available (#1081)** (34282fa)
- **[GPU] Shuffle on the default stream (#1079)** (41957bb)
- **[GPU] GPU broadcast join. (#1062)** (380184d)
- **[GPU] Short circuit shuffle if running on a single GPU (#1078)** (d7fa030)
- **[GPU] Set device memory resource to cuda_async_memory_resource (#1077)** (7c0103d)
- **[GPU] Remove NCCL dependency (#1076)** (0a99b23)
- **[GPU] Make GPU shuffle async to match CPU shuffle (#1070)** (2750ab9)
- **[GPU] GPU bloom filters. (#1055)** (2f72568)
- **[GPU] Replace NCCL with CUDA-aware MPI (#1068)** (d770937)
- **[GPU] Disable pandas metadata in read_parquet (#1067)** (9c48231)
- **[GPU] Use Chunked Parquet Reader and Improve Parquet Read Performance  (#1063)** (bf3878c)
- **Skip Snowflake Iceberg tests (#1066)** (15790f9)
- **Skip SQL database tests (#1065)** (ecdfd9c)
- **[GPU] Avoid GPU calibration with ALWAYS_RUN_ON_GPU (#1060)** (bd3557c)
- **Enable Docker ARM Build (#998)** (eeb32b1)
- **[GPU] Avoid GPU overheads on non-GPU ranks (#1057)** (aa4fb8f)
- **[GPU] Make GPU shuffle data sizes arrive in the same order (#1054)** (fdcccc8)
- **[GPU] Ensure device ID is reset after executor is destroyed (#1053)** (80b295c)
- **[GPU] Use cuDF AST expressions for Parquet filtering (#1051)** (7aac788)
- **Upgrade Duckdb Planner (#1043)** (188e232)
- **Async fixes for join. (#1042)** (eaa0f10)
- **Remove pd.to_datetime() from TPC-H queries (#1046)** (4bcbc35)
- **[GPU] Annotate CPU/GPU directly on plan tree when printing plans (#1045)** (d8d9237)
- **[GPU] GPU CI (#1029)** (9688e49)
- **GPU docs (#1040)** (19d528b)
- **GPU groupby (#1036)** (d3c6e70)
- **Always run operators run on GPU if possible for now (#1037)** (24de56f)
- **2026.2 Release Notes (#1035)** (d9b22be)

_Recap by [Repo Wrapped](https://repowrapped.com/gh/bodo-ai/Bodo?utm_source=github-action)._