Releases · rapidsai/cudf

02 Jul 15:37

raydouglass

v24.06.01

4e59b79

v24.06.01 Latest

Latest

🚨 Breaking Changes

Deprecate Groupby.collect (#15808) @galipremsagar
Raise FileNotFoundError when a literal JSON string that looks like a json filename is passed (#15806) @lithomas1
Support filtered I/O in chunked_parquet_reader and simplify the use of parquet_reader_options (#15764) @mhaseeb123
Raise errors for unsupported operations on certain types (#15712) @galipremsagar
Support DurationType in cudf parquet reader via arrow:schema (#15617) @mhaseeb123
Remove protobuf and use parsed ORC statistics from libcudf (#15564) @bdice
Remove legacy JSON reader from Python (#15538) @bdice
Removing all batching code from parquet writer (#15528) @mhaseeb123
Convert libcudf resource parameters to rmm::device_async_resource_ref (#15507) @harrism
Remove deprecated strings offsets_begin (#15454) @davidwendt
Floating <--> fixed-point conversion must now be called explicitly (#15438) @pmattione-nvidia
Bind read_parquet_metadata API to libcudf instead of pyarrow and extract RowGroup information (#15398) @mhaseeb123
Remove deprecated hash() and spark_murmurhash3_x86_32() (#15375) @davidwendt
Remove empty elements from exploded character-ngrams output (#15371) @davidwendt
[FEA] Performance improvement for mixed left semi/anti join (#15288) @tgujar
Align date_range defaults with pandas, support tz (#15139) @mroeschke

🐛 Bug Fixes

Backport: Use size_t to allow large conditional joins (#16127) (#16133) @bdice
Backport #16045 to 24.06 (#16102) @vyasr
Backport #16038 to 24.06 (#16101) @vyasr
Backport: Fix segfault in conditional join (#16094) (#16100) @bdice
Add patch for incorrect cuco noexcept clauses (#16077) @vyasr
Revert "Fix docs for IO readers and strings_convert" (#15872) @vyasr
Remove problematic call of index setter to unblock dask-cuda CI (#15844) @charlesbluca
Use rapids_cpm_nvtx3 to get same nvtx3 target state as rmm (#15840) @robertmaynard
Return boolean from config_host_memory_resource instead of throwing (#15815) @abellina
Add temporary dask-cudf workaround for categorical sorting (#15801) @rjzamora
Fix row group alignment in ORC writer (#15789) @vuule
Raise error when sorting by categorical column in dask-cudf (#15788) @rjzamora
Upgrade arrow to 16.1 (#15787) @galipremsagar
Add support for PandasArray for pandas<2.1.0 (#15786) @galipremsagar
Limit runtime dependency to libarrow>=16.0.0,<16.1.0a0 (#15782) @pentschev
Fix cat.as_ordered not propogating correct size (#15780) @mroeschke
Handle mixed-like homogeneous types in isin (#15771) @galipremsagar
Fix id_vars and value_vars not accepting string scalars in melt (#15765) @mroeschke
Fix DatetimeIndex.loc for all types of ordering cases (#15761) @galipremsagar
Fix arrow versioning logic (#15755) @vyasr
Avoid running sanitizer on Java test designed to cause an error (#15753) @jlowe
Handle empty dataframe object with index present in setitem of loc (#15752) @galipremsagar
Eliminate circular reference in DataFrame/Series.iloc/loc (#15749) @mroeschke
Cap the absolute row index per pass in parquet chunked reader. (#15735) @nvdbaranec
Fix Index.repeat for datetime64 types (#15722) @galipremsagar
Fix multibyte check for case convert for large strings (#15721) @davidwendt
Fix get_loc to properly fetch results from an index that is in decreasing order (#15719) @galipremsagar
Return same type as the original index for .loc operations (#15717) @galipremsagar
Correct static builds + static arrow (#15715) @robertmaynard
Raise errors for unsupported operations on certain types (#15712) @galipremsagar
Fix ColumnAccessor caching of nrows if empty previously (#15710) @mroeschke
Allow None when nan_as_null=False in column constructor (#15709) @galipremsagar
Refine CudaTest.testCudaException in case throwing wrong type of CudaError under aarch64 (#15706) @sperlingxx
Fix maxima of categorical column (#15701) @rjzamora
Add proxy for inplace operations in cudf.pandas (#15695) @galipremsagar
Make nan_as_null behavior consistent across all APIs (#15692) @galipremsagar
Fix CI s3 api command to fetch latest results (#15687) @galipremsagar
Add NumpyExtensionArray proxy type in cudf.pandas (#15686) @galipremsagar
Properly implement binaryops for proxy types (#15684) @galipremsagar
Fix copy assignment and the comparison operator of rmm_host_allocator (#15677) @vuule
Fix multi-source reading in JSON byte range reader (#15671) @shrshi
Return int64 when pandas compatible mode is turned on for get_indexer (#15659) @galipremsagar
Fix Index contains for error validations and float vs int comparisons (#15657) @galipremsagar
Preserve sub-second data for time scalars in column construction (#15655) @galipremsagar
Check row limit size in cudf::strings::join_strings (#15643) @davidwendt
Enable sorting on column with nulls using query-planning (#15639) @rjzamora
Fix operator precedence problem in Parquet reader (#15638) @etseidl
Fix decoding of dictionary encoded FIXED_LEN_BYTE_ARRAY data in Parquet reader (#15601) @etseidl
Fix debug warnings/errors in from_arrow_device_test.cpp (#15596) @davidwendt
Add "collect" aggregation support to dask-cudf (#15593) @rjzamora
Fix categorical-accessor support and testing in dask-cudf (#15591) @rjzamora
Disable compute-sanitizer usage in CI tests with CUDA<11.6 (#15584) @davidwendt
Preserve RangeIndex.step in to_arrow/from_arrow (#15581) @mroeschke
Ignore new cupy warning (#15574) @vyasr
Add cuda-sanitizer-api dependency for test-cpp matrix 11.4 (#15573) @davidwendt
Allow apply udf to reference global modules in cudf.pandas (#15569) @mroeschke
Fix deprecation warnings for json legacy reader (#15563) @davidwendt
Fix millisecond resampling in cudf Python (#15560) @mroeschke
Rename JSON_READER_OPTION to JSON_READER_OPTION_NVBENCH. (#15553) @bdice
Fix a JNI bug in JSON parsing fixup (#15550) @revans2
Remove conda channel setup from wheel CI image script. (#15539) @bdice
cudf.pandas: Series dt accessor is CombinedDatetimelikeProperties (#15523) @wence-
Fix for some compiler warnings in parquet/page_decode.cuh (#15518) @etseidl
Fix exponent overflow in strings-to-double conversion (#15517) @davidwendt
nanoarrow uses package override for proper pinned versions generation (#15515) @robertmaynard
Remove index name overrides in dask-cudf pyarrow table dispatch (#15514) @charlesbluca
Fix async synchronization issues in json_column.cu (#15497) @karthikeyann
Add new patch to hide more CCCL APIs (#15493) @vyasr
Make improvements in pandas-test reporting (#15485) @galipremsagar
Fixed page data truncation in parquet writer under certain conditions. (#15474) @nvdbaranec
Only use data_type constructor with scale for decimal types (#15472) @wence-
Avoid "p2p" shuffle as a default when dask_cudf is imported (#15469) @rjzamora
Fix debug build errors from to_arrow_device_test.cpp (#15463) @davidwendt
Fix base_normalator::integer_sizeof_fn integer dispatch (#15457) @davidwendt
Allow consumers of static builds to find nanoarrow (#15456) @robertmaynard
Allow jit compilation when using a splayed CUDA toolkit (#15451) @robertmaynard
Handle case of scan aggregation in groupby-transform (#15450) @wence-
Test static builds in CI and fix nanoarrow configure (#15437) @vyasr
Fixes potential race in JSON parser when parsing JSON lines format and when recovering from invalid lines (#15419) @elstehle
Fix errors in chunked ORC writer when no tables were (successfully) written (#15393) @vuule
Support implicit array conversion with query-planning enabled (#15378) @rjzamora
Fix arrow-based round trip of empty dataframes (#15373) @wence-
Remove empty elements from exploded character-ngrams output (#15371) @davidwendt
Remove boundscheck=False setting in cython files (#15362) @wence-
Patch dask-expr var logic in dask-cudf (#15347) @rjzamora
Fix for logical and syntactical errors in libcudf c++ examples (#15346) @mhaseeb123
Disable dask-expr in docs builds. (#15343) @bdice
Apply the cuFile error work around to data_sink as well (#15335) @vuule
Fix parquet predicate filtering with column projection (#15113) @karthikeyann
Check column type equality, handling nested types correctly. (#14531) @bdice

📖 Documentation

Fix docs for IO readers and strings_convert (#15842) @bdice
Update cudf.pandas docs for GA (#15744) @beckernick
Add contributing warning about circular imports (#15691) @er-eis
Update libcudf developer guide for strings offsets column (#15661) @davidwendt
Update developer guide with device_async_resource_ref guidelines (#15562) @harrism
DOC: add pandas intersphinx mapping (#15531) @raybellwaves
rm-dup-doc in frame.py (#15530) @raybellwaves
Update CONTRIBUTING.md to use latest cuda env (#15467) @raybellwaves
Doc: interleave columns pandas compat (#15383) @raybellwaves
Simplified README Examples (#15338) @wkaisertexas
Add debug tips section to libcudf developer guide (#15329) @davidwendt
Fix and clarify notes on result ordering (#13255) @shwina

🚀 New Features

Add JNI bindings for zstd compression of NVCOMP. (#15729) @firestarman
Fix spaces around CSV quoted strings (#15727) @thabetx
Add default pinned pool that falls back to new pinned allocations (#15665) @vuule
Overhaul ops-codeowners coverage (#15660) @raydouglass
Concatenate dictionary of objects along axis=1 (#15623) @er-eis
Construct pylibcudf columns from objects supporting __cuda_array_interface__ (#15615) @brandon-b-miller
Expose some Parquet per-column configuration options via the python API (#15613) @etseidl
Migrate string find operations to pylibcudf (#15604) @brandon-b-miller
Round trip FIXED_LEN_BYTE_ARRAY data properly in Parquet writer (#15600) @etseidl
Reading multi-line JSON in string columns using runtime configurable delimiter (#15556) @shrshi
Remove p...

Contributors

alliepiper, seberg, and 45 other contributors

Assets 2

05 Jun 15:08

raydouglass

v24.06.00

7c706cc

v24.06.00

🚨 Breaking Changes

Deprecate Groupby.collect (#15808) @galipremsagar
Raise FileNotFoundError when a literal JSON string that looks like a json filename is passed (#15806) @lithomas1
Support filtered I/O in chunked_parquet_reader and simplify the use of parquet_reader_options (#15764) @mhaseeb123
Raise errors for unsupported operations on certain types (#15712) @galipremsagar
Support DurationType in cudf parquet reader via arrow:schema (#15617) @mhaseeb123
Remove protobuf and use parsed ORC statistics from libcudf (#15564) @bdice
Remove legacy JSON reader from Python (#15538) @bdice
Removing all batching code from parquet writer (#15528) @mhaseeb123
Convert libcudf resource parameters to rmm::device_async_resource_ref (#15507) @harrism
Remove deprecated strings offsets_begin (#15454) @davidwendt
Floating <--> fixed-point conversion must now be called explicitly (#15438) @pmattione-nvidia
Bind read_parquet_metadata API to libcudf instead of pyarrow and extract RowGroup information (#15398) @mhaseeb123
Remove deprecated hash() and spark_murmurhash3_x86_32() (#15375) @davidwendt
Remove empty elements from exploded character-ngrams output (#15371) @davidwendt
[FEA] Performance improvement for mixed left semi/anti join (#15288) @tgujar
Align date_range defaults with pandas, support tz (#15139) @mroeschke

🐛 Bug Fixes

Revert "Fix docs for IO readers and strings_convert" (#15872) @vyasr
Remove problematic call of index setter to unblock dask-cuda CI (#15844) @charlesbluca
Use rapids_cpm_nvtx3 to get same nvtx3 target state as rmm (#15840) @robertmaynard
Return boolean from config_host_memory_resource instead of throwing (#15815) @abellina
Add temporary dask-cudf workaround for categorical sorting (#15801) @rjzamora
Fix row group alignment in ORC writer (#15789) @vuule
Raise error when sorting by categorical column in dask-cudf (#15788) @rjzamora
Upgrade arrow to 16.1 (#15787) @galipremsagar
Add support for PandasArray for pandas<2.1.0 (#15786) @galipremsagar
Limit runtime dependency to libarrow>=16.0.0,<16.1.0a0 (#15782) @pentschev
Fix cat.as_ordered not propogating correct size (#15780) @mroeschke
Handle mixed-like homogeneous types in isin (#15771) @galipremsagar
Fix id_vars and value_vars not accepting string scalars in melt (#15765) @mroeschke
Fix DatetimeIndex.loc for all types of ordering cases (#15761) @galipremsagar
Fix arrow versioning logic (#15755) @vyasr
Avoid running sanitizer on Java test designed to cause an error (#15753) @jlowe
Handle empty dataframe object with index present in setitem of loc (#15752) @galipremsagar
Eliminate circular reference in DataFrame/Series.iloc/loc (#15749) @mroeschke
Cap the absolute row index per pass in parquet chunked reader. (#15735) @nvdbaranec
Fix Index.repeat for datetime64 types (#15722) @galipremsagar
Fix multibyte check for case convert for large strings (#15721) @davidwendt
Fix get_loc to properly fetch results from an index that is in decreasing order (#15719) @galipremsagar
Return same type as the original index for .loc operations (#15717) @galipremsagar
Correct static builds + static arrow (#15715) @robertmaynard
Raise errors for unsupported operations on certain types (#15712) @galipremsagar
Fix ColumnAccessor caching of nrows if empty previously (#15710) @mroeschke
Allow None when nan_as_null=False in column constructor (#15709) @galipremsagar
Refine CudaTest.testCudaException in case throwing wrong type of CudaError under aarch64 (#15706) @sperlingxx
Fix maxima of categorical column (#15701) @rjzamora
Add proxy for inplace operations in cudf.pandas (#15695) @galipremsagar
Make nan_as_null behavior consistent across all APIs (#15692) @galipremsagar
Fix CI s3 api command to fetch latest results (#15687) @galipremsagar
Add NumpyExtensionArray proxy type in cudf.pandas (#15686) @galipremsagar
Properly implement binaryops for proxy types (#15684) @galipremsagar
Fix copy assignment and the comparison operator of rmm_host_allocator (#15677) @vuule
Fix multi-source reading in JSON byte range reader (#15671) @shrshi
Return int64 when pandas compatible mode is turned on for get_indexer (#15659) @galipremsagar
Fix Index contains for error validations and float vs int comparisons (#15657) @galipremsagar
Preserve sub-second data for time scalars in column construction (#15655) @galipremsagar
Check row limit size in cudf::strings::join_strings (#15643) @davidwendt
Enable sorting on column with nulls using query-planning (#15639) @rjzamora
Fix operator precedence problem in Parquet reader (#15638) @etseidl
Fix decoding of dictionary encoded FIXED_LEN_BYTE_ARRAY data in Parquet reader (#15601) @etseidl
Fix debug warnings/errors in from_arrow_device_test.cpp (#15596) @davidwendt
Add "collect" aggregation support to dask-cudf (#15593) @rjzamora
Fix categorical-accessor support and testing in dask-cudf (#15591) @rjzamora
Disable compute-sanitizer usage in CI tests with CUDA<11.6 (#15584) @davidwendt
Preserve RangeIndex.step in to_arrow/from_arrow (#15581) @mroeschke
Ignore new cupy warning (#15574) @vyasr
Add cuda-sanitizer-api dependency for test-cpp matrix 11.4 (#15573) @davidwendt
Allow apply udf to reference global modules in cudf.pandas (#15569) @mroeschke
Fix deprecation warnings for json legacy reader (#15563) @davidwendt
Fix millisecond resampling in cudf Python (#15560) @mroeschke
Rename JSON_READER_OPTION to JSON_READER_OPTION_NVBENCH. (#15553) @bdice
Fix a JNI bug in JSON parsing fixup (#15550) @revans2
Remove conda channel setup from wheel CI image script. (#15539) @bdice
cudf.pandas: Series dt accessor is CombinedDatetimelikeProperties (#15523) @wence-
Fix for some compiler warnings in parquet/page_decode.cuh (#15518) @etseidl
Fix exponent overflow in strings-to-double conversion (#15517) @davidwendt
nanoarrow uses package override for proper pinned versions generation (#15515) @robertmaynard
Remove index name overrides in dask-cudf pyarrow table dispatch (#15514) @charlesbluca
Fix async synchronization issues in json_column.cu (#15497) @karthikeyann
Add new patch to hide more CCCL APIs (#15493) @vyasr
Make improvements in pandas-test reporting (#15485) @galipremsagar
Fixed page data truncation in parquet writer under certain conditions. (#15474) @nvdbaranec
Only use data_type constructor with scale for decimal types (#15472) @wence-
Avoid "p2p" shuffle as a default when dask_cudf is imported (#15469) @rjzamora
Fix debug build errors from to_arrow_device_test.cpp (#15463) @davidwendt
Fix base_normalator::integer_sizeof_fn integer dispatch (#15457) @davidwendt
Allow consumers of static builds to find nanoarrow (#15456) @robertmaynard
Allow jit compilation when using a splayed CUDA toolkit (#15451) @robertmaynard
Handle case of scan aggregation in groupby-transform (#15450) @wence-
Test static builds in CI and fix nanoarrow configure (#15437) @vyasr
Fixes potential race in JSON parser when parsing JSON lines format and when recovering from invalid lines (#15419) @elstehle
Fix errors in chunked ORC writer when no tables were (successfully) written (#15393) @vuule
Support implicit array conversion with query-planning enabled (#15378) @rjzamora
Fix arrow-based round trip of empty dataframes (#15373) @wence-
Remove empty elements from exploded character-ngrams output (#15371) @davidwendt
Remove boundscheck=False setting in cython files (#15362) @wence-
Patch dask-expr var logic in dask-cudf (#15347) @rjzamora
Fix for logical and syntactical errors in libcudf c++ examples (#15346) @mhaseeb123
Disable dask-expr in docs builds. (#15343) @bdice
Apply the cuFile error work around to data_sink as well (#15335) @vuule
Fix parquet predicate filtering with column projection (#15113) @karthikeyann
Check column type equality, handling nested types correctly. (#14531) @bdice

📖 Documentation

Fix docs for IO readers and strings_convert (#15842) @bdice
Update cudf.pandas docs for GA (#15744) @beckernick
Add contributing warning about circular imports (#15691) @er-eis
Update libcudf developer guide for strings offsets column (#15661) @davidwendt
Update developer guide with device_async_resource_ref guidelines (#15562) @harrism
DOC: add pandas intersphinx mapping (#15531) @raybellwaves
rm-dup-doc in frame.py (#15530) @raybellwaves
Update CONTRIBUTING.md to use latest cuda env (#15467) @raybellwaves
Doc: interleave columns pandas compat (#15383) @raybellwaves
Simplified README Examples (#15338) @wkaisertexas
Add debug tips section to libcudf developer guide (#15329) @davidwendt
Fix and clarify notes on result ordering (#13255) @shwina

🚀 New Features

Add JNI bindings for zstd compression of NVCOMP. (#15729) @firestarman
Fix spaces around CSV quoted strings (#15727) @thabetx
Add default pinned pool that falls back to new pinned allocations (#15665) @vuule
Overhaul ops-codeowners coverage (#15660) @raydouglass
Concatenate dictionary of objects along axis=1 (#15623) @er-eis
Construct pylibcudf columns from objects supporting __cuda_array_interface__ (#15615) @brandon-b-miller
Expose some Parquet per-column configuration options via the python API (#15613) @etseidl
Migrate string find operations to pylibcudf (#15604) @brandon-b-miller
Round trip FIXED_LEN_BYTE_ARRAY data properly in Parquet writer (#15600) @etseidl
Reading multi-line JSON in string columns using runtime configurable delimiter (#15556) @shrshi
Remove public gtest dependency from libcudf conda package (#15534) @robertmaynard
Fea/move to latest nanoarrow (#15526) @robertmaynard
Migrate string case operations to pylibcudf (#15489) @brandon-b-miller
Add Parquet encoding statistics to column chunk metadata (#15452) @etseidl
Implement JNI fo...

Contributors

alliepiper, seberg, and 45 other contributors

Assets 2

24 May 20:20

rapids-bot

v24.08.00a

f5d1c24

[NIGHTLY] v24.08.00 Pre-release

Pre-release

🔗 Links

🚨 Breaking Changes

Remove mr param from write_csv and write_json (#16231) @JayjeetAtGithub
Refactor from_arrow_device/host to use resource_ref (#16160) @harrism
Return FrozenList for Index.names (#16047) @galipremsagar
Add compile option to enable large strings support (#16037) @davidwendt
Rename strings multiple target replace API (#15898) @davidwendt
Pinned vector factory that uses the global pool (#15895) @vuule
Apply clang-tidy autofixes (#15894) @vyasr
Support arrow:schema in Parquet writer to faithfully roundtrip duration types with Arrow (#15875) @mhaseeb123
Expose stream parameter to public rolling APIs (#15865) @srinivasyadav18
Fix large strings handling in nvtext::character_tokenize (#15829) @davidwendt
Remove legacy JSON reader and concurrent_unordered_map.cuh. (#15813) @bdice

🐛 Bug Fixes

Fix logic in to_arrow for empty list column (#16279) @wence-
Add custom name setter and getter for proxy objects in cudf.pandas (#16234) @Matt711
Disable large string support for Java build (#16216) @jlowe
Remove CCCL patch for PR 211. (#16207) @bdice
Add single offset to an empty ListArray in cudf::to_arrow (#16201) @davidwendt
Fix memory_usage when calculating nested list column (#16193) @mroeschke
Support at/iat indexers in cudf.pandas (#16177) @mroeschke
Fix unused-return-value debug build error in from_arrow_stream_test.cpp (#16168) @davidwendt
Fix cudf::strings::replace_multiple hang on empty target (#16167) @davidwendt
Refactor from_arrow_device/host to use resource_ref (#16160) @harrism
interpolate returns new column if no values are interpolated (#16158) @mroeschke
Use provided memory resource for allocating mixed join results. (#16153) @bdice
Run DFG after verify-alpha-spec (#16151) @KyleFromNVIDIA
Use size_t to allow large conditional joins (#16127) @bdice
Allow only scale=0 fixed-point values in fixed_width_column_wrapper (#16120) @davidwendt
Fix pylibcudf Table.num_rows for 0 columns case and add interop to docs (#16108) @lithomas1
Add support for proxy np.flatiter objects (#16107) @Matt711
Ensure cudf objects can astype to any type when empty (#16106) @mroeschke
Support pd.read_pickle and pd.to_pickle in cudf.pandas (#16105) @Matt711
Fix unnecessarily strict check in parquet chunked reader for choosing split locations. (#16099) @nvdbaranec
Fix is_monotonic_* APIs to include nan's (#16085) @galipremsagar
More safely parse CUDA versions when subprocess output is contaminated (#16067) @brandon-b-miller
fast_slow_proxy: Don't import assert_eq at top-level (#16063) @wence-
Prevent bad ColumnAccessor state after .sort_index(axis=1, ignore_index=True) (#16061) @mroeschke
Fix ArrowDeviceArray interface to pass address of event (#16058) @zeroshade
Fix a size overflow bug in hash groupby (#16053) @PointKernel
Fix atomic_ref scope when multiple blocks are updating the same output (#16051) @vuule
Fix initialization error in to_arrow for empty string views (#16033) @wence-
Fix the int32 overflow when computing page fragment sizes for large string columns (#16028) @mhaseeb123
Fix the pool size alignment issue (#16024) @PointKernel
Improve multibyte-split byte-range performance (#16019) @davidwendt
Fix target counting in strings char-parallel replace (#16017) @davidwendt
Support IntervalDtype in cudf.from_pandas (#16014) @mroeschke
Fix memory size in create_byte_range_infos_consecutive (#16012) @davidwendt
Fix Cython typo preventing proper inheritance (#15978) @vyasr
Fix convert_dtypes with convert_integer=False/convert_floating=True (#15964) @mroeschke
Fix nunique for MultiIndex, DataFrame, and all NA case with dropna=False (#15962) @mroeschke
Explicitly build for all GPU architectures (#15959) @vyasr
Preserve column type and class information in more DataFrame operations (#15949) @mroeschke
Add array_interface to cudf.pandas numpy.ndarray proxy (#15936) @mroeschke
Allow tests to be built when stream util is disabled (#15933) @robertmaynard
Fix JSON multi-source reading when total source size exceeds INT_MAX bytes (#15930) @shrshi
Fix dask_cudf.read_parquet regression for legacy timestamp data (#15929) @rjzamora
Fix offsetalator when accessing over 268 million rows (#15921) @davidwendt
Fix debug assert in rowgroup_char_counts_kernel (#15902) @davidwendt
Fix categorical conversion from chunked arrow arrays (#15886) @vyasr
Handling for NaN and inf when converting floating point to fixed point types (#15885) @ttnghia
Manual merge of Branch 24.08 from 24.06 (#15869) @galipremsagar
Avoid unnecessary Index cast in IndexedFrame.index setter (#15843) @charlesbluca
Fix large strings handling in nvtext::character_tokenize (#15829) @davidwendt
Fix multi-replace target count logic for large strings (#15807) @davidwendt
Fix JSON parsing memory corruption - Fix Mixed types nested children removal (#15798) @karthikeyann
Allow anonymous user in devcontainer name. (#15784) @bdice
Add support for additional metaclasses of proxies and use for ExcelWriter (#15399) @vyasr

📖 Documentation

Add docstring for from_dataframe (#16260) @mroeschke
Update libcudf compiler requirements in contributing doc (#16103) @davidwendt
Add libcudf public/detail API pattern to developer guide (#16086) @davidwendt
Explain line profiler and how to know which functions are GPU-accelerated. (#16079) @bdice
cudf.pandas documentation improvement (#15948) @Matt711
Reland "Fix docs for IO readers and strings_convert" (#15872)" (#15941) @lithomas1
Document how to use cudf.pandas in tandem with multiprocessing (#15940) @wence-
DOC: Add documentation for cudf.pandas in the Developer Guide (#15889) @Matt711
Improve options docs (#15888) @bdice
DOC: add linkcode to docs (#15860) @raybellwaves
Update PandasCompat.py to resolve references (#15704) @raybellwaves

🚀 New Features

Publish cudf-polars nightlies (#16213) @lithomas1
Migrate lists/modifying to pylibcudf (#16185) @Matt711
Add missing methods to lists/list_column_view.pxd in pylibcudf (#16175) @Matt711
Migrate pylibcudf lists gathering (#16170) @Matt711
Add groupby_max multi-threaded benchmark (#16154) @srinivasyadav18
Promote has_nested_columns to cudf public API (#16131) @robertmaynard
Promote IO support queries to cudf API (#16125) @robertmaynard
cudf::merge public API now support passing a user stream (#16124) @robertmaynard
Installed cudf header use cudf::allocate_like (#16087) @robertmaynard
cudf-polars string slicing (#16082) @brandon-b-miller
Migrate lists/extract to pylibcudf (#16071) @Matt711
Move common string utilities to public api (#16070) @robertmaynard
stable_distinct public api now has a stream parameter (#16068) @robertmaynard
Add support to ArrowDataSource in SourceInfo (#16050) @lithomas1
Migrate string slice APIs to pylibcudf (#15988) @brandon-b-miller
Migrate lists/contains to pylibcudf (#15981) @Matt711
Remove CCCL 2.2 patches as we now always use 2.5+ (#15969) @robertmaynard
Migrate JSON reader to pylibcudf (#15966) @lithomas1
Add a developer check for proxy objects (#15956) @Matt711
Start migrating I/O writers to pylibcudf (starting with JSON) (#15952) @lithomas1
Kernel copy for pinned memory (#15934) @vuule
Migrate left join and conditional join benchmarks to use nvbench (#15931) @srinivasyadav18
Migrate lists/combine to pylibcudf (#15928) @Matt711
Plumb pylibcudf strings contains_re through cudf_polars (#15918) @brandon-b-miller
Start migrating I/O to pylibcudf (#15899) @lithomas1
Pinned vector factory that uses the global pool (#15895) @vuule
Migrate strings contains operations to pylibcudf (#15880) @brandon-b-miller
Migrate quantile.pxd to pylibcudf (#15874) @lithomas1
Migrate round to pylibcudf (#15863) @lithomas1
Migrate string replace.pxd to pylibcudf (#15839) @lithomas1
Add an Environment Variable for debugging the fast path in cudf.pandas (#15837) @Matt711
Add an option to run cuIO benchmarks with pinned buffers as input (#15830) @vuule
Update pylibcudf testing utilities (#15772) @brandon-b-miller
Migrate string capitalize APIs to pylibcudf (#15503) @brandon-b-miller
Migrate column factories to pylibcudf (#15257) @brandon-b-miller
cuDF/libcudf exponentially weighted moving averages (#9027) @brandon-b-miller

🛠️ Improvements

Introduce version file so we can conditionally handle things in tests (#16280) @wence-
Type & reduce cupy usage (#16277) @mroeschke
Replace is_datetime/timedelta_dtype checks with .kind checks (#16262) @mroeschke
Replace is_bool_type with checking .dtype.kind (#16255) @mroeschke
remove cuco_noexcept.diff (#16254) @trxcllnt
Update contains_tests.cpp to use public cudf::slice (#16253) @davidwendt
Improve the test data for pylibcudf I/O tests (#16247) @lithomas1
Make nvcomp adapter compatible with new version macros (#16245) @vuule
Add Column.strftime/strptime instead of overloading as_string/datetime/timedelta_column (#16243) @mroeschke
Remove temporary functor overloads required by cuco version bump (#16242) @PointKernel
Expose sorted groupby parameters to pylibcudf (#16240) @wence-
Expose reflection to check if casting between two types is supported (#16239) @wence-
Handle nans in groupby-aggregations in polars executor (#16233) @wence-
Remove mr param from write_csv and write_json (#16231) @JayjeetAtGithub
Handler csv reader options in cudf-polars (#16211) @wence-
Add low memory JSON reader for cudf.pandas (#16204) @galipremsagar
Clean up state variables in MultiIndex (#16203) @mroeschke
skip CMake 3.30.0 (#16202) @jameslamb
Assert valid metadata is passed in to_arrow for list_view (#16198) ...

Contributors

seberg, trxcllnt, and 34 other contributors

Assets 2

01 May 12:29

raydouglass

v24.04.01

af9fd84

v24.04.01

🚨 Breaking Changes

Restructure pylibcudf/arrow interop facilities (#15325) @vyasr
Change exceptions thrown by copying APIs (#15319) @vyasr
Change strings_column_view::char_size to return int64 (#15197) @davidwendt
Upgrade to arrow-14.0.2 (#15108) @galipremsagar
Add support for pandas-2.2 in cudf (#15100) @galipremsagar
Deprecate cudf::hashing::spark_murmurhash3_x86_32 (#15074) @davidwendt
Align MultiIndex.get_indexder with pandas 2.2 change (#15059) @mroeschke
Raise an error on import for unsupported GPUs. (#15053) @bdice
Deprecate datelike isin casting strings to dates to match pandas 2.2 (#15046) @mroeschke
Align concat Series name behavior in pandas 2.2 (#15032) @mroeschke
Add future_stack to DataFrame.stack (#15015) @galipremsagar
Deprecate groupby fillna (#15000) @mroeschke
Deprecate replace with categorical columns (#14988) @mroeschke
Deprecate delim_whitespace in read_csv for pandas 2.2 (#14986) @mroeschke
Deprecate parameters similar to pandas 2.2 (#14984) @mroeschke
Add missing atomic operators, refactor atomic operators, move atomic operators to detail namespace. (#14962) @bdice
Add pandas-2.x support in cudf (#14916) @galipremsagar
Use cuco::static_set in the hash-based groupby (#14813) @PointKernel

🐛 Bug Fixes

Fix an issue with creating a series from scalar when dtype='category' (#15476) @galipremsagar
Update pre-commit-hooks to v0.0.3 (#15355) @KyleFromNVIDIA
[BUG][JNI] Trigger MemoryBuffer.onClosed after memory is freed (#15351) @abellina
Fix an issue with multiple short list rowgroups using the Parquet chunked reader. (#15342) @nvdbaranec
Avoid importing dask-expr if "query-planning" config is False (#15340) @rjzamora
Fix gtests/ERROR_TEST errors when run in Debug (#15317) @davidwendt
Fix OOB read in inflate_kernel (#15309) @vuule
Work around a cuFile error when running CSV tests with memcheck (#15293) @vuule
Fix Doxygen upload directory (#15291) @KyleFromNVIDIA
Fix Doxygen check (#15289) @KyleFromNVIDIA
Reintroduce PANDAS_GE_220 import (#15287) @wence-
Fix mean computation for the geometric distribution in the data generator (#15282) @vuule
Fix Parquet decimal64 stats (#15281) @etseidl
Make linking of nvtx3-cpp BUILD_LOCAL_INTERFACE (#15271) @KyleFromNVIDIA
Workaround compute-sanitizer memcheck bug (#15259) @davidwendt
Cleanup hostdevice_vector and add more APIs (#15252) @ttnghia
Fix number of rows in randomly generated lists columns (#15248) @vuule
Fix wrong output for collect_list/collect_set of lists column (#15243) @ttnghia
Fix testchunkedPackTwoPasses to copy from the bounce buffer (#15220) @abellina
Fix accessing .columns by an external API (#15212) @galipremsagar
[JNI] Disable testChunkedPackTwoPasses for now (#15210) @abellina
Update labeler and codeowner configs for CMake files (#15208) @PointKernel
Avoid dict normalization in __dask_tokenize__ (#15187) @rjzamora
Fix memcheck error in distinct inner join (#15164) @PointKernel
Remove unneeded script parameters in test_cpp_memcheck.sh (#15158) @davidwendt
Fix ListColumn.to_pandas() to retain list type (#15155) @galipremsagar
Avoid factorization in MultiIndex.to_pandas (#15150) @mroeschke
Fix GroupBy.get_group and GroupBy.indices (#15143) @wence-
Remove const from range_window_bounds::_extent. (#15138) @mythrocks
DataFrame.columns = ... retains RangeIndex & set dtype (#15129) @mroeschke
Correctly handle output for GroupBy.apply when chunk results are reindexed series (#15109) @brandon-b-miller
Fix Series.groupby.shift with a MultiIndex (#15098) @mroeschke
Fix reductions when DataFrame has MulitIndex columns (#15097) @mroeschke
Fix deprecation warnings for deprecated hash() calls (#15095) @davidwendt
Add support for arrow large_string in cudf (#15093) @galipremsagar
Fix sort_values pytest failure with pandas-2.x regression (#15092) @galipremsagar
Resolve path parsing issues in get_json_object (#15082) @SurajAralihalli
Fix bugs in handling of delta encodings (#15075) @etseidl
Fix is_device_write_preferred in void_sink and user_sink_wrapper (#15064) @vuule
Eliminate duplicate allocation of nested string columns (#15061) @vuule
Raise an error on import for unsupported GPUs. (#15053) @bdice
Align concat Series name behavior in pandas 2.2 (#15032) @mroeschke
Fix Index.difference to handle duplicate values when one of the inputs is empty (#15016) @galipremsagar
Add future_stack to DataFrame.stack (#15015) @galipremsagar
Fix handling of values=None in pylibcudf GroupBy.get_groups (#14998) @shwina
Fix DataFrame.sort_index to respect ignore_index on all axis (#14995) @galipremsagar
Raise for pyarrow array that is tz-aware (#14980) @mroeschke
Direct SeriesGroupBy.aggregate to SeriesGroupBy.agg (#14971) @rjzamora
Respect IntervalDtype and CategoricalDtype objects passed by users (#14961) @mroeschke
unset CUDF_SPILL after a pytest (#14958) @galipremsagar
Fix Null literals to be not parsed as string when mixed types as string is enabled in JSON reader (#14939) @karthikeyann
Fix chunked reads of Parquet delta encoded pages (#14921) @etseidl
Fix reading offset for data stream in ORC reader (#14911) @ttnghia
Enable sanitizer check for a test case testORCReadAndWriteForDecimal128 (#14897) @res-life
Fix dask token normalization (#14829) @rjzamora
Fix 24.04 versions (#14825) @raydouglass
Ensure slow private attrs are maybe proxies (#14380) @mroeschke

📖 Documentation

Ignore DLManagedTensor in the docs build (#15392) @davidwendt
Revert "Temporarily disable docs errors. (#15265)" (#15269) @bdice
Temporarily disable docs errors. (#15265) @bdice
Update developer_guide.md with new guidance on quoted internal includes (#15238) @harrism
Fix broken link for developer guide (#15025) @sanjana098
[DOC] Update typo in docs example of structs_column_wrapper (#14949) @karthikeyann
Update cudf.pandas FAQ. (#14940) @bdice
Optimize doc builds (#14856) @vyasr
Add developer guideline to use east const. (#14836) @bdice
Document how cuDF is pronounced (#14753) @pentschev
Notes convert to Pandas-compat (#12641) @Touutae-lab

🚀 New Features

Address inconsistency in single quote normalization in JSON reader (#15324) @shrshi
Use JNI pinned pool resource with cuIO (#15255) @abellina
Add DELTA_BYTE_ARRAY encoder for Parquet (#15239) @etseidl
Migrate filling operations to pylibcudf (#15225) @brandon-b-miller
[JNI] rmm based pinned pool (#15219) @abellina
Implement zero-copy host buffer source instead of using an arrow implementation (#15189) @vuule
Enable creation of columns from scalar (#15181) @vyasr
Use NVTX from GitHub. (#15178) @bdice
Implement segmented_row_bit_count for computing row sizes by segments of rows (#15169) @ttnghia
Implement search using pylibcudf (#15166) @vyasr
Add distinct left join (#15149) @PointKernel
Add cardinality control for groupby benchs with flat types (#15134) @PointKernel
Add ability to request Parquet encodings on a per-column basis (#15081) @etseidl
Automate include grouping order in .clang-format (#15063) @harrism
Requesting a clean build directory also clears Jitify cache (#15052) @robertmaynard
API for JSON unquoted whitespace normalization (#15033) @shrshi
Implement concatenate, lists.explode, merge, sorting, and stream compaction in pylibcudf (#15011) @vyasr
Implement replace in pylibcudf (#15005) @vyasr
Add distinct key inner join (#14990) @PointKernel
Implement rolling in pylibcudf (#14982) @vyasr
Implement joins in pylibcudf (#14972) @vyasr
Implement scans and reductions in pylibcudf (#14970) @vyasr
Rewrite cudf internals using pylibcudf groupby (#14946) @vyasr
Implement groupby in pylibcudf (#14945) @vyasr
Support casting of Map type to string in JSON reader (#14936) @karthikeyann
POC for whitespace removal in input JSON data using FST (#14931) @shrshi
Support for LZ4 compression in ORC and Parquet (#14906) @vuule
Remove supports_streams from cuDF custom memory resources. (#14857) @harrism
Migrate unary operations to pylibcudf (#14850) @vyasr
Migrate binary operations to pylibcudf (#14821) @vyasr
Add row index and stripe size options to Python ORC chunked writer (#14785) @vuule
Support CUDA 12.2 (#14712) @jameslamb

🛠️ Improvements

Backport: Relax protobuf lower bound to 3.20. (#15506) (#15610) @bdice
Use conda env create --yes instead of --force (#15403) @bdice
Restructure pylibcudf/arrow interop facilities (#15325) @vyasr
Change exceptions thrown by copying APIs (#15319) @vyasr
Enable branch testing for cudf.pandas (#15316) @galipremsagar
Replace black with ruff-format (#15312) @mroeschke
This fixes an NPE when trying to read empty JSON data by adding a new API for missing information (#15307) @revans2
Address poor performance of Parquet string decoding (#15304) @etseidl
Update script input name (#15301) @AyodeAwe
Make test_read_parquet_partitioned_filtered data deterministic (#15296) @mroeschke
Add timeout for cudf.pandas pandas tests (#15284) @galipremsagar
Add upper bound to prevent usage of NumPy 2 (#15283) @bdice
Fix cudf::test::to_host return of host_vector (#15263) @davidwendt
Implement grouped product scan (#15254) @wence-
Add CUDA 12.4 to supported PTX versions (#15247) @brandon-b-miller
Implement DataFrame|Series.squeeze (#15244) @mroeschke
Roll back ipow changes due to register pressure. (#15242) @pmattione-nvidia
Remove create_chars_child_column utility (#15241) @davidwendt
Update dlpack to version 0.8 (#15237) @dantegd
Improve performance in JSON reader when mixed_types_as_string option is enabled (#15236) @shrshi
Remove row conversion code from libcudf (#15234) @ttnghia
Use variable substitution for RAPIDS version in Doxyfile (#15231) @KyleFromNVIDIA
Add ListColumns.to_pandas(arrow_type=) (#15228) @mroeSC...

Contributors

trxcllnt, robertmaynard, and 36 other contributors

Assets 2

10 Apr 15:05

raydouglass

v24.04.00

578c240

v24.04.00

🚨 Breaking Changes

Restructure pylibcudf/arrow interop facilities (#15325) @vyasr
Change exceptions thrown by copying APIs (#15319) @vyasr
Change strings_column_view::char_size to return int64 (#15197) @davidwendt
Upgrade to arrow-14.0.2 (#15108) @galipremsagar
Add support for pandas-2.2 in cudf (#15100) @galipremsagar
Deprecate cudf::hashing::spark_murmurhash3_x86_32 (#15074) @davidwendt
Align MultiIndex.get_indexder with pandas 2.2 change (#15059) @mroeschke
Raise an error on import for unsupported GPUs. (#15053) @bdice
Deprecate datelike isin casting strings to dates to match pandas 2.2 (#15046) @mroeschke
Align concat Series name behavior in pandas 2.2 (#15032) @mroeschke
Add future_stack to DataFrame.stack (#15015) @galipremsagar
Deprecate groupby fillna (#15000) @mroeschke
Deprecate replace with categorical columns (#14988) @mroeschke
Deprecate delim_whitespace in read_csv for pandas 2.2 (#14986) @mroeschke
Deprecate parameters similar to pandas 2.2 (#14984) @mroeschke
Add missing atomic operators, refactor atomic operators, move atomic operators to detail namespace. (#14962) @bdice
Add pandas-2.x support in cudf (#14916) @galipremsagar
Use cuco::static_set in the hash-based groupby (#14813) @PointKernel

🐛 Bug Fixes

Fix an issue with creating a series from scalar when dtype='category' (#15476) @galipremsagar
Update pre-commit-hooks to v0.0.3 (#15355) @KyleFromNVIDIA
[BUG][JNI] Trigger MemoryBuffer.onClosed after memory is freed (#15351) @abellina
Fix an issue with multiple short list rowgroups using the Parquet chunked reader. (#15342) @nvdbaranec
Avoid importing dask-expr if "query-planning" config is False (#15340) @rjzamora
Fix gtests/ERROR_TEST errors when run in Debug (#15317) @davidwendt
Fix OOB read in inflate_kernel (#15309) @vuule
Work around a cuFile error when running CSV tests with memcheck (#15293) @vuule
Fix Doxygen upload directory (#15291) @KyleFromNVIDIA
Fix Doxygen check (#15289) @KyleFromNVIDIA
Reintroduce PANDAS_GE_220 import (#15287) @wence-
Fix mean computation for the geometric distribution in the data generator (#15282) @vuule
Fix Parquet decimal64 stats (#15281) @etseidl
Make linking of nvtx3-cpp BUILD_LOCAL_INTERFACE (#15271) @KyleFromNVIDIA
Workaround compute-sanitizer memcheck bug (#15259) @davidwendt
Cleanup hostdevice_vector and add more APIs (#15252) @ttnghia
Fix number of rows in randomly generated lists columns (#15248) @vuule
Fix wrong output for collect_list/collect_set of lists column (#15243) @ttnghia
Fix testchunkedPackTwoPasses to copy from the bounce buffer (#15220) @abellina
Fix accessing .columns by an external API (#15212) @galipremsagar
[JNI] Disable testChunkedPackTwoPasses for now (#15210) @abellina
Update labeler and codeowner configs for CMake files (#15208) @PointKernel
Avoid dict normalization in __dask_tokenize__ (#15187) @rjzamora
Fix memcheck error in distinct inner join (#15164) @PointKernel
Remove unneeded script parameters in test_cpp_memcheck.sh (#15158) @davidwendt
Fix ListColumn.to_pandas() to retain list type (#15155) @galipremsagar
Avoid factorization in MultiIndex.to_pandas (#15150) @mroeschke
Fix GroupBy.get_group and GroupBy.indices (#15143) @wence-
Remove const from range_window_bounds::_extent. (#15138) @mythrocks
DataFrame.columns = ... retains RangeIndex & set dtype (#15129) @mroeschke
Correctly handle output for GroupBy.apply when chunk results are reindexed series (#15109) @brandon-b-miller
Fix Series.groupby.shift with a MultiIndex (#15098) @mroeschke
Fix reductions when DataFrame has MulitIndex columns (#15097) @mroeschke
Fix deprecation warnings for deprecated hash() calls (#15095) @davidwendt
Add support for arrow large_string in cudf (#15093) @galipremsagar
Fix sort_values pytest failure with pandas-2.x regression (#15092) @galipremsagar
Resolve path parsing issues in get_json_object (#15082) @SurajAralihalli
Fix bugs in handling of delta encodings (#15075) @etseidl
Fix is_device_write_preferred in void_sink and user_sink_wrapper (#15064) @vuule
Eliminate duplicate allocation of nested string columns (#15061) @vuule
Raise an error on import for unsupported GPUs. (#15053) @bdice
Align concat Series name behavior in pandas 2.2 (#15032) @mroeschke
Fix Index.difference to handle duplicate values when one of the inputs is empty (#15016) @galipremsagar
Add future_stack to DataFrame.stack (#15015) @galipremsagar
Fix handling of values=None in pylibcudf GroupBy.get_groups (#14998) @shwina
Fix DataFrame.sort_index to respect ignore_index on all axis (#14995) @galipremsagar
Raise for pyarrow array that is tz-aware (#14980) @mroeschke
Direct SeriesGroupBy.aggregate to SeriesGroupBy.agg (#14971) @rjzamora
Respect IntervalDtype and CategoricalDtype objects passed by users (#14961) @mroeschke
unset CUDF_SPILL after a pytest (#14958) @galipremsagar
Fix Null literals to be not parsed as string when mixed types as string is enabled in JSON reader (#14939) @karthikeyann
Fix chunked reads of Parquet delta encoded pages (#14921) @etseidl
Fix reading offset for data stream in ORC reader (#14911) @ttnghia
Enable sanitizer check for a test case testORCReadAndWriteForDecimal128 (#14897) @res-life
Fix dask token normalization (#14829) @rjzamora
Fix 24.04 versions (#14825) @raydouglass
Ensure slow private attrs are maybe proxies (#14380) @mroeschke

📖 Documentation

Ignore DLManagedTensor in the docs build (#15392) @davidwendt
Revert "Temporarily disable docs errors. (#15265)" (#15269) @bdice
Temporarily disable docs errors. (#15265) @bdice
Update developer_guide.md with new guidance on quoted internal includes (#15238) @harrism
Fix broken link for developer guide (#15025) @sanjana098
[DOC] Update typo in docs example of structs_column_wrapper (#14949) @karthikeyann
Update cudf.pandas FAQ. (#14940) @bdice
Optimize doc builds (#14856) @vyasr
Add developer guideline to use east const. (#14836) @bdice
Document how cuDF is pronounced (#14753) @pentschev
Notes convert to Pandas-compat (#12641) @Touutae-lab

🚀 New Features

Address inconsistency in single quote normalization in JSON reader (#15324) @shrshi
Use JNI pinned pool resource with cuIO (#15255) @abellina
Add DELTA_BYTE_ARRAY encoder for Parquet (#15239) @etseidl
Migrate filling operations to pylibcudf (#15225) @brandon-b-miller
[JNI] rmm based pinned pool (#15219) @abellina
Implement zero-copy host buffer source instead of using an arrow implementation (#15189) @vuule
Enable creation of columns from scalar (#15181) @vyasr
Use NVTX from GitHub. (#15178) @bdice
Implement segmented_row_bit_count for computing row sizes by segments of rows (#15169) @ttnghia
Implement search using pylibcudf (#15166) @vyasr
Add distinct left join (#15149) @PointKernel
Add cardinality control for groupby benchs with flat types (#15134) @PointKernel
Add ability to request Parquet encodings on a per-column basis (#15081) @etseidl
Automate include grouping order in .clang-format (#15063) @harrism
Requesting a clean build directory also clears Jitify cache (#15052) @robertmaynard
API for JSON unquoted whitespace normalization (#15033) @shrshi
Implement concatenate, lists.explode, merge, sorting, and stream compaction in pylibcudf (#15011) @vyasr
Implement replace in pylibcudf (#15005) @vyasr
Add distinct key inner join (#14990) @PointKernel
Implement rolling in pylibcudf (#14982) @vyasr
Implement joins in pylibcudf (#14972) @vyasr
Implement scans and reductions in pylibcudf (#14970) @vyasr
Rewrite cudf internals using pylibcudf groupby (#14946) @vyasr
Implement groupby in pylibcudf (#14945) @vyasr
Support casting of Map type to string in JSON reader (#14936) @karthikeyann
POC for whitespace removal in input JSON data using FST (#14931) @shrshi
Support for LZ4 compression in ORC and Parquet (#14906) @vuule
Remove supports_streams from cuDF custom memory resources. (#14857) @harrism
Migrate unary operations to pylibcudf (#14850) @vyasr
Migrate binary operations to pylibcudf (#14821) @vyasr
Add row index and stripe size options to Python ORC chunked writer (#14785) @vuule
Support CUDA 12.2 (#14712) @jameslamb

🛠️ Improvements

Use conda env create --yes instead of --force (#15403) @bdice
Restructure pylibcudf/arrow interop facilities (#15325) @vyasr
Change exceptions thrown by copying APIs (#15319) @vyasr
Enable branch testing for cudf.pandas (#15316) @galipremsagar
Replace black with ruff-format (#15312) @mroeschke
This fixes an NPE when trying to read empty JSON data by adding a new API for missing information (#15307) @revans2
Address poor performance of Parquet string decoding (#15304) @etseidl
Update script input name (#15301) @AyodeAwe
Make test_read_parquet_partitioned_filtered data deterministic (#15296) @mroeschke
Add timeout for cudf.pandas pandas tests (#15284) @galipremsagar
Add upper bound to prevent usage of NumPy 2 (#15283) @bdice
Fix cudf::test::to_host return of host_vector (#15263) @davidwendt
Implement grouped product scan (#15254) @wence-
Add CUDA 12.4 to supported PTX versions (#15247) @brandon-b-miller
Implement DataFrame|Series.squeeze (#15244) @mroeschke
Roll back ipow changes due to register pressure. (#15242) @pmattione-nvidia
Remove create_chars_child_column utility (#15241) @davidwendt
Update dlpack to version 0.8 (#15237) @dantegd
Improve performance in JSON reader when mixed_types_as_string option is enabled (#15236) @shrshi
Remove row conversion code from libcudf (#15234) @ttnghia
Use variable substitution for RAPIDS version in Doxyfile (#15231) @KyleFromNVIDIA
Add ListColumns.to_pandas(arrow_type=) (#15228) @mroeschke
Treat dask-cudf CI artifacts as pure wheels (#15223) @bdice
Clean...

Contributors

trxcllnt, robertmaynard, and 36 other contributors

Assets 2

27 Feb 15:12

raydouglass

v24.02.02

dd34fdb

v24.02.02

🚨 Breaking Changes

Remove **kwargs from astype (#14765) @mroeschke
Remove mimesis as a testing dependency (#14723) @mroeschke
Update to Dask's shuffle_method kwarg (#14708) @pentschev
Drop Pascal GPU support. (#14630) @bdice
Update to CCCL 2.2.0. (#14576) @bdice
Expunge as_frame conversions in Column algorithms (#14491) @wence-
Deprecate cudf::make_strings_column accepting typed offsets (#14461) @davidwendt
Remove deprecated nvtext::load_merge_pairs_file (#14460) @davidwendt
Include writer code and writerVersion in ORC files (#14458) @vuule
Remove null mask for zero nulls in json readers (#14451) @karthikeyann
REF: Remove **kwargs from to_pandas, raise if nullable is not implemented (#14438) @mroeschke
Consolidate 1D pandas object handling in as_column (#14394) @mroeschke
Move chars column to parent data buffer in strings column (#14202) @karthikeyann
Switch to scikit-build-core (#13531) @vyasr

🐛 Bug Fixes

Bump to nvcomp 3.0.6. (#15128) @bdice
[HOTFIX] Unpin numba<0.58 (#15031) @raydouglass
Exclude tests from builds (#14981) @vyasr
Fix the bounce buffer size in ORC writer (#14947) @vuule
Revert sum/product aggregation to always produce int64_t type (#14907) @SurajAralihalli
Fixed an issue with output chunking computation stemming from input chunking. (#14889) @nvdbaranec
Fix total_byte_size in Parquet row group metadata (#14802) @etseidl
Fix index difference to follow the pandas format (#14789) @amiralimi
Fix shared-workflows repo name (#14784) @raydouglass
Remove unparseable attributes from all nodes (#14780) @vyasr
Refactor and add validation to IntervalIndex.init (#14778) @mroeschke
Work around incompatibilities between V2 page header handling and zStandard compression in Parquet writer (#14772) @etseidl
Fix calls to deprecated strings factory API (#14771) @davidwendt
Fix ptx file discovery in editable installs (#14767) @vyasr
Revise shuffle deprecation to align with dask/dask (#14762) @rjzamora
Enable intermediate proxies to be picklable (#14752) @shwina
Add CUDF_TEST_PROGRAM_MAIN macro to tests lacking it (#14751) @etseidl
Fix CMake args (#14746) @vyasr
Fix logic bug introduced in #14730 (#14742) @wence-
[Java] Choose The Correct RoundingMode For Checking Decimal OutOfBounds (#14731) @razajafri
Fix Groupby.get_group (#14728) @rjzamora
Ensure that all CUDA kernels in cudf have hidden visibility. (#14726) @robertmaynard
Split cuda versions for notebook testing (#14722) @raydouglass
Fix to_numeric not preserving Series index and name (#14718) @mroeschke
Update dask-cudf wheel name (#14713) @raydouglass
Fix strings::contains matching end of string target (#14711) @davidwendt
Update to Dask's shuffle_method kwarg (#14708) @pentschev
Write file-level statistics when writing ORC files with zero rows (#14707) @vuule
Potential fix for peformance regression in #14415 (#14706) @etseidl
Ensure DataFrame column types are preserved during serialization (#14705) @mroeschke
Skip numba test that fails on ARM (#14702) @brandon-b-miller
Allow Z in datetime string parsing in non pandas compat mode (#14701) @mroeschke
Fix nan_as_null not being respected when passing arrow object (#14688) @mroeschke
Fix constructing Series/Index from arrow array and dtype (#14686) @mroeschke
Fix Aggregation Type Promotion: Ensure Unsigned Input Types Result in Unsigned Output for Sum and Multiply (#14679) @SurajAralihalli
Add BaseOffset as a final proxy type to pass instancechecks for offsets against BaseOffset (#14678) @shwina
Add row conversion code from spark-rapids-jni (#14664) @ttnghia
Unconditionally export the CCCL path (#14656) @vyasr
Ensure libcudf searches for our patched version of CCCL first (#14655) @robertmaynard
Constrain CUDA in notebook testing to prevent CUDA 12.1 usage until we have pynvjitlink (#14648) @vyasr
Fix invalid memory access in Parquet reader (#14637) @etseidl
Use column_empty over as_column([]) (#14632) @mroeschke
Add (implicit) handling for torch tensors in is_scalar (#14623) @wence-
Fix astype/fillna not maintaining column subclass and types (#14615) @mroeschke
Remove non-empty nulls in cudf::get_json_object (#14609) @davidwendt
Remove cuda::proclaim_return_type from nested lambda (#14607) @ttnghia
Fix DataFrame.reindex when column reindexing to MultiIndex/RangeIndex (#14605) @mroeschke
Address potential race conditions in Parquet reader (#14602) @etseidl
Fix DataFrame.reindex removing column name (#14601) @mroeschke
Remove unsanitized input test data from copy gtests (#14600) @davidwendt
Fix race detected in Parquet writer (#14598) @etseidl
Correct invalid or missing return types (#14587) @robertmaynard
Fix unsanitized nulls from strings segmented-reduce (#14586) @davidwendt
Upgrade to nvCOMP 3.0.5 (#14581) @davidwendt
Fix unsanitized nulls produced by cudf::clamp APIs (#14580) @davidwendt
Fix unsanitized nulls produced by libcudf dictionary decode (#14578) @davidwendt
Fixes a symbol group lookup table issue (#14561) @elstehle
Drop llvm16 from cuda118-conda devcontainer image (#14526) @charlesbluca
REF: Make DataFrame.from_pandas process by column (#14483) @mroeschke
Improve memory footprint of isin by using contains (#14478) @wence-
Move creation of env.yaml outside the current directory (#14476) @davidwendt
Enable pd.Timestamp objects to be picklable when cudf.pandas is active (#14474) @shwina
Correct dtype of count aggregations on empty dataframes (#14473) @wence-
Avoid DataFrame conversion in MultiIndex.from_pandas (#14470) @mroeschke
JSON writer: avoid default stream use in string_scalar constructors (#14444) @vuule
Fix default stream use in the CSV reader (#14443) @vuule
Preserve DataFrame(columns=).columns dtype during empty-like construction (#14381) @mroeschke
Defer PTX file load to runtime (#13690) @brandon-b-miller

📖 Documentation

Disable parallel build (#14796) @vyasr
Add pylibcudf to the docs (#14791) @vyasr
Describe unpickling expectations when cudf.pandas is enabled (#14693) @shwina
Update CONTRIBUTING for pyproject-only builds (#14653) @vyasr
More doxygen fixes (#14639) @vyasr
Enable doxygen XML generation and fix issues (#14477) @vyasr
Some doxygen improvements (#14469) @vyasr
Remove warning in dask-cudf docs (#14454) @wence-
Update README links with redirects. (#14378) @bdice
Add pip install instructions to README (#13677) @shwina

🚀 New Features

Add ci check for external kernels (#14768) @robertmaynard
JSON single quote normalization API (#14729) @shrshi
Write cuDF version in Parquet "created_by" metadata field (#14721) @etseidl
Implement remaining copying APIs in pylibcudf along with required helper functions (#14640) @vyasr
Don't constrain numba<0.58 (#14616) @brandon-b-miller
Add DELTA_LENGTH_BYTE_ARRAY encoder and decoder for Parquet (#14590) @etseidl
JSON - Parse mixed types as string in JSON reader (#14572) @karthikeyann
JSON quote normalization (#14545) @shrshi
Make DefaultHostMemoryAllocator settable (#14523) @gerashegalov
Implement more copying APIs in pylibcudf (#14508) @vyasr
Include writer code and writerVersion in ORC files (#14458) @vuule
Parquet sub-rowgroup reading. (#14360) @nvdbaranec
Move chars column to parent data buffer in strings column (#14202) @karthikeyann
PARQUET-2261 Size Statistics (#14000) @etseidl
Improve GroupBy JIT error handling (#13854) @brandon-b-miller
Generate unified Python/C++ docs (#13846) @vyasr
Expand JIT groupby test suite (#13813) @brandon-b-miller

🛠️ Improvements

Pin pytest<8 (#14920) @galipremsagar
Move cudf::char_utf8 definition from detail to public header (#14779) @davidwendt
Clean up TimedeltaIndex.__init__ constructor (#14775) @mroeschke
Clean up DatetimeIndex.__init__ constructor (#14774) @mroeschke
Some frame.py typing, move seldom used methods in frame.py (#14766) @mroeschke
Remove **kwargs from astype (#14765) @mroeschke
fix benchmarks compatibility with newer pytest-cases (#14764) @jameslamb
Add pynvjitlink as a dependency (#14763) @brandon-b-miller
Resolve degenerate performance in create_structs_data (#14761) @SurajAralihalli
Simplify ColumnAccessor methods; avoid unnecessary validations (#14758) @mroeschke
Pin pytest-cases<3.8.2 (#14756) @mroeschke
Use _from_data instead of _from_columns for initialzing Frame (#14755) @mroeschke
Consolidate cudf object handling in as_column (#14754) @mroeschke
Reduce execution time of Parquet C++ tests (#14750) @vuule
Implement to_datetime(..., utc=True) (#14749) @mroeschke
Remove usages of rapids-env-update (#14748) @KyleFromNVIDIA
Provide explicit pool size and avoid RMM detail APIs (#14741) @harrism
Implement cudf.MultiIndex.from_arrays (#14740) @mroeschke
Remove unused/single use methods (#14739) @mroeschke
refactor CUDA versions in dependencies.yaml (#14733) @jameslamb
Remove unneeded methods in Column (#14730) @mroeschke
Clean up base column methods (#14725) @mroeschke
Ensure column.fillna signatures are consistent (#14724) @mroeschke
Remove mimesis as a testing dependency (#14723) @mroeschke
Replace as_numerical with as_numerical_column/codes (#14719) @mroeschke
Use offsetalator in gather_chars (#14700) @davidwendt
Use make_strings_children for fill() specialization logic (#14697) @davidwendt
Change io::detail::orc namespace into io::orc::detail (#14696) @ttnghia
Fix call to deprecated factory function (#14695) @davidwendt
Use as_column instead of arange for range like inputs (#14689) @mroeschke
Reorganize ORC reader into multiple files and perform some small fixes to cuIO code (#14665) @ttnghia
Split parquet test into multiple files (#14663) @etseidl
Custom error messages for IO with nonexistent files (#14662) @vuule
Explicitly pass .dtype into is_foo_dtype functions (#14657) @mroeschke
Basic val...

Contributors

trxcllnt, robertmaynard, and 32 other contributors

Assets 2

13 Feb 14:24

raydouglass

v24.02.01

33ffdf5

v24.02.01

🚨 Breaking Changes

Remove **kwargs from astype (#14765) @mroeschke
Remove mimesis as a testing dependency (#14723) @mroeschke
Update to Dask's shuffle_method kwarg (#14708) @pentschev
Drop Pascal GPU support. (#14630) @bdice
Update to CCCL 2.2.0. (#14576) @bdice
Expunge as_frame conversions in Column algorithms (#14491) @wence-
Deprecate cudf::make_strings_column accepting typed offsets (#14461) @davidwendt
Remove deprecated nvtext::load_merge_pairs_file (#14460) @davidwendt
Include writer code and writerVersion in ORC files (#14458) @vuule
Remove null mask for zero nulls in json readers (#14451) @karthikeyann
REF: Remove **kwargs from to_pandas, raise if nullable is not implemented (#14438) @mroeschke
Consolidate 1D pandas object handling in as_column (#14394) @mroeschke
Move chars column to parent data buffer in strings column (#14202) @karthikeyann
Switch to scikit-build-core (#13531) @vyasr

🐛 Bug Fixes

[HOTFIX] Unpin numba<0.58 (#15031) @raydouglass
Exclude tests from builds (#14981) @vyasr
Fix the bounce buffer size in ORC writer (#14947) @vuule
Revert sum/product aggregation to always produce int64_t type (#14907) @SurajAralihalli
Fixed an issue with output chunking computation stemming from input chunking. (#14889) @nvdbaranec
Fix total_byte_size in Parquet row group metadata (#14802) @etseidl
Fix index difference to follow the pandas format (#14789) @amiralimi
Fix shared-workflows repo name (#14784) @raydouglass
Remove unparseable attributes from all nodes (#14780) @vyasr
Refactor and add validation to IntervalIndex.init (#14778) @mroeschke
Work around incompatibilities between V2 page header handling and zStandard compression in Parquet writer (#14772) @etseidl
Fix calls to deprecated strings factory API (#14771) @davidwendt
Fix ptx file discovery in editable installs (#14767) @vyasr
Revise shuffle deprecation to align with dask/dask (#14762) @rjzamora
Enable intermediate proxies to be picklable (#14752) @shwina
Add CUDF_TEST_PROGRAM_MAIN macro to tests lacking it (#14751) @etseidl
Fix CMake args (#14746) @vyasr
Fix logic bug introduced in #14730 (#14742) @wence-
[Java] Choose The Correct RoundingMode For Checking Decimal OutOfBounds (#14731) @razajafri
Fix Groupby.get_group (#14728) @rjzamora
Ensure that all CUDA kernels in cudf have hidden visibility. (#14726) @robertmaynard
Split cuda versions for notebook testing (#14722) @raydouglass
Fix to_numeric not preserving Series index and name (#14718) @mroeschke
Update dask-cudf wheel name (#14713) @raydouglass
Fix strings::contains matching end of string target (#14711) @davidwendt
Update to Dask's shuffle_method kwarg (#14708) @pentschev
Write file-level statistics when writing ORC files with zero rows (#14707) @vuule
Potential fix for peformance regression in #14415 (#14706) @etseidl
Ensure DataFrame column types are preserved during serialization (#14705) @mroeschke
Skip numba test that fails on ARM (#14702) @brandon-b-miller
Allow Z in datetime string parsing in non pandas compat mode (#14701) @mroeschke
Fix nan_as_null not being respected when passing arrow object (#14688) @mroeschke
Fix constructing Series/Index from arrow array and dtype (#14686) @mroeschke
Fix Aggregation Type Promotion: Ensure Unsigned Input Types Result in Unsigned Output for Sum and Multiply (#14679) @SurajAralihalli
Add BaseOffset as a final proxy type to pass instancechecks for offsets against BaseOffset (#14678) @shwina
Add row conversion code from spark-rapids-jni (#14664) @ttnghia
Unconditionally export the CCCL path (#14656) @vyasr
Ensure libcudf searches for our patched version of CCCL first (#14655) @robertmaynard
Constrain CUDA in notebook testing to prevent CUDA 12.1 usage until we have pynvjitlink (#14648) @vyasr
Fix invalid memory access in Parquet reader (#14637) @etseidl
Use column_empty over as_column([]) (#14632) @mroeschke
Add (implicit) handling for torch tensors in is_scalar (#14623) @wence-
Fix astype/fillna not maintaining column subclass and types (#14615) @mroeschke
Remove non-empty nulls in cudf::get_json_object (#14609) @davidwendt
Remove cuda::proclaim_return_type from nested lambda (#14607) @ttnghia
Fix DataFrame.reindex when column reindexing to MultiIndex/RangeIndex (#14605) @mroeschke
Address potential race conditions in Parquet reader (#14602) @etseidl
Fix DataFrame.reindex removing column name (#14601) @mroeschke
Remove unsanitized input test data from copy gtests (#14600) @davidwendt
Fix race detected in Parquet writer (#14598) @etseidl
Correct invalid or missing return types (#14587) @robertmaynard
Fix unsanitized nulls from strings segmented-reduce (#14586) @davidwendt
Upgrade to nvCOMP 3.0.5 (#14581) @davidwendt
Fix unsanitized nulls produced by cudf::clamp APIs (#14580) @davidwendt
Fix unsanitized nulls produced by libcudf dictionary decode (#14578) @davidwendt
Fixes a symbol group lookup table issue (#14561) @elstehle
Drop llvm16 from cuda118-conda devcontainer image (#14526) @charlesbluca
REF: Make DataFrame.from_pandas process by column (#14483) @mroeschke
Improve memory footprint of isin by using contains (#14478) @wence-
Move creation of env.yaml outside the current directory (#14476) @davidwendt
Enable pd.Timestamp objects to be picklable when cudf.pandas is active (#14474) @shwina
Correct dtype of count aggregations on empty dataframes (#14473) @wence-
Avoid DataFrame conversion in MultiIndex.from_pandas (#14470) @mroeschke
JSON writer: avoid default stream use in string_scalar constructors (#14444) @vuule
Fix default stream use in the CSV reader (#14443) @vuule
Preserve DataFrame(columns=).columns dtype during empty-like construction (#14381) @mroeschke
Defer PTX file load to runtime (#13690) @brandon-b-miller

📖 Documentation

Disable parallel build (#14796) @vyasr
Add pylibcudf to the docs (#14791) @vyasr
Describe unpickling expectations when cudf.pandas is enabled (#14693) @shwina
Update CONTRIBUTING for pyproject-only builds (#14653) @vyasr
More doxygen fixes (#14639) @vyasr
Enable doxygen XML generation and fix issues (#14477) @vyasr
Some doxygen improvements (#14469) @vyasr
Remove warning in dask-cudf docs (#14454) @wence-
Update README links with redirects. (#14378) @bdice
Add pip install instructions to README (#13677) @shwina

🚀 New Features

Add ci check for external kernels (#14768) @robertmaynard
JSON single quote normalization API (#14729) @shrshi
Write cuDF version in Parquet "created_by" metadata field (#14721) @etseidl
Implement remaining copying APIs in pylibcudf along with required helper functions (#14640) @vyasr
Don't constrain numba<0.58 (#14616) @brandon-b-miller
Add DELTA_LENGTH_BYTE_ARRAY encoder and decoder for Parquet (#14590) @etseidl
JSON - Parse mixed types as string in JSON reader (#14572) @karthikeyann
JSON quote normalization (#14545) @shrshi
Make DefaultHostMemoryAllocator settable (#14523) @gerashegalov
Implement more copying APIs in pylibcudf (#14508) @vyasr
Include writer code and writerVersion in ORC files (#14458) @vuule
Parquet sub-rowgroup reading. (#14360) @nvdbaranec
Move chars column to parent data buffer in strings column (#14202) @karthikeyann
PARQUET-2261 Size Statistics (#14000) @etseidl
Improve GroupBy JIT error handling (#13854) @brandon-b-miller
Generate unified Python/C++ docs (#13846) @vyasr
Expand JIT groupby test suite (#13813) @brandon-b-miller

🛠️ Improvements

Pin pytest<8 (#14920) @galipremsagar
Move cudf::char_utf8 definition from detail to public header (#14779) @davidwendt
Clean up TimedeltaIndex.__init__ constructor (#14775) @mroeschke
Clean up DatetimeIndex.__init__ constructor (#14774) @mroeschke
Some frame.py typing, move seldom used methods in frame.py (#14766) @mroeschke
Remove **kwargs from astype (#14765) @mroeschke
fix benchmarks compatibility with newer pytest-cases (#14764) @jameslamb
Add pynvjitlink as a dependency (#14763) @brandon-b-miller
Resolve degenerate performance in create_structs_data (#14761) @SurajAralihalli
Simplify ColumnAccessor methods; avoid unnecessary validations (#14758) @mroeschke
Pin pytest-cases<3.8.2 (#14756) @mroeschke
Use _from_data instead of _from_columns for initialzing Frame (#14755) @mroeschke
Consolidate cudf object handling in as_column (#14754) @mroeschke
Reduce execution time of Parquet C++ tests (#14750) @vuule
Implement to_datetime(..., utc=True) (#14749) @mroeschke
Remove usages of rapids-env-update (#14748) @KyleFromNVIDIA
Provide explicit pool size and avoid RMM detail APIs (#14741) @harrism
Implement cudf.MultiIndex.from_arrays (#14740) @mroeschke
Remove unused/single use methods (#14739) @mroeschke
refactor CUDA versions in dependencies.yaml (#14733) @jameslamb
Remove unneeded methods in Column (#14730) @mroeschke
Clean up base column methods (#14725) @mroeschke
Ensure column.fillna signatures are consistent (#14724) @mroeschke
Remove mimesis as a testing dependency (#14723) @mroeschke
Replace as_numerical with as_numerical_column/codes (#14719) @mroeschke
Use offsetalator in gather_chars (#14700) @davidwendt
Use make_strings_children for fill() specialization logic (#14697) @davidwendt
Change io::detail::orc namespace into io::orc::detail (#14696) @ttnghia
Fix call to deprecated factory function (#14695) @davidwendt
Use as_column instead of arange for range like inputs (#14689) @mroeschke
Reorganize ORC reader into multiple files and perform some small fixes to cuIO code (#14665) @ttnghia
Split parquet test into multiple files (#14663) @etseidl
Custom error messages for IO with nonexistent files (#14662) @vuule
Explicitly pass .dtype into is_foo_dtype functions (#14657) @mroeschke
Basic validation in reader benchmarks (#14647) @v...

Contributors

trxcllnt, robertmaynard, and 32 other contributors

Assets 2

12 Feb 21:28

raydouglass

v24.02.00

a48e9fc

v24.02.00

🚨 Breaking Changes

Remove **kwargs from astype (#14765) @mroeschke
Remove mimesis as a testing dependency (#14723) @mroeschke
Update to Dask's shuffle_method kwarg (#14708) @pentschev
Drop Pascal GPU support. (#14630) @bdice
Update to CCCL 2.2.0. (#14576) @bdice
Expunge as_frame conversions in Column algorithms (#14491) @wence-
Deprecate cudf::make_strings_column accepting typed offsets (#14461) @davidwendt
Remove deprecated nvtext::load_merge_pairs_file (#14460) @davidwendt
Include writer code and writerVersion in ORC files (#14458) @vuule
Remove null mask for zero nulls in json readers (#14451) @karthikeyann
REF: Remove **kwargs from to_pandas, raise if nullable is not implemented (#14438) @mroeschke
Consolidate 1D pandas object handling in as_column (#14394) @mroeschke
Move chars column to parent data buffer in strings column (#14202) @karthikeyann
Switch to scikit-build-core (#13531) @vyasr

🐛 Bug Fixes

Exclude tests from builds (#14981) @vyasr
Fix the bounce buffer size in ORC writer (#14947) @vuule
Revert sum/product aggregation to always produce int64_t type (#14907) @SurajAralihalli
Fixed an issue with output chunking computation stemming from input chunking. (#14889) @nvdbaranec
Fix total_byte_size in Parquet row group metadata (#14802) @etseidl
Fix index difference to follow the pandas format (#14789) @amiralimi
Fix shared-workflows repo name (#14784) @raydouglass
Remove unparseable attributes from all nodes (#14780) @vyasr
Refactor and add validation to IntervalIndex.init (#14778) @mroeschke
Work around incompatibilities between V2 page header handling and zStandard compression in Parquet writer (#14772) @etseidl
Fix calls to deprecated strings factory API (#14771) @davidwendt
Fix ptx file discovery in editable installs (#14767) @vyasr
Revise shuffle deprecation to align with dask/dask (#14762) @rjzamora
Enable intermediate proxies to be picklable (#14752) @shwina
Add CUDF_TEST_PROGRAM_MAIN macro to tests lacking it (#14751) @etseidl
Fix CMake args (#14746) @vyasr
Fix logic bug introduced in #14730 (#14742) @wence-
[Java] Choose The Correct RoundingMode For Checking Decimal OutOfBounds (#14731) @razajafri
Fix Groupby.get_group (#14728) @rjzamora
Ensure that all CUDA kernels in cudf have hidden visibility. (#14726) @robertmaynard
Split cuda versions for notebook testing (#14722) @raydouglass
Fix to_numeric not preserving Series index and name (#14718) @mroeschke
Update dask-cudf wheel name (#14713) @raydouglass
Fix strings::contains matching end of string target (#14711) @davidwendt
Update to Dask's shuffle_method kwarg (#14708) @pentschev
Write file-level statistics when writing ORC files with zero rows (#14707) @vuule
Potential fix for peformance regression in #14415 (#14706) @etseidl
Ensure DataFrame column types are preserved during serialization (#14705) @mroeschke
Skip numba test that fails on ARM (#14702) @brandon-b-miller
Allow Z in datetime string parsing in non pandas compat mode (#14701) @mroeschke
Fix nan_as_null not being respected when passing arrow object (#14688) @mroeschke
Fix constructing Series/Index from arrow array and dtype (#14686) @mroeschke
Fix Aggregation Type Promotion: Ensure Unsigned Input Types Result in Unsigned Output for Sum and Multiply (#14679) @SurajAralihalli
Add BaseOffset as a final proxy type to pass instancechecks for offsets against BaseOffset (#14678) @shwina
Add row conversion code from spark-rapids-jni (#14664) @ttnghia
Unconditionally export the CCCL path (#14656) @vyasr
Ensure libcudf searches for our patched version of CCCL first (#14655) @robertmaynard
Constrain CUDA in notebook testing to prevent CUDA 12.1 usage until we have pynvjitlink (#14648) @vyasr
Fix invalid memory access in Parquet reader (#14637) @etseidl
Use column_empty over as_column([]) (#14632) @mroeschke
Add (implicit) handling for torch tensors in is_scalar (#14623) @wence-
Fix astype/fillna not maintaining column subclass and types (#14615) @mroeschke
Remove non-empty nulls in cudf::get_json_object (#14609) @davidwendt
Remove cuda::proclaim_return_type from nested lambda (#14607) @ttnghia
Fix DataFrame.reindex when column reindexing to MultiIndex/RangeIndex (#14605) @mroeschke
Address potential race conditions in Parquet reader (#14602) @etseidl
Fix DataFrame.reindex removing column name (#14601) @mroeschke
Remove unsanitized input test data from copy gtests (#14600) @davidwendt
Fix race detected in Parquet writer (#14598) @etseidl
Correct invalid or missing return types (#14587) @robertmaynard
Fix unsanitized nulls from strings segmented-reduce (#14586) @davidwendt
Upgrade to nvCOMP 3.0.5 (#14581) @davidwendt
Fix unsanitized nulls produced by cudf::clamp APIs (#14580) @davidwendt
Fix unsanitized nulls produced by libcudf dictionary decode (#14578) @davidwendt
Fixes a symbol group lookup table issue (#14561) @elstehle
Drop llvm16 from cuda118-conda devcontainer image (#14526) @charlesbluca
REF: Make DataFrame.from_pandas process by column (#14483) @mroeschke
Improve memory footprint of isin by using contains (#14478) @wence-
Move creation of env.yaml outside the current directory (#14476) @davidwendt
Enable pd.Timestamp objects to be picklable when cudf.pandas is active (#14474) @shwina
Correct dtype of count aggregations on empty dataframes (#14473) @wence-
Avoid DataFrame conversion in MultiIndex.from_pandas (#14470) @mroeschke
JSON writer: avoid default stream use in string_scalar constructors (#14444) @vuule
Fix default stream use in the CSV reader (#14443) @vuule
Preserve DataFrame(columns=).columns dtype during empty-like construction (#14381) @mroeschke
Defer PTX file load to runtime (#13690) @brandon-b-miller

📖 Documentation

Disable parallel build (#14796) @vyasr
Add pylibcudf to the docs (#14791) @vyasr
Describe unpickling expectations when cudf.pandas is enabled (#14693) @shwina
Update CONTRIBUTING for pyproject-only builds (#14653) @vyasr
More doxygen fixes (#14639) @vyasr
Enable doxygen XML generation and fix issues (#14477) @vyasr
Some doxygen improvements (#14469) @vyasr
Remove warning in dask-cudf docs (#14454) @wence-
Update README links with redirects. (#14378) @bdice
Add pip install instructions to README (#13677) @shwina

🚀 New Features

Add ci check for external kernels (#14768) @robertmaynard
JSON single quote normalization API (#14729) @shrshi
Write cuDF version in Parquet "created_by" metadata field (#14721) @etseidl
Implement remaining copying APIs in pylibcudf along with required helper functions (#14640) @vyasr
Don't constrain numba<0.58 (#14616) @brandon-b-miller
Add DELTA_LENGTH_BYTE_ARRAY encoder and decoder for Parquet (#14590) @etseidl
JSON - Parse mixed types as string in JSON reader (#14572) @karthikeyann
JSON quote normalization (#14545) @shrshi
Make DefaultHostMemoryAllocator settable (#14523) @gerashegalov
Implement more copying APIs in pylibcudf (#14508) @vyasr
Include writer code and writerVersion in ORC files (#14458) @vuule
Parquet sub-rowgroup reading. (#14360) @nvdbaranec
Move chars column to parent data buffer in strings column (#14202) @karthikeyann
PARQUET-2261 Size Statistics (#14000) @etseidl
Improve GroupBy JIT error handling (#13854) @brandon-b-miller
Generate unified Python/C++ docs (#13846) @vyasr
Expand JIT groupby test suite (#13813) @brandon-b-miller

🛠️ Improvements

Pin pytest<8 (#14920) @galipremsagar
Move cudf::char_utf8 definition from detail to public header (#14779) @davidwendt
Clean up TimedeltaIndex.__init__ constructor (#14775) @mroeschke
Clean up DatetimeIndex.__init__ constructor (#14774) @mroeschke
Some frame.py typing, move seldom used methods in frame.py (#14766) @mroeschke
Remove **kwargs from astype (#14765) @mroeschke
fix benchmarks compatibility with newer pytest-cases (#14764) @jameslamb
Add pynvjitlink as a dependency (#14763) @brandon-b-miller
Resolve degenerate performance in create_structs_data (#14761) @SurajAralihalli
Simplify ColumnAccessor methods; avoid unnecessary validations (#14758) @mroeschke
Pin pytest-cases<3.8.2 (#14756) @mroeschke
Use _from_data instead of _from_columns for initialzing Frame (#14755) @mroeschke
Consolidate cudf object handling in as_column (#14754) @mroeschke
Reduce execution time of Parquet C++ tests (#14750) @vuule
Implement to_datetime(..., utc=True) (#14749) @mroeschke
Remove usages of rapids-env-update (#14748) @KyleFromNVIDIA
Provide explicit pool size and avoid RMM detail APIs (#14741) @harrism
Implement cudf.MultiIndex.from_arrays (#14740) @mroeschke
Remove unused/single use methods (#14739) @mroeschke
refactor CUDA versions in dependencies.yaml (#14733) @jameslamb
Remove unneeded methods in Column (#14730) @mroeschke
Clean up base column methods (#14725) @mroeschke
Ensure column.fillna signatures are consistent (#14724) @mroeschke
Remove mimesis as a testing dependency (#14723) @mroeschke
Replace as_numerical with as_numerical_column/codes (#14719) @mroeschke
Use offsetalator in gather_chars (#14700) @davidwendt
Use make_strings_children for fill() specialization logic (#14697) @davidwendt
Change io::detail::orc namespace into io::orc::detail (#14696) @ttnghia
Fix call to deprecated factory function (#14695) @davidwendt
Use as_column instead of arange for range like inputs (#14689) @mroeschke
Reorganize ORC reader into multiple files and perform some small fixes to cuIO code (#14665) @ttnghia
Split parquet test into multiple files (#14663) @etseidl
Custom error messages for IO with nonexistent files (#14662) @vuule
Explicitly pass .dtype into is_foo_dtype functions (#14657) @mroeschke
Basic validation in reader benchmarks (#14647) @vuule
Update dependencies.yaml to support CUDA 12.*....

Contributors

trxcllnt, robertmaynard, and 32 other contributors

Assets 2

08 Dec 21:31

raydouglass

v23.12.01

2ce4621

v23.12.01

🚨 Breaking Changes

Raise error in reindex when index is not unique (#14400) @galipremsagar
Expose stream parameter to get_json_object API (#14297) @davidwendt
Refactor cudf_kafka to use skbuild (#14292) @jdye64
Expose stream parameter in public strings convert APIs (#14255) @davidwendt
Upgrade to nvCOMP 3.0.4 (#13815) @vuule

🐛 Bug Fixes

Fix synchronization issue when writing string columns with dictionary to ORC (#14595) @vuule
Update actions/labeler to v4 (#14562) @raydouglass
Fix data corruption when skipping rows (#14557) @etseidl
Fix function name typo in cudf.pandas profiler (#14514) @galipremsagar
Fix intermediate type checking in expression parsing (#14445) @vyasr
Forward merge branch-23.10 into branch-23.12 (#14435) @raydouglass
Remove needs: wheel-build-cudf. (#14427) @bdice
Fix dask dependency in custreamz (#14420) @vyasr
Ensure nvbench initializes nvml context when built statically (#14411) @robertmaynard
Support java AST String literal with desired encoding (#14402) @winningsix
Raise error in reindex when index is not unique (#14400) @galipremsagar
Always build nvbench statically so we don't need to package it (#14399) @robertmaynard
Fix token-count logic in nvtext::tokenize_with_vocabulary (#14393) @davidwendt
Fix as_column(pd.Timestamp/Timedelta, length=) not respecting length (#14390) @mroeschke
cudf.pandas: cuDF subpath checking in module __getattr__ (#14388) @shwina
Fix and disable encoding for nanosecond statistics in ORC writer (#14367) @vuule
Add the new manylinux builds to the build job (#14351) @vyasr
cudf jit parser now supports .pragma instructions with quotes (#14348) @robertmaynard
Fix overflow check in cudf::merge (#14345) @divyegala
Add cramjam (#14344) @vyasr
Enable dask_cudf/io pytests in CI (#14338) @galipremsagar
Temporarily avoid the current build of pydata-sphinx-theme (#14332) @vyasr
Fix host buffer access from device function in the Parquet reader (#14328) @vuule
Run IO tests for Dask-cuDF (#14327) @rjzamora
Fix logical type issues in the Parquet writer (#14322) @vuule
Remove aws-sdk-pinning and revert to arrow 12.0.1 (#14319) @vyasr
test is_valid before reading column data (#14318) @etseidl
Fix gtest validity setting for TextTokenizeTest.Vocabulary (#14312) @davidwendt
Fixes stack context for json lines format that recovers from invalid JSON lines (#14309) @elstehle
Downgrade to Arrow 12.0.0 for aws-sdk-cpp and fix cudf_kafka builds for new CI containers (#14296) @vyasr
fixing thread index overflow issue (#14290) @hyperbolic2346
Fix memset error in nvtext::edit_distance_matrix (#14283) @davidwendt
Changes JSON reader's recovery option's behaviour to ignore all characters after a valid JSON record (#14279) @elstehle
Handle empty string correctly in Parquet statistics (#14257) @etseidl
Fixes behaviour for incomplete lines when recover_with_nulls is enabled (#14252) @elstehle
cudf::detail::pinned_allocator doesn't throw from deallocate (#14251) @robertmaynard
Fix strings replace for adjacent, identical multi-byte UTF-8 character targets (#14235) @davidwendt
Fix the precision when converting a decimal128 column to an arrow array (#14230) @jihoonson
Fixing parquet list of struct interpretation (#13715) @hyperbolic2346

📖 Documentation

Fix io reference in docs. (#14452) @bdice
Update README (#14374) @shwina
Example code for blog on new row comparators (#13795) @divyegala

🚀 New Features

Expose streams in public unary APIs (#14342) @vyasr
Add python tests for Parquet DELTA_BINARY_PACKED encoder (#14316) @etseidl
Update rapids-cmake functions to non-deprecated signatures (#14265) @robertmaynard
Expose streams in public null mask APIs (#14263) @vyasr
Expose streams in binaryop APIs (#14187) @vyasr
Add pylibcudf.Scalar that interoperates with Arrow scalars (#14133) @vyasr
Add decoder for DELTA_BYTE_ARRAY to Parquet reader (#14101) @etseidl
Add DELTA_BINARY_PACKED encoder for Parquet writer (#14100) @etseidl
Add BytePairEncoder class to cuDF (#13891) @davidwendt
Upgrade to nvCOMP 3.0.4 (#13815) @vuule
Use pynvjitlink for CUDA 12+ MVC (#13650) @brandon-b-miller

🛠️ Improvements

Build concurrency for nightly and merge triggers (#14441) @bdice
Cleanup remaining usages of dask dependencies (#14407) @galipremsagar
Update to Arrow 14.0.1. (#14387) @bdice
Remove Cython libcpp wrappers (#14382) @vyasr
Forward-merge branch-23.10 to branch-23.12 (#14372) @bdice
Upgrade to arrow 14 (#14371) @galipremsagar
Fix a pytest typo in test_kurt_skew_error (#14368) @galipremsagar
Use new rapids-dask-dependency metapackage for managing dask versions (#14364) @vyasr
Change nullable() to has_nulls() in cudf::detail::gather (#14363) @divyegala
Split up scan_inclusive.cu to improve its compile time (#14358) @davidwendt
Implement user_datasource_wrapper is_empty() and is_device_read_preferred(). (#14357) @tpn
Added streams to CSV reader and writer api (#14340) @shrshi
Upgrade wheels to use arrow 13 (#14339) @vyasr
Rework nvtext::byte_pair_encoding API (#14337) @davidwendt
Improve performance of nvtext::tokenize_with_vocabulary for long strings (#14336) @davidwendt
Upgrade arrow to 13 (#14330) @galipremsagar
Expose stream parameter in public nvtext replace APIs (#14329) @davidwendt
Drop pyorc dependency and use pandas/pyarrow instead (#14323) @galipremsagar
Avoid pyarrow.fs import for local storage (#14321) @rjzamora
Unpin dask and distributed for 23.12 development (#14320) @galipremsagar
Expose stream parameter in public nvtext tokenize APIs (#14317) @davidwendt
Added streams to JSON reader and writer api (#14313) @shrshi
Minor improvements in source_info (#14308) @vuule
Forward-merge branch-23.10 to branch-23.12 (#14307) @bdice
Add stream parameter to Set Operations (Public List APIs) (#14305) @SurajAralihalli
Expose stream parameter to get_json_object API (#14297) @davidwendt
Sort dictionary data alphabetically in the ORC writer (#14295) @vuule
Expose stream parameter in public strings filter APIs (#14293) @davidwendt
Refactor cudf_kafka to use skbuild (#14292) @jdye64
Update shared-action-workflows references (#14289) @AyodeAwe
Register partd encode dispatch in dask_cudf (#14287) @rjzamora
Update versioning strategy (#14285) @vyasr
Move and rename byte-pair-encoding source files (#14284) @davidwendt
Expose stream parameter in public strings combine APIs (#14281) @davidwendt
Expose stream parameter in public strings contains APIs (#14280) @davidwendt
Add stream parameter to List Sort and Filter APIs (#14272) @SurajAralihalli
Use branch-23.12 workflows. (#14271) @bdice
Refactor LogicalType for Parquet (#14264) @etseidl
Centralize chunked reading code in the parquet reader to reader_impl_chunking.cu (#14262) @nvdbaranec
Expose stream parameter in public strings replace APIs (#14261) @davidwendt
Expose stream parameter in public strings APIs (#14260) @davidwendt
Cleanup of namespaces in parquet code. (#14259) @nvdbaranec
Make parquet schema index type consistent (#14256) @hyperbolic2346
Expose stream parameter in public strings convert APIs (#14255) @davidwendt
Add in java bindings for DataSource (#14254) @revans2
Reimplement cudf::merge for nested types without using comparators (#14250) @divyegala
Add stream parameter to List Manipulation and Operations APIs (#14248) @SurajAralihalli
Expose stream parameter in public strings split/partition APIs (#14247) @davidwendt
Improve contains_column by invoking contains_table (#14238) @PointKernel
Detect and report errors in Parquet header parsing (#14237) @etseidl
Normalizing offsets iterator (#14234) @davidwendt
Forward merge 23.10 into 23.12 (#14231) @galipremsagar
Return error if BOOL8 column-type is used with integers-to-hex (#14208) @davidwendt
Enable indexalator for device code (#14206) @davidwendt
Marginally reduce memory footprint of joins (#14197) @wence-
Add nvtx annotations to spilling-based data movement (#14196) @wence-
Optimize ORC writer for decimal columns (#14190) @vuule
Remove the use of volatile in ORC (#14175) @vuule
Add bytes_per_second to distinct_count of stream_compaction nvbench. (#14172) @Blonck
Add bytes_per_second to transpose benchmark (#14170) @Blonck
cuDF: Build CUDA 12.0 ARM conda packages. (#14112) @bdice
Add bytes_per_second to shift benchmark (#13950) @Blonck
Extract debug_utilities.hpp/cu from column_utilities.hpp/cu (#13720) @ttnghia

Contributors

robertmaynard, tpn, and 26 other contributors

Assets 2

06 Dec 16:26

raydouglass

v23.12.00

c1d3073

v23.12.00

🚨 Breaking Changes

Raise error in reindex when index is not unique (#14400) @galipremsagar
Expose stream parameter to get_json_object API (#14297) @davidwendt
Refactor cudf_kafka to use skbuild (#14292) @jdye64
Expose stream parameter in public strings convert APIs (#14255) @davidwendt
Upgrade to nvCOMP 3.0.4 (#13815) @vuule

🐛 Bug Fixes

Update actions/labeler to v4 (#14562) @raydouglass
Fix data corruption when skipping rows (#14557) @etseidl
Fix function name typo in cudf.pandas profiler (#14514) @galipremsagar
Fix intermediate type checking in expression parsing (#14445) @vyasr
Forward merge branch-23.10 into branch-23.12 (#14435) @raydouglass
Remove needs: wheel-build-cudf. (#14427) @bdice
Fix dask dependency in custreamz (#14420) @vyasr
Ensure nvbench initializes nvml context when built statically (#14411) @robertmaynard
Support java AST String literal with desired encoding (#14402) @winningsix
Raise error in reindex when index is not unique (#14400) @galipremsagar
Always build nvbench statically so we don't need to package it (#14399) @robertmaynard
Fix token-count logic in nvtext::tokenize_with_vocabulary (#14393) @davidwendt
Fix as_column(pd.Timestamp/Timedelta, length=) not respecting length (#14390) @mroeschke
cudf.pandas: cuDF subpath checking in module __getattr__ (#14388) @shwina
Fix and disable encoding for nanosecond statistics in ORC writer (#14367) @vuule
Add the new manylinux builds to the build job (#14351) @vyasr
cudf jit parser now supports .pragma instructions with quotes (#14348) @robertmaynard
Fix overflow check in cudf::merge (#14345) @divyegala
Add cramjam (#14344) @vyasr
Enable dask_cudf/io pytests in CI (#14338) @galipremsagar
Temporarily avoid the current build of pydata-sphinx-theme (#14332) @vyasr
Fix host buffer access from device function in the Parquet reader (#14328) @vuule
Run IO tests for Dask-cuDF (#14327) @rjzamora
Fix logical type issues in the Parquet writer (#14322) @vuule
Remove aws-sdk-pinning and revert to arrow 12.0.1 (#14319) @vyasr
test is_valid before reading column data (#14318) @etseidl
Fix gtest validity setting for TextTokenizeTest.Vocabulary (#14312) @davidwendt
Fixes stack context for json lines format that recovers from invalid JSON lines (#14309) @elstehle
Downgrade to Arrow 12.0.0 for aws-sdk-cpp and fix cudf_kafka builds for new CI containers (#14296) @vyasr
fixing thread index overflow issue (#14290) @hyperbolic2346
Fix memset error in nvtext::edit_distance_matrix (#14283) @davidwendt
Changes JSON reader's recovery option's behaviour to ignore all characters after a valid JSON record (#14279) @elstehle
Handle empty string correctly in Parquet statistics (#14257) @etseidl
Fixes behaviour for incomplete lines when recover_with_nulls is enabled (#14252) @elstehle
cudf::detail::pinned_allocator doesn't throw from deallocate (#14251) @robertmaynard
Fix strings replace for adjacent, identical multi-byte UTF-8 character targets (#14235) @davidwendt
Fix the precision when converting a decimal128 column to an arrow array (#14230) @jihoonson
Fixing parquet list of struct interpretation (#13715) @hyperbolic2346

📖 Documentation

Fix io reference in docs. (#14452) @bdice
Update README (#14374) @shwina
Example code for blog on new row comparators (#13795) @divyegala

🚀 New Features

Expose streams in public unary APIs (#14342) @vyasr
Add python tests for Parquet DELTA_BINARY_PACKED encoder (#14316) @etseidl
Update rapids-cmake functions to non-deprecated signatures (#14265) @robertmaynard
Expose streams in public null mask APIs (#14263) @vyasr
Expose streams in binaryop APIs (#14187) @vyasr
Add pylibcudf.Scalar that interoperates with Arrow scalars (#14133) @vyasr
Add decoder for DELTA_BYTE_ARRAY to Parquet reader (#14101) @etseidl
Add DELTA_BINARY_PACKED encoder for Parquet writer (#14100) @etseidl
Add BytePairEncoder class to cuDF (#13891) @davidwendt
Upgrade to nvCOMP 3.0.4 (#13815) @vuule
Use pynvjitlink for CUDA 12+ MVC (#13650) @brandon-b-miller

🛠️ Improvements

Build concurrency for nightly and merge triggers (#14441) @bdice
Cleanup remaining usages of dask dependencies (#14407) @galipremsagar
Update to Arrow 14.0.1. (#14387) @bdice
Remove Cython libcpp wrappers (#14382) @vyasr
Forward-merge branch-23.10 to branch-23.12 (#14372) @bdice
Upgrade to arrow 14 (#14371) @galipremsagar
Fix a pytest typo in test_kurt_skew_error (#14368) @galipremsagar
Use new rapids-dask-dependency metapackage for managing dask versions (#14364) @vyasr
Change nullable() to has_nulls() in cudf::detail::gather (#14363) @divyegala
Split up scan_inclusive.cu to improve its compile time (#14358) @davidwendt
Implement user_datasource_wrapper is_empty() and is_device_read_preferred(). (#14357) @tpn
Added streams to CSV reader and writer api (#14340) @shrshi
Upgrade wheels to use arrow 13 (#14339) @vyasr
Rework nvtext::byte_pair_encoding API (#14337) @davidwendt
Improve performance of nvtext::tokenize_with_vocabulary for long strings (#14336) @davidwendt
Upgrade arrow to 13 (#14330) @galipremsagar
Expose stream parameter in public nvtext replace APIs (#14329) @davidwendt
Drop pyorc dependency and use pandas/pyarrow instead (#14323) @galipremsagar
Avoid pyarrow.fs import for local storage (#14321) @rjzamora
Unpin dask and distributed for 23.12 development (#14320) @galipremsagar
Expose stream parameter in public nvtext tokenize APIs (#14317) @davidwendt
Added streams to JSON reader and writer api (#14313) @shrshi
Minor improvements in source_info (#14308) @vuule
Forward-merge branch-23.10 to branch-23.12 (#14307) @bdice
Add stream parameter to Set Operations (Public List APIs) (#14305) @SurajAralihalli
Expose stream parameter to get_json_object API (#14297) @davidwendt
Sort dictionary data alphabetically in the ORC writer (#14295) @vuule
Expose stream parameter in public strings filter APIs (#14293) @davidwendt
Refactor cudf_kafka to use skbuild (#14292) @jdye64
Update shared-action-workflows references (#14289) @AyodeAwe
Register partd encode dispatch in dask_cudf (#14287) @rjzamora
Update versioning strategy (#14285) @vyasr
Move and rename byte-pair-encoding source files (#14284) @davidwendt
Expose stream parameter in public strings combine APIs (#14281) @davidwendt
Expose stream parameter in public strings contains APIs (#14280) @davidwendt
Add stream parameter to List Sort and Filter APIs (#14272) @SurajAralihalli
Use branch-23.12 workflows. (#14271) @bdice
Refactor LogicalType for Parquet (#14264) @etseidl
Centralize chunked reading code in the parquet reader to reader_impl_chunking.cu (#14262) @nvdbaranec
Expose stream parameter in public strings replace APIs (#14261) @davidwendt
Expose stream parameter in public strings APIs (#14260) @davidwendt
Cleanup of namespaces in parquet code. (#14259) @nvdbaranec
Make parquet schema index type consistent (#14256) @hyperbolic2346
Expose stream parameter in public strings convert APIs (#14255) @davidwendt
Add in java bindings for DataSource (#14254) @revans2
Reimplement cudf::merge for nested types without using comparators (#14250) @divyegala
Add stream parameter to List Manipulation and Operations APIs (#14248) @SurajAralihalli
Expose stream parameter in public strings split/partition APIs (#14247) @davidwendt
Improve contains_column by invoking contains_table (#14238) @PointKernel
Detect and report errors in Parquet header parsing (#14237) @etseidl
Normalizing offsets iterator (#14234) @davidwendt
Forward merge 23.10 into 23.12 (#14231) @galipremsagar
Return error if BOOL8 column-type is used with integers-to-hex (#14208) @davidwendt
Enable indexalator for device code (#14206) @davidwendt
Marginally reduce memory footprint of joins (#14197) @wence-
Add nvtx annotations to spilling-based data movement (#14196) @wence-
Optimize ORC writer for decimal columns (#14190) @vuule
Remove the use of volatile in ORC (#14175) @vuule
Add bytes_per_second to distinct_count of stream_compaction nvbench. (#14172) @Blonck
Add bytes_per_second to transpose benchmark (#14170) @Blonck
cuDF: Build CUDA 12.0 ARM conda packages. (#14112) @bdice
Add bytes_per_second to shift benchmark (#13950) @Blonck
Extract debug_utilities.hpp/cu from column_utilities.hpp/cu (#13720) @ttnghia

Contributors

robertmaynard, tpn, and 26 other contributors

Assets 2

Releases: rapidsai/cudf

v24.06.01

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

Contributors

v24.06.00

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

Contributors

[NIGHTLY] v24.08.00

🔗 Links

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

🛠️ Improvements

Contributors

v24.04.01

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

🛠️ Improvements

Contributors

v24.04.00

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

🛠️ Improvements

Contributors

v24.02.02

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

🛠️ Improvements

Contributors

v24.02.01

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

🛠️ Improvements

Contributors

v24.02.00

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

🛠️ Improvements

Contributors

v23.12.01

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

🛠️ Improvements

Contributors

v23.12.00

🚨 Breaking Changes

🐛 Bug Fixes

📖 Documentation

🚀 New Features

🛠️ Improvements

Contributors