MCPcopy Create free account
hub / github.com/dask/dask / shuffle

Method shuffle

dask/dataframe/dask_expr/_collection.py:842–952  ·  view source on GitHub ↗

Rearrange DataFrame into new partitions Uses hashing of `on` to map rows to output partitions. After this operation, rows with the same value of `on` will be in the same partition. Parameters ---------- on : str, list of str, or Series, Index, or Dat

(
        self,
        on: str | list | no_default = no_default,  # type: ignore[valid-type]
        ignore_index: bool = False,
        npartitions: int | None = None,
        shuffle_method: str | None = None,
        on_index: bool = False,
        force: bool = False,
        **options,
    )

Source from the content-addressed store, hash-verified

840 return self.partitions[n]
841
842 def shuffle(
843 self,
844 on: str | list | no_default = no_default, # type: ignore[valid-type]
845 ignore_index: bool = False,
846 npartitions: int | None = None,
847 shuffle_method: str | None = None,
848 on_index: bool = False,
849 force: bool = False,
850 **options,
851 ):
852 """Rearrange DataFrame into new partitions
853
854 Uses hashing of `on` to map rows to output partitions. After this
855 operation, rows with the same value of `on` will be in the same
856 partition.
857
858 Parameters
859 ----------
860 on : str, list of str, or Series, Index, or DataFrame
861 Column names to shuffle by.
862 ignore_index : optional
863 Whether to ignore the index. Default is ``False``.
864 npartitions : optional
865 Number of output partitions. The partition count will
866 be preserved by default.
867 shuffle_method : optional
868 Desired shuffle method. Default chosen at optimization time.
869 on_index : bool, default False
870 Whether to shuffle on the index. Mutually exclusive with 'on'.
871 Set this to ``True`` if 'on' is not provided.
872 force : bool, default False
873 This forces the optimizer to keep the shuffle even if the final
874 expression could be further simplified.
875 **options : optional
876 Algorithm-specific options.
877
878 Notes
879 -----
880 This does not preserve a meaningful index/partitioning scheme. This
881 is not deterministic if done in parallel.
882
883 Examples
884 --------
885 >>> df = df.shuffle(df.columns[0]) # doctest: +SKIP
886 """
887 if on is no_default and not on_index: # type: ignore[unreachable]
888 raise TypeError(
889 "Must shuffle on either columns or the index; currently shuffling on "
890 "neither. Pass column(s) to 'on' or set 'on_index' to True."
891 )
892 elif on is not no_default and on_index:
893 raise TypeError(
894 "Cannot shuffle on both columns and the index. Do not pass column(s) "
895 "to 'on' or set 'on_index' to False."
896 )
897
898 # Preserve partition count by default
899 npartitions = npartitions or self.npartitions

Calls 7

is_dask_collectionFunction · 0.90
new_collectionFunction · 0.90
RearrangeByColumnClass · 0.90
get_specified_shuffleFunction · 0.90
anyFunction · 0.85
map_partitionsMethod · 0.45