MCPcopy Create free account

hub / github.com/deepspeedai/DeepSpeed / functions

Functions11,258 in github.com/deepspeedai/DeepSpeed

↓ 2 callersMethod_set_zero_group_parallelism
(self)
deepspeed/runtime/zero/stage3.py:671
↓ 2 callersFunction_shape_numel
(shape)
deepspeed/checkpoint/ds_to_universal.py:454
↓ 2 callersFunction_shape_prod
(values)
deepspeed/module_inject/layers.py:1086
↓ 2 callersFunction_shared_params
(engine)
tests/unit/v1/moe/test_autoep_checkpoint.py:142
↓ 2 callersFunction_single_line
Render untrusted repository text without controls or extra log lines.
ci/tests_fetcher.py:622
↓ 2 callersFunction_source_param_ndim
(param: torch.Tensor | nn.Parameter)
deepspeed/module_inject/auto_ep.py:91
↓ 2 callersFunction_sparse_decode
Gather each token's expert results back and weight them by the gate values.
deepspeed/moe/sharded_moe.py:223
↓ 2 callersFunction_sparse_encode
Place every routed token in its expert slot, returning the [e, c, m] buffer. Inverting the routing map first turns the copy into a single gather.
deepspeed/moe/sharded_moe.py:207
↓ 2 callersMethod_splice_visual_into_text
Replace image placeholder positions in *text_embeds* with *visual_embeds*. This is intentionally architecture-specific. The default raises
deepspeed/sequence/autosp_fusion.py:155
↓ 2 callersFunction_split
Split the tensor along its last dimension and keep the corresponding slice.
deepspeed/compression/basic_layer.py:655
↓ 2 callersFunction_split_affinity
Split a rank's core list into (zf_affinity, pt_affinity): reserve the first ceil(pt_reserved_cores_perc * n) cores for the training thread and giv
deepspeed/runtime/zenflow/zenflow_utils.py:73
↓ 2 callersFunction_split_plan_from_expert_counts
Build a SplitPlan from a global per-expert token histogram (ep_size > 1). A single all-to-all exchanges the whole ``[ep_size, E_local]`` count ma
deepspeed/module_inject/auto_ep_layer.py:95
↓ 2 callersFunction_split_plan_received
Rows of the matrix rank ``ep_rank`` receives from the counts all-to-all.
tests/unit/v1/moe/test_autoep_unit.py:914
↓ 2 callersFunction_storage_id
(tensor)
tests/unit/runtime/zero/test_zero_linear_direct_init.py:28
↓ 2 callersMethod_supported_optims
(self)
deepspeed/runtime/engine.py:1608
↓ 2 callersFunction_supported_sequence_length
(config: Any)
deepspeed/inference/v2/model_implementations/exaone4_5/model.py:72
↓ 2 callersMethod_swap_in_if_offloaded
Read an NVMe-offloaded partition back into memory. The swapper leaves ``ds_tensor.data`` as a 0-dim placeholder while a partition is on NVMe.
deepspeed/ops/adam/zenflow_torch_adam.py:413
↓ 2 callersMethod_swap_out_gradients
(self, parameter, gradient_offsets, gradient_tensors, gradient_swapper)
deepspeed/runtime/swap_tensor/optimizer_utils.py:239
↓ 2 callersMethod_swap_out_optimizer_state
(self, swap_info)
deepspeed/runtime/swap_tensor/partitioned_optimizer_swapper.py:100
↓ 2 callersMethod_swap_out_unpinned_tensors
(self, aio_handle, unpinned_tensors, dest_paths, pinned_buffers)
deepspeed/runtime/swap_tensor/optimizer_utils.py:412
↓ 2 callersMethod_sync_selective_optimizer_lr
(self)
deepspeed/runtime/zenflow/zenflow_stage_1_and_2.py:571
↓ 2 callersMethod_synchronize_tied_weights
(self)
deepspeed/runtime/pipe/module.py:472
↓ 2 callersMethod_take_model_step
(self, lr_kwargs, block_eigenvalue={})
deepspeed/runtime/engine.py:3271
↓ 2 callersFunction_tensor_digest_words
(tensor: torch.Tensor)
deepspeed/moe/ep_tp_dispatch.py:89
↓ 2 callersFunction_test_activation_checkpoint_ordering
(module, expected_ordering, *inputs)
tests/unit/runtime/activation_checkpointing/test_activation_checkpointing.py:87
↓ 2 callersFunction_test_single_mapping_helper
(n_tokens: int, n_experts: int, assigned_exper
tests/unit/inference/v2/kernels/ragged_ops/test_top_k_gating.py:60
↓ 2 callersFunction_to
(v, device)
deepspeed/compile/profilers/graph_profile.py:30
↓ 2 callersFunction_train_moe
(config, hidden_dim, num_chunks, num_steps, use_no_sync, ep_size=1, seed=42)
tests/unit/v1/zero/test_zero_coalesce_grad_reduction.py:276
↓ 2 callersFunction_triton_attention
(qkv, input_mask, layer_past, alibi,
deepspeed/ops/transformer/inference/triton/attention.py:222
↓ 2 callersFunction_triton_packed_flash
(qkv, head_size, mask, sm_scale, causal=False, add_mask=True)
deepspeed/ops/transformer/inference/triton/attention.py:357
↓ 2 callersMethod_unassign_params
(self, tensor_id)
deepspeed/runtime/zero/contiguous_memory_allocator.py:143
↓ 2 callersMethod_under_sources
(self, rel_posix: str)
ci/tests_fetcher.py:314
↓ 2 callersMethod_unfuse_lora_layer
(self, layer_id)
deepspeed/runtime/hybrid_engine.py:139
↓ 2 callersMethod_update_inflight_swap_in
(self, params, swap_in_buffers, inflight_numel)
deepspeed/runtime/swap_tensor/partitioned_param_swapper.py:280
↓ 2 callersMethod_update_scale
(self, has_overflow=False)
deepspeed/runtime/zero/stage3.py:3008
↓ 2 callersMethod_update_scale
(self, skip)
deepspeed/runtime/fp16/fused_optimizer.py:397
↓ 2 callersMethod_update_scale
(self, skip)
deepspeed/runtime/fp16/unfused_optimizer.py:272
↓ 2 callersMethod_valid_stage
(self, stage_id)
deepspeed/runtime/pipe/schedule.py:83
↓ 2 callersFunction_validate_accelerator
(accel_obj)
accelerator/real_accelerator.py:28
↓ 2 callersFunction_validate_config_values
(config_name, config_dict, valid_values)
deepspeed/runtime/model_checkpointing/config.py:22
↓ 2 callersMethod_validate_indices
(self, pp_index, tp_index)
deepspeed/checkpoint/reshape_meg_2d.py:48
↓ 2 callersFunction_validate_tensor_cast_properties
(typed_tensor, byte_tensor)
tests/unit/utils/test_byte_cast.py:17
↓ 2 callersFunction_validate_zero3_model_optim_rank_sets
(model_files, optim_files)
deepspeed/checkpoint/ds_to_universal.py:560
↓ 2 callersFunction_validate_zero3_partitioned_autoep_metadata
(autoep_metadata, require_partitioned=True)
deepspeed/checkpoint/ds_to_universal.py:511
↓ 2 callersFunction_warmup_mmap_file
(path)
deepspeed/runtime/data_pipeline/data_sampling/indexed_dataset.py:331
↓ 2 callersMethod_write
(self, num_bytes)
deepspeed/io/mock_file_writer.py:46
↓ 2 callersMethod_write_from_tensor
(self, buffer_tensor)
deepspeed/io/fast_file_writer.py:188
↓ 2 callersFunction_write_tp_states
(base_dir, param_name, tp_idx, fp32_tensor)
tests/unit/checkpoint/test_autotp_uc_checkpoint.py:216
↓ 2 callersFunction_zero0_baseline_config
()
tests/unit/v1/moe/test_autoep_autotp_grad_parity.py:252
↓ 2 callersFunction_zero2_baseline_config
()
tests/unit/v1/moe/test_autoep_autotp_grad_parity.py:238
↓ 2 callersFunction_zero3_config
(dtype)
tests/unit/runtime/zero/test_zero_late_module_attach.py:34
↓ 2 callersMethod_zero3_consolidated_16bit_state_dict
Get a full non-partitioned state_dict with fp16 weights on cpu. Important: this function must be called on all ranks and not just ran
deepspeed/runtime/engine.py:5410
↓ 2 callersFunction_zero3_rank_from_file
(path)
deepspeed/checkpoint/ds_to_universal.py:460
↓ 2 callersMethod_zero_init_param
(self, param)
deepspeed/runtime/zero/partition_parameters.py:1169
↓ 2 callersFunction_zero_params_with_collective_stubs
(*ds_ids)
tests/unit/v1/compile/test_z3_eager_fallback.py:498
↓ 2 callersFunction_zero_partitioned_param_info
(unpartitioned_numel, world_size)
deepspeed/checkpoint/ds_to_universal.py:447
↓ 2 callersMethodabsolute_name
(self)
op_builder/gds.py:17
↓ 2 callersMethodabsolute_name
(self)
op_builder/xpu/async_io.py:19
↓ 2 callersMethodaccumulate
csrc/aio/common/deepspeed_aio_types.cpp:46
↓ 2 callersFunctionadd_gather_and_release
(graph_id: int, graph: Graph, param_manager, param_nodes: List[Node])
deepspeed/compile/passes/zero3_compile.py:75
↓ 2 callersMethodadd_item_numpy
(self, np_array)
deepspeed/runtime/data_pipeline/data_sampling/indexed_dataset.py:596
↓ 2 callersFunctionadd_pointer_offset
Adds a pointer offset in units of Element
csrc/deepspeed4science/evoformer_attn/iterators/predicated_tile_access_iterator_residual_last.h:289
↓ 2 callersFunctionadd_pre_backward_hook
(hook)
deepspeed/compile/util.py:66
↓ 2 callersFunctionadd_z1_reduce_bw
(gm: GraphModule, graph_id: int, graph_order: List[Tuple[int, bool]], param_manager)
deepspeed/compile/passes/zero1_compile.py:31
↓ 2 callersFunctionadd_z1_reduce_fw
(gm: GraphModule, graph_id: int, profiling_results, param_manager, use_z2=False)
deepspeed/compile/passes/zero1_compile.py:18
↓ 2 callersMethodaio_config
(self)
deepspeed/runtime/engine.py:1424
↓ 2 callersFunctionalign_dense_tensors
(tensor_list, alignment)
deepspeed/runtime/utils.py:991
↓ 2 callersMethodall_gather_object
(self, object_list, obj, group=None)
deepspeed/comm/torch.py:311
↓ 2 callersMethodall_gather_scalar
(self, value, dp_group)
deepspeed/runtime/engine.py:3813
↓ 2 callersMethodall_reduce_coalesced
proxy func to torch.distributed.all_reduce_coalesced, which is included in PyTorch 1.13 and above
deepspeed/comm/torch.py:198
↓ 2 callersFunctionall_reduce_outer_loop
csrc/cpu/comm/shm.cpp:664
↓ 2 callersMethodall_to_all
(self, output_tensor_list, input_tensor_list, group=None, async_op=False)
deepspeed/comm/torch.py:344
↓ 2 callersFunctionall_to_all_single
(output, tensor, output_split_sizes=None, in
deepspeed/comm/comm.py:348
↓ 2 callersFunctionallclose
(x, y)
tests/unit/ops/sparse_attention/test_sparse_attention.py:49
↓ 2 callersMethodallgatherParam
csrc/compile/z3.cpp:119
↓ 2 callersMethodalloc
csrc/aio/py_lib/deepspeed_pin_tensor.cpp:29
↓ 2 callersMethodallocate_blocks
(self, n_blocks: int, cache_group: int = 0)
deepspeed/inference/v2/ragged/ragged_manager.py:205
↓ 2 callersMethodallocate_tensor
(self, size)
deepspeed/runtime/zero/contiguous_memory_allocator.py:51
↓ 2 callersMethodallocate_workspace
(self, size)
deepspeed/model_implementations/transformers/ds_transformer.py:96
↓ 2 callersMethodallreduce_and_copy
(self, small_bucket, dp_group, dp_world_size=None)
deepspeed/runtime/engine.py:3630
↓ 2 callersMethodallreduce_and_scatter
(self, bucket, communication_data_type: torch.dtyp
deepspeed/runtime/zero/stage_1_and_2.py:1255
↓ 2 callersMethodallreduce_bucket
(self, bucket, communication_data_type: torch.dtype,
deepspeed/runtime/zero/stage_1_and_2.py:1771
↓ 2 callersMethodallreduce_no_retain
(self, bucket, dp_group, numel_per_bucket=500000000, dp_world_size=None)
deepspeed/runtime/engine.py:3635
↓ 2 callersMethodallreduce_no_retain
( self, bucket, communication_data_type: torch.dtype, numel_per_bucket=5000000
deepspeed/runtime/zero/stage_1_and_2.py:1848
↓ 2 callersFunctionapply_config_overrides
Apply explicit AutoEP config overrides to a preset. Return the original preset object when there are no overrides. When overrides are present
deepspeed/module_inject/auto_ep_presets/registry.py:119
↓ 2 callersFunctionapply_folding_correction_to_grad_buffer
( folding_spec: ParallelFoldingSpec | None, param, grad, *, tp_group, param_name: str
deepspeed/module_inject/auto_ep_folding.py:474
↓ 2 callersFunctionapply_rotary_pos_emb_backward
(grad_output, freqs_cos, freqs_sin)
deepspeed/sequence/fpdt_layer.py:33
↓ 2 callersFunctionapply_scores_before_experts_if_enabled
Pre-multiply token representations by router scores before expert compute.
deepspeed/module_inject/auto_ep_layer.py:84
↓ 2 callersMethodapply_tensor_parallelism
(self, mp_replace)
deepspeed/module_inject/containers/base.py:226
↓ 2 callersFunctionapply_to_tensors_only
Apply `function` to every Tensor in `value`. Args: functional: The function class to apply. value (Any): Target object to ap
deepspeed/runtime/zero/utils.py:148
↓ 2 callersFunctionassert_all_partitioned
For ZeRO-3, assert every non-persistent param is released after backward. The recompute bug left frozen params gathered after backward, so the re
tests/unit/v1/zero/test_zero_user_backward.py:219
↓ 2 callersFunctionassert_close_for_preferred_dtype
(actual, expected)
tests/unit/model_parallelism/test_autotp_custom_patterns.py:186
↓ 2 callersFunctionassert_group_matches_spec
Ensure cached ``ep_size_N`` rank lists match the requested folding spec.
deepspeed/module_inject/auto_ep_folding.py:496
↓ 2 callersFunctionassert_no_cuda_mismatch
(name="")
op_builder/builder.py:90
↓ 2 callersFunctionassert_persistent_resident
For ZeRO-3, assert every persistent param stays gathered (AVAILABLE) after backward. Persistent frozen params are added to a recompute owner duri
tests/unit/v1/zero/test_zero_user_backward.py:236
↓ 2 callersFunctionassert_tp_payload_consistent
(payload: RoutedAssignmentPayload, *, tp_group, tp_size: int)
deepspeed/moe/ep_tp_dispatch.py:210
↓ 2 callersMethodasync_step
Queue parameter for optimization in the worker process.
deepspeed/runtime/superoffload/superoffload_utils.py:220
↓ 2 callersMethodattention
(self, client_module)
deepspeed/module_inject/containers/vae.py:41
↓ 2 callersMethodattention_quantization
(self)
deepspeed/module_inject/containers/base.py:218
↓ 2 callersFunctionautoep_folding_gradient_reduction_strategy
Classify one folded TP/SP gradient as ``sum``, ``average``, or ``skip``. TP means Tensor Parallel and SP means Sequence Parallel. The parallel mo
deepspeed/module_inject/auto_ep_folding.py:341
← previousnext →1,501–1,600 of 11,258, ranked by callers