Code
Hub
Workspaces
Following
Trending
Connect
MCP
copy
Create free account
hub
/
github.com/deepspeedai/DeepSpeed
/ functions
Functions
11,258 in github.com/deepspeedai/DeepSpeed
⨍
Functions
11,258
◇
Types & classes
1,887
↳
Endpoints
34
↓ 2 callers
Function
_autoep_modules
(engine)
tests/unit/v1/moe/test_autoep_checkpoint.py:122
↓ 2 callers
Method
_autoep_sequence_parallel_world_size
(self)
deepspeed/runtime/engine.py:617
↓ 2 callers
Method
_autoep_zero_optimizer_param_families
(self)
deepspeed/runtime/engine.py:3874
↓ 2 callers
Function
_backfill_missing_profile_metadata
(graph: Graph, profile_complete: bool = True)
deepspeed/compile/profilers/graph_profile.py:86
↓ 2 callers
Method
_backup_cpuinfo
(self)
op_builder/builder.py:438
↓ 2 callers
Function
_bf16_optimizer_stub
(lp, hp_grad)
tests/unit/v1/moe/test_autoep_autotp_bf16_folding_parity.py:15
↓ 2 callers
Function
_blas_linear_helper
(tokens: int, in_channels: int, out_channels: int,
tests/unit/inference/v2/modules/test_blas_linear_module.py:49
↓ 2 callers
Function
_bloom_type_transpose
(input, mp_size)
deepspeed/module_inject/fusedqkv_utils.py:89
↓ 2 callers
Method
_broadcast_model
(self)
deepspeed/runtime/engine.py:1666
↓ 2 callers
Function
_build_average_tensor_optimizer
(monkeypatch, *, copy_streams)
tests/unit/v1/zero/test_overlap_comm_record_stream.py:137
↓ 2 callers
Method
_build_indexes
(self, files: list[Path])
ci/tests_fetcher.py:319
↓ 2 callers
Function
_build_writer
(file_path)
tests/unit/v1/nvme/test_fast_file_writer_fd_close.py:58
↓ 2 callers
Function
_byte_cast_multiple_tensors
(typed_tensor_list)
tests/unit/utils/test_byte_cast.py:30
↓ 2 callers
Function
_byte_cast_single_tensor
(typed_tensor)
tests/unit/utils/test_byte_cast.py:23
↓ 2 callers
Function
_capture_params
(engine)
tests/unit/v1/zero/test_zero2_offload_multi_backward.py:54
↓ 2 callers
Method
_change_recovery_script_permissions
(self, dst)
deepspeed/runtime/engine.py:5318
↓ 2 callers
Method
_choose_module_key
(self, sd)
deepspeed/runtime/state_dict_factory.py:140
↓ 2 callers
Method
_clean_inflight_param_registry
(self)
deepspeed/runtime/zero/partitioned_param_coordinator.py:189
↓ 2 callers
Method
_clear_previous_reduced_grads
(self)
deepspeed/runtime/zero/stage_1_and_2.py:1807
↓ 2 callers
Method
_close_pool
(self, pool, num_procs, force=False)
tests/unit/common.py:358
↓ 2 callers
Function
_collect_autoep_expert_grads
(engine)
tests/unit/v1/moe/test_autoep_grad_parity.py:114
↓ 2 callers
Method
_combine_output_splits
Join the splits of the output into a single result. Args: outputs (List[Any]): The reduced outputs for each output split.
deepspeed/runtime/zero/tiling.py:195
↓ 2 callers
Method
_common_checkpoint_state
(self, module_state_dict, zero_optimizer_state, save_frozen_param)
deepspeed/runtime/engine.py:4830
↓ 2 callers
Function
_compare_optimizers
(model_size, param1, optimizer1, param2, optimizer2)
tests/unit/ops/adam/test_cpu_adam.py:36
↓ 2 callers
Function
_compute
(module, *inputs, do_checkpoint=False)
tests/unit/runtime/activation_checkpointing/test_activation_checkpointing.py:19
↓ 2 callers
Method
_configure_bf16_optimizer
(self, optimizer)
deepspeed/runtime/engine.py:2284
↓ 2 callers
Method
_configure_master_weights
Common validation and dtype selection for ZeRO optimizer master-weight settings. Optionally accepts callables that enforce backend-sp
deepspeed/runtime/base_optimizer.py:467
↓ 2 callers
Method
_configure_train_batch_size
(self)
deepspeed/runtime/config.py:976
↓ 2 callers
Method
_configure_zenflow
Configure ZenFlow optimizer
deepspeed/runtime/zenflow/zenflow_stage_1_and_2.py:88
↓ 2 callers
Method
_configure_zero_optimizer
(self, optimizer)
deepspeed/runtime/engine.py:2307
↓ 2 callers
Method
_consolidated_16bit_state_dict
Consolidate the 16-bit state dictionary.
deepspeed/runtime/engine.py:5384
↓ 2 callers
Method
_convert_to_zero_parameters
(self, param_list)
deepspeed/runtime/zero/partition_parameters.py:1178
↓ 2 callers
Function
_count_deleted_fds
How many fds in /proc/self/fd point at a now-deleted file located under target_dir? Restricting to target_dir avoids false positives from unre
tests/unit/v1/nvme/test_fast_file_writer_fd_close.py:42
↓ 2 callers
Method
_create_checkpoint_file
(self, save_dir, tag, zero_checkpoint)
deepspeed/runtime/engine.py:5140
↓ 2 callers
Method
_create_module_forward_post_hook
(self)
deepspeed/runtime/engine.py:2659
↓ 2 callers
Method
_create_module_forward_pre_hook
(self)
deepspeed/runtime/engine.py:2652
↓ 2 callers
Method
_decode
(self, x, return_dict=True, generator=None)
deepspeed/model_implementations/diffusers/vae.py:34
↓ 2 callers
Function
_define_dc_ops
()
tests/unit/compile/test_list_schedule.py:25
↓ 2 callers
Method
_diff_files
Return (changed, deleted) repo-root-relative paths for ``base_rev..HEAD``. 'changed' = added / modified / renamed-new / copied-new; 'deleted'
ci/tests_fetcher.py:218
↓ 2 callers
Function
_digest_words
(words: torch.Tensor, *, salt: int, modulus: int)
deepspeed/moe/ep_tp_dispatch.py:99
↓ 2 callers
Method
_do_init
(self, path, skip_warmup)
deepspeed/runtime/data_pipeline/data_sampling/indexed_dataset.py:494
↓ 2 callers
Function
_do_io_complete
csrc/aio/common/deepspeed_aio_common.cpp:112
↓ 2 callers
Function
_do_io_submit_block
csrc/aio/common/deepspeed_aio_common.cpp:92
↓ 2 callers
Function
_do_io_submit_singles
csrc/aio/common/deepspeed_aio_common.cpp:71
↓ 2 callers
Function
_do_schedule_without_allgather
(scheduled: List[Node], unscheduled: List[Node], edges: Dict[Node, List[Node]],
deepspeed/compile/list_schedule.py:132
↓ 2 callers
Function
_do_set_z3_leaf_modules
(model: torch.nn.Module, leaf_module_classes: Union[List[Type], List[str]],
deepspeed/utils/z3_leaf_module.py:57
↓ 2 callers
Method
_drain
(self, num_bytes, fd, file_offset, blocking=False)
deepspeed/io/base_io_buffer.py:46
↓ 2 callers
Method
_drain_io_buffer
(self, num_bytes)
deepspeed/io/fast_file_writer.py:132
↓ 2 callers
Function
_eigenvalue_summary_events
(block_eigenvalue, global_samples)
deepspeed/runtime/engine.py:230
↓ 2 callers
Function
_elementwise_flops_compute
(input, other)
deepspeed/profiling/flops_profiler/profiler.py:841
↓ 2 callers
Function
_empty_grad_buffer
(param)
deepspeed/compile/init_z1.py:19
↓ 2 callers
Method
_enable_mem_efficient_linear
(self)
deepspeed/runtime/zero/partition_parameters.py:400
↓ 2 callers
Method
_encode
(self, x, return_dict=True)
deepspeed/model_implementations/diffusers/vae.py:77
↓ 2 callers
Method
_ensure_quantized
(self, tensor: torch.Tensor)
deepspeed/linear/quantization.py:58
↓ 2 callers
Method
_exec_schedule
(self, pipe_schedule)
deepspeed/runtime/pipe/engine.py:1396
↓ 2 callers
Method
_fill_param_grad_accum_attribute
(self, param)
deepspeed/runtime/zero/stage_1_and_2.py:1069
↓ 2 callers
Method
_fini
(self)
deepspeed/io/fast_file_writer.py:113
↓ 2 callers
Method
_flush_buffers_until_complete
(self)
deepspeed/runtime/swap_tensor/async_swapper.py:119
↓ 2 callers
Method
_flush_ready_buffers
(self)
deepspeed/runtime/swap_tensor/async_swapper.py:112
↓ 2 callers
Function
_folded_zero2_tp2_ep4_config
()
tests/unit/v1/moe/test_autoep_autotp_grad_parity.py:211
↓ 2 callers
Method
_forward
(self, sample, timestamp, encoder_hidden_states,
deepspeed/model_implementations/diffusers/unet.py:66
↓ 2 callers
Method
_forward
(self, sample, timestamp, encoder_hidden_states, return_dict=True)
deepspeed/model_implementations/diffusers/vae.py:150
↓ 2 callers
Method
_forward_attention
(self, layer_idx: int, qkv: torch.Tensor, kv_cache: torch.Tensor, ragged_batch_info
deepspeed/inference/v2/model_implementations/exaone4/model.py:145
↓ 2 callers
Function
_fp6_quantized_linear_helper
(tokens: int, in_channels: int, out_channels
tests/unit/inference/v2/modules/test_quantized_linear_module.py:86
↓ 2 callers
Function
_gather
Gather tensors and concatenate along the last dimension.
deepspeed/compression/basic_layer.py:675
↓ 2 callers
Method
_gather_and_compare_params
(self, model, torch_q, torch_o, compare_values=True)
tests/unit/model_parallelism/test_tp_plan_e2e.py:107
↓ 2 callers
Function
_gather_logical_tensor
(tensor, logical_shape, partition_dim,
deepspeed/module_inject/layers.py:1236
↓ 2 callers
Method
_generate
(self, model, tokenizer, prompt)
tests/unit/hybrid_engine/test_he_llama.py:31
↓ 2 callers
Method
_generate
(self, model, tokenizer, prompt)
tests/unit/hybrid_engine/test_he_all.py:31
↓ 2 callers
Function
_get_aio_latencies
csrc/aio/common/deepspeed_aio_common.cpp:60
↓ 2 callers
Function
_get_checkpoint_files
(checkpoint_dir, glob_pattern)
deepspeed/checkpoint/ds_to_universal.py:803
↓ 2 callers
Method
_get_checkpoint_value
(self, key)
deepspeed/checkpoint/deepspeed_checkpoint.py:159
↓ 2 callers
Function
_get_compatible_gpus_v01
We use two heuristics to compute the batch size 1. We use the Lowest Common Multiple of the micro-batches as the base batch size and scale
deepspeed/elasticity/elasticity.py:83
↓ 2 callers
Method
_get_current_buffer
(self)
deepspeed/runtime/swap_tensor/async_swapper.py:161
↓ 2 callers
Function
_get_data_parallel_world_size
Return world size for the data parallel group.
deepspeed/utils/groups.py:771
↓ 2 callers
Method
_get_expert_ckpt_name
(checkpoints_path, layer_id, expert_id, tag, mpu=None)
deepspeed/runtime/engine.py:4180
↓ 2 callers
Function
_get_expert_parallel_ranks
Generate expert parallel and expert data parallel group ranks list. Example - E + M + D parallel world_size = 16 model_degree
deepspeed/utils/groups.py:472
↓ 2 callers
Function
_get_expert_weight
Get expert weight tensor by name, handling both attribute and child module patterns.
deepspeed/moe/ep_repack.py:259
↓ 2 callers
Method
_get_fixture_kwargs
(self, request, func)
tests/unit/common.py:166
↓ 2 callers
Method
_get_flattened_partition
(self, all_partition_states, group=None)
deepspeed/runtime/zero/stage3.py:3174
↓ 2 callers
Method
_get_lean_tensors
(self, padded_flattened_tensor, group_tensors, paddings)
deepspeed/runtime/zero/stage3.py:3044
↓ 2 callers
Function
_get_mem_usage_out_of_torch
()
deepspeed/compile/profilers/graph_profile.py:113
↓ 2 callers
Method
_get_nvtx_domain
(self, domain)
accelerator/cuda_accelerator.py:235
↓ 2 callers
Method
_get_optimizer_ckpt_name
(self, checkpoints_path, tag, expp_rank)
deepspeed/runtime/engine.py:4173
↓ 2 callers
Method
_get_optimizer_loss_scale
(self)
deepspeed/runtime/engine.py:3573
↓ 2 callers
Function
_get_padded_tensor
(src_tensor, size)
deepspeed/runtime/zero/stage_1_and_2.py:97
↓ 2 callers
Method
_get_parallel_write_for_ddp
(self, dp_world_size, dp_rank)
deepspeed/runtime/model_checkpointing/data_parallel_writer_factory.py:195
↓ 2 callers
Function
_get_param_uc_conversion_meta
Return the conversion-facing view of AutoTP UC metadata for a parameter. AutoTP keeps a single parameter-level metadata object with two roles:
deepspeed/module_inject/layers.py:481
↓ 2 callers
Method
_get_parameter_partitions
(self)
deepspeed/runtime/zero/stage3.py:937
↓ 2 callers
Method
_get_rank_zero_ckpt_name
(self, checkpoints_path, tag, mp_rank, dp_rank, bf16_mode)
deepspeed/runtime/engine.py:4131
↓ 2 callers
Method
_get_scale_factor
(self)
deepspeed/runtime/lr_schedules.py:544
↓ 2 callers
Function
_get_send_recv_group
the group id is always the smaller rank unless its a wrap around
deepspeed/runtime/pipe/p2p.py:161
↓ 2 callers
Method
_get_sub_group_partition_count
(self, sub_group_id)
deepspeed/runtime/zero/stage3.py:606
↓ 2 callers
Method
_get_swap_paths
(self, parameters, num_elems)
deepspeed/runtime/swap_tensor/optimizer_utils.py:401
↓ 2 callers
Function
_get_tag_from_path
(path)
deepspeed/runtime/checkpoint_engine/nebula_checkpoint_engine.py:16
↓ 2 callers
Function
_get_test_write_file
(tmpdir, index)
tests/unit/v1/nvme/test_aio.py:52
↓ 2 callers
Method
_get_trainable_parameter_groups
(self)
deepspeed/runtime/zero/stage3.py:645
↓ 2 callers
Function
_get_zero3_model_state_files
(checkpoint_dir)
deepspeed/checkpoint/ds_to_universal.py:782
↓ 2 callers
Function
_git
(cwd: Path, *args: str)
ci/test_torch_latest.py:55
↓ 2 callers
Function
_glm_type_transpose
(input, mp_size)
deepspeed/module_inject/fusedqkv_utils.py:67
← previous
next →
1,301–1,400 of 11,258, ranked by callers