MCPcopy Create free account

hub / github.com/deepspeedai/DeepSpeed / functions

Functions11,258 in github.com/deepspeedai/DeepSpeed

↓ 5 callersMethodis_record_trace
(self)
deepspeed/runtime/zero/partitioned_param_coordinator.py:186
↓ 5 callersMethodkv_ptrs
Pointer to where the list of KV ids associated with a sequence are. If the device Tensor is requested, the Tensor is truncated to the
deepspeed/inference/v2/ragged/ragged_wrapper.py:260
↓ 5 callersFunctionload_with_byte_offset
Loads a fragment from memory
csrc/deepspeed4science/evoformer_attn/iterators/epilogue_predicated_tile_iterator.h:323
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment from memory
csrc/deepspeed4science/evoformer_attn/iterators/predicated_tile_iterator_residual_last.h:812
↓ 5 callersMethodlog_timers
(self)
deepspeed/runtime/swap_tensor/optimizer_utils.py:220
↓ 5 callersFunctionmacs_to_string
(macs, units=None, precision=DEFAULT_PRECISION)
deepspeed/profiling/flops_profiler/profiler.py:1123
↓ 5 callersFunctionmake_cutlass_checkout
(path)
tests/unit/ops/deepspeed4science/test_evoformer_attn_builder.py:16
↓ 5 callersMethodneeds_scaler
Check if this optimizer requires loss scaling for correct backward pass. Returns True if any of the following conditions are met:
deepspeed/runtime/base_optimizer.py:382
↓ 5 callersFunctionparams_to_string
(params_num, units=None, precision=DEFAULT_PRECISION)
deepspeed/profiling/flops_profiler/profiler.py:1169
↓ 5 callersFunctionprint_dist
print message when get_dist_msg() deems it should be logged, see its docstring for details. Use this function instead of `log_dist` when the log
deepspeed/utils/logging.py:126
↓ 5 callersMethodquantize
(self, parameter_group, overflow, eigenvalue_enabled, block_eigenvalue={})
deepspeed/runtime/quantize.py:51
↓ 5 callersFunctionreduce
(tensor, dst, op=ReduceOp.SUM, group=None, async_op=False,
deepspeed/comm/comm.py:595
↓ 5 callersMethodref
Returns a TensorRef to the operand
csrc/deepspeed4science/evoformer_attn/gemm/custom_mma_base.h:118
↓ 5 callersFunctionrequires_offset
csrc/includes/quantization.h:20
↓ 5 callersMethodrun_gpt2_test
(self, test_config, output)
tests/model/Megatron_GPT2/test_common.py:55
↓ 5 callersMethodrun_test
(self, test_config, r_tol)
tests/model/BingBertSquad/BingBertSquad_run_func_test.py:113
↓ 5 callersFunctionrun_tp_layer_fwd_bwd
(tp_size, tp_overlap_comm, column_parallel, use_tp_model_init=False, gather_output=False)
tests/unit/model_parallelism/test_autotp_training.py:431
↓ 5 callersFunctionrun_training_steps
(engine, num_steps=3, seq_len=8, hidden_dim=64)
tests/unit/v1/moe/autoep_test_utils.py:246
↓ 5 callersFunctionsafe_get_full_fp32_param
Assemble and return the fp32 parameter of a low-precision (e.g., fp16) parameter. Args: param (``torch.nn.Parameter``): A model p
deepspeed/utils/tensor_fragment.py:134
↓ 5 callersMethodsave
(self, state_dict, path: str)
deepspeed/runtime/checkpoint_engine/checkpoint_engine.py:32
↓ 5 callersMethodset_dependency
Set dependency can be used for managing dependencies when a mapping is provided in the class definition for the layer. The dep_name h
deepspeed/inference/v2/model_implementations/layer_container_base.py:272
↓ 5 callersMethodset_residual_tile
csrc/deepspeed4science/evoformer_attn/iterators/predicated_tile_access_iterator_residual_last.h:760
↓ 5 callersFunctionshard_tensor_node
(gm: GraphModule, tensor_node: Node)
deepspeed/compile/util.py:596
↓ 5 callersFunctionsingle_all_to_all
(input, scatter_idx, gather_idx, batch_dim_idx, group, async_op=False, handle=None, type=None)
deepspeed/sequence/layer.py:241
↓ 5 callersMethodstep
Update the model parameters. .. note:: This method will be called internally by ZeRO-Offload. DeepSpeed users should
deepspeed/ops/adam/cpu_adam.py:107
↓ 5 callersMethodstep
Not supporting closure.
deepspeed/runtime/zenflow/zenflow_stage_1_and_2.py:750
↓ 5 callersMethodstore_with_pointer_offset
Store a fragment to memory
csrc/deepspeed4science/evoformer_attn/iterators/predicated_tile_iterator_residual_last.h:830
↓ 5 callersFunctiontask_log
(tid, msg, force=False)
deepspeed/nvme/test_ds_aio_utils.py:20
↓ 5 callersFunctiontask_log
(tid, msg, force=False)
csrc/aio/py_test/test_ds_aio_utils.py:20
↓ 5 callersMethodto_layers
(self)
tests/unit/alexnet_model.py:51
↓ 5 callersFunctiontop1gating
Implements Top1Gating on logits.
deepspeed/moe/sharded_moe.py:235
↓ 5 callersMethodtrain_batch
Progress the pipeline to train the next batch of data. The engine will ingest ``self.train_batch_size()`` total samples collectively across al
deepspeed/runtime/pipe/engine.py:341
↓ 5 callersFunctiontrain_cifar
(model, config, num_steps=400, average_dp_losses=True, fp16=True, seed=123)
tests/unit/alexnet_model.py:126
↓ 5 callersFunctiontranspose
(data)
deepspeed/module_inject/utils.py:9
↓ 5 callersMethodunscale_and_clip_grads
(self, grad_groups_flat, total_norm, apply_scale=True)
deepspeed/runtime/fp16/fused_optimizer.py:367
↓ 5 callersFunctionvalidate_folding_metadata
(metadata, *, tp_size,
deepspeed/checkpoint/autoep_universal.py:68
↓ 5 callersMethodwriter
(cls, path, dtype)
deepspeed/runtime/data_pipeline/data_sampling/indexed_dataset.py:375
↓ 5 callersMethodzero_offload_param
(self)
deepspeed/runtime/engine.py:1182
↓ 4 callersMethodInstance
(cls)
deepspeed/ops/transformer/inference/op_binding/workspace.py:32
↓ 4 callersMethodRATIO
csrc/includes/dropout.h:22
↓ 4 callersFunctionRun
csrc/includes/gemm_test.h:123
↓ 4 callersMethod__init__
(self, num_experts=4, ffn_hidden=128, hidden_size=64, intermediate_size=None)
tests/unit/v1/moe/autoep_test_utils.py:56
↓ 4 callersMethod__init__
(self, out=None, bias=None)
tests/unit/runtime/zero/test_zero_context_return.py:28
↓ 4 callersMethod__init__
(self, nf, nx)
tests/unit/compression/test_compression.py:78
↓ 4 callersMethod__release_param
(self, param: Parameter, free_data: bool = True)
deepspeed/runtime/zero/partitioned_param_coordinator.py:624
↓ 4 callersMethod_aligned_size
(self, param)
deepspeed/runtime/zero/partition_parameters.py:1645
↓ 4 callersFunction_assert_params_match
(ref, test, label, tol=5e-5)
tests/unit/v1/zero/test_zero2_offload_multi_backward.py:58
↓ 4 callersMethod_assert_valid_mixed_precision_config
param_dtype, if set, must match the enabled fp16/bf16 mode. The optimizer/master-weight/reduction paths derive the model dtype from t
deepspeed/runtime/engine.py:1443
↓ 4 callersMethod_clear_fp32_optimizer_param_groups
(self)
deepspeed/runtime/zero/stage3.py:3089
↓ 4 callersMethod_configure_expert_parallel
Initialize AutoEP: detect MoE layers, create EP groups, replace with EP-enabled layers.
deepspeed/runtime/engine.py:538
↓ 4 callersMethod_create_zero_config
(self, hidden_dim, leaf_module=None)
tests/unit/runtime/zero/test_zero_leaf_module.py:300
↓ 4 callersFunction_do_parallel_work
(do_work, work_chunks, num_workers)
deepspeed/checkpoint/ds_to_universal.py:384
↓ 4 callersFunction_do_reshape
(src_3d, tgt_3d)
tests/unit/checkpoint/test_reshape_checkpoint.py:9
↓ 4 callersFunction_ds_initialize_for_param_partitioning_testing
(model: Module, cfg: dict)
tests/unit/v1/zero/test_zero.py:497
↓ 4 callersMethod_dump_mapping
(self, data_map, map_tag=None)
deepspeed/checkpoint/deepspeed_checkpoint.py:237
↓ 4 callersMethod_engine
(self, param_dtype=None, buffer_dtype=None, fp16=False, bf16=False)
tests/unit/v1/half_precision/test_mixed_precision_dtype.py:26
↓ 4 callersMethod_engine
(self, module)
tests/unit/v1/half_precision/test_mixed_precision_dtype.py:72
↓ 4 callersFunction_expert_params
(engine)
tests/unit/v1/moe/test_autoep_checkpoint.py:128
↓ 4 callersMethod_flush_gradient_swapper
(self, gradient_swapper)
deepspeed/runtime/swap_tensor/optimizer_utils.py:230
↓ 4 callersFunction_folding_spec
(mp_mode="tp", tp_size=2)
tests/unit/v1/moe/test_autoep_autotp_grad_parity.py:46
↓ 4 callersFunction_full_grad_by_suffix
(engine, suffix)
tests/unit/v1/moe/test_autoep_autotp_grad_parity.py:298
↓ 4 callersMethod_gather_view_and_storage
(self, shard, graph_id, ds_id)
tests/torch_compile/test_deepcompile_z3_release.py:45
↓ 4 callersMethod_get_fp32_grad_state_partition
(self, param, release_swap_buffers)
deepspeed/runtime/zero/stage3.py:2782
↓ 4 callersMethod_get_model_type
Extract model type from module config or class name.
deepspeed/module_inject/auto_tp.py:561
↓ 4 callersMethod_get_param_partition_rank
(self, param)
deepspeed/runtime/zero/stage3.py:597
↓ 4 callersFunction_grad_parity_metrics
(actual, expected)
tests/unit/v1/moe/test_autoep_autotp_grad_parity.py:403
↓ 4 callersFunction_group_meta
Compute per-group m-start (exclusive prefix) and m-size on device (no D2H sync). A single Triton kernel reads ``offs`` and writes both outputs, r
deepspeed/moe/group_gemm_triton.py:276
↓ 4 callersMethod_init_dc
(self)
tests/torch_compile/test_deepcompile_z3_release.py:28
↓ 4 callersFunction_initialize_folded_engine
(*, zero_stage=0, ep_size=2, mixed_precision=True)
tests/unit/v1/moe/test_autoep_autotp_runtime.py:154
↓ 4 callersFunction_make_engine
()
tests/unit/runtime/rollout/test_hybrid_engine_rollout.py:20
↓ 4 callersFunction_make_fwd_graph
()
tests/unit/v1/compile/test_offload_opt_states.py:57
↓ 4 callersFunction_make_node_meta
(node: Node, ds_id: int, comm: bool)
deepspeed/compile/fx.py:112
↓ 4 callersFunction_make_param
(numel, ds_persist=False)
tests/unit/v1/compile/test_selective_gather.py:40
↓ 4 callersFunction_make_tokenizer
()
tests/unit/runtime/rollout/test_hybrid_engine_rollout.py:27
↓ 4 callersFunction_model_config
()
tests/unit/inference/v2/model_implementations/test_exaone4_5.py:25
↓ 4 callersMethod_patch_counts_exchange
Emulate the EP metadata all-to-alls seen by rank ``ep_rank``. Both the legacy per-rank exchange and the per-(rank, local-expert) exch
tests/unit/v1/moe/test_autoep_unit.py:947
↓ 4 callersMethod_reassign_or_swap_out_partitioned_parameters
(self, sub_group_id)
deepspeed/runtime/zero/stage3.py:2531
↓ 4 callersMethod_reentrant_activation_checkpointing
True when the module checkpoints activations with the reentrant function. Reentrant checkpointing needs the first stage's inputs to require g
deepspeed/runtime/pipe/engine.py:898
↓ 4 callersFunction_ref_grouped_mm
Pure-PyTorch reference: out[rows_g] = a[rows_g] @ b[g].
tests/unit/v1/moe/test_group_gemm_triton.py:37
↓ 4 callersMethod_register_param
(self, dc, graph_id, ds_id, shape, persistent=False)
tests/torch_compile/test_deepcompile_z3_release.py:33
↓ 4 callersFunction_rel_names
(repo: TmpRepo, tests)
ci/test_tests_fetcher.py:109
↓ 4 callersMethod_release
(self, view, graph_id, ds_id, n_users, synchronize=True)
tests/torch_compile/test_deepcompile_z3_release.py:54
↓ 4 callersFunction_resolve_expected_grad_dtype
(param)
deepspeed/compile/init_z3.py:75
↓ 4 callersFunction_run_fwd_pass
(monkeypatch, graph, budget_gb="0")
tests/unit/v1/compile/test_offload_opt_states.py:71
↓ 4 callersMethod_run_git
(self, args: list[str])
ci/tests_fetcher.py:191
↓ 4 callersFunction_run_router_grad_boundary
(engine, *, logical_dp_world_size, logical_dp_rank, seed)
tests/unit/v1/moe/test_autoep_autotp_grad_parity.py:285
↓ 4 callersMethod_sample_top_p
Sample from logits with temperature and nucleus (top-p) filtering.
deepspeed/runtime/rollout/hybrid_engine_rollout.py:224
↓ 4 callersFunction_selected_experts_from_counts
Build a routing assignment list whose histogram is exactly ``counts``.
tests/unit/v1/moe/test_autoep_unit.py:920
↓ 4 callersFunction_set_cuda_rng_state
Sets the random number generator state of the current GPU. Arguments: new_state (torch.ByteTensor): The desired state This function i
deepspeed/runtime/activation_checkpointing/checkpointing.py:91
↓ 4 callersMethod_set_fp32_optimizer_param_groups
(self)
deepspeed/runtime/zero/stage3.py:3083
↓ 4 callersMethod_set_param_uc_meta
(self, param, *, partition_ty
deepspeed/module_inject/layers.py:383
↓ 4 callersMethod_should_materialize_tp_partition
(self)
deepspeed/module_inject/layers.py:414
↓ 4 callersMethod_start_timers
(self, timer_names)
deepspeed/runtime/engine.py:3485
↓ 4 callersMethod_toy_model_config
(self, shard_size)
tests/unit/checkpoint/test_mics_optimizer.py:25
↓ 4 callersFunction_tp_consistent_input
(engine, *, seed=1234)
tests/unit/v1/moe/test_autoep_autotp_runtime.py:147
↓ 4 callersMethod_tp_partition
(self, params_list)
deepspeed/module_inject/layers.py:766
↓ 4 callersMethod_validate_autoep_folding_checkpoint_metadata
(state, *,
deepspeed/runtime/engine.py:3897
↓ 4 callersFunction_verify_continuous_increase
(values)
tests/unit/runtime/test_lr_schedulers.py:28
↓ 4 callersFunction_write_stage3_checkpoint
(ckpt_dir, shards_by_param, dp_degree, uc_info)
tests/unit/checkpoint/test_autotp_uc_checkpoint.py:324
↓ 4 callersMethodadd_data
(self, pp_index, tp_index, data)
deepspeed/checkpoint/reshape_meg_2d.py:22
← previousnext →601–700 of 11,258, ranked by callers