MCPcopy Create free account

hub / github.com/deepspeedai/DeepSpeed / functions

Functions11,258 in github.com/deepspeedai/DeepSpeed

↓ 6 callersMethodset_tensor_parallel_config
(self, mp_size, mp_group)
deepspeed/module_inject/auto_tp.py:344
↓ 6 callersMethodsetup_layout
Create layout tensor for the given sequence length Arguments: seq_len: required: an integer determining number of attention head
deepspeed/ops/sparse_attention/sparsity_config.py:31
↓ 6 callersFunctionshard_param
Utility for sharding a parameter. This will return the slice of the parameter that should exist on the given shard_rank given the sharding co
deepspeed/inference/v2/model_implementations/sharding/utils.py:43
↓ 6 callersMethodsimd_width
(self)
op_builder/builder.py:466
↓ 6 callersFunctionsplit_params_into_different_moe_groups_for_optimizer
Split parameters into different MoE groups for optimizer Args: param_groups (Union[Dict[str, Any], Tuple[Dict[str, Any], ...], List[Dict[
deepspeed/moe/utils.py:72
↓ 6 callersMethodsteps_per_print
(self)
deepspeed/runtime/engine.py:1373
↓ 6 callersFunctionstore_with_byte_offset
Stores a fragment to memory
csrc/deepspeed4science/evoformer_attn/iterators/predicated_tile_iterator_atomic.h:679
↓ 6 callersMethodsync_pread
csrc/aio/py_lib/deepspeed_py_io_handle.cpp:329
↓ 6 callersMethodsynchronize
(self)
tests/unit/v1/zero/test_overlap_comm_record_stream.py:45
↓ 6 callersMethodtokens_to_seq
Mapping of token to which sequence it belongs to in the ragged batch. If the device Tensor is requested, the Tensor is truncated to t
deepspeed/inference/v2/ragged/ragged_wrapper.py:240
↓ 6 callersFunctiontopkgating
Implements TopKGating on logits.
deepspeed/moe/sharded_moe.py:434
↓ 6 callersMethodtotal_memory
(self, device_index=None)
accelerator/hpu_accelerator.py:153
↓ 6 callersMethodupdate
Updated the running absmax used to calculate params. Function Arguments : val : The __half2 value to update the running min and max with.
csrc/includes/quantization_utils.h:177
↓ 6 callersFunctionvalidate_aio_operation
csrc/aio/common/deepspeed_aio_common.cpp:325
↓ 6 callersFunctionvalidate_test
(model_w_task, dtype, enable_cuda_graph, enable_triton)
tests/unit/inference/test_inference.py:270
↓ 6 callersMethodversion_dependent_macros
(self)
op_builder/builder.py:762
↓ 6 callersMethodwrite_events
(self, event_list)
deepspeed/monitor/monitor.py:20
↓ 6 callersFunctionzero3_linear_wrap
(input, weight, bias=None)
deepspeed/runtime/zero/linear.py:148
↓ 5 callersMethod__init__
(self, *args)
tests/unit/runtime/zero/test_zero_context.py:24
↓ 5 callersMethod__init__
(self, args, world_info_base64)
deepspeed/launcher/multinode_runner.py:57
↓ 5 callersMethod__init__
(self, )
deepspeed/runtime/domino/transformer.py:46
↓ 5 callersMethod__init__
(self, act_range_momentum=0.95, quant_mode='symmetric')
deepspeed/compression/basic_layer.py:28
↓ 5 callersFunction_allow_dynamo_dynamic_parameter_shapes_for_z3
Acquire process-wide ZeRO-3 Dynamo config ownership and return its release callback.
deepspeed/compile/init_z3.py:30
↓ 5 callersFunction_base_config
(zero_stage, gradient_accumulation_steps=1, cpu_offload=False)
tests/unit/v1/zero/test_zero2_offload_multi_backward.py:17
↓ 5 callersMethod_cast_module_mixed_precision
Cast params to param_dtype; cast buffers only when buffer_dtype is set.
deepspeed/runtime/engine.py:1482
↓ 5 callersMethod_constant_buffered_norm2
(self, input, buffer_size=250000000)
deepspeed/runtime/zero/stage3.py:1766
↓ 5 callersMethod_create_config
(self, world_size, rank)
deepspeed/runtime/model_checkpointing/data_parallel_writer_factory.py:46
↓ 5 callersFunction_cutlass_moe_testing_helper
(tokens: int, in_channels: int, intermediate_d
tests/unit/inference/v2/modules/test_cutlass_moe.py:83
↓ 5 callersMethod_get_fp32_opt_state_partition
(self, param, release_swap_buffers, optim_state_key=None)
deepspeed/runtime/zero/stage3.py:2830
↓ 5 callersMethod_get_norm_groups
(self)
deepspeed/runtime/zero/stage3.py:2334
↓ 5 callersMethod_get_swap_paths
(self, params, must_exist=False)
deepspeed/runtime/swap_tensor/partitioned_param_swapper.py:148
↓ 5 callersMethod_io_aligned_numel
(self, numel)
deepspeed/runtime/swap_tensor/partitioned_param_swapper.py:386
↓ 5 callersFunction_make_internvl_adapter
(world_size=2, rank=0)
deepspeed/sequence/test_autosp.py:512
↓ 5 callersFunction_make_llava_adapter
(world_size=2, rank=0)
deepspeed/sequence/test_autosp.py:400
↓ 5 callersFunction_make_qwen2vl_adapter
(world_size=2, rank=0)
deepspeed/sequence/test_autosp.py:620
↓ 5 callersMethod_mixed_precision_dtypes
Resolve (param_dtype, buffer_dtype) for the module cast. param_dtype follows the enabled fp16/bf16 mode (None if neither). buffer_dty
deepspeed/runtime/engine.py:1462
↓ 5 callersFunction_neg
(graph, arg, name, device_time=0)
tests/unit/compile/test_list_schedule.py:115
↓ 5 callersMethod_optimizer_states_and_gradient_swap_out
(self, sub_group_id, timer_names=None)
deepspeed/runtime/zero/stage3.py:2424
↓ 5 callersMethod_partition_rank
(self, param)
deepspeed/runtime/zero/partition_parameters.py:1659
↓ 5 callersMethod_partitioned_params_swap_out
(self, i)
deepspeed/runtime/zero/stage3.py:1187
↓ 5 callersFunction_payload
()
tests/unit/v1/moe/test_autoep_autotp_dispatch.py:24
↓ 5 callersMethod_post_step
(self, timer_names)
deepspeed/runtime/zero/stage3.py:2512
↓ 5 callersFunction_ragged_embed_test_helper
Helper for embedding test to limit the number of tests to run. Params: embed_dim (int): Model dimension vocab_size (int): Le
tests/unit/inference/v2/kernels/ragged_ops/test_ragged_embed.py:48
↓ 5 callersMethod_read_autotune_table
()
deepspeed/ops/transformer/inference/triton/matmul_ext.py:455
↓ 5 callersMethod_release_params
Helper for `release_[component]` methods. Accepts a list of tuples where the first element is the module param that needs to be delet
deepspeed/module_inject/containers/features/hybrid_engine.py:100
↓ 5 callersFunction_router_grad_model
()
tests/unit/v1/moe/test_autoep_autotp_grad_parity.py:271
↓ 5 callersFunction_scheduled_names
(graph)
tests/unit/compile/test_list_schedule.py:132
↓ 5 callersMethod_test
(self, inputs, class_tmpdir, checkpoint_tag, mp_size, pp_size, mp_resize, pp_resize)
tests/unit/model_parallelism/test_configurable_parallel_pp.py:191
↓ 5 callersFunction_tol
(dtype)
tests/unit/v1/moe/test_group_gemm_triton.py:27
↓ 5 callersMethod_update_model_bit16_weights
(self, group_index)
deepspeed/runtime/zero/stage_1_and_2.py:774
↓ 5 callersFunction_uses_zero3_partitioned_autoep_metadata
(autoep_metadata)
deepspeed/checkpoint/ds_to_universal.py:504
↓ 5 callersFunction_valid_env
(path: Path, **overrides: str)
ci/test_torch_latest.py:40
↓ 5 callersFunction_zero3_config
(memory_efficient_linear)
tests/unit/runtime/zero/test_zero_linear_direct_init.py:16
↓ 5 callersFunction_zero3_functorch_config
()
tests/unit/v1/zero/test_zero_functorch_linear.py:29
↓ 5 callersFunctionadamw
r"""Functional API that performs AdamW algorithm computation. See :class:`~torch.optim.AdamW` for details.
deepspeed/ops/adam/zenflow_torch_adam.py:1044
↓ 5 callersMethodadd_export
(self, key, var)
deepspeed/launcher/multinode_runner.py:37
↓ 5 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
csrc/deepspeed4science/evoformer_attn/iterators/predicated_tile_access_iterator_residual_last.h:764
↓ 5 callersMethodbackward
(ctx, *grads)
deepspeed/runtime/sequence_parallel/ulysses_sp.py:1012
↓ 5 callersFunctionbwc_pipeline_parallel_world_size
Backwards-compatible way of querying the pipeline parallel world size.
deepspeed/utils/bwc.py:81
↓ 5 callersMethodcheck_and_propagate_first_head_layout
If all heads require same sparsity layout, it propagate first head layout to all heads Arguments: layout: required: a tensor of
deepspeed/ops/sparse_attention/sparsity_config.py:48
↓ 5 callersMethodclear
csrc/includes/deepcompile.h:235
↓ 5 callersMethodclear_mask
< Efficiently disables all accesses guarded by mask
csrc/deepspeed4science/evoformer_attn/iterators/predicated_tile_iterator_atomic.h:454
↓ 5 callersFunctioncompare_state_dicts
(state0, state1, expected_mismatch_keys=[])
tests/unit/checkpoint/common.py:88
↓ 5 callersFunctioncompute_split_plan
Compute AllToAllV split sizes for token dispatch/combine. ``num_tokens_per_expert`` may be supplied by the caller to reuse the histogram alre
deepspeed/module_inject/auto_ep_layer.py:135
↓ 5 callersMethodcreate
(cls, zenflow_config)
deepspeed/runtime/zenflow/zenflow_stage_1_and_2.py:82
↓ 5 callersFunctioncreate_config_from_dict
(tmpdir, config_dict)
tests/unit/simple_model.py:291
↓ 5 callersFunctiondeepcompile_z3_forward_context
(engine)
deepspeed/compile/z3_eager_fallback.py:46
↓ 5 callersFunctiondo_aio_operation_overlap
csrc/aio/common/deepspeed_aio_common.cpp:192
↓ 5 callersFunctiondo_aio_operation_sequential
csrc/aio/common/deepspeed_aio_common.cpp:129
↓ 5 callersFunctionds_id
(param)
deepspeed/utils/debug.py:54
↓ 5 callersFunctionflatten
(d, parent_key='', sep='_')
deepspeed/autotuning/tuner/utils.py:55
↓ 5 callersMethodflops_profiler_enabled
(self)
deepspeed/runtime/engine.py:1067
↓ 5 callersMethodforward
(self, x, mask)
tests/unit/runtime/activation_checkpointing/test_activation_checkpointing.py:118
↓ 5 callersMethodfree
(self, buffers)
deepspeed/runtime/swap_tensor/utils.py:220
↓ 5 callersMethodgather
(self, tensor, gather_list, dst, group=None, async_op=False)
deepspeed/comm/ccl.py:131
↓ 5 callersMethodgetDSId
csrc/includes/deepcompile.h:126
↓ 5 callersMethodgetDSTensor
csrc/includes/deepcompile.h:277
↓ 5 callersMethodget_all_ranks_from_group
(self, group)
deepspeed/comm/ccl.py:166
↓ 5 callersFunctionget_backward_inputs
(frame_key=None)
deepspeed/compile/patch_compiled_func.py:106
↓ 5 callersMethodget_bwd_mapping
(self, bw_graph: Graph)
deepspeed/compile/graph_param.py:61
↓ 5 callersMethodget_data_parallel_world_size
The number of pipelines.
deepspeed/runtime/pipe/topology.py:432
↓ 5 callersMethodget_global_rank
(self)
deepspeed/runtime/pipe/topology.py:407
↓ 5 callersMethodget_global_rank
(self, group, group_rank)
deepspeed/comm/torch.py:406
↓ 5 callersFunctionget_gpt2_model
(args_others, mp_size=1)
tests/unit/megatron_model.py:24
↓ 5 callersMethodget_gradient_for_reduction
(self, param)
deepspeed/runtime/zero/stage_1_and_2.py:1082
↓ 5 callersMethodget_lr_ratio
(self)
deepspeed/runtime/lr_schedules.py:874
↓ 5 callersFunctionget_master_port
(base_port=29500, port_range_size=1000)
tests/unit/common.py:47
↓ 5 callersFunctionget_megatron_version
()
tests/unit/megatron_model.py:16
↓ 5 callersMethodget_sequence
Get the sequence descriptor for the given sequence id. If the sequence does not exist, then None is returned.
deepspeed/inference/v2/ragged/ragged_manager.py:125
↓ 5 callersMethodget_sequence_parallel_world_size
(self)
tests/unit/v1/moe/test_autoep_unit.py:286
↓ 5 callersMethodget_slice_parallel_group
(self)
deepspeed/runtime/pipe/topology.py:471
↓ 5 callersMethodget_slice_parallel_world_size
(self)
deepspeed/runtime/pipe/topology.py:465
↓ 5 callersMethodget_swap_paths
(self)
deepspeed/runtime/swap_tensor/utils.py:77
↓ 5 callersMethodgroup_step
(self, group_to_paramlist)
deepspeed/ops/adam/zenflow_torch_adam.py:197
↓ 5 callersMethodhas_all_gather_into_tensor
(self)
deepspeed/comm/torch.py:159
↓ 5 callersMethodinference_all_reduce
(self, tensor, op=ReduceOp.SUM, group=None)
deepspeed/comm/ccl.py:87
↓ 5 callersFunctioninitialize
Initialize the DeepSpeed Engine. Arguments: args: an object containing local_rank and deepspeed_config fields. This is option
deepspeed/__init__.py:93
↓ 5 callersFunctionis_autoep_zero3_partitioned_entry
(entry)
deepspeed/checkpoint/autoep_zero3_metadata.py:34
↓ 5 callersFunctionis_compiling
()
deepspeed/runtime/compiler.py:84
↓ 5 callersMethodis_first_stage
True if this process is in the first stage in the pipeline.
deepspeed/runtime/pipe/engine.py:533
← previousnext →501–600 of 11,258, ranked by callers