MCPcopy Create free account

hub / github.com/NVIDIA/apex / functions

Functions1,387 in github.com/NVIDIA/apex

↓ 115 callersMethodappend
(self, work: torch.distributed.Work)
apex/contrib/optimizers/distributed_fused_adam.py:78
↓ 38 callersMethodparameters
Returns an iterator over optimizer parameters
apex/contrib/optimizers/distributed_fused_adam.py:1194
↓ 28 callersMethodbackward
(ctx, loss_grad)
apex/contrib/transducer/transducer.py:212
↓ 25 callersMethodbackward
(ctx, grad_o)
apex/mlp/mlp.py:22
↓ 24 callersMethodstream
(self)
apex/contrib/torchsched/ops/layer_norm.py:60
↓ 22 callersMethodstep
Performs a single optimization step. Arguments: closure (callable, optional): A closure that reevaluates the model
tests/L0/run_optimizers/test_lamb.py:54
↓ 21 callersMethodget
apex/contrib/csrc/layer_norm/ln.h:157
↓ 20 callersMethodverify_group_norm
( self, tst_func, N=32, C=128, H=256, W=256, G=32,
apex/contrib/test/group_norm/test_group_norm.py:118
↓ 18 callersMethodadd
(self, idx)
apex/contrib/optimizers/distributed_fused_lamb.py:95
↓ 18 callersMethodtest_matches_pytorch
( self, rtol: float | None = None, atol: float | None = None, num_layers: int
apex/contrib/test/optimizers/test_dist_adam.py:122
↓ 16 callersMethod__init__
( self, with_2d=False, )
apex/contrib/sparsity/test/test_permutation_application.py:108
↓ 14 callersMethodstate_dict
Returns a dict containing the current state of this :class:`FP16_Optimizer` instance. This dict contains attributes of :class:`FP16_O
apex/contrib/optimizers/fp16_optimizer.py:184
↓ 13 callersMethodzero_grad
(self)
apex/contrib/optimizers/fused_lamb.py:104
↓ 12 callersFunctiondiv_up
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:646
↓ 12 callersMethodrun_transducer_joint
(self, for_vector_kernel, pack_output, relu, dropout)
apex/contrib/test/transducer/test_transducer_joint.py:74
↓ 12 callersMethodstep
Performs a single optimization step. Arguments: closure (callable, optional): A closure that reevaluates the model
apex/contrib/optimizers/fused_sgd.py:130
↓ 11 callersFunction_cast_if_autocast_enabled
(*args)
apex/_autocast_utils.py:22
↓ 11 callersMethodupdate
(self, alphafold: nn.Module)
apex/contrib/test/openfold_triton/test_fused_adam_swa.py:46
↓ 11 callersMethodzero_grad
(self)
apex/optimizers/fused_sgd.py:129
↓ 10 callersFunctioncheck_args
csrc/layer_norm_cuda.cpp:21
↓ 10 callersMethodgen_single_type_test
( self, param_type=torch.float, device="cuda", *, skip_assert: bool = False )
tests/L0/run_optimizers/test_fused_optimizer.py:61
↓ 10 callersMethodgen_single_type_test
(self, param_type=torch.float, device="cuda")
tests/L0/run_optimizers/test_lamb.py:207
↓ 10 callersMethodparameter
Get optimizer parameter Can either accept two ints or one DistributedFusedAdam.ParameterFragment. Arguments: par
apex/contrib/optimizers/distributed_fused_adam.py:1198
↓ 9 callersMethodgen_param_optim
(self, tensors, options, tst_options=None)
tests/L0/run_optimizers/test_fused_optimizer.py:20
↓ 9 callersFunctiongenerateStrides
apex/contrib/csrc/conv_bias_relu/conv_bias_relu.cpp:72
↓ 9 callersFunctiongetConvFusionString
TODO: better name
apex/contrib/csrc/bottleneck/bottleneck.cpp:351
↓ 9 callersFunctiongetConvFusionString
apex/contrib/csrc/conv_bias_relu/conv_bias_relu.cpp:103
↓ 9 callersMethodget_max_diff
(self, ref_param, tst_param)
tests/L0/run_optimizers/test_fused_optimizer.py:50
↓ 9 callersMethodload_state_dict
(self, state_dict)
apex/optimizers/fused_novograd.py:119
↓ 8 callersMethodinit_params
Initialize optimizer state for parameters Ignores parameters that have already been initialized. Arguments: params (iter
apex/contrib/optimizers/distributed_fused_adam.py:1227
↓ 8 callersMethodwait
(self)
apex/contrib/optimizers/distributed_fused_adam.py:82
↓ 7 callersMethoddtypes
Datatypes for the bucket's compute and communication
apex/contrib/optimizers/distributed_fused_adam.py:477
↓ 7 callersMethodgen_grad
(self, ref_param, tst_param)
tests/L0/run_optimizers/test_fused_optimizer.py:38
↓ 7 callersFunctionmake_grouped
(A)
apex/contrib/sparsity/permutation_search_kernels/permutation_utilities.py:194
↓ 7 callersFunctionread_from_smem
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:209
↓ 7 callersFunctionrun_dconv
apex/contrib/csrc/bottleneck/bottleneck.cpp:810
↓ 7 callersFunctionsum_after_2_to_4
(matrix)
apex/contrib/sparsity/permutation_search_kernels/permutation_utilities.py:57
↓ 7 callersFunctionuse_gpu
(initial_override=True)
apex/contrib/sparsity/permutation_search_kernels/permutation_utilities.py:24
↓ 7 callersFunctionwrite_to_smem
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:258
↓ 6 callersMethodbackward
(ctx, grad_y)
apex/contrib/groupbn/batch_norm.py:80
↓ 6 callersMethodgen_grad
(self, ref_param, tst_param)
tests/L0/run_optimizers/test_lamb.py:195
↓ 6 callersMethodgen_param_optim
(self, tensors, lamb_option)
tests/L0/run_optimizers/test_lamb.py:183
↓ 6 callersFunctiongenerateStrides
apex/contrib/csrc/bottleneck/bottleneck.cpp:66
↓ 6 callersFunctionget_stream_name
Generate CUDA Stream name from stream index number. Args: stream_idx: Non-negative index number. 0 refers to the default stream, others r
apex/contrib/torchsched/inductor/_utils.py:30
↓ 6 callersMethodinit
apex/contrib/csrc/bottleneck/bottleneck.cpp:2603
↓ 6 callersFunctionrun_conv_scale_bias_add_activation
apex/contrib/csrc/bottleneck/bottleneck.cpp:437
↓ 6 callersMethodtest_matches_pytorch
Make sure PyTorch and Apex gradient clipping produce same results
apex/contrib/test/clip_grad/test_clip_grad.py:52
↓ 6 callersFunctionto_float
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:81
↓ 5 callersFunction_multi_tensor_copy
Copy between corresponding buffers Uses fused copy kernel if possible.
apex/contrib/optimizers/distributed_fused_adam.py:167
↓ 5 callersMethod_test_fused_layer_norm
( self, batch_size, contiguous, elementwise_affine, mixed_fused,
tests/L0/run_fused_layer_norm/test_fused_layer_norm.py:23
↓ 5 callersMethod_test_fused_rms_norm
( self, batch_size, contiguous, elementwise_affine, mixed_fused,
tests/L0/run_fused_layer_norm/test_fused_layer_norm.py:97
↓ 5 callersFunctionadd
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:301
↓ 5 callersMethodgrad_sync
Ensure that all gradients are synchronized
apex/contrib/optimizers/distributed_fused_adam.py:2123
↓ 5 callersMethodinit_model_for_pruning
Call this method to modify your model to take advantage of sparse matrix multiplication. Note that this call alone only augments the model wit
apex/contrib/sparsity/asp.py:41
↓ 5 callersMethodinit_optimizer_for_pruning
Call this method to monkey patch optimizer step function so that masks can be applied to gradients and weights during training. You mu
apex/contrib/sparsity/asp.py:271
↓ 5 callersFunctionldg
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:103
↓ 5 callersFunctionmake_models
( num_layers: int, size: int, *, lr: float = 0.1, adam_w_mode: bool = True, model_dtyp
apex/contrib/test/optimizers/test_dist_adam.py:33
↓ 5 callersFunctionread_from_gmem
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:191
↓ 5 callersMethodrun_transducer_loss
(self, scalar_t, fuse_softmax_backward, packed_input, for_vector_kernel)
apex/contrib/test/transducer/test_transducer_loss.py:83
↓ 5 callersMethodset_stream
(self, stream: torch.cuda.Stream)
apex/contrib/torchsched/ops/layer_norm.py:49
↓ 5 callersMethodskip_if_v2_not_supported
(self)
apex/contrib/test/group_norm/test_group_norm.py:349
↓ 5 callersFunctionsupports_custom_op
()
apex/normalization/fused_layer_norm.py:17
↓ 5 callersMethodupdate
(self, val, n=1)
tests/L1/common/main_amp.py:565
↓ 5 callersFunctionzero_array
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:281
↓ 4 callersFunctionMaxSharedMemoryPerMultiprocessor
apex/contrib/csrc/groupbn/cuda_utils.h:10
↓ 4 callersMethod__init__
(self, halo_ex)
apex/contrib/bottleneck/halo_exchangers.py:204
↓ 4 callersFunction_coalescing_manager
(group, device, reqs)
apex/contrib/optimizers/distributed_fused_adam.py:64
↓ 4 callersFunction_devices_match
Whether two PyTorch devices are equivalent
apex/contrib/optimizers/distributed_fused_adam.py:149
↓ 4 callersFunction_prep_inputs
(batch_size, normalized_shape, dtype)
tests/L0/run_fused_layer_norm/test_fused_layer_norm.py:11
↓ 4 callersMethodcompute_sparse_masks
Call this method to enable sparsity. If init(...) was called with allow_recompute_mask=False AND sparsity is disabled, pruned field can be Non
apex/contrib/sparsity/asp.py:316
↓ 4 callersFunctiongetFwdConvOutputDim
apex/contrib/csrc/bottleneck/bottleneck.cpp:89
↓ 4 callersFunctiongetFwdConvOutputDim
apex/contrib/csrc/conv_bias_relu/conv_bias_relu.cpp:95
↓ 4 callersFunctionget_cudnn_manager
Get the CuDNN front-end context manager. Returns: CuDNNManager: Global CuDNN manager.
apex/contrib/torchsched/ops/layer_norm.py:67
↓ 4 callersMethodget_v2_hw_c_list
()
apex/contrib/test/group_norm/test_group_norm.py:332
↓ 4 callersFunctiongroup_norm_nhwc_fprop
( x: torch.Tensor, G: int, weight: torch.Tensor, bias: torch.Tensor, eps: float, act:
apex/contrib/group_norm/group_norm.py:49
↓ 4 callersFunctioninter_block_sync
It is expected that all threads in the CTA enter this function!
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:651
↓ 4 callersFunctionldg_stream
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:107
↓ 4 callersFunctionlog2_ceil
csrc/megatron/scaled_masked_softmax.h:62
↓ 4 callersFunctionmake_params
Construct parameters with random configurations
apex/contrib/test/clip_grad/test_clip_grad.py:13
↓ 4 callersFunctionmultiply
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:311
↓ 4 callersMethodparam_sync
Ensure that all parameters are synchronized
apex/contrib/optimizers/distributed_fused_adam.py:2137
↓ 4 callersFunctionparse
()
examples/imagenet/main_amp.py:38
↓ 4 callersFunctionrun_dconv_drelu_dscale
apex/contrib/csrc/bottleneck/bottleneck.cpp:696
↓ 4 callersMethodsave_graph_to_json
This function is used to save the graph into JSON file for inspection.
apex/contrib/sparsity/permutation_lib.py:2004
↓ 4 callersFunctionstg_stream
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:154
↓ 4 callersMethodtest_checkpoint
Test state_dict and load_state_dict functions Two models are constructed, possibly on different process groups. One of the models is
apex/contrib/test/optimizers/test_dist_adam.py:492
↓ 4 callersFunctionto_python_float
(scalar_tensor: torch.Tensor)
examples/imagenet/main_amp.py:19
↓ 4 callersFunctionwrite_to_gmem
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:239
↓ 4 callersMethodzero_grad
(self)
apex/optimizers/fused_adam.py:139
↓ 3 callersMethod__init__
( self, normalized_shape, eps=1e-5, elementwise_affine=True, memory_ef
apex/normalization/fused_layer_norm.py:901
↓ 3 callersMethod_apply_state_scale
Compute and apply scaling factor for scaled optimizer state The scaling factor is chosen to maximize the dynamic range while avoiding
apex/contrib/optimizers/distributed_fused_adam.py:2833
↓ 3 callersMethod_check_params_shard_dtypes
Make sure local shards of parameters are in expected datatypes The Adam kernel only supports floating-point datatypes. If we want to
apex/contrib/optimizers/distributed_fused_adam.py:2776
↓ 3 callersFunction_coalescing_manager_append_work
Add asynchronous request to coalescing manager
apex/contrib/optimizers/distributed_fused_adam.py:103
↓ 3 callersMethod_finish_bucket_grad_sync
Wait for any gradient synchronizations that are in progress
apex/contrib/optimizers/distributed_fused_adam.py:1965
↓ 3 callersMethod_init_grad_buffer
Allocate contiguous buffer for grad buckets
apex/contrib/optimizers/distributed_fused_adam.py:1155
↓ 3 callersMethod_init_param_state
Initialize optimizer state for a parameter
apex/contrib/optimizers/distributed_fused_adam.py:1346
↓ 3 callersMethod_start_bucket_param_sync
Synchronize parameter buckets Parameter synchronization is asynchronous. Involves all-gather over distributed process group. Assumes
apex/contrib/optimizers/distributed_fused_adam.py:2031
↓ 3 callersFunctionbwd_dx
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:1449
↓ 3 callersFunctionbwd_update
apex/contrib/csrc/groupbn/nhwc_batch_norm_kernel.h:1435
↓ 3 callersMethodcodegen_buffers_record_stream
Generate data structure for recording steam on return tensors before program exit. Args: buffers: Names of buffers that need to b
apex/contrib/torchsched/inductor/wrapper.py:271
next →1–100 of 1,387, ranked by callers