MCPcopy Create free account

hub / github.com/NVIDIA/apex / types & classes

Types & classes224 in github.com/NVIDIA/apex

↓ 5 callersClassMLP
Launch MLP in C++ Args: mlp_sizes (list of int): MLP sizes. Example: [1024,1024,1024] will create 2 MLP layers with shape 1024x1024
apex/mlp/mlp.py:33
↓ 4 callersClassFusedAdam
Implements Adam algorithm. Currently GPU-only. Requires Apex to be installed via ``python setup.py install --cuda_ext --cpp_ext``. It has be
apex/contrib/optimizers/fused_adam.py:9
↓ 4 callersClassFusedLayerNorm
r"""Applies Layer Normalization over a mini-batch of inputs as described in the paper `Layer Normalization`_ . Currently only runs on cuda()
apex/normalization/fused_layer_norm.py:724
↓ 4 callersClassFusedRMSNorm
r"""Applies RMS Normalization over a mini-batch of inputs Currently only runs on cuda() tensors. .. math:: y = \frac{x}{\mathrm{RMS}
apex/normalization/fused_layer_norm.py:841
↓ 4 callersClassPeerMemoryPool
apex/contrib/peer_memory/peer_memory.py:6
↓ 3 callersClassGroupNorm
Optimized GroupNorm for NHWC layout with optional Swish/SiLU fusion. There are two version of CUDA kernels under the hood: one pass and two p
apex/contrib/group_norm/group_norm.py:210
↓ 2 callersClassAverageMeter
Computes and stores the average and current value
tests/L1/common/main_amp.py:553
↓ 2 callersClassAverageMeter
Computes and stores the average and current value
examples/imagenet/main_amp.py:452
↓ 2 callersClassBottleneck
apex/contrib/bottleneck/bottleneck.py:153
↓ 2 callersClassCudaEventSym
Symbolic representation of CUDA Events in the Inductor scheduling phase. Args: factory: The CUDAEventFactory that generate this event.
apex/contrib/torchsched/inductor/event.py:31
↓ 2 callersClassHaloExchangerPeer
apex/contrib/bottleneck/halo_exchangers.py:146
↓ 2 callersClassMixedFusedLayerNorm
apex/normalization/fused_layer_norm.py:959
↓ 2 callersClassMixedFusedRMSNorm
apex/normalization/fused_layer_norm.py:1000
↓ 2 callersClassPeerHaloExchanger1d
apex/contrib/peer_memory/peer_halo_exchanger_1d.py:5
↓ 2 callersClassdata_prefetcher
tests/L1/common/main_amp.py:352
↓ 2 callersClassdata_prefetcher
examples/imagenet/main_amp.py:244
↓ 1 callersClassAlphaFoldSWA
AlphaFold SWA (Stochastic Weight Averaging) module wrapper.
apex/contrib/test/openfold_triton/test_fused_adam_swa.py:31
↓ 1 callersClassArgs
apex/contrib/sparsity/test/toy_problem.py:96
↓ 1 callersClassArgs
apex/contrib/sparsity/test/checkpointing_test_part1.py:104
↓ 1 callersClassArgs
apex/contrib/sparsity/test/checkpointing_test_reference.py:103
↓ 1 callersClassArgs
apex/contrib/sparsity/test/checkpointing_test_part2.py:86
↓ 1 callersClassBNModel
apex/contrib/test/cudnn_gbn/test_cudnn_gbn_with_two_gpus.py:59
↓ 1 callersClassBNModelRef
apex/contrib/test/cudnn_gbn/test_cudnn_gbn_with_two_gpus.py:39
↓ 1 callersClassCUDAStreamPool
A pool managing reusable CUDA streams to optimize GPU operations. Attributes: pool_size (int): The maximum number of CUDA streams managed
apex/contrib/torchsched/inductor/_utils.py:43
↓ 1 callersClassCuDNNManager
CuDNN fronted context manager. Notice: CuDNN handle must be created after distributed process group initialization.
apex/contrib/torchsched/ops/layer_norm.py:18
↓ 1 callersClassCudaEventFactory
A factory that managements CUDA event creations and materializations. This factory maintains internal states to ensure that created cuda events g
apex/contrib/torchsched/inductor/event.py:158
↓ 1 callersClassDecompositionsWrapper
A wrapper class for handling decompositions in model compilation. This class extends the `_TorchCompileInductorWrapper` to include additional
apex/contrib/torchsched/backend.py:177
↓ 1 callersClassDiscriminator
examples/dcgan/main_amp.py:164
↓ 1 callersClassDistributedFusedAdam
r"""Adam optimizer with ZeRO algorithm. Currently GPU-only. Requires Apex to be installed via ``python setup.py install --cuda_ext --cpp_ext
apex/contrib/optimizers/distributed_fused_adam.py:269
↓ 1 callersClassDistributedFusedLAMB
Implements LAMB algorithm. Currently GPU-only. Requires Apex to be installed via ``pip install -v --no-cache-dir --global-option="--cpp_ext"
apex/contrib/optimizers/distributed_fused_lamb.py:27
↓ 1 callersClassEnterCudaStreamContextLine
Enter a context executed by respective CUDA Stream and insert necessary syncs. Attributes: wrapper: The code-gen wrapper of the current c
apex/contrib/torchsched/inductor/wrapper.py:88
↓ 1 callersClassEnterDeviceContextManagerWithStreamInfoLine
Enter a CUDA device context and allocate required side streams. Note: - The number of allocated streams is controlled by :attr:`torchsche
apex/contrib/torchsched/inductor/wrapper.py:43
↓ 1 callersClassExitCudaStreamContextLine
Generate code to exit the current stream context. Note: Most attributes and checking logics of this class have been moved to :met
apex/contrib/torchsched/inductor/wrapper.py:126
↓ 1 callersClassExitDeviceContextManagerWithStreamInfoLine
Exit a CUDA device context and release allocated streams.
apex/contrib/torchsched/inductor/wrapper.py:74
↓ 1 callersClassFastLayerNorm
apex/contrib/layer_norm/layer_norm.py:45
↓ 1 callersClassFusedAdamSWA
apex/contrib/openfold_triton/fused_adam_swa.py:210
↓ 1 callersClassGPUTimer
apex/contrib/test/layer_norm/test_fast_layer_norm.py:15
↓ 1 callersClassGenerator
examples/dcgan/main_amp.py:122
↓ 1 callersClassMHA_test
MultiheadAttention modules are unique, we need to check permutations for input and ouput projections
apex/contrib/sparsity/test/test_permutation_application.py:550
↓ 1 callersClassModel
tests/distributed/DDP/ddp_race_condition_test.py:26
↓ 1 callersClassModel
tests/L0/run_optimizers/test_adam.py:16
↓ 1 callersClassModelFoo
apex/contrib/test/optimizers/test_distributed_fused_lamb.py:29
↓ 1 callersClassMultiStreamWrapperCodegen
Wrapper code generator for graph scheduling.
apex/contrib/torchsched/inductor/wrapper.py:141
↓ 1 callersClassMultiTensorApply
apex/multi_tensor_apply/multi_tensor_apply.py:1
↓ 1 callersClassSimpleModel
apex/contrib/test/optimizers/test_dist_adam.py:19
↓ 1 callersClassSpatialBottleneck
apex/contrib/bottleneck/bottleneck.py:830
↓ 1 callersClassToyModel
apex/contrib/examples/nccl_allocator/toy_ddp.py:12
↓ 1 callersClassTransducerJoint
Transducer joint Detail of this loss function can be found in: Sequence Transduction with Recurrent Neural Networks Arguments:
apex/contrib/transducer/transducer.py:6
↓ 1 callersClassTransducerLoss
Transducer loss Detail of this loss function can be found in: Sequence Transduction with Recurrent Neural Networks Arguments:
apex/contrib/transducer/transducer.py:88
↓ 1 callersClass_CoalescingManager
apex/contrib/optimizers/distributed_fused_adam.py:74
↓ 1 callersClass_CudaEventRecordLine
apex/contrib/torchsched/inductor/event.py:128
↓ 1 callersClass_CudaEventWaitLine
apex/contrib/torchsched/inductor/event.py:142
↓ 1 callersClassalready_sparse
if weights are already sparse, permutations should be skipped
apex/contrib/sparsity/test/test_permutation_application.py:761
↓ 1 callersClassconv_1d
1D convolutions in isolation and with siblings
apex/contrib/sparsity/test/test_permutation_application.py:105
↓ 1 callersClasscoparent_poison
A single coparent that cannot permute along K poisons all other coparents in its group
apex/contrib/sparsity/test/test_permutation_application.py:387
↓ 1 callersClassdepthwise_child_is_sibling
The child of a depthwise convolution should act as a sibling
apex/contrib/sparsity/test/test_permutation_application.py:427
↓ 1 callersClassdifferent_grouped_convs
Convolutions with different group sizes need to use the GCD of the input channel counts if siblings
apex/contrib/sparsity/test/test_permutation_application.py:297
↓ 1 callersClassgrouped_convs
Stack of 2d convolutions with different types of grouped convolutions
apex/contrib/sparsity/test/test_permutation_application.py:145
↓ 1 callersClassmodule_attribute
Attributes of some module must be permuted if they feed some operation that is permuted
apex/contrib/sparsity/test/test_permutation_application.py:470
↓ 1 callersClassone_sparse_sibling
If only one of two siblings is sparse, both need to be permuted
apex/contrib/sparsity/test/test_permutation_application.py:576
↓ 1 callersClasssiblings_poison
A single sibling that cannot permute along C poisons all other siblings in its group
apex/contrib/sparsity/test/test_permutation_application.py:347
↓ 1 callersClasssimple_convs
Stack of 2d convolutions with different normalization and activation functions
apex/contrib/sparsity/test/test_permutation_application.py:22
↓ 1 callersClasssimple_forks_joins
Some simple residual connections to test collecting parameters into a single group. Four sections: input, blocka + residual, blockb + blockc, output
apex/contrib/sparsity/test/test_permutation_application.py:213
↓ 1 callersClasssquare_attribute
Attributes with multiple dimensions matching the permutation length should only be permuted along the correct dimension
apex/contrib/sparsity/test/test_permutation_application.py:521
↓ 1 callersClassswa_avg_fn
Averaging function for EMA with configurable decay rate (Supplementary '1.11.7 Evaluator setup').
apex/contrib/test/openfold_triton/test_fused_adam_swa.py:56
↓ 1 callersClasstest_concat
If concats are along the channel dimension (dim1 of NCHW), downstream layers can still be permuted despite C!=parentK
apex/contrib/sparsity/test/test_permutation_application.py:616
↓ 1 callersClasstest_flatten_module
flatten modules may change the effective channel count, typically by collapsing N,C,H,W into N,C*H*W before a classifier
apex/contrib/sparsity/test/test_permutation_application.py:709
↓ 1 callersClasstest_flatten_op
flatten ops may change the effective channel count, typically by collapsing N,C,H,W into N,C*H*W before a classifier
apex/contrib/sparsity/test/test_permutation_application.py:680
↓ 1 callersClasstest_trace_failure
make sure tracing failures are handled gracefully
apex/contrib/sparsity/test/test_permutation_application.py:739
ClassASP
apex/contrib/sparsity/asp.py:29
ClassAdamMathType
apex/contrib/openfold_triton/fused_adam_swa.py:48
ClassAdamTest
tests/L0/run_optimizers/test_adam.py:50
ClassAdd
csrc/megatron/scaled_upper_triang_masked_softmax.h:92
ClassAdd
csrc/megatron/scaled_masked_softmax.h:69
ClassAdd
csrc/megatron/generic_scaled_masked_softmax.h:30
ClassAtomicCounter
apex/contrib/optimizers/distributed_fused_lamb.py:87
ClassBatchNorm2d_NHWC
apex/contrib/groupbn/batch_norm.py:289
ClassBf16
apex/contrib/csrc/group_norm/traits.h:60
ClassBf16IOBf16W
apex/contrib/csrc/group_norm/traits.h:137
ClassBf16IOFp16W
///////////////////////////////////////////////////////////////////////////////////////////////
apex/contrib/csrc/group_norm/traits.h:130
ClassBf16IOFp32W
apex/contrib/csrc/group_norm/traits.h:144
ClassBottleneckFunction
apex/contrib/bottleneck/bottleneck.py:79
ClassBuildExtensionSeparateDir
setup.py:867
ClassBwdParams
apex/contrib/csrc/layer_norm/ln.h:70
ClassBwdRegistrar
apex/contrib/csrc/layer_norm/ln.h:176
ClassClipGradNormTest
apex/contrib/test/clip_grad/test_clip_grad.py:46
ClassComputeType2Key
apex/contrib/csrc/layer_norm/ln.h:149
ClassConvBiasMaskReLU_
apex/contrib/conv_bias_relu/conv_bias_relu.py:31
ClassConvBiasReLU_
apex/contrib/conv_bias_relu/conv_bias_relu.py:9
ClassConvBias_
apex/contrib/conv_bias_relu/conv_bias_relu.py:53
ClassConvFrozenScaleBiasReLU_
apex/contrib/conv_bias_relu/conv_bias_relu.py:75
ClassDenseNoBiasFunc
apex/fused_dense/fused_dense.py:25
ClassDeprecatedFeatureWarning
apex/__init__.py:31
ClassDistributedTestBase
apex/distributed_testing/distributed_test_base.py:23
ClassFP16_Optimizer
:class:`FP16_Optimizer` A cutdown version of apex.fp16_utils.FP16_Optimizer. Designed only to wrap apex.contrib.optimizers.FusedAdam, FusedSG
apex/contrib/optimizers/fp16_optimizer.py:6
ClassFastLayerNormFN
apex/contrib/layer_norm/layer_norm.py:8
ClassFocalLoss
apex/contrib/focal_loss/focal_loss.py:5
ClassFocalLossTest
apex/contrib/test/focal_loss/test_focal_loss.py:24
ClassFp16
apex/contrib/csrc/group_norm/traits.h:34
ClassFp16IOBf16W
apex/contrib/csrc/group_norm/traits.h:115
next →1–100 of 224, ranked by callers