MCPcopy Create free account

hub / github.com/JieShibo/MoLE / functions

Functions2,110 in github.com/JieShibo/MoLE

↓ 1 callersMethodset_input_tensor
See megatron.legacy.model.transformer.set_input_tensor()
pretrain/megatron/legacy/model/t5_model.py:103
↓ 1 callersMethodset_is_first_microbatch
Sets the is_first_microbatch flag if it exists. When this flag is set, TE modules will update their fp8 parameter cache.
pretrain/megatron/core/transformer/module.py:90
↓ 1 callersFunctionset_jit_fusion_options
Set PyTorch JIT layer fusion options.
pretrain/megatron/training/initialize.py:306
↓ 1 callersMethodset_layer_number
Set the layer number for the router.
pretrain/megatron/core/transformer/moe/router.py:87
↓ 1 callersMethodset_loss_scale
set the scale of the aux loss. Args: scale (torch.Tensor): The scale value to set. Please ensure that the scale passed in
pretrain/megatron/core/transformer/moe/moe_utils.py:97
↓ 1 callersMethodset_special_tokens
Add a list of additional tokens to the encoder. The additional tokens are indexed starting from the last index of the current
pretrain/megatron/training/tokenizer/gpt2_tokenization.py:181
↓ 1 callersFunctionsetup_model_and_optimizer
Setup model and optimizer.
pretrain/megatron/training/training.py:484
↓ 1 callersMethodsharded_param_state_dp_zero
Naive implementation which reuses gather/scatter from the legacy ckpt format. During saving, gathers the parameters state on DP rank 0 and sa
pretrain/megatron/core/optimizer/distrib_optimizer.py:852
↓ 1 callersMethodsharded_param_state_fs_bucket_space
Sharded state dict where each noncontiguous buffer is a separate ShardedTensor. Results in fully parallel save and load without any inter-pro
pretrain/megatron/core/optimizer/distrib_optimizer.py:881
↓ 1 callersMethodsharded_state_dict
( self, prefix: str = '', sharded_offsets: tuple = (), metadata: Optional[dict] = None )
pretrain/megatron/core/transformer/vanillamlp.py:128
↓ 1 callersMethodsharded_state_dict
(self, prefix='', sharded_offsets=(), metadata=None)
pretrain/megatron/core/transformer/moe/experts.py:167
↓ 1 callersFunctionsharded_tensor_to_torch_sharded_tensor
Convert MCore ShardedTensor to PyT ShardedTensor. PyT requires information about all chunks. NOTE: this function assumes regular (grid) sharding
pretrain/megatron/core/dist_checkpointing/strategies/torch.py:78
↓ 1 callersMethodsignals_received
(self)
pretrain/megatron/training/dist_signal_handler.py:54
↓ 1 callersFunctionsinkhorn
(cost, tol=0.0001)
pretrain/megatron/legacy/model/transformer.py:162
↓ 1 callersFunctionsinkhorn
Sinkhorn based MoE routing function
pretrain/megatron/core/transformer/moe/moe_utils.py:43
↓ 1 callersMethodsinkhorn_load_balancing
Apply sinkhorn routing to the logits tensor. Args: logits (torch.Tensor): The logits tensor. Returns: torch.
pretrain/megatron/core/transformer/moe/router.py:107
↓ 1 callersFunctionsplit_tensor_along_last_dim
Split a tensor along its last dimension. Args: tensor: input tensor. num_partitions: number of partitions to split t
pretrain/megatron/core/tensor_parallel/utils.py:11
↓ 1 callersFunctionsplit_tensor_into_1d_equal_chunks
Break a tensor into equal 1D chunks across tensor parallel ranks. Returns a Tensor or View with this rank's portion of the data. Ar
pretrain/megatron/core/tensor_parallel/utils.py:37
↓ 1 callersMethodstate_dict
(self)
pretrain/megatron/core/optimizer/optimizer.py:529
↓ 1 callersMethodswap_key_value_dict
swap between batches
pretrain/megatron/core/inference_params.py:12
↓ 1 callersFunctionswitch_load_balancing_loss_func
Calculate the auxiliary loss for better load balacing. Please refer to the Switch Transformer paper (https://arxiv.org/abs/2101.03961) for detail
pretrain/megatron/core/transformer/moe/moe_utils.py:8
↓ 1 callersFunctiont5_extended_attention_mask
(attention_mask_list)
pretrain/megatron/legacy/model/t5_model.py:19
↓ 1 callersFunctiont5_extended_attention_mask
(attention_mask_list: List[Tensor])
pretrain/megatron/core/models/T5/t5_model.py:413
↓ 1 callersFunctiontest_allmasked_softmax_backward
()
pretrain/megatron/legacy/fused_kernels/tests/test_fused_kernels.py:345
↓ 1 callersFunctiontest_allmasked_softmax_forward
()
pretrain/megatron/legacy/fused_kernels/tests/test_fused_kernels.py:329
↓ 1 callersFunctiontest_broadcast_data
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_data.py:14
↓ 1 callersFunctiontest_column_parallel_linear
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_layers.py:173
↓ 1 callersFunctiontest_cross_entropy
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_cross_entropy.py:45
↓ 1 callersFunctiontest_cuda_rng_tracker
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_random.py:73
↓ 1 callersFunctiontest_fused_softmax
()
pretrain/megatron/legacy/fused_kernels/tests/test_fused_kernels.py:24
↓ 1 callersFunctiontest_fused_upper_triangle_mask_softmax
()
pretrain/megatron/legacy/fused_kernels/tests/test_fused_kernels.py:122
↓ 1 callersFunctiontest_get_tensor_model_parallel_src_rank
(tensor_model_parallel_size_)
pretrain/megatron/legacy/mpu/tests/test_initialize.py:49
↓ 1 callersFunctiontest_initialize_affine_weight
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_layers.py:94
↓ 1 callersFunctiontest_initialize_model_parallel
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_initialize.py:11
↓ 1 callersFunctiontest_layer_norm
()
pretrain/megatron/legacy/fused_kernels/tests/test_fused_kernels.py:222
↓ 1 callersFunctiontest_load_fused_kernels
()
pretrain/megatron/legacy/fused_kernels/tests/test_fused_kernels.py:12
↓ 1 callersFunctiontest_masked_softmax_backward
()
pretrain/megatron/legacy/fused_kernels/tests/test_fused_kernels.py:308
↓ 1 callersFunctiontest_masked_softmax_forward
()
pretrain/megatron/legacy/fused_kernels/tests/test_fused_kernels.py:293
↓ 1 callersFunctiontest_model_parallel_cuda_manual_seed
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_random.py:144
↓ 1 callersFunctiontest_parallel_embedding
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_layers.py:16
↓ 1 callersFunctiontest_parallel_self_attention
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_layers.py:348
↓ 1 callersFunctiontest_parallel_transformer_layer
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_layers.py:436
↓ 1 callersFunctiontest_row_parallel_linear
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_layers.py:240
↓ 1 callersFunctiontest_set_cuda_rng_state
(tensor_model_parallel_size)
pretrain/megatron/legacy/mpu/tests/test_random.py:11
↓ 1 callersMethodtie_embeddings_and_output_weights_state_dict
Ties the embedding and output weights in a given sharded state dict. Args: sharded_state_dict (ShardedStateDict): state dict with
pretrain/megatron/core/models/common/language_module/language_module.py:160
↓ 1 callersMethodtoken_permutation
Dispatch tokens to experts. Args: tokens (torch.Tensor): Input tokens. indices (torch.Tensor): indices tensor.
pretrain/megatron/core/transformer/moe/token_dispatcher.py:26
↓ 1 callersMethodtoken_unpermutation
Restores the expert output to its original ordering. Args: expert_output (torch.Tensor): The output tensor from the expert models
pretrain/megatron/core/transformer/moe/token_dispatcher.py:41
↓ 1 callersMethodtokenize
Tokenize a string.
pretrain/megatron/training/tokenizer/gpt2_tokenization.py:236
↓ 1 callersFunctiontorch_cross_entropy
(batch_size, seq_length, vocab_size, logits_scale, seed)
pretrain/megatron/legacy/mpu/tests/test_cross_entropy.py:16
↓ 1 callersMethodtrack_and_report_progress
Utility function for tracking progress
pretrain/megatron/legacy/indexer.py:67
↓ 1 callersFunctiontrack_moe_metrics
( loss_scale, iteration, writer, wandb_writer=None, total_loss_dict=None, per_layer_logging=False )
pretrain/megatron/core/transformer/moe/moe_utils.py:193
↓ 1 callersFunctiontrain
Train the model function.
pretrain/megatron/training/training.py:876
↓ 1 callersMethodtrain
Train index. Args: config (RetroPreprocessingConfig): Retro preprocessing config.
pretrain/megatron/core/datasets/retro/index/indexes/faiss_base.py:77
↓ 1 callersFunctiontrain_index
Entry point for training the index. We select whether to train a new index, or validate an existing index. Args: config (RetroPrepro
pretrain/megatron/core/datasets/retro/index/build.py:207
↓ 1 callersFunctiontrain_on_embeddings
Train index on embedded DB chunks. Args: config (RetroPreprocessingConfig): Retro preprocessing config.
pretrain/megatron/core/datasets/retro/index/build.py:159
↓ 1 callersFunctiontrain_step
Single training step.
pretrain/megatron/training/training.py:528
↓ 1 callersFunctiontraining_log
Log training information such as losses, timing, ....
pretrain/megatron/training/training.py:594
↓ 1 callersMethodunwrap
(self)
pretrain/megatron/core/dist_checkpointing/mapping.py:224
↓ 1 callersMethodupdate
(self, consumed_samples, consistency_check)
pretrain/megatron/training/microbatches.py:74
↓ 1 callersMethodupdate
(self, consumed_samples, consistency_check)
pretrain/megatron/training/microbatches.py:127
↓ 1 callersMethodupdate_center
Update center used for teacher output.
pretrain/megatron/legacy/model/vision/dino.py:73
↓ 1 callersFunctionupdate_chunk_counts
Set n_chunks_train & n_chunks sampled for each individual DB. Args: config (RetroPreprocessingConfig): Retro preprocessing config.
pretrain/megatron/core/datasets/retro/db/build.py:396
↓ 1 callersMethodupdate_momentum
(self, iteration)
pretrain/megatron/legacy/model/vision/dino.py:286
↓ 1 callersFunctionupdate_train_iters
(args)
pretrain/megatron/training/training.py:302
↓ 1 callersFunctionvalidate
Validation method for validating saved neighbor IDs. Args: f (h5py.File): File containing save neighbor IDs.
pretrain/megatron/core/datasets/retro/query/query.py:300
↓ 1 callersFunctionvalidate_args
(args, defaults={})
pretrain/megatron/training/arguments.py:145
↓ 1 callersFunctionvalidate_yaml
(args, defaults={})
pretrain/megatron/training/yaml_arguments.py:40
↓ 1 callersFunctionvocab_parallel_cross_entropy
Performs cross entropy loss when logits are split across tensor parallel ranks Args: vocab_parallel_logits: logits split across tens
pretrain/megatron/core/tensor_parallel/cross_entropy.py:129
↓ 1 callersMethodvocab_range_from_global_vocab_size
( global_vocab_size: int, rank: int, world_size: int )
pretrain/megatron/core/tensor_parallel/utils.py:107
↓ 1 callersMethodvocab_range_from_per_partition_vocab_size
( per_partition_vocab_size: int, rank, world_size: int )
pretrain/megatron/core/tensor_parallel/utils.py:99
↓ 1 callersFunctionwindow_reverse
Args: windows: (num_windows*B, window_size, window_size, C) window_size (int): Window size H (int): Height of image
pretrain/megatron/legacy/model/vision/esvit_swin_backbone.py:59
↓ 1 callersFunctionwindow_reverse
Args: windows: (num_windows*B, window_size, window_size, C) window_size (int): Window size H (int): Height of image
pretrain/megatron/legacy/model/vision/swin_backbone.py:54
↓ 1 callersFunctionwrite_args_to_tensorboard
Write arguments to tensorboard.
pretrain/megatron/training/initialize.py:297
↓ 1 callersFunctionz_loss_func
Encourages the router's logits to remain small to enhance stability. Please refer to the ST-MoE paper (https://arxiv.org/pdf/2202.08906.pdf) for d
pretrain/megatron/core/transformer/moe/moe_utils.py:28
↓ 1 callersMethodzero_grad_buffer
Zeros out all grad buffers. Needs to be called at the beginning of each training iteration.
pretrain/megatron/core/distributed/distributed_data_parallel.py:244
↓ 1 callersMethodzero_parameters
Zero out all parameters in embedding.
pretrain/megatron/legacy/model/language_model.py:185
FunctionPYBIND11_MODULE
pretrain/megatron/core/datasets/helpers.cpp:759
Method__call__
(self, img)
pretrain/megatron/legacy/data/vit_dataset.py:24
Method__call__
(self, img)
pretrain/megatron/legacy/data/vit_dataset.py:43
Method__call__
(self, input)
pretrain/megatron/legacy/data/vit_dataset.py:74
Method__call__
(self, input)
pretrain/megatron/legacy/data/vit_dataset.py:140
Method__call__
(self, image)
pretrain/megatron/legacy/data/vit_dataset.py:206
Method__call__
Define call method for ImageNetPolicy class.
pretrain/megatron/legacy/data/autoaugment.py:103
Method__call__
Define call method for SubPolicy class.
pretrain/megatron/legacy/data/autoaugment.py:310
Method__call__
Call timer with name and log level.
pretrain/megatron/core/timers.py:171
Method__call__
Invocation of the forward methods. Note that self.inference_params is being modified by the forward step.
pretrain/megatron/inference/text_generation/forward_step.py:40
Method__del__
Clean up the object
pretrain/megatron/core/datasets/indexed_dataset.py:302
Method__del__
Clean up the object
pretrain/megatron/core/datasets/indexed_dataset.py:400
Method__exit__
(self, type, value, tb)
pretrain/megatron/training/dist_signal_handler.py:72
Method__getitem__
Get an ICT example of a pseudo-query and the block of text from which it was extracted
pretrain/megatron/legacy/data/ict_dataset.py:77
Method__getitem__
(self, idx)
pretrain/megatron/legacy/data/orqa_wiki_dataset.py:150
Method__getitem__
(self, idx)
pretrain/megatron/legacy/data/multimodal_dataset.py:41
Method__getitem__
Args: index (int): Index Returns: tuple: (sample, target) where target is class_index of the target class.
pretrain/megatron/legacy/data/image_folder.py:207
Method__getitem__
Get the data associated with an indexed sample.
pretrain/megatron/legacy/data/realm_dataset_utils.py:105
Method__getitem__
(self, idx)
pretrain/megatron/legacy/data/data_samplers.py:117
Method__getitem__
Get the data associated with an indexed sample.
pretrain/megatron/legacy/data/biencoder_dataset_utils.py:115
Method__getitem__
Return a sample that contains a dummy image, text sequence and the associated labels and cost and attention masks. Args: idx (int
pretrain/megatron/core/datasets/multimodal_dataset.py:42
Method__getitem__
Abstract method implementation Args: idx (int): The index into the dataset Returns: Dict[str, Union[int, nu
pretrain/megatron/core/datasets/t5_dataset.py:93
Method__getitem__
(self, idx: int)
pretrain/megatron/core/datasets/blended_dataset.py:91
Method__getitem__
Abstract method implementation Args: idx (int): The index into the dataset Returns: Dict[str, Union[int, nu
pretrain/megatron/core/datasets/bert_dataset.py:81
← previousnext →901–1,000 of 2,110, ranked by callers