MCPcopy Create free account

hub / github.com/OpenSparseLLMs/Linear-MoE / functions

Functions1,498 in github.com/OpenSparseLLMs/Linear-MoE

Method_collate
(x)
linear_moe/lm_evaluate.py:151
Method_compare
(srcs, tgts, names, parallelism)
linear_moe/model/mixtral/transformer/attention.py:429
Method_compare
(srcs, tgts, names, parallelism)
linear_moe/model/qwen2/transformer/attention.py:429
Method_compare
(srcs, tgts, names, parallelism)
linear_moe/model/llama3/transformer/attention.py:425
Method_convert_id_to_token
Converts an id to a token, special tokens included
linear_moe/tokenizer/tokenization_qwen_vl.py:290
Method_convert_id_to_token
Converts an index (integer) in a token (str) using the vocab.
linear_moe/tokenizer/tokenization_baichuan.py:101
Method_convert_id_to_token
Converts an index (integer) in a token (str) using the vocab.
linear_moe/tokenizer/tokenization_yi.py:114
Method_convert_token_to_id
Converts a token to an id using the vocab, special tokens included
linear_moe/tokenizer/tokenization_qwen_vl.py:296
Method_convert_token_to_id
Converts a token (str) in an id using the vocab.
linear_moe/tokenizer/tokenization_baichuan.py:97
Method_convert_token_to_id
Converts a token (str) in an id using the vocab.
linear_moe/tokenizer/tokenization_yi.py:110
Method_create_model
( self, pretrained: str, **kwargs, )
linear_moe/lm_evaluate.py:63
Function_cross_entropy_forward_step
Simple forward step with cross-entropy loss.
linear_moe/finetune_utils.py:63
Method_decode
( self, token_ids: Union[int, List[int]], skip_special_tokens: bool = False, e
linear_moe/tokenizer/tokenization_qwen_vl.py:313
Method_decode_imgurl
(img_token_ids)
linear_moe/tokenizer/tokenization_qwen_vl.py:323
Method_encode_imgurl
(img_tokens)
linear_moe/tokenizer/tokenization_qwen_vl.py:252
Method_encode_vl_info
(tokens)
linear_moe/tokenizer/tokenization_qwen_vl.py:341
Function_fwd_inter_kernel
( Q, Out, KV, n: tl.constexpr, d: tl.constexpr, e: tl.constexpr, BLOCK: tl.constex
linear_moe/sequence_modeling/lasp2/lasp2_with_mask_triton_op.py:245
Function_fwd_intra_kernel
( Q, K, V, Out, n: tl.constexpr, d: tl.constexpr, e: tl.constexpr, BLOCK: tl.c
linear_moe/sequence_modeling/lasp2/lasp2_with_mask_triton_op.py:156
Function_fwd_kernel
( Q, Out, KV, n: tl.constexpr, d: tl.constexpr, e: tl.constexpr, BLOCK: tl.constex
linear_moe/sequence_modeling/lasp2/lasp2_without_mask_triton_op.py:156
Function_fwd_m_calculate
( K, V, KV, n: tl.constexpr, d: tl.constexpr, e: tl.constexpr, BLOCK: tl.constexpr
linear_moe/sequence_modeling/lasp2/lasp2_with_mask_triton_op.py:11
Function_fwd_m_calculate
( K, V, KV, n: tl.constexpr, d: tl.constexpr, e: tl.constexpr, BLOCK: tl.constexpr
linear_moe/sequence_modeling/lasp2/lasp2_without_mask_triton_op.py:11
Function_fwd_m_cumsum
( KV, d: tl.constexpr, e: tl.constexpr, NUM_BLOCK: tl.constexpr, D_FBLOCK: tl.constexpr,
linear_moe/sequence_modeling/lasp2/lasp2_with_mask_triton_op.py:80
Function_fwd_m_cumsum
( KV, d: tl.constexpr, e: tl.constexpr, NUM_BLOCK: tl.constexpr, D_FBLOCK: tl.constexpr,
linear_moe/sequence_modeling/lasp2/lasp2_without_mask_triton_op.py:80
Function_fwd_m_update
( KV, GKV, d: tl.constexpr, e: tl.constexpr, NUM_BLOCK: tl.constexpr, D_FBLOCK: tl.con
linear_moe/sequence_modeling/lasp2/lasp2_with_mask_triton_op.py:114
Function_fwd_m_update
( KV, GKV, d: tl.constexpr, e: tl.constexpr, NUM_BLOCK: tl.constexpr, D_FBLOCK: tl.con
linear_moe/sequence_modeling/lasp2/lasp2_without_mask_triton_op.py:114
Function_init_weights
( module, n_layer, initializer_range=0.02, # Now only used for embedding layer. rescale_preno
linear_moe/sequence_modeling/ssm.py:28
Function_init_weights
( module, n_layer, initializer_range=0.02, # Now only used for embedding layer. rescale_preno
linear_moe/sequence_modeling/mamba2/mamba_block.py:28
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/hgrn2/hgrn2.py:57
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/rwkv7/rwkv7.py:65
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/gated_deltanet/gated_deltanet.py:79
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/based/based.py:47
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/deltanet/deltanet.py:86
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/rwkv6/rwkv6.py:61
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/basic_linear_attention/basic_linear_attention.py:104
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/rebased/rebased.py:53
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/lasp2/lasp2.py:53
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/gla/gla.py:63
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/lightning_attention/lightning_attention.py:70
Method_initialize_weights
(self, module: torch.nn.Module)
linear_moe/sequence_modeling/retention/retention.py:68
Function_l2_norm_bwd_kernel
( X, # pointer to the input # Y, # pointer to the output to be recomputed DY, # pointer to the o
linear_moe/model/common_modules/l2norm.py:64
Function_l2_norm_fwd_1pass_kernel
( X, # pointer to the input Y, # pointer to the output stride_x_row, # how much to increase the
linear_moe/model/common_modules/l2norm.py:22
Function_layer_norm_bwd_kernel
( X, # pointer to the input W, # pointer to the weights B, # pointer to the biases Y, # po
linear_moe/model/common_modules/layernorm.py:212
Function_layer_norm_fwd_1pass_kernel
( X, # pointer to the input Y, # pointer to the output W, # pointer to the weights B, # po
linear_moe/model/common_modules/layernorm.py:69
Method_set_cos_sin_cache
(self, seq_len, device, dtype)
linear_moe/model/deepseek_v2/yarn_rotary_pos_embedding.py:120
Method_sinkhorn_activation
(logits)
linear_moe/model/mixtral/moe/router.py:129
Method_sinkhorn_activation
(logits)
linear_moe/model/qwen2/moe/router.py:204
Method_sinkhorn_activation
(logits)
linear_moe/model/deepseek_v2/moe/router_old.py:130
Method_sinkhorn_activation
(logits)
linear_moe/model/deepseek_v2/moe/router.py:204
Method_tokenize
Converts a string in a sequence of tokens (string), using the tokenizer. Split in words for word-based vocabulary or sub-words for su
linear_moe/tokenizer/tokenization_qwen_vl.py:304
Method_tokenize
Returns a tokenized string.
linear_moe/tokenizer/tokenization_baichuan.py:93
Functionadd_ckpt_args
(parser)
toolkits/model_checkpoints_convertor/llama/hf2mcore.py:726
Functionadd_ckpt_args
(parser)
toolkits/model_checkpoints_convertor/mistral/hf2mcore.py:522
Functionadd_ckpt_args
(parser)
toolkits/model_checkpoints_convertor/qwen/hf2megablocks_qwen1.5.py:605
Functionadd_extra_args
(parser)
toolkits/model_checkpoints_convertor/deepseek/hf2mcore_deepseek_v2_moe.py:419
Functionadd_extra_args
(parser)
toolkits/model_checkpoints_convertor/qwen/hf2mcore_qwen1.5_moe.py:399
Functionadd_extra_args
(parser)
toolkits/model_checkpoints_convertor/qwen/hf2mcore_qwen1.5_dense_mha_to_moe.py:225
Functionadd_extra_args
(parser)
toolkits/model_checkpoints_convertor/qwen/hf2mcore_qwen2_dense_and_moe_gqa.py:879
Functionadd_extra_args
(parser)
toolkits/model_checkpoints_convertor/qwen/hf2mcore_qwen1.5_dense_mha.py:291
Methodadd_tokentype_embeddings
Add token-type embedding. This function is provided so we can add token-type embeddings in case the pretrained model does not have it.
linear_moe/model/llama3/language_model.py:195
Methodallocate_inference_cache
(self, batch_size, max_seqlen, dtype=None)
linear_moe/sequence_modeling/ssm.py:193
Methodallocate_inference_cache
(self, batch_size, max_seqlen, dtype=None)
linear_moe/sequence_modeling/mamba2/mamba_block.py:171
Methodallocate_inference_cache
(self, batch_size, max_seqlen, dtype=None)
linear_moe/sequence_modeling/mamba2/mamba_mixer.py:475
Methodallocate_inference_cache
(self, batch_size, max_seqlen, dtype=None)
linear_moe/sequence_modeling/mamba2/mamba_layer.py:79
Methodappend_message
(self, role, message)
linear_moe/data/llava/conversation.py:120
Functionapply_rotary_emb
Arguments: x: (batch_size, seqlen, nheads, headdim) if cu_seqlens is None else (total_seqlen, nheads, headdim) cos, s
linear_moe/model/common_modules/rotary.py:100
Functionapply_rotary_emb_torch
x: (batch_size, seqlen, nheads, headdim) cos, sin: (seqlen, rotary_dim / 2) or (batch_size, seqlen, rotary_dim / 2)
linear_moe/model/common_modules/rotary.py:22
Functionassert_grouped_gemm_is_available
()
linear_moe/model/mixtral/moe/grouped_gemm_util.py:23
Methodbackward
(ctx, do)
linear_moe/sequence_modeling/lasp2/lasp2_with_mask_triton_op.py:997
Methodbackward
(ctx, do)
linear_moe/sequence_modeling/lasp2/lasp2_without_mask_triton_op.py:719
Methodbackward
Compute and scale the gradient for auxiliary loss.. Args: grad_output (torch.Tensor): The gradient of the output. Return
linear_moe/model/mixtral/moe/moe_utils.py:97
Methodbackward
(ctx, dy, *args)
linear_moe/model/common_modules/layernorm.py:432
Methodbackward
(ctx, dout, *args)
linear_moe/model/common_modules/layernorm.py:730
Methodbackward
(ctx, dy, *args)
linear_moe/model/common_modules/l2norm.py:188
Methodbackward
(ctx, dout)
linear_moe/model/common_modules/activations.py:36
Methodbackward
(ctx, dy)
linear_moe/model/common_modules/activations.py:144
Methodbackward
(ctx, dout)
linear_moe/model/common_modules/activations.py:179
Methodbackward
(ctx, grad_output)
linear_moe/model/common_modules/activations.py:225
Methodbackward
(ctx, grad_output)
linear_moe/model/common_modules/activations.py:263
Methodbackward
(ctx, grad_output)
linear_moe/model/common_modules/activations.py:296
Methodbackward
(ctx, dout)
linear_moe/model/common_modules/activations.py:346
Methodbackward
(ctx, dout, *args)
linear_moe/model/common_modules/activations.py:371
Methodbackward
(ctx, do)
linear_moe/model/common_modules/rotary.py:76
Functionbeam_search_and_post_process
Run beam search and post-process outputs, i.e., detokenize, move to cpu and convert to list. Args: model (torch.nn.Module): The m
linear_moe/generation/api.py:214
Functionbias_dropout_add_fused_inference
(x: torch.Tensor, bias: Optional[torch.Tensor],
linear_moe/model/llama3/transformer_legacy.py:856
Functionbias_dropout_add_fused_train
(x: torch.Tensor, bias: Optional[torch.Tensor],
linear_moe/model/llama3/transformer_legacy.py:848
Functionbuild_finetune_dataset
(dataset)
linear_moe/data/__init__.py:39
Methodbuild_inputs_with_special_tokens
(self, token_ids_0, token_ids_1=None)
linear_moe/tokenizer/tokenization_baichuan.py:152
Methodbuild_layer
(layer_spec, layer_number)
linear_moe/model/mixtral/transformer_block.py:164
Methodbuild_layer
(layer_spec, layer_number)
linear_moe/model/qwen2/transformer_block.py:164
Methodbuild_layer
(layer_spec, layer_number)
linear_moe/model/deepseek_v2/transformer_block.py:160
Methodbuild_layer
(layer_number)
linear_moe/model/llama3/transformer_legacy.py:1495
Functionbuild_pretrain_dataset_from_idxmap
Build train, valid, and test datasets for pretraining a LLAMA model on mmap format data. Args: data_prefix (str): common prefix added
linear_moe/data/__init__.py:116
Functioncheck_mg_eg_forward
(mgmodel, hgmodel, mgargs)
toolkits/model_checkpoints_convertor/llama/hf2mcore.py:596
Functioncheck_tokenizer_is_same
(hgtokenizer, mgtokenizer)
toolkits/model_checkpoints_convertor/llama/hf2mcore.py:708
Functioncheckpoint
(func)
linear_moe/utils.py:17
Methodcheckpoint_handler
(forward_func)
linear_moe/model/mixtral/transformer_block.py:244
Methodcheckpoint_handler
(forward_func)
linear_moe/model/mixtral/hybrid/hybrid_transformer_block.py:236
Methodcheckpoint_handler
(forward_func)
linear_moe/model/qwen2/transformer_block.py:244
Methodcheckpoint_handler
(forward_func)
linear_moe/model/qwen2/hybrid/hybrid_transformer_block.py:236
Methodcheckpoint_handler
(forward_func)
linear_moe/model/deepseek_v2/transformer_block.py:235
← previousnext →901–1,000 of 1,498, ranked by callers