MCPcopy Create free account

hub / github.com/VITA-MLLM/VITA-Audio / functions

Functions566 in github.com/VITA-MLLM/VITA-Audio

↓ 88 callersMethodget
Retrieve a specified number of elements from the buffer in a thread-safe manner. Parameters: - length (int): The number of e
web/queue.py:42
↓ 81 callersMethodsize
Get the current size of the queue in a thread-safe manner. Parameters: - None Returns: - int: The number of
web/queue.py:156
↓ 57 callersMethodfrom_pretrained
(model:str=None, **kwargs)
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:729
↓ 32 callersMethoddecode
(self, audio_tokens, **kwargs)
vita_audio/tokenizer_snac.py:118
↓ 23 callersMethodencode
(self, audio_path, **kwargs)
vita_audio/tokenizer_snac.py:99
↓ 17 callersMethodapply_to_role
(self, role, **kwargs)
vita_audio/tokenizer_snac.py:138
↓ 11 callersFunctionget_audio_tokenizer
(model_name_or_path, model_type, flow_path=None, rank=None)
vita_audio/tokenizer.py:81
↓ 9 callersMethod__init__
(self, config)
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:90
↓ 9 callersMethod__init__
(self, config)
vita_audio/models/qwen2_v4_48_3/modeling_qwen2.py:45
↓ 9 callersMethod__init__
(self, config)
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:87
↓ 7 callersFunctionadd_audio_input_contiguous
(input_ids, audio_paths, tokenizer, audio_tokenizer)
vita_audio/data/processor/audio_processor.py:87
↓ 7 callersMethodkeys
(self)
evaluation/compute-cer.py:246
↓ 7 callersMethodprocess_audios
(self, audio_path, is_discrete=False, is_contiguous=False, **kwargs)
vita_audio/data/processor/audio_processor.py:35
↓ 6 callersMethod__init__
(self, *args, **kwargs)
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:344
↓ 6 callersMethodencode
Frontend + Encoder. Note that this method is used by asr_inference.py Args: speech: (Batch, Length, ...) speec
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:787
↓ 6 callersMethodprocess_video
(self, video_file_or_dir, max_num_frame=8, max_fps=1)
vita_audio/data/processor/image_processor.py:135
↓ 5 callersMethodget_input_embeddings
(self)
vita_audio/models/qwen2_v4_48_3/modeling_qwen2.py:495
↓ 5 callersMethodload_model
(self)
vita_audio/data/processor/audio_processor.py:31
↓ 5 callersMethodprocess_images_with_subpatch
(self, img_or_path)
vita_audio/data/processor/image_processor.py:224
↓ 5 callersMethodrun_infer
( self, audio_path=None, prompt_audio_path=None, stream_stride=4, max_
tools/inference_sts.py:453
↓ 4 callersMethodget_vocab
(self)
vita_audio/models/qwen2_v4_48_3/tokenization_qwen2.py:215
↓ 4 callersMethodinterrupt
(self)
web/parms.py:74
↓ 4 callersMethodmaybe_init_ret
(self, source, force=False)
vita_audio/data/dataset_deepseek.py:45
↓ 4 callersMethodmaybe_init_ret
(self, source, force=False)
vita_audio/data/dataset_hunyuan.py:45
↓ 4 callersMethodmaybe_init_ret
(self, source, force=False)
vita_audio/data/dataset_cosyvoice2.py:45
↓ 4 callersMethodmaybe_init_ret
(self, source, force=False)
vita_audio/data/dataset_llama3.py:45
↓ 4 callersMethodmaybe_init_ret
(self, source, force=False)
vita_audio/data/dataset_qwen2.py:45
↓ 4 callersMethodprocess_images
(self, img_or_path_list)
vita_audio/data/processor/image_processor.py:179
↓ 4 callersMethodrun_infer_stream
( self, audio_path=None, prompt_audio_path=None, stream_stride=4, max_
tools/inference_sts.py:580
↓ 3 callersMethod_prepare_mtp_for_generation
( self, mtp_inference_mode, max_new_tokens, )
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:1243
↓ 3 callersMethodadd_ret
(self, ret, source)
vita_audio/data/dataset_deepseek.py:87
↓ 3 callersMethodadd_ret
(self, ret, source)
vita_audio/data/dataset_hunyuan.py:87
↓ 3 callersMethodadd_ret
(self, ret, source)
vita_audio/data/dataset_cosyvoice2.py:87
↓ 3 callersMethodadd_ret
(self, ret, source)
vita_audio/data/dataset_llama3.py:87
↓ 3 callersMethodadd_ret
(self, ret, source)
vita_audio/data/dataset_qwen2.py:87
↓ 3 callersFunctionextract_token_ids_as_int
(text)
tools/inference_sts.py:174
↓ 3 callersFunctionis_wav
(file_path)
web_demo.py:26
↓ 3 callersMethodoutput_size
(self)
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:619
↓ 3 callersMethodreset
(self)
web/parms.py:58
↓ 3 callersMethodreset_states
(self)
web/vad.py:56
↓ 3 callersFunctionresize
(img_or_path: str, size: Tuple[int, int], format="JPEG")
vita_audio/data/utils.py:27
↓ 2 callersMethod__init__
( self, model_name_or_path, audio_tokenizer_path, audio_tokenizer_type, flow_path=None )
tools/inference_sts.py:191
↓ 2 callersMethod__init__
(self, config)
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modular_qwen2.py:31
↓ 2 callersMethod__init__
(self, config)
vita_audio/models/qwen2_v4_48_3/modular_qwen2.py:31
↓ 2 callersMethod__init__
(self, config)
vita_audio/models/qwen2_mtp_v4_48_3/modular_qwen2.py:31
↓ 2 callersFunctionapply_rotary_pos_emb
Applies Rotary Position Embedding to the query and key tensors. Args: q (`torch.Tensor`): The query tensor. k (`torch.Tensor`): T
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:112
↓ 2 callersFunctionapply_rotary_pos_emb
Applies Rotary Position Embedding to the query and key tensors. Args: q (`torch.Tensor`): The query tensor. k (`torch.Tensor`): T
vita_audio/models/qwen2_v4_48_3/modeling_qwen2.py:67
↓ 2 callersFunctionapply_rotary_pos_emb
Applies Rotary Position Embedding to the query and key tensors. Args: q (`torch.Tensor`): The query tensor. k (`torch.Tensor`): T
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:109
↓ 2 callersMethodapply_to_role
(self, role, **kwargs)
vita_audio/data/processor/audio_processor.py:74
↓ 2 callersMethodbuild_model
(**kwargs)
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:1120
↓ 2 callersFunctioncharacterize
(string)
evaluation/compute-cer.py:16
↓ 2 callersFunctioncharacterize
(string)
evaluation/compute-wer.py:15
↓ 2 callersMethodcluster
(self, data)
evaluation/compute-cer.py:235
↓ 2 callersMethodcluster
(self, data)
evaluation/compute-wer.py:228
↓ 2 callersMethoddecode
( self, token_ids, skip_special_tokens: bool = False, clean_up_tokenization_sp
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/tokenization_qwen2.py:291
↓ 2 callersFunctionfind_audio_segments_regex
Find all substrings between <|begin_of_audio|> and <|end_of_audio|> using regex. Args: text (str): The input string to search throug
tools/inference_sts.py:159
↓ 2 callersMethodforward_attention
Compute attention context vector. Args: value (torch.Tensor): Transformed value (#batch, n_head, time2, d_k). scores
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:244
↓ 2 callersMethodforward_fsmn
(self, inputs, mask, mask_shfit_chunk=None)
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:197
↓ 2 callersMethodforward_qkv
Transform query, key and value. Args: query (torch.Tensor): Query tensor (#batch, time1, size). key (torch.Tensor): K
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:215
↓ 2 callersMethodget_input_embeddings
(self)
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:543
↓ 2 callersFunctionget_pairs
Return set of symbol pairs in a word. Word is represented as tuple of symbols (symbols being variable-length strings).
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/tokenization_qwen2.py:69
↓ 2 callersFunctionget_pairs
Return set of symbol pairs in a word. Word is represented as tuple of symbols (symbols being variable-length strings).
vita_audio/models/qwen2_v4_48_3/tokenization_qwen2.py:69
↓ 2 callersFunctionget_pairs
Return set of symbol pairs in a word. Word is represented as tuple of symbols (symbols being variable-length strings).
vita_audio/models/qwen2_mtp_v4_48_3/tokenization_qwen2.py:69
↓ 2 callersFunctionhas_audio
(sample)
vita_audio/data/dataset_deepseek.py:815
↓ 2 callersFunctionhas_audio
(sample)
vita_audio/data/dataset_qwen2.py:830
↓ 2 callersMethodis_empty
Check if the queue is empty in a thread-safe manner. Parameters: - None Returns: - bool: True if the queue
web/queue.py:143
↓ 2 callersFunctionis_list_in_string
(text, candidate)
evaluation/compute-acc-of-contain.py:20
↓ 2 callersMethodload_data
(self)
vita_audio/data/dataset_base.py:120
↓ 2 callersFunctionload_json
(data_file, output_dir)
vita_audio/data/dataset_base.py:312
↓ 2 callersMethodload_model
(self)
vita_audio/tokenizer_sensevoice_glm4voice.py:109
↓ 2 callersMethodload_model
(self)
vita_audio/tokenizer_glm4voice.py:104
↓ 2 callersMethodload_model
(self)
vita_audio/tokenizer_sensevoice_sparktts.py:91
↓ 2 callersMethodload_model
(self)
vita_audio/tokenizer_cosyvoice2.py:85
↓ 2 callersMethodload_model
(self)
vita_audio/tokenizer_snac.py:87
↓ 2 callersFunctionmain
()
tools/finetune_sts_v4_48_3.py:278
↓ 2 callersMethodmaybe_init_ret
(self, source, force=False)
vita_audio/data/dataset_mistral.py:39
↓ 2 callersMethodmtp_forward
( self, mtp_idx, input_ids: torch.LongTensor = None, hidden_states: torch.Tens
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:872
↓ 2 callersMethodmtp_forward
( self, mtp_idx, input_ids: torch.LongTensor = None, hidden_states: torch.Tens
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:819
↓ 2 callersFunctionnormalize
sentence, ignore_words are both in unicode
evaluation/compute-cer.py:67
↓ 2 callersFunctionnormalize
sentence, ignore_words are both in unicode
evaluation/compute-wer.py:64
↓ 2 callersMethodput
Receives tokens, decodes them, and logger.infos them to stdout as soon as they form entire words.
web_demo_stream.py:176
↓ 2 callersMethodrelease
(self, obj)
web/pool.py:50
↓ 2 callersFunctionrepeat_kv
This is the equivalent of torch.repeat_interleave(x, dim=1, repeats=n_rep). The hidden states go from (batch, num_key_value_heads, seqlen, he
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:139
↓ 2 callersFunctionrepeat_kv
This is the equivalent of torch.repeat_interleave(x, dim=1, repeats=n_rep). The hidden states go from (batch, num_key_value_heads, seqlen, he
vita_audio/models/qwen2_v4_48_3/modeling_qwen2.py:94
↓ 2 callersFunctionrepeat_kv
This is the equivalent of torch.repeat_interleave(x, dim=1, repeats=n_rep). The hidden states go from (batch, num_key_value_heads, seqlen, he
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:136
↓ 2 callersMethodreset_dialog
(self)
web/vad.py:170
↓ 2 callersFunctionrotate_half
Rotates half the hidden dims of the input.
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:105
↓ 2 callersFunctionrotate_half
Rotates half the hidden dims of the input.
vita_audio/models/qwen2_v4_48_3/modeling_qwen2.py:60
↓ 2 callersFunctionrotate_half
Rotates half the hidden dims of the input.
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:102
↓ 2 callersFunctionrun_infer_stream
(audio_tensor, sid)
web_demo_stream.py:244
↓ 2 callersFunctionwidth
(string)
evaluation/compute-cer.py:250
↓ 2 callersFunctionwidth
(string)
evaluation/compute-wer.py:243
↓ 1 callersFunctionForCausalLMLoss
( logits, labels, vocab_size: int, num_items_in_batch: int = None, ignore_index: int = -100, **kwargs )
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:56
↓ 1 callersFunctionForCausalLMLoss
( logits, labels, vocab_size: int, num_items_in_batch: int = None, ignore_index: int = -100, **kwargs )
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:53
↓ 1 callersMethod__len__
(self)
vita_audio/data/dataset_base.py:250
↓ 1 callersMethod_calc_ctc_loss
( self, encoder_out: torch.Tensor, encoder_out_lens: torch.Tensor, ys_pad: tor
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:827
↓ 1 callersMethod_calc_rich_ce_loss
( self, encoder_out: torch.Tensor, ys_pad: torch.Tensor, )
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_sensevoice.py:844
↓ 1 callersMethod_dynamic_frequency_update
dynamic RoPE layers should recompute `inv_freq` in the following situations: 1 - growing beyond the cached sequence length (allow sca
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:348
↓ 1 callersMethod_dynamic_frequency_update
dynamic RoPE layers should recompute `inv_freq` in the following situations: 1 - growing beyond the cached sequence length (allow sca
vita_audio/models/qwen2_v4_48_3/modeling_qwen2.py:303
↓ 1 callersMethod_dynamic_frequency_update
dynamic RoPE layers should recompute `inv_freq` in the following situations: 1 - growing beyond the cached sequence length (allow sca
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:345
next →1–100 of 566, ranked by callers