MCPcopy Create free account

hub / github.com/boson-ai/higgs-audio / functions

Functions316 in github.com/boson-ai/higgs-audio

Functiondisable_deepspeed_ulysses
Disable deepspeed ulysses (sequence parallelism) if it is enabled
boson_multimodal/model/higgs_audio/utils.py:740
Functiondrop_tokens
(input_, dim=0, group=None, grad_scale=1)
boson_multimodal/model/higgs_audio/utils.py:690
Functionema_inplace
(moving_avg, new, decay: float)
boson_multimodal/audio_processing/quantization/core_vq_lsx_version.py:58
Methodencode
(self, audio_path_or_wv, sr=None, loudness_normalize=False, loudness_threshold=-23.0)
boson_multimodal/audio_processing/higgs_audio_tokenizer.py:237
Methodencode
Encode a given input tensor with the specified sample rate at the given bandwidth. The RVQ encode method sets the appropriate number of quanti
boson_multimodal/audio_processing/quantization/vq.py:104
Methodencode
(self, x)
boson_multimodal/audio_processing/quantization/core_vq_lsx_version.py:344
Methodencode
(self, x: torch.Tensor, n_q: tp.Optional[int] = None)
boson_multimodal/audio_processing/quantization/core_vq_lsx_version.py:407
Methodencode
(self, x)
boson_multimodal/audio_processing/quantization/core_vq.py:279
Methodencode
(self, x: torch.Tensor, n_q: tp.Optional[int] = None)
boson_multimodal/audio_processing/quantization/core_vq.py:342
Functionencode_base64_content_from_file
Encode a content from a local file to base64 format.
boson_multimodal/serve/utils.py:27
Functionextract_generation_prompt_from_input_tokens
Extract the generation prompt and reference answer from the input tokens. For example: Input Text = '<|begin_of_text|><|start_header_id|>use
boson_multimodal/dataset/chatml_dataset.py:455
Methodforward
( self, hidden_states: torch.Tensor, causal_mask: torch.Tensor, position_ids:
boson_multimodal/model/higgs_audio/cuda_graph_runner.py:106
Methodforward
(ctx, input_, group)
boson_multimodal/model/higgs_audio/utils.py:565
Methodforward
(ctx, input_, dim, group, grad_scale)
boson_multimodal/model/higgs_audio/utils.py:654
Methodforward
(ctx, input_, dim, group, grad_scale)
boson_multimodal/model/higgs_audio/utils.py:676
Methodforward
Args: hidden_states (`torch.Tensor` of shape `(batch_size, seq_len, hidden_size)`): Hidden states from the LLM co
boson_multimodal/model/higgs_audio/audio_head.py:39
Methodforward
(self, audio_features)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:170
Methodforward
r""" Args: input_features (`torch.LongTensor` of shape `(batch_size, feature_size, sequence_length)`): Float value
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:229
Methodforward
Args: hidden_states (`torch.FloatTensor`): input to the layer of shape `(batch, seq_len, embed_dim)` attention_mask (
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:430
Methodforward
Forward pass for the Higgs-Audio model. Args: input_ids (:obj:`torch.LongTensor`): The input ids of the prompt. I
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:1142
Methodforward
Forward pass for the split embedding wrapper. :param input_ids: Tensor of shape [batch_size, seq_len] with indices in [0..original_vo
boson_multimodal/model/higgs_audio/custom_modules.py:46
Methodforward
(self, input_tensor)
boson_multimodal/model/higgs_audio/custom_modules.py:135
Methodforward
(self, raw_audio, sampling_rate=16000, return_tensors="pt")
boson_multimodal/audio_processing/higgs_audio_tokenizer.py:34
Methodforward
(self, x: torch.Tensor, bw: int)
boson_multimodal/audio_processing/higgs_audio_tokenizer.py:209
Methodforward
Args: x (Tensor): Float tensor variable with the shape (B, C, T). Returns: Tensor: Float tensor variable wit
boson_multimodal/audio_processing/semantic_module.py:46
Methodforward
(self, x)
boson_multimodal/audio_processing/semantic_module.py:80
Methodforward
Args: x (Tensor): Float tensor variable with the shape (B, C, T). Returns: Tensor: Float tensor variable wit
boson_multimodal/audio_processing/semantic_module.py:114
Methodforward
(self, x)
boson_multimodal/audio_processing/semantic_module.py:143
Methodforward
(self, x)
boson_multimodal/audio_processing/semantic_module.py:186
Methodforward
(self, x)
boson_multimodal/audio_processing/semantic_module.py:225
Methodforward
(self, z)
boson_multimodal/audio_processing/semantic_module.py:277
Methodforward
Residual vector quantization on the given input tensor. Args: x (torch.Tensor): Input tensor. sample_rate (int): Sampl
boson_multimodal/audio_processing/quantization/vq.py:74
Methodforward
(ctx, tensor)
boson_multimodal/audio_processing/quantization/ddp_utils.py:39
Methodforward
(self, *inputs, **kwargs)
boson_multimodal/audio_processing/quantization/ddp_utils.py:73
Methodforward
(self, x)
boson_multimodal/audio_processing/quantization/core_vq_lsx_version.py:252
Methodforward
(self, x)
boson_multimodal/audio_processing/quantization/core_vq_lsx_version.py:356
Methodforward
(self, x, n_q: tp.Optional[int] = None)
boson_multimodal/audio_processing/quantization/core_vq_lsx_version.py:387
Methodforward
(self, x)
boson_multimodal/audio_processing/quantization/core_vq.py:198
Methodforward
(self, x)
boson_multimodal/audio_processing/quantization/core_vq.py:291
Methodforward
(self, x, n_q: tp.Optional[int] = None)
boson_multimodal/audio_processing/quantization/core_vq.py:322
Methodforward
(self, x)
boson_multimodal/audio_processing/descriptaudiocodec/dac/model/dac.py:35
Methodforward
(self, x)
boson_multimodal/audio_processing/descriptaudiocodec/dac/model/dac.py:60
Methodforward
(self, x)
boson_multimodal/audio_processing/descriptaudiocodec/dac/model/dac.py:90
Methodforward
(self, x)
boson_multimodal/audio_processing/descriptaudiocodec/dac/model/dac.py:112
Methodforward
(self, x)
boson_multimodal/audio_processing/descriptaudiocodec/dac/model/dac.py:148
Methodforward
Model forward pass Parameters ---------- audio_data : Tensor[B x 1 x T] Audio data to encode sample_rate
boson_multimodal/audio_processing/descriptaudiocodec/dac/model/dac.py:271
Methodforward
(self, x)
boson_multimodal/audio_processing/descriptaudiocodec/dac/nn/layers.py:32
Methodforward
Quantized the input tensor using a fixed codebook and returns the corresponding codebook vectors Parameters ----------
boson_multimodal/audio_processing/descriptaudiocodec/dac/nn/quantize.py:34
Methodforward
Quantized the input tensor using a fixed set of `n` codebooks and returns the corresponding codebook vectors Parameters ------
boson_multimodal/audio_processing/descriptaudiocodec/dac/nn/quantize.py:122
Methodfreeze_audio_encoder_proj
(self)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:2138
Methodfreeze_audio_tower
(self)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:2133
Methodfreeze_llm
(self, freeze_embed=True, freeze_embed_until_idx: Optional[int] = None)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:2143
Methodfreeze_text_head
Freeze the final text head
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:2173
Methodfrom_latents
Given the unquantized latents, reconstruct the continuous representation after quantization. Parameters ---------- la
boson_multimodal/audio_processing/descriptaudiocodec/dac/nn/quantize.py:213
Functionfull_to_half_width
Convert full-width punctuation to half-width in a given string.
boson_multimodal/serve/utils.py:204
Functiongather_tokens
(input_, dim=0, group=None, grad_scale=1)
boson_multimodal/model/higgs_audio/utils.py:697
Methodgenerate
The generate function in huggingface generally follows these steps: for sample_step in 1, 2, 3, 4, 5, ... ...
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:1933
Methodgenerate_delta_stream
Generate audio from a chatml sample. Args: chat_ml_sample: A chatml sample. max_new_tokens: The maximum numbe
boson_multimodal/serve/serve_engine.py:428
Functionget_commit_hash
()
boson_multimodal/audio_processing/quantization/ddp_utils.py:63
Methodget_input_embeddings
(self)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:223
Functionget_interleaved_dialogue_input_sample
()
examples/serve_engine/input_samples.py:14
Methodget_last_layer
(self)
boson_multimodal/audio_processing/higgs_audio_tokenizer.py:153
Functionget_logger
(cfg, name=None)
boson_multimodal/audio_processing/quantization/ddp_utils.py:28
Functionget_sequence_data_parallel_group
()
boson_multimodal/model/higgs_audio/utils.py:599
Functionget_sequence_data_parallel_rank
()
boson_multimodal/model/higgs_audio/utils.py:595
Functionget_sequence_data_parallel_world_size
()
boson_multimodal/model/higgs_audio/utils.py:591
Functionget_timestamp
()
boson_multimodal/audio_processing/quantization/ddp_utils.py:59
Functionget_voice_clone_input_sample
()
examples/serve_engine/input_samples.py:59
Functionget_zero_shot_input_sample
()
examples/serve_engine/input_samples.py:38
Functioninit_weights
(m)
boson_multimodal/audio_processing/descriptaudiocodec/dac/model/dac.py:18
Functionis_only_punctuation
(text: str)
boson_multimodal/serve/utils.py:153
Functionlaplace_smoothing
(x, n_categories: int, epsilon: float = 1e-5)
boson_multimodal/audio_processing/quantization/core_vq_lsx_version.py:62
Methodmax_score_sample
(self)
boson_multimodal/dataset/chatml_dataset.py:281
Methodmerge
Merges a list of ChatMLDatasetSample instances, inserting eos_token_id and ignore_index between them, and adjusting offsets for audio_ids_start and au
boson_multimodal/dataset/chatml_dataset.py:129
Methodmerge_weights_from_checkpoint
(cls, checkpoint_dir: str, merged_output_dir: str, *model_args, **kwargs)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:2186
Methodmin_score_sample
(self)
boson_multimodal/dataset/chatml_dataset.py:286
Methodnum_audios
(self)
boson_multimodal/dataset/chatml_dataset.py:48
Methodnum_codebooks
(self)
boson_multimodal/audio_processing/higgs_audio_tokenizer.py:146
Methodpadding
(self)
boson_multimodal/audio_processing/descriptaudiocodec/dac/model/base.py:57
Methodparameter_count_per_component
Count the number of parameters per component in the model. HiggsAudio has the following main components: audio_tower: For mapping
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:2025
Functionpcm16_to_target_format
( np_audio: np.ndarray, sample_rate: int, bit_depth: int, channels: int, format: str,
boson_multimodal/serve/utils.py:35
Functionprepare_chatml_dataframe
(df, tokenizer, num_process=16)
boson_multimodal/dataset/chatml_dataset.py:502
Functionrandom_uuid
()
boson_multimodal/serve/utils.py:15
Functionrank
()
boson_multimodal/audio_processing/quantization/distrib.py:14
Functionremove_bracket
(text: str)
boson_multimodal/serve/utils.py:86
Functionremove_emoji
(text: str)
boson_multimodal/serve/utils.py:179
Functionremove_repeated_punctuations
(text, punctuations)
boson_multimodal/serve/utils.py:197
Functionreplace_blank
(text: str)
boson_multimodal/serve/utils.py:68
Functionreplace_corner_mark
(text: str)
boson_multimodal/serve/utils.py:79
Functionrope_decorator
(rope_func=None)
boson_multimodal/model/higgs_audio/utils.py:513
Methodsampling_rate
(self)
boson_multimodal/audio_processing/higgs_audio_tokenizer.py:142
Functionsequence_chunking_per_rank
Slice the inputs to create chuncks per the sequence parallel rank. This is used for the context parallel training. Args: sp_size (`i
boson_multimodal/model/higgs_audio/utils.py:704
Methodset_delay_pattern
(self)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:910
Methodset_encode_audio_in_tokens
(self)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:2130
Methodset_input_embeddings
(self, value: nn.Module)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:226
Methodset_num_activation_checkpointing_layers
(self, num_layers)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:907
Functionset_random_seed
(seed)
boson_multimodal/audio_processing/quantization/ddp_utils.py:17
Methodset_skip_audio_tower
(self)
boson_multimodal/model/higgs_audio/modeling_higgs_audio.py:2126
Functionsp_group
(self)
boson_multimodal/model/higgs_audio/utils.py:467
Functionsp_rank
(self)
boson_multimodal/model/higgs_audio/utils.py:459
← previousnext →201–300 of 316, ranked by callers