MCPcopy Create free account

hub / github.com/AIGC-Audio/AudioGPT / functions

Functions2,752 in github.com/AIGC-Audio/AudioGPT

↓ 4 callersMethodfit
(self, model)
NeuralSeq/utils/pl_utils.py:477
↓ 4 callersMethodforward
(self, audio, text)
text_to_audio/Make_An_Audio/ldm/modules/encoders/CLAP/clap.py:85
↓ 4 callersFunctionget_ckpt_path
(name, root, check=False)
text_to_audio/Make_An_Audio/ldm/util.py:128
↓ 4 callersFunctionget_cont_lf0
(f0, frame_period=5.0)
NeuralSeq/utils/cwt.py:46
↓ 4 callersMethodget_first_stage_encoding
(self, encoder_posterior)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio_inpaint.py:161
↓ 4 callersMethodget_fold_unfold
:param x: img of size (bs, c, h, w) :return: n img crops of size (n, bs, c, kernel_size[0], kernel_size[1])
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio.py:242
↓ 4 callersMethodget_fold_unfold
:param x: img of size (bs, c, h, w) :return: n img crops of size (n, bs, c, kernel_size[0], kernel_size[1])
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:599
↓ 4 callersMethodget_fold_unfold
:param x: img of size (bs, c, h, w) :return: n img crops of size (n, bs, c, kernel_size[0], kernel_size[1])
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio_inpaint.py:220
↓ 4 callersMethodget_input
(self, batch, k)
text_to_audio/Make_An_Audio/ldm/models/autoencoder.py:124
↓ 4 callersMethodget_input
(self, batch, k)
text_to_audio/Make_An_Audio/ldm/models/autoencoder_multi.py:82
↓ 4 callersMethodget_last_layer
(self)
text_to_audio/Make_An_Audio/ldm/models/autoencoder.py:250
↓ 4 callersMethodget_last_layer
(self)
text_to_audio/Make_An_Audio/ldm/models/autoencoder_multi.py:155
↓ 4 callersFunctionisimage
(x)
text_to_audio/Make_An_Audio/ldm/util.py:80
↓ 4 callersFunctionmean_with_lens
features: [N, T, ...] (assume the second dimension represents length) lens: [N,]
audio_to_text/captioning/models/utils.py:39
↓ 4 callersFunctionmove_data_to_device
(x, device)
audio_detection/audio_infer/pytorch/pytorch_utils.py:7
↓ 4 callersMethodnorm_spec
(self, x)
NeuralSeq/modules/diff/shallow_diffusion_tts.py:279
↓ 4 callersMethodplot_pitch
(self, batch_idx, sample, model_out)
NeuralSeq/tasks/tts/fs2.py:303
↓ 4 callersMethodplot_wav
(self, batch_idx, gt_wav, wav_out, is_mel=False, gt_f0=None, f0=None, name=None)
NeuralSeq/tasks/svs/diffspeech_task.py:112
↓ 4 callersMethodpredict_start_from_noise
(self, x_t, t, noise)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:214
↓ 4 callersFunctionpreprocess_wav
Applies the preprocessing operations used in training the Speaker Encoder to a waveform either on disk or in memory. The waveform will be re
NeuralSeq/data_gen/tts/emotion/audio.py:13
↓ 4 callersMethodq_posterior
(self, x_start, x_t, t)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:220
↓ 4 callersFunctionread_hdf5
Read hdf5 dataset. Args: hdf5_name (str): Filename of hdf5 file. hdf5_path (str): Dataset name in hdf5 file. Return:
NeuralSeq/modules/parallel_wavegan/utils/utils.py:39
↓ 4 callersMethodreduce_distributed_output
(self, output, num_gpus)
NeuralSeq/utils/pl_utils.py:1050
↓ 4 callersMethodremove_weight_norm
Remove weight normalization module from all of the layers.
NeuralSeq/modules/parallel_wavegan/models/parallel_wavegan.py:173
↓ 4 callersMethodsample_log
(self,cond,batch_size,ddim, ddim_steps,**kwargs)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:1233
↓ 4 callersMethodsample_log
(self,cond,batch_size,ddim, ddim_steps,**kwargs)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio_inpaint.py:923
↓ 4 callersMethodshared_step
(self, batch, t=None)
text_to_audio/Make_An_Audio/ldm/models/diffusion/classifier.py:179
↓ 4 callersFunctionstft
Perform STFT and convert to magnitude spectrogram. Args: x (Tensor): Input signal tensor (B, T). fft_size (int): FFT size.
NeuralSeq/modules/parallel_wavegan/losses/stft_loss.py:12
↓ 4 callersMethodtest_step
(self, sample, batch_idx)
NeuralSeq/tasks/base_task.py:210
↓ 4 callersMethodtrain
(self, mode)
audio_to_text/captioning/models/encoder.py:576
↓ 4 callersMethodvocode
(self, spec)
text_to_audio/Make_An_Audio/vocoder/bigvgan/models.py:406
↓ 4 callersMethodwav2spec
(wav_fn, return_linear=False)
NeuralSeq/vocoders/pwg.py:106
↓ 3 callersMethod__init__
(self, in_dims=80)
NeuralSeq/modules/diff/net.py:82
↓ 3 callersMethod__init__
(self, in_dim=80, out_dim=256, kernel=5, n_layers=3, strides=None)
NeuralSeq/modules/fastspeech/pe.py:8
↓ 3 callersMethod__init__
Initialize Conv2d module.
NeuralSeq/modules/parallel_wavegan/layers/upsample.py:50
↓ 3 callersMethod__init__
Initialize STFT loss module.
NeuralSeq/modules/parallel_wavegan/losses/stft_loss.py:79
↓ 3 callersMethod__init__
(self, date=None, chntext=None)
NeuralSeq/utils/text_norm.py:508
↓ 3 callersMethod__init__
(self)
NeuralSeq/tasks/svs/diffsinger_task.py:31
↓ 3 callersMethod__init__
(self, sample_rate, window_size, hop_size, mel_bins, fmin, fmax, classes_num, out_emb)
text_to_audio/Make_An_Audio/wav_evaluation/models/audio.py:108
↓ 3 callersMethod__init__
(self, # audio audioenc_name: str, sample_rate: int,
text_to_audio/Make_An_Audio/wav_evaluation/models/clap.py:56
↓ 3 callersMethod__init__
(self, num_features, logdet=False, affine=True, allow_reverse_init=False)
text_to_audio/Make_An_Audio/ldm/modules/discriminator/model.py:6
↓ 3 callersMethod__init__
(self, use_dropout=True)
text_to_audio/Make_An_Audio/ldm/modules/losses_audio/lpaps.py:19
↓ 3 callersMethod__init__
(self, sample_rate, window_size, hop_size, mel_bins, fmin, fmax, classes_num, out_emb)
text_to_audio/Make_An_Audio/ldm/modules/encoders/CLAP/audio.py:108
↓ 3 callersMethod__init__
(self, # audio audioenc_name: str, sample_rate: int,
text_to_audio/Make_An_Audio/ldm/modules/encoders/CLAP/clap.py:55
↓ 3 callersMethod__init__
(self, ddconfig, lossconfig, n_embed, embe
text_to_audio/Make_An_Audio/ldm/models/autoencoder.py:15
↓ 3 callersMethod__init__
(self, unet_config, timesteps=1000, beta_schedule="linear",
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:45
↓ 3 callersMethod_generic_batch_inference
r"""Process audio and/or text per batch
text_to_audio/Make_An_Audio/wav_evaluation/models/CLAPWrapper.py:223
↓ 3 callersMethod_generic_batch_inference
r"""Process audio and/or text per batch
text_to_audio/Make_An_Audio/ldm/modules/encoders/CLAP/CLAPWrapper.py:204
↓ 3 callersMethod_get_denoise_row_from_list
(self, samples, desc='', force_no_decoder_quantization=False)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio.py:145
↓ 3 callersMethod_get_item
(self, index)
NeuralSeq/tasks/tts/fs2_utils.py:53
↓ 3 callersFunctionact
(x, activation)
sound_extraction/model/modules.py:472
↓ 3 callersFunctionadd_blur
(img, sf=4)
text_to_audio/Make_An_Audio/ldm/modules/image_degradation/bsrgan_light.py:325
↓ 3 callersMethodadd_dur
:param dur_input: [B, T_txt, H] :param mel2ph: [B, T_mel] :param txt_tokens: [B, T_txt] :param ret: :return:
NeuralSeq/modules/fastspeech/fs2.py:140
↓ 3 callersMethodadd_dur_loss
:param dur_pred: [B, T], float, log scale :param mel2ph: [B, T] :param txt_tokens: [B, T] :param losses: :ret
NeuralSeq/tasks/svs/diffsinger_task.py:351
↓ 3 callersMethodadd_f0_loss
(self, p_pred, f0, uv, losses, nonpadding)
NeuralSeq/tasks/tts/fs2.py:252
↓ 3 callersMethodadd_item
(self, item)
NeuralSeq/utils/indexed_datasets.py:47
↓ 3 callersMethodattention
(self, x: torch.Tensor, attn_mask: Optional[torch.Tensor] = None)
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/model.py:276
↓ 3 callersMethodbackward
(self, loss, optimizer)
NeuralSeq/tasks/base_task.py:338
↓ 3 callersMethodbuild_dataloader
(self, dataset, shuffle, max_tokens=None, max_sentences=None, required_batch_size_mul
NeuralSeq/tasks/tts/tts_base.py:92
↓ 3 callersMethodbuild_dataloader
(self, dataset, shuffle, max_tokens=None, max_sentences=None, required_batch_size_mul
NeuralSeq/tasks/tts/tts.py:49
↓ 3 callersMethodbuild_dataloader
(self, dataset, shuffle, max_sentences, endless=False)
NeuralSeq/tasks/vocoder/vocoder_base.py:37
↓ 3 callersFunctioncheckpoint
Evaluate a function without caching intermediate activations, allowing for reduced memory at the expense of extra compute in the backward pas
text_to_audio/Make_An_Audio/ldm/modules/diffusionmodules/util.py:102
↓ 3 callersMethodcopy_trainer_model_properties
(self, model)
NeuralSeq/utils/pl_utils.py:776
↓ 3 callersFunctioncount_params
(model, verbose=False)
text_to_audio/Make_An_Audio/ldm/util.py:104
↓ 3 callersFunctioncreate_logging
(log_dir, filemode)
audio_detection/audio_infer/utils/utilities.py:34
↓ 3 callersFunctioncreate_window
(window_size, channel)
NeuralSeq/modules/commons/ssim.py:324
↓ 3 callersFunctiond_prime
(auc)
audio_detection/audio_infer/utils/utilities.py:112
↓ 3 callersFunctiondefault
(val, d)
text_to_audio/Make_An_Audio/ldm/modules/attention.py:19
↓ 3 callersMethoddenorm_spec
(self, x)
NeuralSeq/modules/diff/shallow_diffusion_tts.py:282
↓ 3 callersMethodencode_first_stage
(self, x)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio.py:472
↓ 3 callersMethodencode_first_stage
(self, x)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:824
↓ 3 callersFunctionfind_contiguous_regions
Find contiguous regions from bool valued numpy.array. Copy of https://dcase-repo.github.io/dcase_util/_modules/dcase_util/data/decisions.html#Deci
audio_detection/target_sound_detection/src/utils.py:34
↓ 3 callersFunctionfreeze_batch_norm_2d
Converts all `BatchNorm2d` and `SyncBatchNorm` layers of provided module into `FrozenBatchNorm2d`. If `module` is itself an instance of eithe
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/utils.py:42
↓ 3 callersFunctionfspecial
python code from: https://github.com/ronaldosena/imagens-medicas-2/blob/40171a6c259edec7827a6693a93955de2bd39e76/Aulas/aula_2_-_uniform_filte
text_to_audio/Make_An_Audio/ldm/modules/image_degradation/bsrgan_light.py:210
↓ 3 callersFunctionfspecial
python code from: https://github.com/ronaldosena/imagens-medicas-2/blob/40171a6c259edec7827a6693a93955de2bd39e76/Aulas/aula_2_-_uniform_filte
text_to_audio/Make_An_Audio/ldm/modules/image_degradation/bsrgan.py:210
↓ 3 callersMethodgenerate_square_subsequent_mask
(self, max_length)
audio_to_text/captioning/models/decoder.py:645
↓ 3 callersMethodget_first_stage_encoding
(self, encoder_posterior)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio.py:157
↓ 3 callersMethodget_first_stage_encoding
(self, encoder_posterior)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:540
↓ 3 callersFunctionget_incremental_state
Helper for getting incremental state for an nn.Module.
NeuralSeq/utils/tts_utils.py:48
↓ 3 callersMethodget_input
(self, batch, k)
text_to_audio/Make_An_Audio/ldm/models/autoencoder.py:365
↓ 3 callersFunctionget_mel2ph
(tg_fn, ph, mel, hparams)
NeuralSeq/data_gen/tts/data_gen_utils.py:274
↓ 3 callersMethodget_pos_embed
(self, word2word, x2word)
NeuralSeq/modules/syntaspeech/syntaspeech.py:259
↓ 3 callersFunctionget_pretrained_url
(model: str, tag: str)
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/pretrained.py:102
↓ 3 callersMethodget_text_embeddings
r"""Load list of class labels and return text embeddings
text_to_audio/Make_An_Audio/wav_evaluation/models/CLAPWrapper.py:177
↓ 3 callersMethodget_weight
(self, device, reverse)
NeuralSeq/modules/GenerSpeech/model/glow_modules.py:225
↓ 3 callersMethodget_weight
(self, device, reverse)
NeuralSeq/modules/commons/normalizing_flow/glow_modules.py:167
↓ 3 callersMethodget_weighting
(self, h, w, Ly, Lx, device)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio.py:226
↓ 3 callersMethodget_weighting
(self, h, w, Ly, Lx, device)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:583
↓ 3 callersMethodget_weighting
(self, h, w, Ly, Lx, device)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio_inpaint.py:204
↓ 3 callersMethodin_proj_v
(self, value)
NeuralSeq/modules/commons/transformer.py:442
↓ 3 callersMethodin_proj_v
(self, value)
NeuralSeq/modules/commons/common_layers.py:443
↓ 3 callersFunctioninit_bn
Initialize a Batchnorm layer.
audio_to_text/captioning/models/encoder.py:25
↓ 3 callersFunctioninit_bn
Initialize a Batchnorm layer.
audio_detection/target_sound_detection/src/models.py:145
↓ 3 callersMethodinit_from_ckpt
(self, path, ignore_keys=list(), only_model=False)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:184
↓ 3 callersFunctioninit_layer
Initialize a Linear or Convolutional layer.
audio_to_text/captioning/models/encoder.py:16
↓ 3 callersFunctionload_ckpt
(cur_model, ckpt_base_dir, model_name='model', force=True, strict=True)
NeuralSeq/utils/ckpt_utils.py:28
↓ 3 callersMethodload_pretrained
(self, pretrained)
audio_to_text/captioning/models/encoder.py:426
↓ 3 callersFunctionload_state_dict
(checkpoint_path: str, map_location="cpu", skip_params=True)
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/factory.py:51
↓ 3 callersFunctionmax_with_lens
features: [N, T, ...] (assume the second dimension represents length) lens: [N,]
audio_to_text/captioning/models/utils.py:63
↓ 3 callersFunctionmean_flat
Take the mean over all non-batch dimensions.
text_to_audio/Make_An_Audio/ldm/modules/diffusionmodules/util.py:192
← previousnext →201–300 of 2,752, ranked by callers