MCPcopy Create free account

hub / github.com/AIGC-Audio/AudioGPT / functions

Functions2,752 in github.com/AIGC-Audio/AudioGPT

↓ 2 callersFunctioncopy_tensor
(src, dst)
NeuralSeq/utils/__init__.py:49
↓ 2 callersMethodcopy_to
(self, model)
text_to_audio/Make_An_Audio/ldm/modules/ema.py:46
↓ 2 callersFunctioncount_flops_attn
A counter for the `thop` package to count the operations in an attention operation. Meant to be used like: macs, params = thop.pr
text_to_audio/Make_An_Audio/ldm/modules/diffusionmodules/openaimodel.py:327
↓ 2 callersFunctioncreate_model
( amodel_name: str, tmodel_name: str, pretrained: str = "", precision: str = "fp32", devic
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/factory.py:67
↓ 2 callersFunctioncreate_system
根据数字系统类型返回创建相应的数字系统,默认为 mid NUMBERING_TYPES = ['low', 'mid', 'high']: 中文数字系统类型 low: '兆' = '亿' * '十' = $10^{9}$, '京' = '兆' * '十', et
NeuralSeq/utils/text_norm.py:191
↓ 2 callersFunctioncubic
(x)
text_to_audio/Make_An_Audio/ldm/modules/image_degradation/utils_image.py:700
↓ 2 callersFunctioncut_dialogue_history
(history_memory, keep_last_n_words = 500)
audio-chatgpt.py:77
↓ 2 callersFunctioncwt2f0
(cwt_spec, mean, std, cwt_scales)
NeuralSeq/utils/cwt.py:135
↓ 2 callersMethoddecode
(self, quant)
text_to_audio/Make_An_Audio/ldm/models/autoencoder.py:107
↓ 2 callersMethoddecode
(self, z)
text_to_audio/Make_An_Audio/ldm/models/autoencoder.py:351
↓ 2 callersMethoddecode
(self, z)
text_to_audio/Make_An_Audio/ldm/models/autoencoder_multi.py:68
↓ 2 callersFunctiondefault
(val, d)
NeuralSeq/modules/diff/diffusion.py:23
↓ 2 callersFunctiondefault
(val, d)
NeuralSeq/modules/diff/shallow_diffusion_tts.py:24
↓ 2 callersMethoddefault_collate
r"""Puts each data field into a tensor with outer dimension batch size
text_to_audio/Make_An_Audio/wav_evaluation/models/CLAPWrapper.py:73
↓ 2 callersMethoddefault_collate
r"""Puts each data field into a tensor with outer dimension batch size
text_to_audio/Make_An_Audio/ldm/modules/encoders/CLAP/CLAPWrapper.py:71
↓ 2 callersMethoddelta_border
:param h: height :param w: width :return: normalized distance to image border, wtith min distance = 0 at border and
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio.py:212
↓ 2 callersMethoddelta_border
:param h: height :param w: width :return: normalized distance to image border, wtith min distance = 0 at border and
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:569
↓ 2 callersMethoddelta_border
:param h: height :param w: width :return: normalized distance to image border, wtith min distance = 0 at border and
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio_inpaint.py:190
↓ 2 callersMethoddigit2chntext
(self)
NeuralSeq/utils/text_norm.py:447
↓ 2 callersMethoddisplacements2warpfield
(self, displacements, seq_length)
mono2binaural/src/warping.py:97
↓ 2 callersFunctiondownload_pretrained
(url: str, root: str = os.path.expanduser("~/.cache/clip"))
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/pretrained.py:111
↓ 2 callersMethodema_scope
(self, context=None)
text_to_audio/Make_An_Audio/ldm/models/autoencoder.py:64
↓ 2 callersFunctionembed_frames_batch
Computes embeddings for a batch of mel spectrogram. :param frames_batch: a batch mel of spectrogram as a numpy array of float32 of shape
NeuralSeq/data_gen/tts/emotion/inference.py:43
↓ 2 callersMethodencode
(self, x)
NeuralSeq/modules/GenerSpeech/model/prosody_util.py:33
↓ 2 callersMethodevaluate
Run evaluation code. :param model: PT model :param dataloaders: list of PT dataloaders :param max_batches: Scalar :pa
NeuralSeq/utils/pl_utils.py:1146
↓ 2 callersMethodevaluate
Forward evaluation data and calculate statistics. Args: data_loader: object Returns: statistics: dict,
audio_detection/audio_infer/pytorch/evaluate.py:15
↓ 2 callersFunctionevaluate_annotation
(key2refs, scorer)
audio_to_text/captioning/utils/eval_round_robin.py:8
↓ 2 callersFunctionevaluate_prediction
(key2pred, key2refs, scorer)
audio_to_text/captioning/utils/eval_round_robin.py:30
↓ 2 callersMethodexample_run
(cls, inp)
NeuralSeq/inference/svs/base_svs_infer.py:235
↓ 2 callersFunctionexists
(x)
NeuralSeq/modules/diff/diffusion.py:19
↓ 2 callersFunctionexists
(x)
NeuralSeq/modules/diff/shallow_diffusion_tts.py:20
↓ 2 callersFunctionexists
(val)
text_to_audio/Make_An_Audio/ldm/modules/attention.py:11
↓ 2 callersMethodexpand_f0_ph
(f0, mel2ph)
NeuralSeq/tasks/tts/fs2.py:501
↓ 2 callersMethodexpand_queue
(self, queue)
audio_detection/audio_infer/utils/data_generator.py:208
↓ 2 callersMethodextract
(self, caption_file: str, model, output, dev: bool)
audio_to_text/captioning/utils/bert/create_sent_embedding.py:27
↓ 2 callersMethodfind_in_interval
(self, n)
text_to_audio/Make_An_Audio/ldm/lr_scheduler.py:52
↓ 2 callersMethodforward
(self, x)
NeuralSeq/modules/GenerSpeech/model/prosody_util.py:221
↓ 2 callersMethodforward
:param x: [B, T, H] :return: [B, T, H]
NeuralSeq/modules/commons/conv.py:100
↓ 2 callersMethodforward_dur
:param dur_input: [B, T_txt, H] :param mel2ph: [B, T_mel] :param txt_tokens: [B, T_txt] :param ret: :return:
NeuralSeq/modules/syntaspeech/syntaspeech.py:234
↓ 2 callersMethodforward_features
(self, x, longer_idx = None)
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/htsat.py:774
↓ 2 callersFunctiongather_features
( audio_features, text_features, audio_features_mlp=None, text_features_mlp=N
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/loss.py:15
↓ 2 callersMethodget_attn_stats
(self, attn, sample, logging_outputs, prefix='')
NeuralSeq/tasks/tts/ps.py:109
↓ 2 callersMethodget_audio_embeddings
r"""Load list of audio files and return a audio embeddings
text_to_audio/Make_An_Audio/wav_evaluation/models/CLAPWrapper.py:184
↓ 2 callersFunctionget_avg_stats
(workspace, bgn_iter, fin_iter, interval_iter, filename, data_type)
audio_detection/audio_infer/utils/plot_statistics.py:1524
↓ 2 callersMethodget_conditioning
(self, batch, k=None)
text_to_audio/Make_An_Audio/ldm/models/diffusion/classifier.py:133
↓ 2 callersFunctionget_diagonal_focus_rate
attn: bx x L_t x L_s attn_ks: shape: tensor with shape [batch_size], input_lens/output_lens diagonal: y=k*x (k=attn_ks, x:output, y:inpu
NeuralSeq/utils/tts_utils.py:108
↓ 2 callersMethodget_embedding
Build sinusoidal embeddings. This matches the implementation in tensor2tensor, but differs slightly from the description in Section 3
NeuralSeq/modules/commons/transformer.py:31
↓ 2 callersMethodget_embedding
Build sinusoidal embeddings. This matches the implementation in tensor2tensor, but differs slightly from the description in Section 3
NeuralSeq/modules/commons/common_layers.py:105
↓ 2 callersFunctionget_focus_rate
attn: bs x L_t x L_s
NeuralSeq/utils/tts_utils.py:73
↓ 2 callersFunctionget_hop_size
(hparams)
NeuralSeq/utils/audio.py:20
↓ 2 callersFunctionget_id_sets
(csv_path)
audio_detection/audio_infer/utils/create_black_list.py:23
↓ 2 callersMethodget_input
(self, batch, k)
text_to_audio/Make_An_Audio/ldm/models/diffusion/classifier.py:124
↓ 2 callersMethodget_input
(self, batch, k, return_first_stage_outputs=False, force_c_encode=False, cond_key=None, retu
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm.py:652
↓ 2 callersMethodget_last_layer
(self)
text_to_audio/Make_An_Audio/ldm/models/autoencoder.py:428
↓ 2 callersMethodget_lr
(self)
NeuralSeq/utils/training_utils.py:26
↓ 2 callersFunctionget_pairs
Return set of symbol pairs in a word. Word is represented as tuple of symbols (symbols being variable-length strings).
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/tokenizer.py:44
↓ 2 callersFunctionget_phone_coverage_rate
attn: bs x L_t x L_s
NeuralSeq/utils/tts_utils.py:88
↓ 2 callersMethodget_plot_dur_info
(self, sample, model_out)
NeuralSeq/tasks/tts/ps_adv.py:271
↓ 2 callersFunctionget_symbol
(char, system)
NeuralSeq/utils/text_norm.py:234
↓ 2 callersMethodget_text_embeddings
r"""Load list of class labels and return text embeddings
text_to_audio/Make_An_Audio/ldm/modules/encoders/CLAP/CLAPWrapper.py:167
↓ 2 callersFunctiongroup_dict_by_key
(cond, d)
text_to_audio/Make_An_Audio/ldm/modules/x_transformer.py:93
↓ 2 callersFunctiongroup_hidden_by_segs
:param h: [B, T, H] :param seg_ids: [B, T] :return: h_ph: [B, T_ph, H]
NeuralSeq/modules/syntaspeech/syntactic_graph_encoder.py:16
↓ 2 callersFunctiongroup_hidden_by_segs
:param h: [B, T, H] :param seg_ids: [B, T] :return: h_ph: [B, T_ph, H]
NeuralSeq/utils/tts_utils.py:357
↓ 2 callersFunctiongroupby_prefix_and_trim
(prefix, d)
text_to_audio/Make_An_Audio/ldm/modules/x_transformer.py:110
↓ 2 callersFunctionimage_transform
( image_size: int, is_train: bool, mean=(0.48145466, 0.4578275, 0.40821073), s
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/transform.py:9
↓ 2 callersMethodin_proj_k
(self, key)
NeuralSeq/modules/commons/transformer.py:432
↓ 2 callersMethodin_proj_k
(self, key)
NeuralSeq/modules/commons/common_layers.py:433
↓ 2 callersMethodin_proj_q
(self, query)
NeuralSeq/modules/commons/transformer.py:423
↓ 2 callersMethodin_proj_q
(self, query)
NeuralSeq/modules/commons/common_layers.py:424
↓ 2 callersMethodinit_from_ckpt
(self, path, ignore_keys=list(), only_model=False)
text_to_audio/Make_An_Audio/ldm/models/diffusion/ddpm_audio.py:1101
↓ 2 callersMethodinit_layer
Initialize a Linear or Convolutional layer.
audio_detection/target_sound_detection/src/models.py:731
↓ 2 callersMethodinit_optimizers
(self, optimizers)
NeuralSeq/utils/pl_utils.py:492
↓ 2 callersMethodinitialize
(self, input)
text_to_audio/Make_An_Audio/ldm/modules/discriminator/model.py:17
↓ 2 callersMethodinput_to_batch
(self, item)
NeuralSeq/inference/svs/base_svs_infer.py:200
↓ 2 callersFunctionkaiser_sinc_filter1d
(cutoff, half_width, kernel_size)
text_to_audio/Make_An_Audio/vocoder/bigvgan/alias_free_torch/filter.py:28
↓ 2 callersMethodkl
(self, other=None)
text_to_audio/Make_An_Audio/ldm/modules/distributions/distributions.py:39
↓ 2 callersFunctionlist_models
enumerate available model architectures based on config files
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/factory.py:247
↓ 2 callersFunctionload_ckpt
(cur_model, ckpt_base_dir, prefix_in_ckpt='model', force=True, strict=True)
NeuralSeq/utils/__init__.py:178
↓ 2 callersFunctionload_data_preprocessor
()
NeuralSeq/tasks/tts/tts_utils.py:37
↓ 2 callersMethodload_from_file
load network parameters from model_file :param model_file: file containing the model parameters
mono2binaural/src/utils.py:39
↓ 2 callersFunctionload_model
(config_path, checkpoint_path)
NeuralSeq/vocoders/hifigan.py:17
↓ 2 callersFunctionload_pwg_model
(config_path, checkpoint_path, stats_path)
NeuralSeq/vocoders/pwg.py:16
↓ 2 callersMethodlog_best_model
(self, model, loss, epoch, optimizer, metrics_dict)
text_to_audio/Make_An_Audio/ldm/modules/losses_audio/vggishish/logger.py:76
↓ 2 callersMethodlog_epoch_loss
(self, loss, epoch, phase)
text_to_audio/Make_An_Audio/ldm/modules/losses_audio/vggishish/logger.py:54
↓ 2 callersMethodlog_epoch_metrics
(self, metrics_dict, epoch, phase)
text_to_audio/Make_An_Audio/ldm/modules/losses_audio/vggishish/logger.py:58
↓ 2 callersMethodlog_iter_loss
(self, loss, iter, phase)
text_to_audio/Make_An_Audio/ldm/modules/losses_audio/vggishish/logger.py:51
↓ 2 callersMethodlog_metrics
Logs the metric dict passed in. :param metrics: :param grad_norm_dic:
NeuralSeq/utils/pl_utils.py:917
↓ 2 callersMethodlog_param_num
(self, model)
text_to_audio/Make_An_Audio/ldm/modules/losses_audio/vggishish/logger.py:45
↓ 2 callersMethodlog_test_metrics
(self, metrics_dict, hparams_dict, best_epoch)
text_to_audio/Make_An_Audio/ldm/modules/losses_audio/vggishish/logger.py:64
↓ 2 callersFunctionmake_ddim_sampling_parameters
(alphacums, ddim_timesteps, eta, verbose=True)
text_to_audio/Make_An_Audio/ldm/modules/diffusionmodules/util.py:63
↓ 2 callersFunctionmake_ddim_timesteps
(ddim_discr_method, num_ddim_timesteps, num_ddpm_timesteps, verbose=True)
text_to_audio/Make_An_Audio/ldm/modules/diffusionmodules/util.py:46
↓ 2 callersFunctionmd5_hash
(path)
text_to_audio/Make_An_Audio/ldm/util.py:43
↓ 2 callersFunctionmerge_audio
(audio_path_1, audio_path_2)
audio-chatgpt.py:92
↓ 2 callersFunctionmkdir
(path)
text_to_audio/Make_An_Audio/ldm/modules/image_degradation/utils_image.py:153
↓ 2 callersFunctionnorm_cdf
(x)
text_to_audio/Make_An_Audio/ldm/modules/encoders/open_clap/htsat.py:169
↓ 2 callersFunctionnorm_f0
(f0, uv, hparams)
NeuralSeq/utils/pitch_utils.py:34
↓ 2 callersFunctionnorm_scale
(Wavelet_lf0)
NeuralSeq/utils/cwt.py:72
↓ 2 callersFunctionnormalize_energy_torch
If the signal is almost empty(determined by threshold), if will only be divided by 2**15 :param audio: 1d waveform, 2**15 :param alpha: t
sound_extraction/utils/create_mixtures.py:54
↓ 2 callersFunctionnormalize_tensor
(x, eps=1e-10)
text_to_audio/Make_An_Audio/ldm/modules/losses_audio/lpaps.py:138
↓ 2 callersMethodon_epoch_end
(self, epoch, logs=None)
NeuralSeq/utils/pl_utils.py:327
← previousnext →401–500 of 2,752, ranked by callers