MCPcopy Create free account

hub / github.com/FunAudioLLM/ThinkSound / types & classes

Types & classes175 in github.com/FunAudioLLM/ThinkSound

↓ 12 callersClassAuralossLoss
ThinkSound/training/losses/losses.py:70
↓ 10 callersClassChannelLastConv1d
ThinkSound/models/blocks.py:341
↓ 9 callersClassConvMLP
ThinkSound/models/blocks.py:385
↓ 8 callersClassValueLoss
ThinkSound/training/losses/losses.py:16
↓ 7 callersClassLayerNorm
ThinkSound/models/transformer.py:173
↓ 6 callersClassLocalDatasetConfig
ThinkSound/data/dataset.py:113
↓ 6 callersClassResidualUnit
ThinkSound/models/autoencoders.py:39
↓ 6 callersClassSelfAttention1d
ThinkSound/models/blocks.py:34
↓ 5 callersClassMLP
ThinkSound/models/blocks.py:351
↓ 5 callersClassMono
ThinkSound/data/utils.py:327
↓ 5 callersClassMultiResolutionSTFTLoss
Multi resolution STFT loss module. See [Yamamoto et al., 2019](https://arxiv.org/abs/1910.11480) Args: fft_sizes (list): List of FFT
ThinkSound/training/losses/auraloss.py:449
↓ 5 callersClassPhaseFlipper
Randomly invert the phase of a signal
ThinkSound/data/utils.py:319
↓ 5 callersClassStereo
ThinkSound/data/utils.py:331
↓ 4 callersClassAttention
ThinkSound/models/transformer.py:271
↓ 4 callersClassFOA
ThinkSound/data/utils.py:345
↓ 4 callersClassMMDitSingleBlock
ThinkSound/models/transformer_layers.py:135
↓ 4 callersClassPattern
Base implementation of a pattern over a sequence with multiple codebooks. The codebook pattern consists in a layout, defining for each sequence s
ThinkSound/models/codebook_patterns.py:19
↓ 3 callersClassAveragePooling
data_utils/ext/synchformer/motionformer.py:386
↓ 3 callersClassDiffusionAttnUnet1D
ThinkSound/models/diffusion.py:408
↓ 3 callersClassMultiLoss
ThinkSound/training/losses/losses.py:84
↓ 2 callersClassAdaRMSNorm
ThinkSound/models/blocks.py:211
↓ 2 callersClassContinuousLocalTransformer
ThinkSound/models/local_attention.py:14
↓ 2 callersClassDataModule
ThinkSound/data/datamodule.py:35
↓ 2 callersClassDiTUncondWrapper
ThinkSound/models/diffusion.py:708
↓ 2 callersClassDiTWrapper
ThinkSound/models/diffusion.py:522
↓ 2 callersClassDiffusionModelWrapper
ThinkSound/models/diffusion.py:42
↓ 2 callersClassDiffusionTransformer
ThinkSound/models/dit.py:13
↓ 2 callersClassDividedAttention
data_utils/ext/synchformer/vit_helper.py:36
↓ 2 callersClassFeedForward
ThinkSound/models/transformer.py:221
↓ 2 callersClassFourierFeatures
ThinkSound/models/blocks.py:84
↓ 2 callersClassGLU
ThinkSound/models/transformer.py:196
↓ 2 callersClassMultiModalDataset
ThinkSound/data/dataset.py:558
↓ 2 callersClassPadCrop_Normalized_T
ThinkSound/data/utils.py:23
↓ 2 callersClassRotaryEmbedding
ThinkSound/models/transformer.py:89
↓ 2 callersClassSTFTMagnitudeLoss
STFT magnitude loss module. See [Arik et al., 2018](https://arxiv.org/abs/1808.06719) and [Engel et al., 2020](https://arxiv.org/abs/2001.046
ThinkSound/training/losses/auraloss.py:183
↓ 2 callersClassSynchformer
data_utils/ext/synchformer/synchformer.py:10
↓ 2 callersClassTemporalTransformerEncoderLayer
Aggregates temporal dimension with attention. Also used with pos emb as global aggregation in both streams.
data_utils/ext/synchformer/motionformer.py:368
↓ 2 callersClassTimestepEmbedder
Embeds scalar timesteps into vector representations.
ThinkSound/models/embeddings.py:43
↓ 2 callersClassUNet1DCondWrapper
ThinkSound/models/diffusion.py:303
↓ 2 callersClassUNet1DUncondWrapper
ThinkSound/models/diffusion.py:354
↓ 1 callersClassAbsolutePositionalEmbedding
ThinkSound/models/transformer.py:44
↓ 1 callersClassAudioAutoencoder
ThinkSound/models/autoencoders.py:230
↓ 1 callersClassAudioDataset
ThinkSound/data/dataset.py:353
↓ 1 callersClassAudio_Text
data_utils/v2a_utils/audio_text_dataset.py:27
↓ 1 callersClassAudiocraftCompressionPretransform
ThinkSound/models/pretransforms.py:194
↓ 1 callersClassAutoencoderPretransform
ThinkSound/models/pretransforms.py:28
↓ 1 callersClassCLAPAudioConditioner
ThinkSound/models/conditioners.py:439
↓ 1 callersClassCLAPTextConditioner
ThinkSound/models/conditioners.py:356
↓ 1 callersClassCLIPConditioner
ThinkSound/models/conditioners.py:236
↓ 1 callersClassCLIPTextConditioner
ThinkSound/models/conditioners.py:603
↓ 1 callersClassConformerModule
ThinkSound/models/transformer.py:554
↓ 1 callersClassContinuousTransformer
ThinkSound/models/transformer.py:702
↓ 1 callersClassCrossAttention
ThinkSound/models/transformer_layers.py:89
↓ 1 callersClassDACDecoderWrapper
ThinkSound/models/autoencoders.py:217
↓ 1 callersClassDACEncoderWrapper
ThinkSound/models/autoencoders.py:194
↓ 1 callersClassDACRVQBottleneck
ThinkSound/models/bottleneck.py:208
↓ 1 callersClassDACRVQVAEBottleneck
ThinkSound/models/bottleneck.py:261
↓ 1 callersClassDAU1DCondWrapper
ThinkSound/models/diffusion.py:374
↓ 1 callersClassDecoderBlock
ThinkSound/models/autoencoders.py:83
↓ 1 callersClassDiffusionAutoencoder
ThinkSound/models/autoencoders.py:569
↓ 1 callersClassDiffusionCondDemoCallback
ThinkSound/training/diffusion.py:462
↓ 1 callersClassDiffusionCondTrainingWrapper
Wrapper for training a conditional audio diffusion model.
ThinkSound/training/diffusion.py:45
↓ 1 callersClassDownsample1d
ThinkSound/models/blocks.py:111
↓ 1 callersClassEncoderBlock
ThinkSound/models/autoencoders.py:64
↓ 1 callersClassExceptionCallback
train.py:21
↓ 1 callersClassFIRFilter
FIR pre-emphasis filtering module. Args: filter_type (str): Shape of the desired FIR filter ("hp", "fd", "aw"). Default: "hp" coe
ThinkSound/training/losses/auraloss.py:76
↓ 1 callersClassFSQBottleneck
ThinkSound/models/bottleneck.py:313
↓ 1 callersClassFeaturesUtils
extract_latents.py:62
↓ 1 callersClassFeaturesUtils
data_utils/v2a_utils/feature_utils_224.py:52
↓ 1 callersClassFeaturesUtils
data_utils/v2a_utils/feature_utils_224_audio.py:54
↓ 1 callersClassFinalBlock
ThinkSound/models/transformer_layers.py:259
↓ 1 callersClassIntConditioner
ThinkSound/models/conditioners.py:298
↓ 1 callersClassJointBlock
ThinkSound/models/transformer_layers.py:211
↓ 1 callersClassL1Loss
ThinkSound/training/losses/losses.py:25
↓ 1 callersClassL2Bottleneck
ThinkSound/models/bottleneck.py:129
↓ 1 callersClassLatentDataset
ThinkSound/data/dataset.py:268
↓ 1 callersClassLocalWebDatasetConfig
ThinkSound/data/dataset.py:938
↓ 1 callersClassMMDiTWrapper
ThinkSound/models/diffusion.py:575
↓ 1 callersClassMMmodule
ThinkSound/models/mmdit.py:27
↓ 1 callersClassMSELoss
ThinkSound/training/losses/losses.py:44
↓ 1 callersClassMetaCLIPTextConditioner
ThinkSound/models/conditioners.py:694
↓ 1 callersClassMlp
data_utils/ext/synchformer/vit_helper.py:189
↓ 1 callersClassModelConfigEmbedderCallback
train.py:25
↓ 1 callersClassMotionFormer
This class serves three puposes: 1. Renames the class to MotionFormer. 2. Downloads the cfg from the original repo and patche
data_utils/ext/synchformer/motionformer.py:31
↓ 1 callersClassMultiConditioner
A module that applies multiple conditioners to an input dictionary based on the keys Args: conditioners: a dictionary of conditioner
ThinkSound/models/conditioners.py:889
↓ 1 callersClassNumberConditioner
Conditioner that takes a list of floats, normalizes them for a given range, and returns a list of embeddings
ThinkSound/models/conditioners.py:321
↓ 1 callersClassOobleckDecoder
ThinkSound/models/autoencoders.py:150
↓ 1 callersClassOobleckEncoder
ThinkSound/models/autoencoders.py:116
↓ 1 callersClassPQMFPretransform
ThinkSound/models/pretransforms.py:111
↓ 1 callersClassPadCrop
ThinkSound/data/utils.py:9
↓ 1 callersClassPadCrop_DualVideo_Normalized_T
ThinkSound/data/utils.py:258
↓ 1 callersClassPadCrop_Video_Hiera_Normalized_T
ThinkSound/data/utils.py:189
↓ 1 callersClassPadCrop_Video_Image_Normalized_T
ThinkSound/data/utils.py:131
↓ 1 callersClassPadCrop_Video_Normalized_T
ThinkSound/data/utils.py:73
↓ 1 callersClassPhonemeConditioner
A conditioner that turns text into phonemes and embeds them using a lookup table Only works for English text Args: output_dim: t
ThinkSound/models/conditioners.py:745
↓ 1 callersClassPreprocessedConditions
ThinkSound/models/mmdit.py:19
↓ 1 callersClassPretrainedDACPretransform
ThinkSound/models/pretransforms.py:133
↓ 1 callersClassPretransformConditioner
A conditioner that uses a pretransform's encoder for conditioning Args: pretransform: an instantiated pretransform to use for condit
ThinkSound/models/conditioners.py:859
↓ 1 callersClassProfiler
ThinkSound/training/diffusion.py:28
↓ 1 callersClassProfiler
ThinkSound/models/diffusion.py:18
next →1–100 of 175, ranked by callers