MCPcopy Create free account

hub / github.com/antgroup/echomimic_v3 / functions

Functions301 in github.com/antgroup/echomimic_v3

↓ 28 callersMethodfrom_pretrained
(cls, pretrained_model_path, additional_kwargs={})
src/wan_vae.py:688
↓ 11 callersMethodset_timesteps
Sets the discrete timesteps used for the diffusion chain (to be run before inference). Args: num_inference_steps (`int`):
src/fm_solvers.py:226
↓ 10 callersMethod_sigma_to_alpha_sigma_t
(self, sigma)
src/fm_solvers.py:333
↓ 9 callersMethod__init__
(self, dim, out_dim, patch_size, eps=1e-6)
src/wan_transformer3d_audio_2512.py:972
↓ 9 callersMethod__init__
(self, dim, out_dim, patch_size, eps=1e-6)
src/wan_transformer3d_audio.py:922
↓ 8 callersMethod__init__
(self, dim, channel_first=True, images=True, bias=False)
src/wan_vae.py:47
↓ 8 callersMethod_sigma_to_alpha_sigma_t
(self, sigma)
src/fm_solvers_unipc.py:272
↓ 8 callersFunctionfilter_kwargs
(cls, kwargs)
src/utils.py:12
↓ 8 callersFunctionrope_params
(max_seq_len, dim, theta=10000)
src/wan_transformer3d_audio_2512.py:495
↓ 8 callersFunctionrope_params
(max_seq_len, dim, theta=10000)
src/wan_transformer3d_audio.py:435
↓ 7 callersMethod__init__
(self, dim, mid_dim)
src/wan_image_encoder.py:98
↓ 7 callersFunctionattention
( q, k, v, q_lens=None, k_lens=None, dropout_p=0., softmax_scale=None, q_scale
src/wan_transformer3d_audio.py:165
↓ 7 callersFunctionget_image_to_video_latent3
(validation_image_start, validation_image_end, video_length, sample_size)
src/utils.py:377
↓ 6 callersMethod__init__
(self, dim, eps=1e-6)
src/wan_text_encoder.py:45
↓ 6 callersFunctionattention
( q, k, v, q_lens=None, k_lens=None, dropout_p=0., softmax_scale=None, q_scale
src/wan_transformer3d_audio_2512.py:241
↓ 6 callersFunctionhalf
(x)
src/wan_transformer3d_audio_2512.py:102
↓ 6 callersFunctionhalf
(x)
src/wan_transformer3d_audio.py:80
↓ 5 callersFunctionfp16_clamp
(x)
src/wan_text_encoder.py:14
↓ 4 callersMethodclear_cache
(self)
src/wan_vae.py:592
↓ 4 callersMethodenable_riflex
( self, k = 6, L_test = 66, L_test_scale = 4.886, )
src/wan_transformer3d_audio.py:1128
↓ 4 callersMethodenable_teacache
( self, coefficients, num_steps: int, rel_l1_thresh: float, num_skip_s
src/wan_transformer3d_audio.py:1113
↓ 4 callersMethodencode
( self, extract_features, attention_mask=None, mask_time_indices=None,
src/wav2vec2.py:97
↓ 4 callersFunctionflash_attention
q: [B, Lq, Nq, C1]. k: [B, Lk, Nk, C1]. v: [B, Lk, Nk, C2]. Nq must be divisible by Nk. q_lens
src/wan_transformer3d_audio.py:45
↓ 4 callersFunctionget_mask_coord
(image_path)
src/face_detect.py:8
↓ 4 callersFunctionget_teacache_coefficients
(model_name)
src/cache_utils.py:5
↓ 4 callersFunctionsave_videos_grid
(videos: torch.Tensor, path: str, rescale=False, n_rows=6, fps=12, imageio_backend=True, color_transfer_post_p
src/utils.py:54
↓ 3 callersMethoddecode
(self, z: torch.Tensor, return_dict: bool = True)
src/wan_vae.py:680
↓ 3 callersFunctionflash_attention
q: [B, Lq, Nq, C1]. k: [B, Lk, Nk, C1]. v: [B, Lk, Nk, C2]. Nq must be divisible by Nk. q_lens
src/wan_transformer3d_audio_2512.py:67
↓ 3 callersFunctionretrieve_timesteps
Calls the scheduler's `set_timesteps` method and retrieves timesteps from the scheduler after the call. Handles custom timesteps. Any kwargs
src/pipeline_wan_fun_inpaint_audio.py:45
↓ 3 callersFunctionretrieve_timesteps
Calls the scheduler's `set_timesteps` method and retrieves timesteps from the scheduler after the call. Handles custom timesteps. Any kwargs
src/pipeline_wan_fun_inpaint_audio_2512.py:48
↓ 3 callersMethodscale_model_input
Ensures interchangeability with schedulers that need to scale the denoising model input depending on the current timestep. Ar
src/fm_solvers.py:800
↓ 3 callersMethodstep
Predict the sample from the previous timestep by reversing the SDE. This function propagates the sample with the multistep DPMSolver.
src/fm_solvers.py:706
↓ 2 callersMethod__init__
(self, vocab_size=250002, max_seq_len=514, type_size=1,
src/wan_xlm_roberta.py:81
↓ 2 callersMethod_get_t5_prompt_embeds
( self, prompt: Union[str, List[str]] = None, num_videos_per_prompt: int = 1,
src/pipeline_wan_fun_inpaint_audio.py:193
↓ 2 callersMethod_get_t5_prompt_embeds
( self, prompt: Union[str, List[str]] = None, num_videos_per_prompt: int = 1,
src/pipeline_wan_fun_inpaint_audio_2512.py:197
↓ 2 callersMethod_threshold_sample
"Dynamic thresholding: At each sampling step we set s to a certain percentile absolute pixel value in xt0 (the prediction of x_0 at t
src/fm_solvers.py:292
↓ 2 callersMethod_threshold_sample
"Dynamic thresholding: At each sampling step we set s to a certain percentile absolute pixel value in xt0 (the prediction of x_0 at t
src/fm_solvers_unipc.py:230
↓ 2 callersFunctioncolor_transfer
Transfer color distribution from of sc, referred to dc. Args: sc (numpy.ndarray): input image to be transfered. dc (numpy.nd
src/utils.py:26
↓ 2 callersMethodcompute_rel_l1_distance
(prev: torch.Tensor, cur: torch.Tensor)
src/cache_utils.py:62
↓ 2 callersFunctioncount_conv3d
(model)
src/wan_vae.py:481
↓ 2 callersMethoddecode_latents
(self, latents: torch.Tensor)
src/pipeline_wan_fun_inpaint_audio.py:380
↓ 2 callersMethoddecode_latents
(self, latents: torch.Tensor)
src/pipeline_wan_fun_inpaint_audio_2512.py:379
↓ 2 callersMethodencode
(self, x, scale)
src/wan_vae.py:522
↓ 2 callersMethodencode_prompt
r""" Encodes the prompt into text encoder hidden states. Args: prompt (`str` or `List[str]`, *optional*):
src/pipeline_wan_fun_inpaint_audio.py:237
↓ 2 callersMethodforward
(self, x)
src/wan_vae.py:57
↓ 2 callersMethodforward
(self, x)
src/wan_image_encoder.py:108
↓ 2 callersFunctionget_mean_and_std
(img)
src/utils.py:38
↓ 2 callersFunctionget_sampling_sigmas
(sampling_steps, shift)
src/fm_solvers.py:22
↓ 2 callersMethodindex_for_timestep
(self, timestep, schedule_timesteps=None)
src/fm_solvers.py:679
↓ 2 callersMethodindex_for_timestep
(self, timestep, schedule_timesteps=None)
src/fm_solvers_unipc.py:628
↓ 2 callersFunctionlinear_interpolation
(features, seq_len)
src/wav2vec2.py:19
↓ 2 callersMethodreset
(self)
src/cache_utils.py:67
↓ 2 callersFunctionrope_apply
(x, grid_sizes, freqs)
src/wan_transformer3d_audio_2512.py:583
↓ 2 callersFunctionrope_apply
(x, grid_sizes, freqs)
src/wan_transformer3d_audio.py:523
↓ 1 callersMethod__init__
(self, in_dim, out_dim, kernel_size, stride, num_residual_blocks=1)
src/wan_camera_adapter.py:5
↓ 1 callersFunction_clip
(pretrained=False, pretrained_name=None, model_cls=XLMRobertaCLIP, return_transf
src/wan_image_encoder.py:436
↓ 1 callersMethod_decode
(self, zs)
src/wan_vae.py:670
↓ 1 callersMethod_encode
(self, x: torch.Tensor)
src/wan_vae.py:650
↓ 1 callersMethod_init_step_index
Initialize the step_index counter for the scheduler.
src/fm_solvers.py:693
↓ 1 callersMethod_init_step_index
Initialize the step_index counter for the scheduler.
src/fm_solvers_unipc.py:643
↓ 1 callersMethod_norm
(self, x)
src/wan_transformer3d_audio_2512.py:628
↓ 1 callersMethod_norm
(self, x)
src/wan_transformer3d_audio.py:568
↓ 1 callersMethod_relative_position_bucket
(self, rel_pos)
src/wan_text_encoder.py:235
↓ 1 callersFunction_video_vae
Autoencoder3d adapted from Stable Diffusion 1.x, 2.x and XL.
src/wan_vae.py:602
↓ 1 callersFunctionaudio_mask_attention
( q, k, v, q_lens=None, k_lens=None, dropout_p=0., softmax_scale=None, q_scale
src/wan_transformer3d_audio_2512.py:432
↓ 1 callersFunctionaudio_mask_attention
( q, k, v, q_lens=None, k_lens=None, dropout_p=0., softmax_scale=None, q_scale
src/wan_transformer3d_audio.py:357
↓ 1 callersMethodcheck_inputs
( self, prompt, height, width, negative_prompt, callback_on_st
src/pipeline_wan_fun_inpaint_audio.py:406
↓ 1 callersMethodcheck_inputs
( self, prompt, height, width, negative_prompt, callback_on_st
src/pipeline_wan_fun_inpaint_audio_2512.py:400
↓ 1 callersFunctionclip_xlm_roberta_vit_h_14
( pretrained=False, pretrained_name='open-clip-xlm-roberta-large-vit-huge-14', **kwarg
src/wan_image_encoder.py:473
↓ 1 callersMethodconvert_model_output
Convert the model output to the corresponding type the DPMSolver/DPMSolver++ algorithm needs. DPM-Solver is designed to discretize an
src/fm_solvers.py:341
↓ 1 callersMethodconvert_model_output
r""" Convert the model output to the corresponding type the UniPC algorithm needs. Args: model_output (`torch.Tensor`):
src/fm_solvers_unipc.py:279
↓ 1 callersMethoddecode
(self, z, scale)
src/wan_vae.py:553
↓ 1 callersMethoddpm_solver_first_order_update
One step for the first-order DPMSolver (equivalent to DDIM). Args: model_output (`torch.Tensor`): The dir
src/fm_solvers.py:415
↓ 1 callersMethodenable_multi_gpus_inference
(self,)
src/wan_transformer3d_audio.py:1155
↓ 1 callersMethodencode
( self, x: torch.Tensor, return_dict: bool = True )
src/wan_vae.py:659
↓ 1 callersMethodencode_prompt
r""" Encodes the prompt into text encoder hidden states. Args: prompt (`str` or `List[str]`, *optional*):
src/pipeline_wan_fun_inpaint_audio_2512.py:241
↓ 1 callersFunctionextract_audio_features
Extract audio features using Wav2Vec.
app_mm.py:123
↓ 1 callersFunctionextract_audio_features
Extract audio features using Wav2Vec.
infer_preview.py:118
↓ 1 callersFunctionextract_audio_features
Extract audio features using Wav2Vec.
app.py:112
↓ 1 callersMethodforward
r""" Args: x(Tensor): Shape [B, L1, C] e(Tensor): Shape [B, C]
src/wan_transformer3d_audio_2512.py:987
↓ 1 callersMethodforward
r""" Args: x(Tensor): Shape [B, L1, C] e(Tensor): Shape [B, C]
src/wan_transformer3d_audio.py:937
↓ 1 callersFunctionget_1d_rotary_pos_embed_riflex
RIFLEx: Precompute the frequency tensor for complex exponentials (cis) with given dimensions. This function calculates a frequency tensor wi
src/wan_transformer3d_audio_2512.py:506
↓ 1 callersFunctionget_1d_rotary_pos_embed_riflex
RIFLEx: Precompute the frequency tensor for complex exponentials (cis) with given dimensions. This function calculates a frequency tensor wi
src/wan_transformer3d_audio.py:446
↓ 1 callersFunctionget_ai2v_audio
( audio, vae_scale, audio_window, device, dtype )
src/wan_transformer3d_audio_2512.py:45
↓ 1 callersFunctionget_audio_embed
(mel_input, wav2vec_feature_extractor, audio_encoder, video_length, sr=16000, fps=25, device='cpu')
infer_flash.py:134
↓ 1 callersFunctionget_audio_mask
ip_mask B, n_q
src/wan_transformer3d_audio.py:408
↓ 1 callersFunctionget_file_path
Helper function to find the file path with multiple extensions.
infer_preview.py:152
↓ 1 callersFunctionget_image_to_video_latent2
(validation_image_start, validation_image_end, video_length, sample_size)
src/utils.py:278
↓ 1 callersFunctionget_ip_mask
(coords)
app_mm.py:150
↓ 1 callersFunctionget_ip_mask
(coords)
infer_preview.py:144
↓ 1 callersFunctionget_ip_mask
(coords)
app.py:139
↓ 1 callersFunctionget_sample_size
(pil_img, sample_size)
infer_flash.py:110
↓ 1 callersFunctionget_sample_size
Calculate the sample size based on the input image dimensions.
app_mm.py:133
↓ 1 callersFunctionget_sample_size
Calculate the sample size based on the input image dimensions.
infer_preview.py:127
↓ 1 callersFunctionget_sample_size
Calculate the sample size based on the input image dimensions.
app.py:122
↓ 1 callersFunctionload_wav2vec_models
Load Wav2Vec models for audio feature extraction.
app_mm.py:115
↓ 1 callersFunctionload_wav2vec_models
Load Wav2Vec models for audio feature extraction.
infer_preview.py:110
↓ 1 callersFunctionload_wav2vec_models
Load Wav2Vec models for audio feature extraction.
app.py:104
↓ 1 callersFunctionloudness_norm
(audio_array, sr=16000, lufs=-23)
infer_flash.py:151
↓ 1 callersFunctionmain
()
infer_flash.py:159
next →1–100 of 301, ranked by callers