MCPcopy Create free account

hub / github.com/bytedance/LatentSync / functions

Functions512 in github.com/bytedance/LatentSync

↓ 1 callersFunction_lws_processor
()
latentsync/utils/audio.py:68
↓ 1 callersMethod_main_loop
(self, audio_features: Tensor, tokens: Tensor)
latentsync/whisper/whisper/decoding.py:591
↓ 1 callersFunction_no_grad_trunc_normal_
(tensor, mean, std, a, b)
latentsync/trepa/third_party/VideoMAEv2/videomaev2_finetune.py:21
↓ 1 callersFunction_ntuple
(n)
latentsync/trepa/third_party/VideoMAEv2/videomaev2_finetune.py:79
↓ 1 callersMethod_validate_indices
Validate int64 integers and convert negative integers to positive by backward search
latentsync/utils/av_reader.py:147
↓ 1 callersMethod_verify_options
(self, options: DecodingOptions)
latentsync/whisper/whisper/decoding.py:499
↓ 1 callersFunctionadd_segment
( *, start: float, end: float, encoder_embeddings )
latentsync/whisper/whisper/transcribe.py:88
↓ 1 callersFunctionadjust_offset
(video_input: str, video_output: str, av_offset: int, fps: int = 25)
preprocess/sync_av.py:42
↓ 1 callersMethodalign_warp_face
(self, img, landmarks3, smooth=True)
latentsync/utils/affine_transform.py:25
↓ 1 callersFunctionbounding_box_iou
(boxA, boxB)
eval/syncnet_detect.py:238
↓ 1 callersFunctionbuild_tokenizer
(name: str = "gpt2")
latentsync/whisper/whisper/tokenizer.py:274
↓ 1 callersFunctioncalc_pdist
(feat1, feat2, vshift=10)
eval/syncnet/syncnet_eval.py:21
↓ 1 callersFunctioncheck_ffmpeg_installed
()
latentsync/utils/util.py:270
↓ 1 callersMethodcheck_inputs
(self, height, width, callback_steps)
latentsync/pipelines/lipsync_pipeline.py:163
↓ 1 callersFunctioncheck_mem
(cuda_device)
tools/occupy_gpu.py:21
↓ 1 callersMethodcleanup_caching
Clean up any resources or hooks after decoding is finished
latentsync/whisper/whisper/decoding.py:127
↓ 1 callersMethodclose
(self)
preprocess/remove_incorrect_affined.py:52
↓ 1 callersFunctioncombine_video_audio
(video_frames, video_input_path, video_output_path, process_temp_dir)
preprocess/affine_transform.py:38
↓ 1 callersFunctioncompression_ratio
(text)
latentsync/whisper/whisper/utils.py:26
↓ 1 callersFunctioncompute_fvd
(feats_fake: np.ndarray, feats_real: np.ndarray)
eval/fvd.py:10
↓ 1 callersFunctioncount_total_videos_time
(fileslist_path: str)
tools/count_total_videos_time.py:5
↓ 1 callersFunctioncreate_args
( video_path: str, audio_path: str, output_path: str, inference_steps: int, guidance_scale: float, seed: i
gradio_app.py:56
↓ 1 callersMethodcrop_audio_window
(self, original_mel, start_index)
latentsync/data/unet_dataset.py:64
↓ 1 callersMethodcrop_audio_window
(self, original_mel, start_index)
latentsync/data/syncnet_dataset.py:58
↓ 1 callersMethodcrop_overlap_audio_window
(self, audio_feat, start_index)
latentsync/whisper/audio2feature.py:142
↓ 1 callersMethodcrop_video
(self, track, cropfile, frames_dir, frame_rate, temp_dir, video_dir, crop_scale=0.4)
eval/syncnet_detect.py:168
↓ 1 callersFunctioncuda_to_int
Convert the string with format "cuda:X" to integer X.
latentsync/utils/face_detector.py:72
↓ 1 callersFunctiondata_processing_pipeline
( total_num_workers, per_gpu_num_workers, resolution, sync_conf_threshold, temp_dir, input_dir )
preprocess/data_processing_pipeline.py:29
↓ 1 callersFunctiondecode
Decode locations from predictions using priors to undo the encoding we did for offset regression at train time. Args: loc (tensor): lo
eval/detectors/s3fd/box_utils.py:42
↓ 1 callersMethoddecode_latents
(self, latents)
latentsync/pipelines/lipsync_pipeline.py:140
↓ 1 callersMethoddetect_face
(self, image)
eval/eval_fvd.py:30
↓ 1 callersMethoddetect_face
(self, frames_dir, facedet_scale=0.25)
eval/syncnet_detect.py:150
↓ 1 callersMethoddetect_face
(self, image)
preprocess/filter_high_resolution.py:44
↓ 1 callersMethoddetect_face
(self, image)
preprocess/remove_incorrect_affined.py:28
↓ 1 callersMethoddetect_faces
(self, image, conf_th=0.8, scales=[1])
eval/detectors/s3fd/__init__.py:29
↓ 1 callersFunctiondetect_shot
(video_input, output_dir)
preprocess/detect_shot.py:35
↓ 1 callersMethoddetect_video
(self, video_path)
preprocess/filter_high_resolution.py:63
↓ 1 callersMethoddetect_video
(self, video_path)
preprocess/remove_incorrect_affined.py:39
↓ 1 callersFunctiondownload_videos
(num_workers, video_urls, video_paths)
tools/download_web_videos.py:38
↓ 1 callersFunctiondownload_weights
(url, dest)
predict.py:13
↓ 1 callersMethoddraw
(self, save_path, plot_val=True)
eval/draw_syncnet_lines.py:31
↓ 1 callersFunctiondrop_path
Adapted from timm codebase
latentsync/trepa/third_party/VideoMAEv2/videomaev2_finetune.py:91
↓ 1 callersFunctioneval_fvd
(real_videos_dir: str, fake_videos_dir: str)
eval/eval_fvd.py:87
↓ 1 callersMethodevaluate
(self, video_path, temp_dir="temp", batch_size=20, vshift=15)
eval/syncnet/syncnet_eval.py:47
↓ 1 callersFunctionexact_div
(x, y)
latentsync/whisper/whisper/utils.py:5
↓ 1 callersFunctionextract_vid
(video_url)
tools/download_web_videos.py:43
↓ 1 callersMethodfeature2chunks
(self, feature_array, fps)
latentsync/whisper/audio2feature.py:88
↓ 1 callersFunctionfilter_high_resolution_multiprocessing
(input_dir, output_dir, resolution, num_workers)
preprocess/filter_high_resolution.py:96
↓ 1 callersFunctionfilter_video
(video_input, video_out, resolution)
preprocess/filter_high_resolution.py:76
↓ 1 callersMethodfinalize
Finalize search and return the final candidate sequences Parameters ---------- tokens : Tensor, shape = (n_audio, n_group, cu
latentsync/whisper/whisper/decoding.py:228
↓ 1 callersMethodforward_aud
(self, x)
eval/syncnet/syncnet.py:88
↓ 1 callersMethodforward_features
(self, x, mask)
latentsync/trepa/third_party/VideoMAEv2/videomaev2_pretrain.py:123
↓ 1 callersMethodforward_lip
(self, x)
eval/syncnet/syncnet.py:98
↓ 1 callersMethodforward_lipfeat
(self, x)
eval/syncnet/syncnet.py:107
↓ 1 callersFunctiongather_loss
(loss, device)
latentsync/utils/util.py:235
↓ 1 callersFunctiongather_paths
(input_dir, output_dir)
tools/move_files_recur.py:22
↓ 1 callersFunctiongather_paths
(input_dir, output_dir)
preprocess/segment_videos.py:23
↓ 1 callersFunctiongather_paths
(input_dir, output_dir)
preprocess/filter_visual_quality.py:30
↓ 1 callersFunctiongather_paths
(input_dir, output_dir)
preprocess/detect_shot.py:23
↓ 1 callersFunctiongather_paths
(input_dir, output_dir)
preprocess/sync_av.py:28
↓ 1 callersFunctiongather_paths
(input_dir, output_dir)
preprocess/resample_fps_hz.py:24
↓ 1 callersFunctiongather_video_paths
(input_dir, paths)
latentsync/utils/util.py:253
↓ 1 callersFunctiongather_video_paths
(input_dir, output_dir, resolution)
preprocess/filter_high_resolution.py:25
↓ 1 callersFunctiongather_video_paths
(input_dir, output_dir)
preprocess/affine_transform.py:26
↓ 1 callersMethodgetTensor
Returns a tensor of the video frames at the given index. Args: index: The index of the video frames to return.
latentsync/trepa/utils/data_utils.py:267
↓ 1 callersMethodget_all
Get all the stored features as NumPy Array. Returns: Concatenation of the stored features.
latentsync/trepa/utils/metric_utils.py:106
↓ 1 callersFunctionget_down_block
( down_block_type, num_layers, in_channels, out_channels, temb_channels, add_downsampl
latentsync/models/unet_blocks.py:11
↓ 1 callersMethodget_frames
(self, video_reader: VideoReader)
latentsync/data/unet_dataset.py:69
↓ 1 callersMethodget_frames
(self, video_reader: VideoReader)
latentsync/data/syncnet_dataset.py:63
↓ 1 callersFunctionget_position_angle_vec
(position)
latentsync/trepa/third_party/VideoMAEv2/videomaev2_finetune.py:361
↓ 1 callersFunctionget_up_block
( up_block_type, num_layers, in_channels, out_channels, prev_output_channel, temb_chan
latentsync/models/unet_blocks.py:82
↓ 1 callersFunctionget_video_fps
(video_path: str)
preprocess/resample_fps_hz.py:36
↓ 1 callersFunctioninference_video_from_fileslist
( video_fileslist: str, audio_fileslist: str, output_dir: str, unet_config_path: str, ckpt
eval/inference_videos.py:21
↓ 1 callersMethodinstall_kv_cache_hooks
The `MultiHeadAttention` module optionally accepts `kv_cache` which stores the key and value tensors calculated for the previous posi
latentsync/whisper/whisper/model.py:256
↓ 1 callersMethodinterpolate_pos_encoding
(self, t)
latentsync/trepa/third_party/VideoMAEv2/videomaev2_finetune.py:480
↓ 1 callersFunctionis_image_file
(filename)
latentsync/trepa/utils/data_utils.py:29
↓ 1 callersFunctionload_audio
Open an audio file and read as mono waveform, resampling as necessary Parameters ---------- file: str The audio file to open
latentsync/whisper/whisper/audio.py:22
↓ 1 callersMethodload_video_frames
Loads all the video frames under the dataroot and returns a list of all the video frames. Args: dataroot: The root direc
latentsync/trepa/utils/data_utils.py:237
↓ 1 callersFunctionload_videomae_model
(device, ckpt_path=None, with_cp=False)
latentsync/trepa/third_party/VideoMAEv2/utils.py:44
↓ 1 callersFunctionlog_mel_spectrogram
Compute the log-Mel spectrogram of Parameters ---------- audio: Union[str, np.ndarray, torch.Tensor], shape = (*) The path t
latentsync/whisper/whisper/audio.py:92
↓ 1 callersMethodloop_video
(self, whisper_chunks: list, video_frames: np.ndarray)
latentsync/pipelines/lipsync_pipeline.py:281
↓ 1 callersFunctionmain
(input_dir, output_dir)
tools/move_files_recur.py:36
↓ 1 callersFunctionmain
(input_dir, fig_path)
tools/plot_videos_time_distribution.py:33
↓ 1 callersFunctionmain
(urls_txt_path, output_dir, num_workers)
tools/download_web_videos.py:66
↓ 1 callersFunctionmain
()
tools/occupy_gpu.py:46
↓ 1 callersFunctionmain
()
eval/fvd.py:48
↓ 1 callersFunctionmain
(config)
eval/eval_syncnet_acc.py:27
↓ 1 callersFunctionmain
()
eval/eval_sync_conf.py:44
↓ 1 callersFunctionmain
(config)
scripts/train_syncnet.py:39
↓ 1 callersFunctionmain
(config)
scripts/train_unet.py:60
↓ 1 callersFunctionmel_filters
load the mel filterbank matrix for projecting STFT into a Mel spectrogram. Allows decoupling librosa dependency; saved using: np.sav
latentsync/whisper/whisper/audio.py:77
↓ 1 callersFunctionnms
Apply non-maximum suppression at test time to avoid detecting too many overlapping bounding boxes for a given object. Args: boxes: (te
eval/detectors/s3fd/box_utils.py:63
↓ 1 callersFunctionnms_
Courtesy of Ross Girshick [https://github.com/rbgirshick/py-faster-rcnn/blob/master/lib/nms/py_cpu_nms.py]
eval/detectors/s3fd/box_utils.py:8
↓ 1 callersFunctionnum_frames
Compute number of time frames of spectrogram
latentsync/utils/audio.py:83
↓ 1 callersFunctionone_step_sampling
(ddim_scheduler, pred_noise, timesteps, x_t)
latentsync/utils/util.py:168
↓ 1 callersMethodoutput
(result: Union[str, int])
latentsync/whisper/whisper/normalizers/english.py:171
↓ 1 callersFunctionpad_or_trim
Pad or trim the audio array to N_SAMPLES, as expected by the encoder.
latentsync/whisper/whisper/audio.py:52
↓ 1 callersMethodpaste_surrounding_pixels_back
(decoded_latents, pixel_values, masks, device, weight_dtype)
latentsync/pipelines/lipsync_pipeline.py:237
↓ 1 callersFunctionplot_histogram
(data, fig_path)
tools/plot_videos_time_distribution.py:20
↓ 1 callersMethodpostprocess
(self, s: str)
latentsync/whisper/whisper/normalizers/english.py:410
← previousnext →101–200 of 512, ranked by callers