MCPcopy Create free account

hub / github.com/VITA-MLLM/VITA / functions

Functions1,650 in github.com/VITA-MLLM/VITA

↓ 5 callersFunctionexpand2square
(pil_img, background_color)
vita/util/data_utils_video_audio_neg_patch.py:782
↓ 5 callersFunctionexpand2square
(pil_img, background_color)
vita/util/data_utils_video_audio_neg_patch_fo.py:780
↓ 5 callersFunctionexpand2square
(pil_img, background_color)
vita/util/data_utils_video_patch_audio.py:641
↓ 5 callersFunctionexpand2square
(pil_img, background_color)
vita/util/data_utils_video_audio_neg_frameCat.py:525
↓ 5 callersFunctionexpand2square
(pil_img, background_color)
vita/util/data_utils_video_audio.py:326
↓ 5 callersFunctionexpand2square
(pil_img, background_color)
vita/util/data_utils_video_audio_patch_sf.py:640
↓ 5 callersFunctionget_dimension_rating
(data_path)
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:10
↓ 5 callersFunctiongpt_key_set
()
VLMEvalKit/vlmeval/smp/vlm.py:139
↓ 5 callersMethodmessage_to_promptimg
(self, message, dataset=None)
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer2d5.py:134
↓ 5 callersFunctionmodel_gen
(model, text, images, need_bos=True, padding=False, beams=3, max_token=500)
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer2_4KHD.py:60
↓ 5 callersFunctionmodel_gen
(model, text, images, need_bos=True, padding=False, beams=3, max_token=500)
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer2d5.py:62
↓ 5 callersFunctionpreprocess_multimodal
( sources: Sequence[str], data_args: DataArguments, image_token_num=1, audio_lens: int = 0 )
vita/util/data_utils_video_audio.py:39
↓ 5 callersFunctionproxy_set
(s)
VLMEvalKit/vlmeval/smp/misc.py:95
↓ 5 callersMethodrun
Run the speech decoder process. Parameters: - hidden (torch.Tensor): The output for embedding layer of the language model.
vita/model/vita_tts/decoder/llm2tts.py:114
↓ 4 callersFunctionanls_compute
(groundtruth, prediction, threshold=0.5)
VLMEvalKit/vlmeval/dataset/mmlongbench.py:102
↓ 4 callersFunctionbuild_dataset
(dataset_name, **kwargs)
VLMEvalKit/vlmeval/dataset/__init__.py:155
↓ 4 callersFunctiondisable_torch_init
Disable the redundant torch default initialization to accelerate model creation.
vita/util/utils.py:14
↓ 4 callersMethoddump_image
(self, origin_line)
VLMEvalKit/vlmeval/dataset/dude.py:88
↓ 4 callersFunctionencode_image_file_to_base64
(image_path, target_size=-1)
VLMEvalKit/vlmeval/smp/vlm.py:96
↓ 4 callersFunctionexpand2square
(pil_img, background_color)
vita/util/mm_utils.py:16
↓ 4 callersMethodframe_paths
(self, video, num_frames=8)
VLMEvalKit/vlmeval/dataset/video_base.py:47
↓ 4 callersFunctionget_length_grouped_indices
(lengths, batch_size, world_size, generator=None, merge=True)
vita/train/vita_trainer.py:100
↓ 4 callersFunctionhit_calculate
(result, dataset_name, anls_threshold=0.5)
VLMEvalKit/vlmeval/dataset/utils/vqa_eval.py:160
↓ 4 callersMethodinfer
(self, xs, buffer, buffer_index, buffer_out, pe_index)
vita/model/vita_tts/encoder/transformer.py:267
↓ 4 callersMethodinterrupt
(self)
web_demo/vita_html/web/parms.py:33
↓ 4 callersFunctionload_env
()
VLMEvalKit/vlmeval/smp/misc.py:154
↓ 4 callersMethodload_model
(self)
web_demo/wakeup_and_vad/wakeup_and_vad.py:128
↓ 4 callersFunctionnormalize
(x)
VLMEvalKit/vlmeval/dataset/utils/tablevqabench.py:209
↓ 4 callersFunctionparse_file
(s)
VLMEvalKit/vlmeval/smp/file.py:280
↓ 4 callersMethodpool_feats
(self, x, out_size)
vita/model/vita_arch.py:146
↓ 4 callersFunctionprocess_answer
(answer)
VLMEvalKit/vlmeval/dataset/utils/vqa_eval.py:239
↓ 4 callersFunctionrepeat
Repeat module N times. :param int N: repeat time :param function fn: function to generate module :return: repeated modules :rtype: Mu
vita/model/multimodal_encoder/whale/module/component/transformer.py:24
↓ 4 callersMethoduse_custom_prompt
(self, dataset)
VLMEvalKit/vlmeval/vlm/ovis.py:44
↓ 3 callersFunctionMMLongBench_auxeval
(model, line)
VLMEvalKit/vlmeval/dataset/mmlongbench.py:341
↓ 3 callersMethodadd_arguments
Add TDNN common arguments.
vita/model/multimodal_encoder/whale/module/component/mamba.py:85
↓ 3 callersMethodbuild_prompt_default
(self, message, add_brief=False, add_yes_or_no=False)
VLMEvalKit/vlmeval/vlm/idefics.py:93
↓ 3 callersFunctioncalc_aAcc
(data)
VLMEvalKit/vlmeval/dataset/utils/yorn.py:67
↓ 3 callersFunctioncalc_fAcc
(data)
VLMEvalKit/vlmeval/dataset/utils/yorn.py:51
↓ 3 callersFunctioncalc_qAcc
(data)
VLMEvalKit/vlmeval/dataset/utils/yorn.py:59
↓ 3 callersFunctionconcat_images
(image_list, max_concat=1, column_num=1)
VLMEvalKit/vlmeval/dataset/mmlongbench.py:263
↓ 3 callersFunctiond2df
(D)
VLMEvalKit/vlmeval/smp/misc.py:115
↓ 3 callersFunctionextract_characters_regex
(s)
VLMEvalKit/vlmeval/dataset/utils/videomme.py:118
↓ 3 callersMethodfeature_select
(self, image_forward_outs)
vita/model/multimodal_encoder/siglip/siglip_encoder.py:30
↓ 3 callersMethodgenerate_vanilla
(self, image_path, text)
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer2_4KHD.py:138
↓ 3 callersMethodgenerate_vanilla
(self, image_path, text)
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer2.py:117
↓ 3 callersMethodgenerate_vanilla
(self, image_path, text)
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer2d5.py:193
↓ 3 callersMethodget_index
(self, bound, fps, max_frame, first_idx=0)
VLMEvalKit/vlmeval/dataset/mvbench.py:190
↓ 3 callersMethodinfer_data
( self, data_sample, system=' ', question_prompt='', # add in the end of question ans
VLMEvalKit/vlmeval/vlm/video_llm/videochat2.py:302
↓ 3 callersFunctionmaybe_zero_3
(param, ignore_status=False, name=None)
vita/train/train.py:98
↓ 3 callersFunctionmcq_vanilla_eval
(model, data, meta, nproc, result_file, dataset_name=None)
VLMEvalKit/vlmeval/dataset/utils/multiple_choice.py:337
↓ 3 callersFunctionpost_check
(line, prefetch=False)
VLMEvalKit/vlmeval/dataset/utils/mathvista.py:74
↓ 3 callersFunctionpost_check
(line, prefetch=False)
VLMEvalKit/vlmeval/dataset/utils/mathv.py:100
↓ 3 callersMethodread_video
(self, video_path, bound=None)
VLMEvalKit/vlmeval/vlm/video_llm/videochat2.py:228
↓ 3 callersMethodremove_weight_norm
(self)
vita/model/vita_tts/decoder/ticodec/models.py:516
↓ 3 callersFunctionreport_acc
(df)
VLMEvalKit/vlmeval/dataset/utils/multiple_choice.py:68
↓ 3 callersMethodreset
(self)
web_demo/vita_html/web/parms.py:16
↓ 3 callersMethodreset_states
(self)
web_demo/wakeup_and_vad/wakeup_and_vad.py:53
↓ 3 callersMethodsave_video_frames
(self, imgs, video_name, frames)
VLMEvalKit/vlmeval/dataset/mvbench.py:241
↓ 3 callersFunctionsubsequent_mask
Create mask for subsequent steps (size, size). This mask is used only in decoder which works in an auto-regressive mode. This means the curre
vita/model/vita_tts/masks.py:153
↓ 3 callersMethodsupported_datasets
(cls)
VLMEvalKit/vlmeval/dataset/__init__.py:82
↓ 2 callersFunctionHD_transform
(img, im_num=36, id_scale=1.5)
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer2d5.py:29
↓ 2 callersFunctionMISSING
(lvl)
VLMEvalKit/vlmeval/tools.py:132
↓ 2 callersFunctionRUN
(lvl, model)
VLMEvalKit/vlmeval/tools.py:287
↓ 2 callersFunctionYOrN_Extraction
(output)
VLMEvalKit/vlmeval/dataset/utils/yorn.py:185
↓ 2 callersMethod__init__
( self, enc_out_dim: int = 512, llm_embed_dim: int = 4096, kernel_size: int =
vita/model/vita_tts/adapter.py:11
↓ 2 callersMethod__init__
(self, args)
vita/model/vita_tts/encoder/subsampling.py:86
↓ 2 callersMethod__init__
( self, enc_out_dim: int = 512, llm_embed_dim: int = 4096, kernel_size: int =
vita/model/multimodal_encoder/whale/adapter.py:7
↓ 2 callersMethod__init__
(self, args)
vita/model/multimodal_encoder/whale/module/component/subsampling.py:57
↓ 2 callersMethod_generate_one_step
Generates the model's next output based on the current input and state. Parameters: - inputs: The input tensor containing th
vita/model/vita_tts/audioLLM.py:387
↓ 2 callersFunction_get_rawvideo_dec
(video_path, image_processor, max_frames=MAX_IMAGE_LENGTH, image_resolution=224, video_framerate=1, s=None, e=
videomme/yt_video_inference_qa.py:185
↓ 2 callersFunction_get_rawvideo_dec
( video_path, image_processor, max_frames=32, min_frames=4, image_resolution=384, vide
vita/util/data_utils_video_audio_patch.py:579
↓ 2 callersFunction_get_rawvideo_dec
( video_path, image_processor, max_frames=32, min_frames=4, image_resolution=384, vide
vita/util/data_utils_video_audio_neg_patch.py:721
↓ 2 callersFunction_get_rawvideo_dec
( video_path, image_processor, max_frames=32, min_frames=4, image_resolution=384, vide
vita/util/data_utils_video_audio_neg_patch_fo.py:719
↓ 2 callersFunction_get_rawvideo_dec
( video_path, image_processor, max_frames=32, min_frames=4, image_resolution=384, vide
vita/util/data_utils_video_patch_audio.py:587
↓ 2 callersFunction_get_rawvideo_dec
( video_path, image_processor, max_frames=MAX_IMAGE_LENGTH, min_frames=MIN_IMAGE_LENGTH, i
vita/util/data_utils_video_audio_neg_frameCat.py:442
↓ 2 callersFunction_get_rawvideo_dec
( video_path, image_processor, max_frames=32, min_frames=4, image_resolution=384, vide
vita/util/data_utils_video_audio.py:265
↓ 2 callersFunction_get_rawvideo_dec
( video_path, image_processor, max_frames=32, min_frames=4, image_resolution=384, vide
vita/util/data_utils_video_audio_patch_sf.py:579
↓ 2 callersMethod_merge_multimodal_embeddings
Merge `vision_embeddings` into `input_embeds` by overwriting the positions in `input_embeds` corresponding to placeholder image token
web_demo/vllm_tools/vllm_file/qwen2.py:1064
↓ 2 callersFunction_parse_text
(text)
web_demo/web_ability_demo.py:209
↓ 2 callersMethod_process
(self, formatted_messages, formatted_images)
VLMEvalKit/vlmeval/vlm/idefics.py:86
↓ 2 callersFunction_to_float
(text: str)
VLMEvalKit/vlmeval/dataset/utils/vqa_eval.py:194
↓ 2 callersFunctionacc
(key, mode='normal')
VLMEvalKit/vlmeval/dataset/utils/yorn.py:17
↓ 2 callersMethodadd_arguments
Add Subsampling common arguments.
vita/model/vita_tts/encoder/subsampling.py:77
↓ 2 callersFunctionadd_encoder_args
Add Encoder common arguments.
vita/model/vita_tts/encoder/encoder.py:12
↓ 2 callersFunctionadd_optional_chunk_mask
Apply optional mask for encoder. Args: xs (torch.Tensor): padded input, (B, L, D), L for max length mask (torch.Tensor): mask fo
vita/model/vita_tts/masks.py:59
↓ 2 callersFunctionadd_optional_chunk_mask
( xs: torch.Tensor, masks: torch.Tensor, use_dynamic_chunk: bool, use_dynamic_left_chunk: bool
vita/model/multimodal_encoder/whale/utils.py:105
↓ 2 callersMethodanswer
(self, conv, model, img_list, do_sample=True, max_new_tokens=500, num_beams=1, min_length=1, top_p=0.9,
VLMEvalKit/vlmeval/vlm/video_llm/videochat2.py:269
↓ 2 callersMethodask
(self, text, conv)
VLMEvalKit/vlmeval/vlm/video_llm/videochat2.py:241
↓ 2 callersFunctionassign_args_from_dict
(args, dict, prefix_key=None)
vita/model/vita_tts/encoder/encoder.py:36
↓ 2 callersFunctionassign_args_from_dict
(args, dict, prefix_key=None)
vita/model/multimodal_encoder/whale/module/encoder/encoder.py:45
↓ 2 callersFunctionbuild_audio_encoder
(audio_encoder_config, **kwargs)
vita/model/multimodal_encoder/builder.py:46
↓ 2 callersFunctionbuild_choices
(item)
VLMEvalKit/vlmeval/dataset/utils/multiple_choice.py:224
↓ 2 callersFunctionbuild_prompt
(line)
VLMEvalKit/vlmeval/dataset/utils/llavabench.py:32
↓ 2 callersMethodbuild_prompt
(self, line)
VLMEvalKit/vlmeval/dataset/image_vqa.py:410
↓ 2 callersFunctionbuild_transform
(input_size)
VLMEvalKit/vlmeval/vlm/mmalaya.py:72
↓ 2 callersMethodbuild_video_prompt
(self, prompt, dataset=None, max_nframe=64)
VLMEvalKit/vlmeval/vlm/internvl_chat.py:210
↓ 2 callersFunctionbuild_vision_projector
(config, delay_load=False, **kwargs)
vita/model/multimodal_projector/builder.py:154
↓ 2 callersFunctionbuild_vision_tower
(vision_tower_cfg, **kwargs)
vita/model/multimodal_encoder/builder.py:14
↓ 2 callersFunctioncal_f1_score
(y_true, y_pred)
VLMEvalKit/vlmeval/dataset/utils/yorn.py:103
↓ 2 callersFunctioncheck_ans
(pred, gt)
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:32
← previousnext →101–200 of 1,650, ranked by callers