Code
Hub
Workspaces
Following
Trending
Connect
MCP
copy
Create free account
hub
/
github.com/VITA-MLLM/VITA
/ types & classes
Types & classes
320 in github.com/VITA-MLLM/VITA
⨍
Functions
1,650
◇
Types & classes
320
↳
Endpoints
1
↓ 13 callers
Class
KeywordsStoppingCriteria
vita/util/mm_utils.py:121
↓ 9 callers
Class
Conversation
A class that keeps all conversation history.
vita/conversation.py:18
↓ 6 callers
Class
Quantizer_module
vita/model/vita_tts/decoder/ticodec/models.py:525
↓ 5 callers
Class
DiscriminatorP
vita/model/vita_tts/decoder/ticodec/models.py:257
↓ 5 callers
Class
GroupScale
Rescales the input PIL.Image to the given 'size'. 'size' will be the size of the smaller edge. For example, if height > width, then image wil
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:150
↓ 4 callers
Class
CrossEntropyLoss
vita/model/vita_tts/decoder/decoder.py:16
↓ 3 callers
Class
Block
vita/model/multimodal_encoder/eva_clip/eva_vit.py:430
↓ 3 callers
Class
DiscriminatorS
vita/model/vita_tts/decoder/ticodec/models.py:337
↓ 3 callers
Class
DropPath
Drop paths (Stochastic Depth) per sample (when applied in main path of residual blocks).
vita/model/multimodal_encoder/eva_clip/eva_vit.py:181
↓ 2 callers
Class
Attention
vita/model/multimodal_encoder/eva_clip/eva_vit.py:263
↓ 2 callers
Class
GroupCenterCrop
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:106
↓ 2 callers
Class
GroupNormalize
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:134
↓ 2 callers
Class
InternRMSNorm
vita/model/multimodal_encoder/internvit/modeling_intern_vit.py:34
↓ 2 callers
Class
NumberValue
VLMEvalKit/vlmeval/dataset/utils/tablevqabench.py:284
↓ 2 callers
Class
Qwen2Model
web_demo/vllm_tools/vllm_file/qwen2.py:571
↓ 2 callers
Class
Stack
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:405
↓ 2 callers
Class
StoppingCriteriaSub
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer.py:10
↓ 2 callers
Class
StreamToLogger
Fake file-like stream object that redirects writes to a logger instance.
vita/util/utils.py:68
↓ 2 callers
Class
ToTorchFormatTensor
Converts a PIL.Image (RGB) or numpy.ndarray (H x W x C) in the range [0, 255] to a torch.FloatTensor of shape (C x H x W) in the range [0.0, 1.0]
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:424
↓ 2 callers
Class
VADIterator
web_demo/wakeup_and_vad/wakeup_and_vad.py:12
↓ 2 callers
Class
llm2TTS
vita/model/vita_tts/decoder/llm2tts.py:17
↓ 1 callers
Class
AttrDict
vita/model/vita_tts/decoder/ticodec/vqvae.py:10
↓ 1 callers
Class
AudioLLM
vita/model/vita_tts/audioLLM.py:20
↓ 1 callers
Class
CLIPVisionCfg
vita/model/multimodal_encoder/eva_clip/eva_vit.py:877
↓ 1 callers
Class
CLIPVisionTower
vita/model/multimodal_encoder/clip/clip_encoder.py:6
↓ 1 callers
Class
CNNAdapter
vita/model/vita_tts/adapter.py:10
↓ 1 callers
Class
CNNAdapter
vita/model/multimodal_encoder/whale/adapter.py:6
↓ 1 callers
Class
CNNSubsampling
vita/model/vita_tts/adapter.py:72
↓ 1 callers
Class
CNNSubsampling
vita/model/multimodal_encoder/whale/adapter.py:68
↓ 1 callers
Class
COCO_Caption_Scorer
VLMEvalKit/vlmeval/dataset/image_caption.py:5
↓ 1 callers
Class
Conv2dSubsampling4
Convolutional 2D subsampling (to 1/4 length). Args: idim (int): Input dimension. odim (int): Output dimension. dropout_ra
vita/model/vita_tts/encoder/subsampling.py:15
↓ 1 callers
Class
Conv2dSubsampling4
Convolutional 2D subsampling (to 1/4 length). Args: idim (int): Input dimension. odim (int): Output dimension. dropout_ra
vita/model/multimodal_encoder/whale/module/component/subsampling.py:15
↓ 1 callers
Class
CustomMCQDataset
VLMEvalKit/vlmeval/dataset/image_mcq.py:629
↓ 1 callers
Class
CustomTextMCQDataset
VLMEvalKit/vlmeval/dataset/text_mcq.py:112
↓ 1 callers
Class
CustomVQADataset
VLMEvalKit/vlmeval/dataset/image_vqa.py:473
↓ 1 callers
Class
DataCollatorForSupervisedDataset
Collate examples for supervised fine-tuning.
vita/util/data_utils_video_audio_patch.py:1248
↓ 1 callers
Class
DataCollatorForSupervisedDataset
Collate examples for supervised fine-tuning.
vita/util/data_utils_video_audio_neg_patch.py:1390
↓ 1 callers
Class
DataCollatorForSupervisedDataset
Collate examples for supervised fine-tuning.
vita/util/data_utils_video_audio_neg_patch_fo.py:1390
↓ 1 callers
Class
DataCollatorForSupervisedDataset
Collate examples for supervised fine-tuning.
vita/util/data_utils_video_patch_audio.py:1298
↓ 1 callers
Class
DataCollatorForSupervisedDataset
Collate examples for supervised fine-tuning.
vita/util/data_utils_video_audio_neg_frameCat.py:1127
↓ 1 callers
Class
DataCollatorForSupervisedDataset
Collate examples for supervised fine-tuning.
vita/util/data_utils_video_audio.py:754
↓ 1 callers
Class
DataCollatorForSupervisedDataset
Collate examples for supervised fine-tuning.
vita/util/data_utils_video_audio_patch_sf.py:1254
↓ 1 callers
Class
DateValue
VLMEvalKit/vlmeval/dataset/utils/tablevqabench.py:339
↓ 1 callers
Class
DownloadProgressBar
VLMEvalKit/vlmeval/smp/file.py:184
↓ 1 callers
Class
EVAVisionTransformer
Vision Transformer with support for patch or hybrid CNN input stage
vita/model/multimodal_encoder/eva_clip/eva_vit.py:591
↓ 1 callers
Class
Encoder
vita/model/vita_tts/decoder/ticodec/models.py:429
↓ 1 callers
Class
Eva2LargePlusEncoder
vita/model/multimodal_encoder/eva_clip/eva_vit.py:945
↓ 1 callers
Class
EvaClipImageTrainProcessor
vita/model/multimodal_encoder/eva_clip/eva_clip_processors.py:34
↓ 1 callers
Class
EvaClipVisionTower
vita/model/multimodal_encoder/eva_clip/eva_clip_encoder.py:8
↓ 1 callers
Class
FlashAttention
Implement the scaled dot product attention with softmax. Arguments --------- softmax_scale: The temperature to use for the softmax att
vita/model/multimodal_encoder/internvit/flash_attention.py:16
↓ 1 callers
Class
Generator
vita/model/vita_tts/decoder/ticodec/models.py:169
↓ 1 callers
Class
GlobalCMVN
vita/model/vita_tts/encoder/cmvn.py:7
↓ 1 callers
Class
GlobalCMVN
vita/model/multimodal_encoder/whale/cmvn.py:7
↓ 1 callers
Class
GlobalParams
web_demo/vita_html/web/parms.py:7
↓ 1 callers
Class
GlobalTokenEncoder
vita/model/vita_tts/decoder/ticodec/models.py:22
↓ 1 callers
Class
GroupRandomCrop
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:50
↓ 1 callers
Class
IdentityMap
vita/model/multimodal_projector/builder.py:12
↓ 1 callers
Class
InternAttention
Multi-headed attention from 'Attention Is All You Need' paper
vita/model/multimodal_encoder/internvit/modeling_intern_vit.py:125
↓ 1 callers
Class
InternMLP
vita/model/multimodal_encoder/internvit/modeling_intern_vit.py:205
↓ 1 callers
Class
InternViTVisionTower
vita/model/multimodal_encoder/internvit/internvit_encoder.py:8
↓ 1 callers
Class
InternVisionConfig
r""" This is the configuration class to store the configuration of a [`InternVisionModel`]. It is used to instantiate a vision encoder accordi
web_demo/vllm_tools/qwen2p5_model_weight_file/configuration_intern_vit.py:15
↓ 1 callers
Class
InternVisionEmbeddings
vita/model/multimodal_encoder/internvit/modeling_intern_vit.py:68
↓ 1 callers
Class
InternVisionEncoder
Transformer encoder consisting of `config.num_hidden_layers` self attention layers. Each layer is a [`InternEncoderLayer`]. Args:
vita/model/multimodal_encoder/internvit/modeling_intern_vit.py:256
↓ 1 callers
Class
InternVisionEncoderLayer
vita/model/multimodal_encoder/internvit/modeling_intern_vit.py:220
↓ 1 callers
Class
InternVisionModel
vita/model/multimodal_encoder/internvit/modeling_intern_vit.py:321
↓ 1 callers
Class
LDPBlock
vita/model/multimodal_projector/builder.py:75
↓ 1 callers
Class
LDPNetProjector
vita/model/multimodal_projector/builder.py:105
↓ 1 callers
Class
LLM2TTSCodecAR
E2E module. Args: idim (int): dimension of inputs odim (int): dimension of outputs args (namespace): argument Namespace c
vita/model/vita_tts/decoder/decoder.py:32
↓ 1 callers
Class
LazySupervisedDataset
Dataset for supervised fine-tuning.
vita/util/data_utils_video_audio_patch.py:685
↓ 1 callers
Class
LazySupervisedDataset
Dataset for supervised fine-tuning.
vita/util/data_utils_video_audio_neg_patch.py:827
↓ 1 callers
Class
LazySupervisedDataset
Dataset for supervised fine-tuning.
vita/util/data_utils_video_audio_neg_patch_fo.py:825
↓ 1 callers
Class
LazySupervisedDataset
Dataset for supervised fine-tuning.
vita/util/data_utils_video_patch_audio.py:731
↓ 1 callers
Class
LazySupervisedDataset
Dataset for supervised fine-tuning.
vita/util/data_utils_video_audio_neg_frameCat.py:559
↓ 1 callers
Class
LazySupervisedDataset
Dataset for supervised fine-tuning.
vita/util/data_utils_video_audio.py:371
↓ 1 callers
Class
LazySupervisedDataset
Dataset for supervised fine-tuning.
vita/util/data_utils_video_audio_patch_sf.py:685
↓ 1 callers
Class
LengthGroupedSampler
r""" Sampler that samples indices in a way that groups together features of the dataset of roughly the same length while keeping a bit of rand
vita/train/vita_trainer.py:117
↓ 1 callers
Class
LinearAdapter
vita/model/vita_tts/adapter.py:59
↓ 1 callers
Class
LinearAdapter
vita/model/multimodal_encoder/whale/adapter.py:54
↓ 1 callers
Class
MMAlaya2
This implementation fine-tunes 20 LoRA modules based on the InternVL-Chat-V1-5 model. The fine-tuned LoRA modules are then merged with the In
VLMEvalKit/vlmeval/vlm/mmalaya.py:193
↓ 1 callers
Class
MambaBlock
vita/model/multimodal_encoder/whale/module/component/mamba.py:22
↓ 1 callers
Class
MambaSSM
vita/model/multimodal_encoder/whale/module/component/mamba.py:83
↓ 1 callers
Class
Minigpt
vita/model/multimodal_projector/builder.py:24
↓ 1 callers
Class
Mlp
vita/model/multimodal_encoder/eva_clip/eva_vit.py:195
↓ 1 callers
Class
MultiHeadedAttention
Multi-Head Attention layer. :param int n_head: the number of head s :param int n_feat: the number of features :param float dropout_rate:
vita/model/vita_tts/encoder/attention.py:268
↓ 1 callers
Class
MultiHeadedAttention
Multi-Head Attention layer. :param int n_head: the number of head s :param int n_feat: the number of features :param float dropout_rate:
vita/model/multimodal_encoder/whale/module/layer/attention.py:273
↓ 1 callers
Class
MultiSequential
Multi-input multi-output torch.nn.Sequential.
vita/model/vita_tts/encoder/transformer.py:27
↓ 1 callers
Class
MultiSequential
Multi-input multi-output torch.nn.Sequential.
vita/model/multimodal_encoder/whale/module/component/transformer.py:35
↓ 1 callers
Class
OpenAIWrapper
VLMEvalKit/vlmeval/api/gpt.py:32
↓ 1 callers
Class
PCMQueue
web_demo/vita_html/web/queue.py:12
↓ 1 callers
Class
PatchDropout
https://arxiv.org/abs/2212.00794
vita/model/multimodal_encoder/eva_clip/eva_vit.py:123
↓ 1 callers
Class
PatchEmbed
Image to Patch Embedding
vita/model/multimodal_encoder/eva_clip/eva_vit.py:525
↓ 1 callers
Class
Quantizer
vita/model/vita_tts/decoder/ticodec/models.py:540
↓ 1 callers
Class
Qwen2Attention
web_demo/vllm_tools/vllm_file/qwen2.py:432
↓ 1 callers
Class
Qwen2DecoderLayer
web_demo/vllm_tools/vllm_file/qwen2.py:509
↓ 1 callers
Class
Qwen2ImagePixelInputs
web_demo/vllm_tools/vllm_file/qwen2.py:72
↓ 1 callers
Class
Qwen2MLP
web_demo/vllm_tools/vllm_file/qwen2.py:402
↓ 1 callers
Class
Qwen2MultiModalAudioProjector
web_demo/vllm_tools/vllm_file/qwen2.py:803
↓ 1 callers
Class
Qwen2MultiModalVisionProjector
web_demo/vllm_tools/vllm_file/qwen2.py:788
↓ 1 callers
Class
RelPositionalEncoding
Relative positional encoding module. See : Appendix B in https://arxiv.org/abs/1901.02860 Args: d_model (int): Embedding dimension.
web_demo/vllm_tools/vllm_file/whale.py:274
↓ 1 callers
Class
RelativePositionBias
vita/model/multimodal_encoder/eva_clip/eva_vit.py:550
next →
1–100 of 320, ranked by callers