MCPcopy Create free account

hub / github.com/HumanMLLM/ViSpeak / types & classes

Types & classes292 in github.com/HumanMLLM/ViSpeak

↓ 13 callersClassKeywordsStoppingCriteria
vispeak/util/mm_utils.py:128
↓ 9 callersClassConversation
A class that keeps all conversation history.
vispeak/conversation.py:18
↓ 6 callersClassQuantizer_module
vispeak/model/vita_tts/decoder/ticodec/models.py:525
↓ 5 callersClassCardinal
CARDINAL类
audio_eval/cn_tn.py:756
↓ 5 callersClassDiscriminatorP
vispeak/model/vita_tts/decoder/ticodec/models.py:257
↓ 4 callersClassChineseNumberUnit
中文数字/数位字符 每个字符除繁简体外还有一个额外的大写字符 e.g. '陆' 和 '陸'
audio_eval/cn_tn.py:412
↓ 4 callersClassGroupScale
Rescales the input PIL.Image to the given 'size'. 'size' will be the size of the smaller edge. For example, if height > width, then image wil
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:150
↓ 3 callersClassBlock
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:430
↓ 3 callersClassDiscriminatorS
vispeak/model/vita_tts/decoder/ticodec/models.py:337
↓ 3 callersClassDropPath
Drop paths (Stochastic Depth) per sample (when applied in main path of residual blocks).
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:181
↓ 2 callersClassCrossEntropyLoss
vispeak/model/vita_tts/decoder/decoder.py:16
↓ 2 callersClassDigit
DIGIT类
audio_eval/cn_tn.py:771
↓ 2 callersClassFraction
FRACTION类
audio_eval/cn_tn.py:821
↓ 2 callersClassInternRMSNorm
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:34
↓ 2 callersClassNumberValue
VLMEvalKit/vlmeval/dataset/utils/tablevqabench.py:284
↓ 2 callersClassStoppingCriteriaSub
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer.py:10
↓ 2 callersClassStreamToLogger
Fake file-like stream object that redirects writes to a logger instance.
vispeak/util/utils.py:68
↓ 2 callersClassTelePhone
TELEPHONE类
audio_eval/cn_tn.py:787
↓ 2 callersClassTextNorm
audio_eval/cn_tn.py:1066
↓ 1 callersClassAttention
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:263
↓ 1 callersClassAttrDict
vispeak/model/vita_tts/decoder/ticodec/vqvae.py:10
↓ 1 callersClassAudioDataset
audio_eval/eval_asr.py:75
↓ 1 callersClassAudioLLM
vispeak/model/vita_tts/audioLLM.py:20
↓ 1 callersClassBasicTextNormalizer
audio_eval/whisper_normalizer/basic.py:56
↓ 1 callersClassBenchmark
StreamingBench/src/benchmark/Benchmark.py:1
↓ 1 callersClassCLIPVisionCfg
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:877
↓ 1 callersClassCLIPVisionTower
vispeak/model/multimodal_encoder/clip/clip_encoder.py:6
↓ 1 callersClassCNNAdapter
vispeak/model/vita_tts/adapter.py:10
↓ 1 callersClassCNNAdapter
vispeak/model/multimodal_encoder/whale/adapter.py:6
↓ 1 callersClassCNNSubsampling
vispeak/model/vita_tts/adapter.py:72
↓ 1 callersClassCNNSubsampling
vispeak/model/multimodal_encoder/whale/adapter.py:68
↓ 1 callersClassCOCO_Caption_Scorer
VLMEvalKit/vlmeval/dataset/image_caption.py:5
↓ 1 callersClassChineseNumberDigit
中文数字字符
audio_eval/cn_tn.py:448
↓ 1 callersClassConv2dSubsampling4
Convolutional 2D subsampling (to 1/4 length). Args: idim (int): Input dimension. odim (int): Output dimension. dropout_ra
vispeak/model/vita_tts/encoder/subsampling.py:15
↓ 1 callersClassConv2dSubsampling4
Convolutional 2D subsampling (to 1/4 length). Args: idim (int): Input dimension. odim (int): Output dimension. dropout_ra
vispeak/model/multimodal_encoder/whale/module/component/subsampling.py:15
↓ 1 callersClassCustomMCQDataset
VLMEvalKit/vlmeval/dataset/image_mcq.py:629
↓ 1 callersClassCustomTextMCQDataset
VLMEvalKit/vlmeval/dataset/text_mcq.py:112
↓ 1 callersClassCustomVQADataset
VLMEvalKit/vlmeval/dataset/image_vqa.py:473
↓ 1 callersClassDataCollatorForSupervisedDataset
Collate examples for supervised fine-tuning.
vispeak/util/data_utils.py:1000
↓ 1 callersClassDate
DATE类
audio_eval/cn_tn.py:839
↓ 1 callersClassDateValue
VLMEvalKit/vlmeval/dataset/utils/tablevqabench.py:339
↓ 1 callersClassDownloadProgressBar
VLMEvalKit/vlmeval/smp/file.py:184
↓ 1 callersClassEVAVisionTransformer
Vision Transformer with support for patch or hybrid CNN input stage
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:591
↓ 1 callersClassEncoder
vispeak/model/vita_tts/decoder/ticodec/models.py:429
↓ 1 callersClassEnglishNumberNormalizer
Convert any spelled-out numbers into arabic numbers, while handling: - remove any commas - keep the suffixes such as: `1960s`, `274th`,
audio_eval/whisper_normalizer/english.py:12
↓ 1 callersClassEnglishSpellingNormalizer
Applies British-American spelling mappings as listed in [1]. [1] https://www.tysto.com/uk-us-spelling-list.html
audio_eval/whisper_normalizer/english.py:450
↓ 1 callersClassEnglishTextNormalizer
audio_eval/whisper_normalizer/english.py:465
↓ 1 callersClassEva2LargePlusEncoder
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:945
↓ 1 callersClassEvaClipImageTrainProcessor
vispeak/model/multimodal_encoder/eva_clip/eva_clip_processors.py:34
↓ 1 callersClassEvaClipVisionTower
vispeak/model/multimodal_encoder/eva_clip/eva_clip_encoder.py:8
↓ 1 callersClassEvalViSpeak
OVO-Bench/models/ViSpeak.py:23
↓ 1 callersClassEvaluationTokenizer
A generic evaluation-time tokenizer, which leverages built-in tokenizers in sacreBLEU (https://github.com/mjpost/sacrebleu). It additionally provi
audio_eval/evaluate_tokenizer.py:9
↓ 1 callersClassFlashAttention
Implement the scaled dot product attention with softmax. Arguments --------- softmax_scale: The temperature to use for the softmax att
vispeak/model/multimodal_encoder/internvit/flash_attention.py:16
↓ 1 callersClassGenerator
vispeak/model/vita_tts/decoder/ticodec/models.py:169
↓ 1 callersClassGlobalCMVN
vispeak/model/vita_tts/encoder/cmvn.py:7
↓ 1 callersClassGlobalCMVN
vispeak/model/multimodal_encoder/whale/cmvn.py:7
↓ 1 callersClassGlobalTokenEncoder
vispeak/model/vita_tts/decoder/ticodec/models.py:22
↓ 1 callersClassGroupCenterCrop
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:106
↓ 1 callersClassGroupNormalize
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:134
↓ 1 callersClassGroupRandomCrop
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:50
↓ 1 callersClassIdentityMap
vispeak/model/multimodal_projector/builder.py:12
↓ 1 callersClassInferenceSampler
audio_eval/eval_asr.py:133
↓ 1 callersClassInternAttention
Multi-headed attention from 'Attention Is All You Need' paper
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:125
↓ 1 callersClassInternMLP
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:205
↓ 1 callersClassInternViTVisionTower
vispeak/model/multimodal_encoder/internvit/internvit_encoder.py:8
↓ 1 callersClassInternVisionEmbeddings
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:68
↓ 1 callersClassInternVisionEncoder
Transformer encoder consisting of `config.num_hidden_layers` self attention layers. Each layer is a [`InternEncoderLayer`]. Args:
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:256
↓ 1 callersClassInternVisionEncoderLayer
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:220
↓ 1 callersClassLDPBlock
vispeak/model/multimodal_projector/builder.py:75
↓ 1 callersClassLDPNetProjector
vispeak/model/multimodal_projector/builder.py:105
↓ 1 callersClassLLM2TTSCodecAR
E2E module. Args: idim (int): dimension of inputs odim (int): dimension of outputs args (namespace): argument Namespace c
vispeak/model/vita_tts/decoder/decoder.py:32
↓ 1 callersClassLazySupervisedDataset
Dataset for supervised fine-tuning.
vispeak/util/data_utils.py:678
↓ 1 callersClassLengthGroupedSampler
r""" Sampler that samples indices in a way that groups together features of the dataset of roughly the same length while keeping a bit of rand
vispeak/train/vispeak_trainer.py:117
↓ 1 callersClassLinearAdapter
vispeak/model/vita_tts/adapter.py:59
↓ 1 callersClassLinearAdapter
vispeak/model/multimodal_encoder/whale/adapter.py:54
↓ 1 callersClassMMAlaya2
This implementation fine-tunes 20 LoRA modules based on the InternVL-Chat-V1-5 model. The fine-tuned LoRA modules are then merged with the In
VLMEvalKit/vlmeval/vlm/mmalaya.py:193
↓ 1 callersClassMambaBlock
vispeak/model/multimodal_encoder/whale/module/component/mamba.py:22
↓ 1 callersClassMambaSSM
vispeak/model/multimodal_encoder/whale/module/component/mamba.py:83
↓ 1 callersClassMathSymbol
用于中文数字系统的数学符号 (繁/简体), e.g. positive = ['正', '正'] negative = ['负', '負'] point = ['点', '點']
audio_eval/cn_tn.py:492
↓ 1 callersClassMinigpt
vispeak/model/multimodal_projector/builder.py:24
↓ 1 callersClassMlp
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:195
↓ 1 callersClassModel
StreamingBench/src/model/modelclass.py:1
↓ 1 callersClassModelArguments
VLMEvalKit/vlmeval/vlm/vispeak/vispeak_qwen2.py:32
↓ 1 callersClassMoney
MONEY类
audio_eval/cn_tn.py:897
↓ 1 callersClassMultiHeadedAttention
Multi-Head Attention layer. :param int n_head: the number of head s :param int n_feat: the number of features :param float dropout_rate:
vispeak/model/vita_tts/encoder/attention.py:268
↓ 1 callersClassMultiHeadedAttention
Multi-Head Attention layer. :param int n_head: the number of head s :param int n_feat: the number of features :param float dropout_rate:
vispeak/model/multimodal_encoder/whale/module/layer/attention.py:273
↓ 1 callersClassMultiSequential
Multi-input multi-output torch.nn.Sequential.
vispeak/model/vita_tts/encoder/transformer.py:27
↓ 1 callersClassMultiSequential
Multi-input multi-output torch.nn.Sequential.
vispeak/model/multimodal_encoder/whale/module/component/transformer.py:35
↓ 1 callersClassNumberSystem
中文数字系统
audio_eval/cn_tn.py:485
↓ 1 callersClassOVOBenchOfflineScore
OVO-Bench/utils/OVOBenchScore.py:8
↓ 1 callersClassOpenAIWrapper
VLMEvalKit/vlmeval/api/gpt.py:32
↓ 1 callersClassPatchDropout
https://arxiv.org/abs/2212.00794
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:123
↓ 1 callersClassPatchEmbed
Image to Patch Embedding
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:525
↓ 1 callersClassPercentage
PERCENTAGE类
audio_eval/cn_tn.py:920
↓ 1 callersClassQuantizer
vispeak/model/vita_tts/decoder/ticodec/models.py:540
↓ 1 callersClassRelativePositionBias
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:550
↓ 1 callersClassSPP
vispeak/model/multimodal_projector/builder.py:114
↓ 1 callersClassSiglipVisionTower
vispeak/model/multimodal_encoder/siglip/siglip_encoder.py:8
↓ 1 callersClassSiglipVisionTowerS2
vispeak/model/multimodal_encoder/siglip/siglip_encoder.py:82
↓ 1 callersClassStack
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:405
next →1–100 of 292, ranked by callers