Code
Hub
Workspaces
Following
Trending
Connect
MCP
copy
Create free account
hub
/
github.com/HumanMLLM/ViSpeak
/ types & classes
Types & classes
292 in github.com/HumanMLLM/ViSpeak
⨍
Functions
1,447
◇
Types & classes
292
↓ 13 callers
Class
KeywordsStoppingCriteria
vispeak/util/mm_utils.py:128
↓ 9 callers
Class
Conversation
A class that keeps all conversation history.
vispeak/conversation.py:18
↓ 6 callers
Class
Quantizer_module
vispeak/model/vita_tts/decoder/ticodec/models.py:525
↓ 5 callers
Class
Cardinal
CARDINAL类
audio_eval/cn_tn.py:756
↓ 5 callers
Class
DiscriminatorP
vispeak/model/vita_tts/decoder/ticodec/models.py:257
↓ 4 callers
Class
ChineseNumberUnit
中文数字/数位字符 每个字符除繁简体外还有一个额外的大写字符 e.g. '陆' 和 '陸'
audio_eval/cn_tn.py:412
↓ 4 callers
Class
GroupScale
Rescales the input PIL.Image to the given 'size'. 'size' will be the size of the smaller edge. For example, if height > width, then image wil
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:150
↓ 3 callers
Class
Block
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:430
↓ 3 callers
Class
DiscriminatorS
vispeak/model/vita_tts/decoder/ticodec/models.py:337
↓ 3 callers
Class
DropPath
Drop paths (Stochastic Depth) per sample (when applied in main path of residual blocks).
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:181
↓ 2 callers
Class
CrossEntropyLoss
vispeak/model/vita_tts/decoder/decoder.py:16
↓ 2 callers
Class
Digit
DIGIT类
audio_eval/cn_tn.py:771
↓ 2 callers
Class
Fraction
FRACTION类
audio_eval/cn_tn.py:821
↓ 2 callers
Class
InternRMSNorm
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:34
↓ 2 callers
Class
NumberValue
VLMEvalKit/vlmeval/dataset/utils/tablevqabench.py:284
↓ 2 callers
Class
StoppingCriteriaSub
VLMEvalKit/vlmeval/vlm/xcomposer/xcomposer.py:10
↓ 2 callers
Class
StreamToLogger
Fake file-like stream object that redirects writes to a logger instance.
vispeak/util/utils.py:68
↓ 2 callers
Class
TelePhone
TELEPHONE类
audio_eval/cn_tn.py:787
↓ 2 callers
Class
TextNorm
audio_eval/cn_tn.py:1066
↓ 1 callers
Class
Attention
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:263
↓ 1 callers
Class
AttrDict
vispeak/model/vita_tts/decoder/ticodec/vqvae.py:10
↓ 1 callers
Class
AudioDataset
audio_eval/eval_asr.py:75
↓ 1 callers
Class
AudioLLM
vispeak/model/vita_tts/audioLLM.py:20
↓ 1 callers
Class
BasicTextNormalizer
audio_eval/whisper_normalizer/basic.py:56
↓ 1 callers
Class
Benchmark
StreamingBench/src/benchmark/Benchmark.py:1
↓ 1 callers
Class
CLIPVisionCfg
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:877
↓ 1 callers
Class
CLIPVisionTower
vispeak/model/multimodal_encoder/clip/clip_encoder.py:6
↓ 1 callers
Class
CNNAdapter
vispeak/model/vita_tts/adapter.py:10
↓ 1 callers
Class
CNNAdapter
vispeak/model/multimodal_encoder/whale/adapter.py:6
↓ 1 callers
Class
CNNSubsampling
vispeak/model/vita_tts/adapter.py:72
↓ 1 callers
Class
CNNSubsampling
vispeak/model/multimodal_encoder/whale/adapter.py:68
↓ 1 callers
Class
COCO_Caption_Scorer
VLMEvalKit/vlmeval/dataset/image_caption.py:5
↓ 1 callers
Class
ChineseNumberDigit
中文数字字符
audio_eval/cn_tn.py:448
↓ 1 callers
Class
Conv2dSubsampling4
Convolutional 2D subsampling (to 1/4 length). Args: idim (int): Input dimension. odim (int): Output dimension. dropout_ra
vispeak/model/vita_tts/encoder/subsampling.py:15
↓ 1 callers
Class
Conv2dSubsampling4
Convolutional 2D subsampling (to 1/4 length). Args: idim (int): Input dimension. odim (int): Output dimension. dropout_ra
vispeak/model/multimodal_encoder/whale/module/component/subsampling.py:15
↓ 1 callers
Class
CustomMCQDataset
VLMEvalKit/vlmeval/dataset/image_mcq.py:629
↓ 1 callers
Class
CustomTextMCQDataset
VLMEvalKit/vlmeval/dataset/text_mcq.py:112
↓ 1 callers
Class
CustomVQADataset
VLMEvalKit/vlmeval/dataset/image_vqa.py:473
↓ 1 callers
Class
DataCollatorForSupervisedDataset
Collate examples for supervised fine-tuning.
vispeak/util/data_utils.py:1000
↓ 1 callers
Class
Date
DATE类
audio_eval/cn_tn.py:839
↓ 1 callers
Class
DateValue
VLMEvalKit/vlmeval/dataset/utils/tablevqabench.py:339
↓ 1 callers
Class
DownloadProgressBar
VLMEvalKit/vlmeval/smp/file.py:184
↓ 1 callers
Class
EVAVisionTransformer
Vision Transformer with support for patch or hybrid CNN input stage
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:591
↓ 1 callers
Class
Encoder
vispeak/model/vita_tts/decoder/ticodec/models.py:429
↓ 1 callers
Class
EnglishNumberNormalizer
Convert any spelled-out numbers into arabic numbers, while handling: - remove any commas - keep the suffixes such as: `1960s`, `274th`,
audio_eval/whisper_normalizer/english.py:12
↓ 1 callers
Class
EnglishSpellingNormalizer
Applies British-American spelling mappings as listed in [1]. [1] https://www.tysto.com/uk-us-spelling-list.html
audio_eval/whisper_normalizer/english.py:450
↓ 1 callers
Class
EnglishTextNormalizer
audio_eval/whisper_normalizer/english.py:465
↓ 1 callers
Class
Eva2LargePlusEncoder
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:945
↓ 1 callers
Class
EvaClipImageTrainProcessor
vispeak/model/multimodal_encoder/eva_clip/eva_clip_processors.py:34
↓ 1 callers
Class
EvaClipVisionTower
vispeak/model/multimodal_encoder/eva_clip/eva_clip_encoder.py:8
↓ 1 callers
Class
EvalViSpeak
OVO-Bench/models/ViSpeak.py:23
↓ 1 callers
Class
EvaluationTokenizer
A generic evaluation-time tokenizer, which leverages built-in tokenizers in sacreBLEU (https://github.com/mjpost/sacrebleu). It additionally provi
audio_eval/evaluate_tokenizer.py:9
↓ 1 callers
Class
FlashAttention
Implement the scaled dot product attention with softmax. Arguments --------- softmax_scale: The temperature to use for the softmax att
vispeak/model/multimodal_encoder/internvit/flash_attention.py:16
↓ 1 callers
Class
Generator
vispeak/model/vita_tts/decoder/ticodec/models.py:169
↓ 1 callers
Class
GlobalCMVN
vispeak/model/vita_tts/encoder/cmvn.py:7
↓ 1 callers
Class
GlobalCMVN
vispeak/model/multimodal_encoder/whale/cmvn.py:7
↓ 1 callers
Class
GlobalTokenEncoder
vispeak/model/vita_tts/decoder/ticodec/models.py:22
↓ 1 callers
Class
GroupCenterCrop
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:106
↓ 1 callers
Class
GroupNormalize
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:134
↓ 1 callers
Class
GroupRandomCrop
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:50
↓ 1 callers
Class
IdentityMap
vispeak/model/multimodal_projector/builder.py:12
↓ 1 callers
Class
InferenceSampler
audio_eval/eval_asr.py:133
↓ 1 callers
Class
InternAttention
Multi-headed attention from 'Attention Is All You Need' paper
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:125
↓ 1 callers
Class
InternMLP
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:205
↓ 1 callers
Class
InternViTVisionTower
vispeak/model/multimodal_encoder/internvit/internvit_encoder.py:8
↓ 1 callers
Class
InternVisionEmbeddings
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:68
↓ 1 callers
Class
InternVisionEncoder
Transformer encoder consisting of `config.num_hidden_layers` self attention layers. Each layer is a [`InternEncoderLayer`]. Args:
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:256
↓ 1 callers
Class
InternVisionEncoderLayer
vispeak/model/multimodal_encoder/internvit/modeling_intern_vit.py:220
↓ 1 callers
Class
LDPBlock
vispeak/model/multimodal_projector/builder.py:75
↓ 1 callers
Class
LDPNetProjector
vispeak/model/multimodal_projector/builder.py:105
↓ 1 callers
Class
LLM2TTSCodecAR
E2E module. Args: idim (int): dimension of inputs odim (int): dimension of outputs args (namespace): argument Namespace c
vispeak/model/vita_tts/decoder/decoder.py:32
↓ 1 callers
Class
LazySupervisedDataset
Dataset for supervised fine-tuning.
vispeak/util/data_utils.py:678
↓ 1 callers
Class
LengthGroupedSampler
r""" Sampler that samples indices in a way that groups together features of the dataset of roughly the same length while keeping a bit of rand
vispeak/train/vispeak_trainer.py:117
↓ 1 callers
Class
LinearAdapter
vispeak/model/vita_tts/adapter.py:59
↓ 1 callers
Class
LinearAdapter
vispeak/model/multimodal_encoder/whale/adapter.py:54
↓ 1 callers
Class
MMAlaya2
This implementation fine-tunes 20 LoRA modules based on the InternVL-Chat-V1-5 model. The fine-tuned LoRA modules are then merged with the In
VLMEvalKit/vlmeval/vlm/mmalaya.py:193
↓ 1 callers
Class
MambaBlock
vispeak/model/multimodal_encoder/whale/module/component/mamba.py:22
↓ 1 callers
Class
MambaSSM
vispeak/model/multimodal_encoder/whale/module/component/mamba.py:83
↓ 1 callers
Class
MathSymbol
用于中文数字系统的数学符号 (繁/简体), e.g. positive = ['正', '正'] negative = ['负', '負'] point = ['点', '點']
audio_eval/cn_tn.py:492
↓ 1 callers
Class
Minigpt
vispeak/model/multimodal_projector/builder.py:24
↓ 1 callers
Class
Mlp
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:195
↓ 1 callers
Class
Model
StreamingBench/src/model/modelclass.py:1
↓ 1 callers
Class
ModelArguments
VLMEvalKit/vlmeval/vlm/vispeak/vispeak_qwen2.py:32
↓ 1 callers
Class
Money
MONEY类
audio_eval/cn_tn.py:897
↓ 1 callers
Class
MultiHeadedAttention
Multi-Head Attention layer. :param int n_head: the number of head s :param int n_feat: the number of features :param float dropout_rate:
vispeak/model/vita_tts/encoder/attention.py:268
↓ 1 callers
Class
MultiHeadedAttention
Multi-Head Attention layer. :param int n_head: the number of head s :param int n_feat: the number of features :param float dropout_rate:
vispeak/model/multimodal_encoder/whale/module/layer/attention.py:273
↓ 1 callers
Class
MultiSequential
Multi-input multi-output torch.nn.Sequential.
vispeak/model/vita_tts/encoder/transformer.py:27
↓ 1 callers
Class
MultiSequential
Multi-input multi-output torch.nn.Sequential.
vispeak/model/multimodal_encoder/whale/module/component/transformer.py:35
↓ 1 callers
Class
NumberSystem
中文数字系统
audio_eval/cn_tn.py:485
↓ 1 callers
Class
OVOBenchOfflineScore
OVO-Bench/utils/OVOBenchScore.py:8
↓ 1 callers
Class
OpenAIWrapper
VLMEvalKit/vlmeval/api/gpt.py:32
↓ 1 callers
Class
PatchDropout
https://arxiv.org/abs/2212.00794
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:123
↓ 1 callers
Class
PatchEmbed
Image to Patch Embedding
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:525
↓ 1 callers
Class
Percentage
PERCENTAGE类
audio_eval/cn_tn.py:920
↓ 1 callers
Class
Quantizer
vispeak/model/vita_tts/decoder/ticodec/models.py:540
↓ 1 callers
Class
RelativePositionBias
vispeak/model/multimodal_encoder/eva_clip/eva_vit.py:550
↓ 1 callers
Class
SPP
vispeak/model/multimodal_projector/builder.py:114
↓ 1 callers
Class
SiglipVisionTower
vispeak/model/multimodal_encoder/siglip/siglip_encoder.py:8
↓ 1 callers
Class
SiglipVisionTowerS2
vispeak/model/multimodal_encoder/siglip/siglip_encoder.py:82
↓ 1 callers
Class
Stack
VLMEvalKit/vlmeval/dataset/utils/mvbench.py:405
next →
1–100 of 292, ranked by callers