↓ 2 callersFunctionrepeat_kv This is the equivalent of torch.repeat_interleave(x, dim=1, repeats=n_rep). The hidden states go from (batch, num_key_value_heads, seqlen, he
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:139
↓ 2 callersFunctionrepeat_kv This is the equivalent of torch.repeat_interleave(x, dim=1, repeats=n_rep). The hidden states go from (batch, num_key_value_heads, seqlen, he
vita_audio/models/qwen2_v4_48_3/modeling_qwen2.py:94
↓ 2 callersFunctionrepeat_kv This is the equivalent of torch.repeat_interleave(x, dim=1, repeats=n_rep). The hidden states go from (batch, num_key_value_heads, seqlen, he
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:136
↓ 1 callersFunctionForCausalLMLoss(
logits, labels, vocab_size: int, num_items_in_batch: int = None, ignore_index: int = -100, **kwargs
)
vita_audio/models/qwen2_mtp_sensevoice_v4_48_3/modeling_qwen2.py:56
↓ 1 callersFunctionForCausalLMLoss(
logits, labels, vocab_size: int, num_items_in_batch: int = None, ignore_index: int = -100, **kwargs
)
vita_audio/models/qwen2_mtp_v4_48_3/modeling_qwen2.py:53