MCPcopy Create free account

hub / github.com/brontoguana/krasis / functions

Functions12,380 in github.com/brontoguana/krasis

↓ 2 callersFunctioncreate_marlin_buffers
Create properly-shaped Marlin buffers with random data.
tests/archive/test_marlin_gpu1.py:25
↓ 2 callersFunctioncreate_prefill_engine_for_server
( store: &mut GpuDecodeStore, max_context_tokens: usize, )
src/server.rs:237
↓ 2 callersFunctioncutlassGetStatusString
Convert cutlass status to status strings
src/cuda/flash_attn/cutlass/cutlass/cutlass.h:62
↓ 2 callersMethoddata
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:81
↓ 2 callersMethoddecode_step
Full decode step — runs entire layer loop in Rust. Replaces the Python step() method. One Python call per token. output_ptr: *mut f32 [vocab_size] —
src/decode.rs:3120
↓ 2 callersFunctiondelete_hqq_attention_pending_manifest
( model_path: str, cache_profile: Optional[str] = HQQ_CACHE_PROFILE_BASELINE, nbits: int = 4,
python/krasis/attention_backend.py:1032
↓ 2 callersFunctiondepth
src/cuda/flash_attn/cutlass/cute/layout.hpp:623
↓ 2 callersFunctiondequant_f16
(data: &[u8], n: usize)
src/gguf.rs:547
↓ 2 callersFunctiondequant_q4_0
Q4_0: 32 elements per block, 18 bytes = fp16 d + 16 bytes qs (nibbles). Each element = d * (nibble - 8).
src/gguf.rs:635
↓ 2 callersFunctiondequant_q4_k
Dequantize Q4_K blocks to FP32. Q4_K block (144 bytes per 256 elements): fp16 d (super-block scale), fp16 dmin (super-block min), 12 bytes packed 6-b
src/gguf.rs:681
↓ 2 callersFunctiondequant_q5_0
Q5_0: 32 elements per block, 22 bytes = fp16 d + 4 bytes qh + 16 bytes qs (nibbles). Each element = d * ((qs_nibble | (qh_bit << 4)) - 16).
src/gguf.rs:599
↓ 2 callersFunctiondequant_q5_k
Dequantize Q5_K blocks to FP32. Q5_K block (176 bytes per 256 elements): fp16 d, fp16 dmin, 12 bytes scales, 32 bytes qh (5th bits), 128 bytes qs (4-
src/gguf.rs:740
↓ 2 callersFunctiondequant_q6_k
Dequantize Q6_K blocks to FP32. Q6_K block (210 bytes per 256 elements): 128 bytes ql (lower 4 bits), 64 bytes qh (upper 2 bits), 16 bytes scales (in
src/gguf.rs:813
↓ 2 callersFunctiondequant_q8_0
Q8_0 block: fp16 scale + 32 x int8 weights. Block size = 32.
src/gguf.rs:574
↓ 2 callersFunctiondequantize_mxfp4_to_bf16
Dequantize MXFP4 blocks + E8M0 scales to BF16 for a single projection. # Arguments - `blocks`: contiguous [out_features, num_blocks, 16] u8 (each byt
src/weights/mod.rs:6225
↓ 2 callersFunctiondestination_for_repo
(models_dir: str, repo_id: str)
python/krasis/hf_downloader.py:444
↓ 2 callersFunctiondetect_linear_attention
Detect whether this model uses linear attention (recurrent state) layers. Models with linear attention (e.g. Gated DeltaNet in Qwen3-Coder-Next)
tests/reference_test.py:117
↓ 2 callersFunctiondetect_repetition
Detect if text contains repeating patterns (sign of degeneration). Returns True if a pattern of length >= min_pattern_len repeats >= min_repeats
tests/archive/test_v2lite_thorough.py:62
↓ 2 callersFunctiondiagnose_solve_scale_drift
(name: str, chunk: torch.Tensor, zero: torch.Tensor, scale_seed: torch.Tensor)
tests/test_hqq_rust_quantizer.py:1937
↓ 2 callersMethoddispatch_matmul_ext
Dispatch matmul to correct INT4/INT8 kernel (uses provided buffers).
src/decode.rs:1497
↓ 2 callersFunctiondomain_distribute
src/cuda/flash_attn/cutlass/cute/layout.hpp:1449
↓ 2 callersFunctiondrain_vram_pressure_for_state
( state: &mut ServerState, reason: &str, force_measure: bool, )
src/server.rs:131
↓ 2 callersFunctionemit_reference_generation_trace
( context: str, contract: Dict[str, Any], artifact_path: str, *, max_new_tokens: Optional[
tests/reference_contract.py:1279
↓ 2 callersMethodenable_mask
Clears the predicate set efficiently
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator_triangular_matrix.h:758
↓ 2 callersMethodenable_mask
Clears the predicate set efficiently
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator_2dthreadtile.h:739
↓ 2 callersMethodend_epilogue
Called after all steps have been completed
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/fusion/visitor_2x.hpp:121
↓ 2 callersFunctionepilogue_predication
src/cuda/flash_attn/cutlass/cute/algorithm/cooperative_gemm.hpp:58
↓ 2 callersFunctionerror_metrics_from_diff
(source: torch.Tensor, diff: torch.Tensor)
tests/hqq_attention_diff.py:527
↓ 2 callersFunctionevaluate_ppl_with_skip
Run PPL evaluation, skipping attention for specified layers. When a layer is skipped, its attention output is set to zeros so the residual st
perplexity/ablate_attention.py:37
↓ 2 callersFunctionexchange_sort
src/cuda/flash_attn/cutlass/cute/int_tuple.hpp:586
↓ 2 callersMethodexpect_transaction
Performs an expected transaction bytes increment without doing an arrive operation
src/cuda/flash_attn/cutlass/cutlass/arch/barrier.h:553
↓ 2 callersFunctionexpert_byte_sizes
(config: &ModelConfig, group_size: usize, num_bits: u8)
src/weights/mod.rs:1748
↓ 2 callersFunctionexpert_forward_bf16
BF16 reference expert forward: SiLU(gate(x)) * up(x) → down().
tests/test_rust_vs_python.py:77
↓ 2 callersFunctionextract_bundle
(path: Path, contracts: dict[str, dict[str, object]])
scripts/build_sidecars.py:442
↓ 2 callersFunctionextract_token
Extract a token string from tokenizer_config.json. Handles both `"bos_token": "<s>"` and `"bos_token": {"content": "<s>", ...}`.
src/chat_template.rs:296
↓ 2 callersFunctionfast_acos
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:823
↓ 2 callersFunctionfast_log
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:904
↓ 2 callersFunctionfast_silu_mul_avx2
(gate: &[f32], up: &[f32], output: &mut [f32], n: usize)
src/decode.rs:2053
↓ 2 callersFunctionfence_view_async_tmem_store
src/cuda/flash_attn/cutlass/cutlass/arch/barrier.h:894
↓ 2 callersFunctionfield_stage
(layer: int, tensor: str)
tests/hqq_attention_diff.py:3507
↓ 2 callersFunctionfile_mtime
(path: &str)
build.rs:120
↓ 2 callersFunctionfind_log2
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:245
↓ 2 callersFunctionfind_reference_data
Find reference data JSON for a model.
tests/reference_test.py:155
↓ 2 callersMethodfind_shared_expert_tensors
Find shared expert tensor names for a given layer.
src/gguf.rs:517
↓ 2 callersFunctionflash_attn_cu_files
()
scripts/build_sidecars.py:150
↓ 2 callersFunctionfloat_exmy_base<T, Derived>
Floating point conversion
src/cuda/flash_attn/cutlass/cutlass/exmy_base.h:1023
↓ 2 callersFunctionformat_completion
Format a complete (non-streaming) chat completion response.
src/server.rs:648
↓ 2 callersFunctionformat_contract_report
(reference_contract: Dict[str, Any], runtime_contract: Dict[str, Any], validation: Dict[str, List[str]])
tests/reference_contract.py:1089
↓ 2 callersFunctionfpclassify
src/cuda/flash_attn/cutlass/cutlass/half.h:531
↓ 2 callersMethodfront
src/cuda/flash_attn/cutlass/cutlass/array.h:385
↓ 2 callersFunctiongather_ranges
(t: torch.Tensor, ranges: list[tuple[int, int]])
tests/hqq_attention_diff.py:2064
↓ 2 callersFunctiongcd_cxx11
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:194
↓ 2 callersFunctiongemm
src/cuda/flash_attn/cutlass/cute/algorithm/cooperative_gemm.hpp:519
↓ 2 callersFunctiongenerate
Generate with or without CUDA graph.
tests/test_la_graph.py:30
↓ 2 callersFunctiongenerate
Generate with original or inplace recurrent forward.
tests/archive/test_la_inplace.py:30
↓ 2 callersFunctiongenerate
Generate text from a prompt and return (text, stats_dict).
tests/archive/run_long_prompt_test.py:24
↓ 2 callersFunctiongenerate_hf_reference
Generate reference logits using HuggingFace transformers.
tests/archive/test_v2lite_thorough.py:176
↓ 2 callersFunctiongenerate_weight_perm_int4
()
tests/test_marlin_attn_shapes.rs:149
↓ 2 callersFunctiongenerate_weight_perm_int8
()
tests/test_marlin_attn_shapes.rs:176
↓ 2 callersMethodgetParams
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/conv2d_tile_iterator.h:102
↓ 2 callersFunctiongetVersionBuild
src/cuda/flash_attn/cutlass/cutlass/version.h:64
↓ 2 callersMethodget_1d_coord
src/cuda/flash_attn/cutlass/cute/layout.hpp:272
↓ 2 callersMethodget_abs_max
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/epilogue_with_absmax.h:190
↓ 2 callersMethodget_callbacks
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/fusion/visitor_2x.hpp:134
↓ 2 callersMethodget_current_target
(self)
src/cuda/fla/compile_kernels.py:307
↓ 2 callersFunctionget_current_work_cta_m_n_in_cluster
Given raster order and current work tile linear index, reset cta m and n index in the cluster.
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:732
↓ 2 callersFunctionget_err_ratio
(x, y)
src/cuda/fla/utils.py:88
↓ 2 callersMethodget_grid_dims
Returns the grid extents in thread blocks to launch
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/params_universal_base.h:216
↓ 2 callersMethodget_hier_coord
src/cuda/flash_attn/cutlass/cute/layout.hpp:250
↓ 2 callersMethodget_layoutA_MK
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:430
↓ 2 callersMethodget_layoutB_NK
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:467
↓ 2 callersMethodget_layoutC_MN
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:397
↓ 2 callersMethodget_layoutC_TV
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:414
↓ 2 callersMethodget_mask
Gets the mask
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator_triangular_matrix.h:770
↓ 2 callersMethodget_mask
Gets the mask
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator_2dthreadtile.h:751
↓ 2 callersFunctionget_model_details
(repo_or_url: str)
python/krasis/hf_downloader.py:437
↓ 2 callersMethodget_ptr_aux_output_abs_max
src/cuda/flash_attn/cutlass/cutlass/epilogue/thread/linear_combination_generic_with_scaling.h:305
↓ 2 callersMethodget_ptr_output_abs_max
src/cuda/flash_attn/cutlass/cutlass/epilogue/thread/linear_combination_generic_with_scaling.h:300
↓ 2 callersFunctionget_scale_min_k4
(j: usize, scales: &[u8])
src/gguf_kernels.rs:640
↓ 2 callersFunctionget_scale_min_k4_ptr
(j: usize, scales: *const u8)
src/gguf_kernels.rs:652
↓ 2 callersMethodget_shape_C
Get C extents. fprop: C extents array contains [N,Z,P,Q,K]. Turn that into ((Q,P,Z,N), (K)) dgrad: C extents array contains [N,D,H,W,C]. Turn that int
src/cuda/flash_attn/cutlass/cutlass/conv/convnd_problem_shape.hpp:457
↓ 2 callersMethodget_shared_expert_unified
Backward compat: returns CPU shared expert ref.
src/weights/mod.rs:4294
↓ 2 callersFunctionget_slice
src/cuda/flash_attn/cutlass/cute/atom/copy_atom.hpp:370
↓ 2 callersFunctionget_stride
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/ell_predicated_tile_access_iterator.h:402
↓ 2 callersFunctionget_strided_dgrad_tile_m
////////////////////////////////////////////////////////////////////////////////////////////
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/threadblock_swizzle.h:53
↓ 2 callersFunctionget_tensor_decision
Get the quantization decision for a specific tensor. For v2 templates: input projections get "int4" (AWQ-scaled at load time), output project
python/krasis/awq_calibrate.py:968
↓ 2 callersFunctionget_tma_swizzle_base
src/cuda/flash_attn/cutlass/cute/atom/copy_traits_sm90_tma_swizzle.hpp:110
↓ 2 callersFunctionget_tma_swizzle_bits
src/cuda/flash_attn/cutlass/cute/atom/copy_traits_sm90_tma_swizzle.hpp:72
↓ 2 callersMethodget_work_k_tile_count
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/sm100_tile_scheduler.hpp:499
↓ 2 callersFunctiongithub_asset
(token: str | None, release: dict[str, object], name: str)
scripts/build_sidecars.py:590
↓ 2 callersFunctiongithub_repo
()
scripts/build_sidecars.py:505
↓ 2 callersFunctiongmem_wait
Wait until we have at least one completed global fetch stage
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/mma_multistage.h:485
↓ 2 callersFunctiongmem_wait
Wait until we have at least one completed global fetch stage
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/mma_pipelined.h:274
↓ 2 callersFunctiongpu_mem
()
tests/archive/decode_only_bench.py:43
↓ 2 callersFunctiongpu_mem
()
tests/archive/bench_qwen235b_pp2.py:47
↓ 2 callersFunctiongpu_mem
Return per-GPU allocated MB.
tests/archive/run_hcs_benchmark.py:52
↓ 2 callersFunctiongpu_mem
()
tests/archive/run_hcs_hybrid_bench.py:55
↓ 2 callersFunctiongpu_mem
()
tests/archive/bench_prefill_10k.py:44
↓ 2 callersFunctiongpu_mem
()
tests/archive/investigate_decode_regression.py:41
↓ 2 callersMethodgridDim
Commonly used utility functions
src/cuda/flash_attn/cutlass/cutlass/cluster_launch.hpp:95
← previousnext →1,701–1,800 of 12,380, ranked by callers