Code
Hub
Workspaces
Following
Trending
Connect
MCP
copy
Create free account
hub
/
github.com/brontoguana/krasis
/ functions
Functions
12,380 in github.com/brontoguana/krasis
⨍
Functions
12,380
◇
Types & classes
7,782
↳
Endpoints
10
↓ 2 callers
Function
group_diagnostics
(row_idx: int, group_idx: int)
tests/hqq_attention_diff.py:1385
↓ 2 callers
Function
group_events
(events: list[dict[str, Any]])
tests/hqq_attention_diff.py:3541
↓ 2 callers
Function
handle_chat_completion
Handle /v1/chat/completions request.
src/server.rs:868
↓ 2 callers
Function
handle_prefill_logits
Handle /v1/internal/prefill_logits endpoint. Runs a full prefill pass and extracts top-k logprobs at sampled positions.
src/server.rs:1571
↓ 2 callers
Function
handle_reference_test
Handle /v1/internal/reference_test endpoint. Accepts raw input_token_ids, runs greedy prefill + decode, returns output tokens with logprobs. Used for
src/server.rs:1724
↓ 2 callers
Method
has_merged_experts
Check if this GGUF uses merged expert tensors (ffn_gate_exps).
src/gguf.rs:512
↓ 2 callers
Function
hf_auth_status
()
python/krasis/hf_downloader.py:259
↓ 2 callers
Function
hqq46_auto_budget_bytes_from_mib
(value: Optional[int])
python/krasis/attention_backend.py:572
↓ 2 callers
Function
hqq46_tensor_nbits
(tensor_name: str)
python/krasis/attention_backend.py:343
↓ 2 callers
Function
hqq_search_cuda_tensor_ptr
( input_ptr: usize, rows: usize, cols: usize, group_size: usize, nbits: usize, global_
src/hqq.rs:481
↓ 2 callers
Function
implicit_gemm_k_iterations_per_channel
src/cuda/flash_attn/cutlass/cutlass/conv/conv2d_problem_size.h:493
↓ 2 callers
Method
inf_with_sign
src/cuda/flash_attn/cutlass/cutlass/exmy_base.h:603
↓ 2 callers
Function
infer_capture_profile
(reference: Dict[str, Any], tokenizer: Any)
tests/reference_contract.py:361
↓ 2 callers
Function
info
(msg: str)
tests/reference_inventory.py:53
↓ 2 callers
Method
init_params
Initialize params member
src/cuda/flash_attn/cutlass/cutlass/gemm/device/gemm_universal_base.h:207
↓ 2 callers
Method
inverse
Computes the inverse of a 2-by-2 matrix given the matrix's determinant
src/cuda/flash_attn/cutlass/cutlass/matrix.h:3240
↓ 2 callers
Function
inverse_marlin_repack
Convert Marlin-packed INT4 weights to standard GPTQ packed format. Args: w_marlin: [..., K//16, 2*N] int32 (Marlin format, any leading ba
python/krasis/triton_moe.py:73
↓ 2 callers
Function
inverse_scale_permute
Un-permute Marlin scales to standard GPTQ format. Args: s_marlin: [..., K//gs, N] bf16 (Marlin permuted + transposed scales) Returns
python/krasis/triton_moe.py:136
↓ 2 callers
Method
is_C_load_needed
src/cuda/flash_attn/cutlass/cutlass/epilogue/fusion/sm90_visitor_compute_tma_warpspecialized.hpp:143
↓ 2 callers
Function
is_auxiliary_prefix
(prefix: str)
python/krasis/config.py:224
↓ 2 callers
Method
is_f16c_supported
src/cuda/flash_attn/cutlass/cutlass/half.h:149
↓ 2 callers
Method
is_full_attention_layer
True if this layer uses standard full attention (GQA/MLA).
python/krasis/config.py:779
↓ 2 callers
Method
is_producer_load_needed
src/cuda/flash_attn/cutlass/cutlass/epilogue/fusion/sm90_visitor_compute_tma_warpspecialized.hpp:138
↓ 2 callers
Method
is_producer_load_needed
src/cuda/flash_attn/cutlass/cutlass/epilogue/collective/sm90_epilogue_tma_warpspecialized.hpp:413
↓ 2 callers
Method
is_reduction_unit
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:115
↓ 2 callers
Function
is_same_row_or_col
src/cuda/flash_attn/cutlass/cutlass/pipeline/sm100_pipeline.hpp:382
↓ 2 callers
Method
is_source1_needed
src/cuda/flash_attn/cutlass/cutlass/epilogue/thread/linear_combination_tensor_broadcast.hpp:200
↓ 2 callers
Method
is_source_needed
src/cuda/flash_attn/cutlass/cutlass/epilogue/collective/sm70_epilogue_vectorized.hpp:262
↓ 2 callers
Function
is_stacked_experts
Detect stacked expert format (Qwen3.5, Mistral 4). These models store all experts in stacked 3D tensors per layer: experts.gate_up_proj [E, 2*inter, h
src/weights/mod.rs:6198
↓ 2 callers
Function
is_valid
src/cuda/flash_attn/cutlass/cute/util/type_traits.hpp:262
↓ 2 callers
Method
is_zero
src/cuda/flash_attn/cutlass/cutlass/epilogue/fusion/sm90_visitor_load_tma_warpspecialized.hpp:313
↓ 2 callers
Function
isfinite
src/cuda/flash_attn/cutlass/cutlass/tfloat32.h:208
↓ 2 callers
Function
isfinite
src/cuda/flash_attn/cutlass/cutlass/half.h:510
↓ 2 callers
Function
isinf
src/cuda/flash_attn/cutlass/cutlass/half.h:521
↓ 2 callers
Function
judge_overall
Return overall PASS, WARN, or FAIL.
tests/reference_test.py:767
↓ 2 callers
Function
judge_prompt
Return PASS, WARN, or FAIL for a single prompt. Primary metric: prefill top-10 containment (same input, only quant differs). Secondary: first
tests/reference_test.py:727
↓ 2 callers
Function
kill_server
Kill server process group.
tests/validate_model.py:317
↓ 2 callers
Function
kill_stale_krasis_processes
Kill any lingering krasis Python processes (excluding ourselves).
tests/release_test.py:641
↓ 2 callers
Method
kn
Obtains a Coord<2> from GemmCoord
src/cuda/flash_attn/cutlass/cutlass/gemm_coord.h:186
↓ 2 callers
Function
l2_normalize_expand_avx2
( src: &[f32], dst: &mut [f32], src_offset: usize, dk: usize, nv: usize, hr: usize, scale: f32,
src/decode.rs:4470
↓ 2 callers
Function
l2norm
(x, dim=-1, eps=1e-6)
tests/hf_layer0_compare.py:22
↓ 2 callers
Function
lcm
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:183
↓ 2 callers
Function
lcm_cxx11
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:204
↓ 2 callers
Function
linear_attention_recurrent_avx2
( state: &mut [f32], q: &[f32], k: &[f32], v: &[f32], g: &[f32], beta: &[f32],
src/decode.rs:1653
↓ 2 callers
Function
list_available_references
List all available reference data directories.
tests/reference_test.py:172
↓ 2 callers
Function
load
Loads a fragment
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/shared_load_iterator.h:210
↓ 2 callers
Method
load_ab_init
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:638
↓ 2 callers
Method
load_ab_tail
Perform a Producer Epilogue to prevent early exit of ctas in a Cluster
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:896
↓ 2 callers
Function
load_and_quantize_weight
Load a BF16 or FP8 weight tensor and quantize it to INT4. For FP8 tensors, loads the per-tensor `weight_scale_inv` and dequantizes to BF16 first.
src/weights/mod.rs:6547
↓ 2 callers
Function
load_and_time
(label, threshold, expert_divisor=-1)
tests/archive/test_qwen3_next_decode_timing.py:15
↓ 2 callers
Function
load_bf16_tensor_as_f32
Load BF16 bias tensor from safetensors as f32 slice.
src/weights/mod.rs:6287
↓ 2 callers
Function
load_case_map
(path: Path)
tests/hqq_attention_diff.py:346
↓ 2 callers
Function
load_case_summary_doc
(path: Path)
tests/hqq_attention_diff.py:327
↓ 2 callers
Method
load_config
Parse TOML config and return all model × config combinations.
python/krasis/suite.py:148
↓ 2 callers
Method
load_final_norm
Load final RMSNorm weight (BF16, tiny).
python/krasis/weight_loader.py:228
↓ 2 callers
Function
load_hqq_attention_pending_manifest
( model_path: str, cache_profile: Optional[str] = HQQ_CACHE_PROFILE_BASELINE, nbits: int = 4,
python/krasis/attention_backend.py:1006
↓ 2 callers
Function
load_init
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_array_warpspecialized.hpp:501
↓ 2 callers
Method
load_lm_head
Load LM head weight. Returns (weight_int8, scale) if INT8, or plain BF16 tensor if BF16.
python/krasis/weight_loader.py:282
↓ 2 callers
Function
load_model
Load V2-Lite with specified threshold.
tests/test_gpu_decode.py:25
↓ 2 callers
Function
load_model
Load V2-Lite with dual-format cache.
tests/archive/test_v2lite_dual_format.py:19
↓ 2 callers
Function
load_model
Load Qwen3-Coder-Next PP=2 with active_only + static_pin.
tests/archive/test_qwen3_next_gpu_decode.py:22
↓ 2 callers
Function
load_mxfp4_layer_experts
Load all experts for a single layer from MXFP4 format, dequantize to BF16, then quantize to INT4/INT8. MXFP4 stores all experts in a single tensor pe
src/weights/mod.rs:6366
↓ 2 callers
Function
load_prompt_text
Load a full Gutenberg prompt by index (0-5). Cached after first load.
tests/prompt_utils.py:25
↓ 2 callers
Method
load_sf
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:916
↓ 2 callers
Method
load_sf_init
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:704
↓ 2 callers
Method
load_sf_tail
Perform a Producer Epilogue to prevent early exit of ctas in a Cluster
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:977
↓ 2 callers
Function
load_stacked_layer_experts
Load all experts from stacked 3D tensors (Qwen3.5/Mistral 4 format) and quantize. Stacked format: experts.gate_up_proj [E, 2*inter, hidden], experts.
src/weights/mod.rs:6997
↓ 2 callers
Method
load_tail
Perform a Producer Epilogue to prevent early exit of blocks in a Cluster
src/cuda/flash_attn/cutlass/cutlass/conv/collective/sm90_implicit_gemm_gmma_ss_warpspecialized.hpp:634
↓ 2 callers
Function
load_template
Load a template for a model, if one exists. Searches: 0. Explicit override from KRASIS_AWQ_TEMPLATE_PATH 1. Bundled templates in template
python/krasis/awq_calibrate.py:899
↓ 2 callers
Function
load_with_pointer_offset
Load
src/cuda/flash_attn/cutlass/cutlass/epilogue/warp/tile_iterator_volta_tensor_op.h:216
↓ 2 callers
Function
log_build_timing
(label: &str, elapsed: std::time::Duration)
build.rs:86
↓ 2 callers
Method
make_fragment_C
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:130
↓ 2 callers
Function
make_layout_like
src/cuda/flash_attn/cutlass/cute/layout.hpp:439
↓ 2 callers
Function
make_linspace
(start: f32, end: f32, steps: usize)
src/hqq.rs:62
↓ 2 callers
Function
make_ordered_layout
src/cuda/flash_attn/cutlass/cute/layout.hpp:423
↓ 2 callers
Function
make_swizzle_strides
src/cuda/flash_attn/cutlass/cute/swizzle_layout.hpp:175
↓ 2 callers
Function
make_tiled_copy_C_atom
src/cuda/flash_attn/cutlass/cute/atom/copy_atom.hpp:481
↓ 2 callers
Function
make_tiled_mma
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:575
↓ 2 callers
Function
make_tma_copy_C_sm90
src/cuda/flash_attn/cutlass/cute/atom/copy_traits_sm90_tma.hpp:1531
↓ 2 callers
Function
make_tmem_warp_partitioner
src/cuda/flash_attn/cutlass/cute/atom/copy_traits_sm100.hpp:296
↓ 2 callers
Function
manual_mla_attention
Manual MLA attention using torch.matmul (no FlashInfer). Args: q_nope_absorbed: [M, H, D_ckv] query with w_kc absorbed q_pe: [M,
tests/archive/test_attn_verify.py:21
↓ 2 callers
Function
marlin_cache_config_hash
( config_str: &str, gpu_bits: u8, expert_int4_calib_mode: ExpertInt4CalibMode, expert_int4_cal
src/weights/mod.rs:1516
↓ 2 callers
Function
marlin_expert_byte_sizes
Compute per-expert byte sizes for Marlin GPU format. For INT4: pack_factor=8, packed shape [K/16, 2*N] per tile. For INT8: pack_factor=4, packed shape
src/weights/mod.rs:1584
↓ 2 callers
Function
matmul_bf16_scalar
Scalar BF16 matrix-vector multiply (reference). Computes output[n] = sum_k(weight[n, k] * activation[k]).
src/kernel/avx2.rs:27
↓ 2 callers
Function
matmul_int4_marlin_scalar
Scalar Marlin-native INT4 matmul for correctness verification. Reads Marlin-packed `[K/16, 2*N]` weights and Marlin-permuted `[K/gs, N]` scales. Comp
src/kernel/avx2.rs:1884
↓ 2 callers
Function
matmul_int4_transposed_integer_parallel_tiled
Parallel tiled INT4 matmul. Data in [N/TILE, K/8, TILE] layout.
src/kernel/avx2.rs:1406
↓ 2 callers
Function
matmul_int4_transposed_scalar
Scalar INT4 transposed matmul (correctness reference). Weight layout: packed[K/8, N], scales[K/group_size, N]. Computes output[n] = sum_k(activation[
src/kernel/avx2.rs:818
↓ 2 callers
Function
matmul_int8_integer
Safe wrapper for the AVX2 integer INT8 matmul kernel.
src/kernel/avx2.rs:708
↓ 2 callers
Function
matmul_int8_transposed_integer_parallel_tiled
Parallel tiled INT8 matmul. Data in [N/TILE, K, TILE] layout.
src/kernel/avx2.rs:1762
↓ 2 callers
Function
matvec_q4_0_scalar
(weights: &[u8], input: &[f32], n: usize, k: usize, output: &mut [f32])
src/gguf_kernels.rs:552
↓ 2 callers
Function
matvec_q8_0_scalar
(weights: &[u8], input: &[f32], n: usize, k: usize, output: &mut [f32])
src/gguf_kernels.rs:534
↓ 2 callers
Function
measure_gpu_memory
Return (allocated_mb, reserved_mb).
tests/archive/test_quant_config.py:38
↓ 2 callers
Method
mk
Obtains a Coord<2> from GemmCoord
src/cuda/flash_attn/cutlass/cutlass/gemm_coord.h:168
↓ 2 callers
Function
mla_attn_dot_fp16_avx2
( q: &[f32], cache: &[u16], dim: usize, )
src/decode.rs:4847
↓ 2 callers
Function
moe_worker
Background worker for async MoE computation. Owns its own scratch buffers. Processes batches of tokens sequentially, with expert-level parallelism wi
src/moe.rs:1394
↓ 2 callers
Method
next_u32
(&mut self)
src/decode.rs:4929
↓ 2 callers
Function
norm
src/cuda/flash_attn/cutlass/cutlass/quaternion.h:410
↓ 2 callers
Function
normalize_hqq_auto_budget_pct
(value: Optional[float], attention_quant: str = "hqq_auto")
python/krasis/attention_backend.py:583
↓ 2 callers
Function
normalize_selected_gpus
(raw: Optional[str])
tests/reference_contract.py:903
↓ 2 callers
Function
nullspace
src/cuda/flash_attn/cutlass/cute/layout.hpp:1506
← previous
next →
1,801–1,900 of 12,380, ranked by callers