MCPcopy Create free account

hub / github.com/brontoguana/krasis / functions

Functions12,380 in github.com/brontoguana/krasis

↓ 5 callersFunctionhash_bytes
(data)
tools/expert_dedup_analysis.py:77
↓ 5 callersFunctionhqq_cache_algorithm_for_nbits
(nbits: int)
python/krasis/attention_backend.py:356
↓ 5 callersFunctioni32_bytes
Convert tensor to raw INT32 bytes.
tests/generate_kernel_test_data.py:42
↓ 5 callersFunctionimplicit_gemm_k_iterations
Determine the number of gemm_k iterations for conv2d problem using implicit gemm algorithm
src/cuda/flash_attn/cutlass/cutlass/conv/conv3d_problem_size.h:372
↓ 5 callersMethodinverse
src/cuda/flash_attn/cutlass/cutlass/layout/tensor_op_multiplicand_sm70.h:376
↓ 5 callersMethodinverse
src/cuda/flash_attn/cutlass/cutlass/layout/tensor_op_multiplicand_sm75.h:303
↓ 5 callersMethodis_moe_layer
(self, layer_idx: int)
python/krasis/config.py:811
↓ 5 callersMethodis_numa
Whether NUMA-aware placement is meaningful (>1 node).
src/numa.rs:197
↓ 5 callersFunctionleft_inverse
src/cuda/flash_attn/cutlass/cute/layout.hpp:1322
↓ 5 callersMethodload
Loads a fragment from memory at the location pointed to by the iterator.
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm80.h:2359
↓ 5 callersFunctionload_with_pointer_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm80.h:301
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm80.h:2366
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment from memory
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_iterator_tensor_op_sm70.h:507
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment from memory
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator.h:1459
↓ 5 callersMethodload_with_pointer_offset
Loads a fragment
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_iterator_pitch_linear.h:339
↓ 5 callersFunctionmake_tensor
(rows: int, cols: int, start: int)
tests/test_hqq_attention_carry.py:15
↓ 5 callersFunctionmake_tensor_like
src/cuda/flash_attn/cutlass/cute/tensor_impl.hpp:422
↓ 5 callersFunctionmake_tma_atom
src/cuda/flash_attn/cutlass/cute/atom/copy_traits_sm90_tma.hpp:1330
↓ 5 callersFunctionmatmul_int4_parallel
Parallel AVX2 INT4 matmul — splits output rows across rayon threads. Each thread processes a chunk of rows independently. The activation vector is sh
src/kernel/avx2.rs:537
↓ 5 callersFunctionmax
src/cuda/flash_attn/cutlass/cute/numeric/math.hpp:48
↓ 5 callersMethodmin_free_mb
Get the minimum free VRAM observed on a device (in MB).
src/vram_monitor.rs:684
↓ 5 callersFunctionminimal_manifest
(complete=True, profile=HQQ_CACHE_PROFILE_BASELINE)
tests/test_hqq_cache_profile.py:57
↓ 5 callersFunctionminmax_affine_params
(source: torch.Tensor)
tests/hqq_attention_diff.py:779
↓ 5 callersFunctionmoe_forward_gguf
Scratch buffers for GGUF expert computation (reused across calls). Full MoE forward using native GGUF weights (no dequant→requant). Uses raw GGUF blo
src/moe.rs:1045
↓ 5 callersFunctionprefetch_expert_configurable
(expert: &UnifiedExpertWeights, stride: usize, hint: u8)
src/moe.rs:1211
↓ 5 callersMethodproduct
Returns the product of all elements
src/cuda/flash_attn/cutlass/cutlass/coord.h:320
↓ 5 callersFunctionprofile_decode
Profile decode with per-token timing after prefill.
tests/archive/test_v2lite_profile.py:98
↓ 5 callersFunctionquantize_hqq4_tensor_rust
Experimental Rust shadow implementation for HQQ4 tensor quantization.
python/krasis/attention_backend.py:1684
↓ 5 callersFunctionquantizer_candidate_metrics
( name: str, source: torch.Tensor, input_group: torch.Tensor, *, scale: float, zero: f
tests/hqq_attention_diff.py:1001
↓ 5 callersMethodraw
Obtains raw bits
src/cuda/flash_attn/cutlass/cutlass/tfloat32.h:161
↓ 5 callersFunctionrelatively_equal_float
src/cuda/flash_attn/cutlass/cutlass/relatively_equal.h:57
↓ 5 callersMethodrows
(&self)
src/weights/mod.rs:269
↓ 5 callersMethodset_kgroup_index
Notify the iterator which k-group it is currently pointing to. This does not advance the iterator. Rather, it overrides its internal tracking with co
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm80.h:2427
↓ 5 callersFunctionset_params_alibi
src/cuda/flash_attn/fa2/flash_api_reference.cpp:331
↓ 5 callersMethodset_sign_bit
src/cuda/flash_attn/cutlass/cutlass/exmy_base.h:523
↓ 5 callersMethodsetup_gpu_decode_store
Create and configure a GpuDecodeStore for Rust-native GPU decode. Registers all attention weights, MoE layers, RoPE tables, and KV cache.
python/krasis/model.py:6607
↓ 5 callersFunctionshared_load
src/cuda/flash_attn/cutlass/cutlass/arch/memory_sm75.h:231
↓ 5 callersMethodsignificand_hidden_bits
src/cuda/flash_attn/cutlass/cutlass/exmy_base.h:629
↓ 5 callersFunctionsize
src/cuda/flash_attn/cutlass/cute/int_tuple.hpp:265
↓ 5 callersFunctionslice
src/cuda/flash_attn/cutlass/cute/layout.hpp:689
↓ 5 callersFunctionsource_prefix_candidates
(layer: int, family: str)
tests/hqq_attention_diff.py:2261
↓ 5 callersFunctionstage_hqq
(store: GpuDecodeStore, layer_idx: int, tensor_name: str, weight: torch.Tensor)
tests/test_hqq_fused_branch_runtime.py:147
↓ 5 callersMethodstart
(self, startup_timeout: float = 12.0)
python/krasis/ssh_tunnel.py:117
↓ 5 callersMethodstart
(self, name: str)
tests/hqq_attention_diff.py:35
↓ 5 callersMethodstore_init
src/cuda/flash_attn/cutlass/cutlass/epilogue/collective/detail.hpp:370
↓ 5 callersFunctionstore_with_pointer_offset
Store
src/cuda/flash_attn/cutlass/cutlass/epilogue/warp/tile_iterator_simt.h:195
↓ 5 callersFunctionstore_with_pointer_offset
Store
src/cuda/flash_attn/cutlass/cutlass/epilogue/warp/tile_iterator_tensor_op_mixed.h:203
↓ 5 callersMethodstore_with_pointer_offset
Store a fragment to memory
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_iterator_tensor_op_sm70.h:519
↓ 5 callersMethodstore_with_pointer_offset
Store a fragment to memory
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator.h:1477
↓ 5 callersMethodstore_with_pointer_offset
Stores a fragment
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_iterator_pitch_linear.h:357
↓ 5 callersFunctionstrided_dgrad_tile_m_per_filter
/////////////////////////////////////////////////////////////////////////////////////////////// Strided dgrad helper functions
src/cuda/flash_attn/cutlass/cutlass/conv/conv2d_problem_size.h:617
↓ 5 callersFunctiontensor_bytes
(shapes: list[tuple[str, int, int]], nbits_for_name)
python/krasis/vram_budget.py:482
↓ 5 callersMethodtensormaps_cp_fence_release
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm90_mma_array_tma_gmma_ss_warpspecialized.hpp:735
↓ 5 callersMethodtest_wait
src/cuda/flash_attn/cutlass/cutlass/arch/barrier.h:357
↓ 5 callersFunctiontimed_run
(args: list[str], label: str)
scripts/build_sidecars.py:123
↓ 5 callersMethodtop_k
Number of experts selected per token (top-k).
src/moe.rs:2102
↓ 5 callersFunctiontrait_ratio
src/cuda/flash_attn/cutlass/cute/numeric/integral_ratio.hpp:271
↓ 5 callersFunctiontx
(t)
tests/release_test.py:1541
↓ 5 callersFunctionvalidate_hqq_attention_tensors
( tensors: Dict[str, torch.Tensor], *, expected_nbits: Optional[int] = None, expected_axis: in
python/krasis/attention_backend.py:1259
↓ 5 callersMethodvisit
src/cuda/flash_attn/cutlass/cutlass/epilogue/fusion/sm90_visitor_tma_warpspecialized.hpp:597
↓ 5 callersFunctionwrite_manifest
(model_path: Path, profile: str, manifest: dict, nbits: int = 4)
tests/test_hqq_cache_profile.py:82
↓ 4 callersFunctionOffsetBytes
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:712
↓ 4 callersFunction_candidate_metrics
( name: str, source: torch.Tensor, input_rows: torch.Tensor, *, scale: float, zero: fl
python/krasis/hqq_self_calibrate.py:2685
↓ 4 callersFunction_csv_filter
(value: Optional[str])
python/krasis/hqq_self_calibrate.py:4024
↓ 4 callersFunction_dense_artifact_tensor_specs_from_real_case
( real_case: dict, selected_tensor_names: list[str] | None = None, )
tests/test_hqq_rust_quantizer.py:214
↓ 4 callersMethod_estimate_attention_vram
Estimate total GPU VRAM used by attention weights across all layers. Returns bytes consumed by all attention-related parameters: norm
python/krasis/model.py:3661
↓ 4 callersMethod_export_kv_to_rust
Set KV cache position for Rust decode after prefill. The KV cache is shared — Rust prefill already wrote FP8 data into the same GPU b
python/krasis/model.py:8591
↓ 4 callersFunction_family_identity
(event: Optional[Dict[str, Any]])
tests/trace_diff.py:744
↓ 4 callersFunction_family_member_key
(event: Dict[str, Any])
tests/trace_diff.py:755
↓ 4 callersFunction_find_nvidia_smi
Find nvidia-smi, including the WSL2 driver mount.
python/krasis/setup.py:59
↓ 4 callersFunction_format_timing
Format server timing into compact single-line format. Example: (pp: 21t 682ms 31t/s - tg 980t>12t 12s 89.2t/s)
python/krasis/chat.py:628
↓ 4 callersFunction_format_tokens
Format token count as e.g. '72K' or '1.2M'.
python/krasis/launcher.py:1081
↓ 4 callersFunction_get_nvcc_version
Get nvcc major.minor version, or None if not found.
python/krasis/setup.py:200
↓ 4 callersFunction_gpu_cleanup
Force GPU cleanup between tests to avoid stale CUDA state.
tests/archive/test_gpu_prefill.py:261
↓ 4 callersFunction_headline
Print a compact headline that stays readable when stdout is log-prefixed.
python/krasis/server.py:131
↓ 4 callersMethod_init_state
Initialize conv and recurrent states to zero.
python/krasis/linear_attention.py:222
↓ 4 callersFunction_load_activation_rows
(trace_path: str, *, layer: int, tensor: str)
python/krasis/hqq_self_calibrate.py:2480
↓ 4 callersFunction_load_trace_events
( path: Path, *, include_events: Optional[set[str]], include_components: Optional[set[str]],
tests/trace_diff.py:644
↓ 4 callersFunction_minmax_params
(source: torch.Tensor)
python/krasis/hqq_self_calibrate.py:2512
↓ 4 callersFunction_normalized_manifest_entry
(entry: dict)
tests/test_hqq_rust_quantizer.py:1060
↓ 4 callersFunction_normalized_runtime_entry
(entry: dict)
tests/test_hqq_rust_quantizer.py:1066
↓ 4 callersMethod_pack_expert_matrix
(self, weight_cpu: torch.Tensor)
tests/decode_harness.py:201
↓ 4 callersFunction_pack_hqq_quant
(quant: torch.Tensor, nbits: int)
python/krasis/attention_backend.py:1211
↓ 4 callersFunction_package_version
(package_name: str)
tests/generate_reference.py:123
↓ 4 callersFunction_parse_ranked_pairs
(text: Any)
tests/trace_diff.py:269
↓ 4 callersFunction_real_model_quant_cfg
()
tests/test_hqq_fused_branch_runtime.py:395
↓ 4 callersFunction_require_hqq4_rust_symbols
()
python/krasis/attention_backend.py:67
↓ 4 callersFunction_select_int8_exception_candidates
( candidate_report: Dict[str, Any], *, requested_groups: Optional[set[Tuple[int, str, int]]],
python/krasis/hqq_self_calibrate.py:1979
↓ 4 callersMethod_set_interactive_attention_quant
Apply an interactive attention preset. The TUI exposes four simple attention presets: HQQ4, HQQ4+10%, HQQ6, and HQQ6+10%.
python/krasis/launcher.py:1399
↓ 4 callersMethod_stream_attn_prefetch
Start async DMA for a layer's attention weights on the DMA stream. Call this while CPU is busy with MoE experts. Before using the pre
python/krasis/model.py:4320
↓ 4 callersFunction_trace_enabled
(component: str)
tests/reference_contract.py:1128
↓ 4 callersFunction_truncate_ansi
Truncate text to a visible width while preserving ANSI color codes.
python/krasis/launcher.py:87
↓ 4 callersFunction_validate_int8_exception_byte_cap
(value: Optional[int])
python/krasis/hqq_self_calibrate.py:1925
↓ 4 callersFunction_vram_ledger_enabled
()
python/krasis/model.py:208
↓ 4 callersFunction_write_real_artifact_layer_with_writer
( model: KrasisModel, real_case: dict, *, writer, selected_tensor_names: list[str] | None
tests/test_hqq_rust_quantizer.py:418
↓ 4 callersMethodaccum
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:1079
↓ 4 callersMethodaccum_init
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized_blockwise_scaling.hpp:823
↓ 4 callersFunctionadd_byte_offset_
Adds a pointer offset in units of element
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/conv2d_dgrad_output_gradient_tile_access_iterator_optimized.h:673
↓ 4 callersFunctionadd_pointer_offset
Adds a pointer offset in units of Element
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_access_iterator_tensor_op_sm80.h:163
↓ 4 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/ell_predicated_tile_iterator.h:808
← previousnext →901–1,000 of 12,380, ranked by callers