MCPcopy Create free account

hub / github.com/brontoguana/krasis / functions

Functions12,380 in github.com/brontoguana/krasis

↓ 21 callersMethodvalid
Returns true if the current coordinate is within the filter tensor W
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/conv3d_fprop_filter_tile_access_iterator_optimized.h:227
↓ 20 callersFunction_fail
(msg: str)
tests/test_network.py:40
↓ 20 callersFunction_to_device
Transfer tensor to target device, using CPU bounce if P2P is broken.
python/krasis/model.py:155
↓ 20 callersFunctionconst_max
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:758
↓ 20 callersFunctioncutlass_get_smem_pointer
CUTLASS helper to get SMEM pointer
src/cuda/flash_attn/cutlass/cutlass/arch/memory_sm75.h:63
↓ 20 callersFunctiondispatch_matmul_free
Free-function dispatch for quantized matmul (avoids &self borrow conflict).
src/decode.rs:1902
↓ 20 callersFunctionflatten
src/cuda/flash_attn/cutlass/cute/layout.hpp:529
↓ 20 callersMethodget
Returns a pointer
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_scale_bias_vector_access_iterator.h:129
↓ 20 callersMethodget
src/cuda/flash_attn/cutlass/cute/pointer_base.hpp:157
↓ 20 callersFunctionload_with_byte_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator.h:385
↓ 20 callersFunctionlog_memory_usage
Log current memory usage from /proc/self/status and /proc/meminfo. Called from diagnostic logging points to track memory consumption.
src/syscheck.rs:227
↓ 20 callersMethodmake_fragment_B
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:172
↓ 20 callersFunctionmake_tiled_copy_impl
src/cuda/flash_attn/cutlass/cute/atom/copy_atom.hpp:438
↓ 20 callersFunctionquantize_activation_int16
Quantize BF16 activation vector to INT16 with per-group scales. For each group: scale = max(|x|) / 32767, int16[i] = round(x[i] / scale). INT16 per-g
src/kernel/avx2.rs:234
↓ 20 callersFunctionreal
src/cuda/flash_attn/cutlass/cutlass/complex.h:395
↓ 20 callersMethodsize_bits
src/cuda/marlin/scalar_type.hpp:157
↓ 20 callersFunctionstore_with_pointer_offset
Stores a fragment to memory at the location pointed to by the iterator
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_simt_tile_iterator.h:507
↓ 19 callersFunction_html_escape
(s: str)
tests/reference_test.py:85
↓ 19 callersMethodcopy
src/cuda/flash_attn/cutlass/cute/arch/copy.hpp:55
↓ 19 callersFunctioncp_async4
src/cuda/marlin/marlin_vendor_common.h:102
↓ 19 callersFunctionfast_divmod
* Find quotient and remainder using device-side intrinsics */
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:275
↓ 19 callersFunctionfast_sqrt
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:859
↓ 19 callersMethodget
Value
src/cuda/flash_attn/cutlass/cute/container/array_subbyte.hpp:161
↓ 19 callersMethodget_flat_coord
src/cuda/flash_attn/cutlass/cute/layout.hpp:261
↓ 19 callersFunctioninfo
(msg: str)
tests/validate_model.py:105
↓ 19 callersMethodinit_masks
src/cuda/flash_attn/cutlass/cutlass/pipeline/sm100_pipeline.hpp:150
↓ 19 callersMethodisnan
Is NaN implementation
src/cuda/flash_attn/cutlass/cutlass/float8.h:190
↓ 19 callersFunctionmake_identity_layout
src/cuda/flash_attn/cutlass/cute/layout.hpp:483
↓ 19 callersFunctionmake_tiled_copy_A
src/cuda/flash_attn/cutlass/cute/atom/copy_atom.hpp:452
↓ 19 callersFunctionprint
src/cuda/flash_attn/cutlass/cute/atom/copy_atom.hpp:648
↓ 19 callersMethodtensor_info
Get metadata for a tensor.
src/weights/safetensors_io.rs:155
↓ 18 callersMethod_dense_mlp_tensor
Extract a BF16 tensor from a dense MLP weight value. Dense MLP weights are either plain BF16 tensors or (weight_int8, scale) tuples f
python/krasis/model.py:6595
↓ 18 callersFunction_format_metric_pct
(value: Optional[float], digits: int = 0)
tests/reference_test.py:777
↓ 18 callersFunction_parse_float
(text: Any)
tests/trace_diff.py:205
↓ 18 callersMethodadd
(self, pos: int, row_hash: str, values: List[float])
python/krasis/hqq_self_calibrate.py:239
↓ 18 callersFunctionadd_pointer_offset
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/shared_load_iterator.h:67
↓ 18 callersFunctionblock_id_in_cluster
Returns the relative dim3 block rank local to the cluster.
src/cuda/flash_attn/cutlass/cute/arch/cluster_sm90.hpp:126
↓ 18 callersFunctionget_gmem_layout
src/cuda/flash_attn/cutlass/cutlass/detail/collective/mixed_input_utils.hpp:519
↓ 18 callersMethodinverse
src/cuda/flash_attn/cutlass/cutlass/layout/matrix.h:114
↓ 18 callersMethodisfinite
Is finite implementation
src/cuda/flash_attn/cutlass/cutlass/float8.h:176
↓ 18 callersFunctionload_prequantized_weight
Load a pre-quantized INT4 weight directly (compressed-tensors format). Reads weight_packed (I32), weight_scale (BF16), weight_shape (I32[2]).
src/weights/mod.rs:6488
↓ 18 callersMethodmake_fragment_A
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:146
↓ 18 callersMethodraw
Accesses raw internal state
src/cuda/flash_attn/cutlass/cutlass/half.h:463
↓ 18 callersMethodretile_S
src/cuda/flash_attn/cutlass/cute/atom/copy_atom.hpp:416
↓ 18 callersFunctiontake
src/cuda/flash_attn/cutlass/cute/layout.hpp:507
↓ 17 callersFunctionblocked_product
src/cuda/flash_attn/cutlass/cute/layout.hpp:1748
↓ 17 callersMethodflush
Flush any remaining buffered tokens (end of stream).
src/server.rs:45
↓ 17 callersMethodis_source_needed
Determine if the source is needed. May return false if
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/epilogue_with_broadcast.h:123
↓ 17 callersFunctionload_hqq_attention_manifest
( model_path: str, cache_profile: Optional[str] = HQQ_CACHE_PROFILE_BASELINE, nbits: int = 4,
python/krasis/attention_backend.py:756
↓ 17 callersFunctionmake_zip_tensor
src/cuda/flash_attn/cutlass/cute/tensor_zip.hpp:157
↓ 17 callersMethodmoe_forward
Run MoE forward for a single token on one layer. Args: moe_layer_idx: 0-based MoE layer index (skipping dense layers) activation_bf16: BF16 activatio
src/moe.rs:1968
↓ 17 callersFunctionnorm
src/cuda/flash_attn/cutlass/cutlass/complex.h:481
↓ 17 callersFunctionpartition_fragment_C
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:621
↓ 17 callersFunctionreport_event
Record a named event with current VRAM snapshot. No-op if reporting is not enabled.
src/vram_monitor.rs:392
↓ 17 callersMethodresize
Changes the size of the view without affecting pointer or layout
src/cuda/flash_attn/cutlass/cutlass/tensor_view.h:170
↓ 17 callersFunctionstride
src/cuda/flash_attn/cutlass/cute/tensor_impl.hpp:541
↓ 16 callersFunctionabs
src/cuda/flash_attn/cutlass/cute/numeric/math.hpp:65
↓ 16 callersFunctionalignment_for_swizzle
src/cuda/flash_attn/cutlass/cutlass/detail/layout.hpp:418
↓ 16 callersMethodclear
Efficient clear method
src/cuda/flash_attn/cutlass/cutlass/array.h:360
↓ 16 callersFunctionconj
src/cuda/flash_attn/cutlass/cutlass/complex.h:568
↓ 16 callersMethodfill
src/cuda/flash_attn/cutlass/cute/container/array.hpp:171
↓ 16 callersFunctionget_logical_ptr
src/cuda/flash_attn/cutlass/cutlass/detail/collective/mixed_input_utils.hpp:493
↓ 16 callersFunctionget_nonswizzle_portion
src/cuda/flash_attn/cutlass/cute/swizzle_layout.hpp:152
↓ 16 callersFunctionimag
src/cuda/flash_attn/cutlass/cutlass/complex.h:404
↓ 16 callersFunctionis_awq_scaled_tensor
Check if a tensor should have AWQ per-channel scaling applied.
python/krasis/awq_calibrate.py:1002
↓ 16 callersMethodload_init
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized.hpp:495
↓ 16 callersFunctionload_with_byte_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_complex_tensor_op_tile_iterator_sm80.h:253
↓ 16 callersFunctionmake_tiled_copy_B
src/cuda/flash_attn/cutlass/cute/atom/copy_atom.hpp:461
↓ 16 callersMethodmma
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized.hpp:648
↓ 16 callersMethodstride
Returns the stride of the layout
src/cuda/flash_attn/cutlass/cutlass/layout/tensor_op_multiplicand_sm80.h:127
↓ 16 callersFunctionsynclog_emit_tma_load
src/cuda/flash_attn/cutlass/cutlass/arch/synclog.hpp:813
↓ 16 callersFunctionthr_size
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:690
↓ 16 callersFunctiontile_shape
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:672
↓ 15 callersFunctionRematerializeBlockIdxZ
Helper to rematerialize block Idx. Reduces register liveness.
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/index_remat.h:78
↓ 15 callersFunction_emit_real_model_timing
(payload: dict)
python/krasis/attention_backend.py:370
↓ 15 callersFunction_fused_add_rmsnorm
In-place fused add + RMSNorm. hidden = RMSNorm(hidden + residual, weight, eps) residual = hidden + residual (before norm)
python/krasis/layer.py:38
↓ 15 callersMethod_message_screen
(self, title: str, lines: List[str], *, wait: bool = True)
python/krasis/launcher.py:1779
↓ 15 callersFunction_pass
(msg: str)
tests/test_network.py:36
↓ 15 callersFunction_utc_now
()
python/krasis/hqq_self_calibrate.py:125
↓ 15 callersFunctionassert_raises_contains
(fn, text: str)
tests/test_hqq_cache_profile.py:89
↓ 15 callersFunctioncanonical_lane_idx
Returns a lane index in the warp. The threads in warp may not be convergent
src/cuda/flash_attn/cutlass/cutlass/cutlass.h:114
↓ 15 callersFunctionceil_div
src/cuda/flash_attn/cutlass/cute/layout.hpp:1608
↓ 15 callersFunctioncheck_cuda_status
src/cuda/flash_attn/cutlass/cutlass/experimental/distributed/device/detail.hpp:46
↓ 15 callersFunctioncomposition
src/cuda/flash_attn/cutlass/cute/swizzle_layout.hpp:299
↓ 15 callersFunctionconstruct
src/cuda/flash_attn/cutlass/cute/algorithm/tuple_algorithms.hpp:626
↓ 15 callersMethodcount
src/cuda/flash_attn/cutlass/cutlass/pipeline/sm90_pipeline.hpp:198
↓ 15 callersFunctioncp_async_fence
Establishes an ordering w.r.t previously issued cp.async instructions. Does not block.
src/cuda/flash_attn/cutlass/cute/arch/copy_sm80.hpp:162
↓ 15 callersFunctiondiv_ceil
src/cuda/marlin/marlin_vendor_common.h:86
↓ 15 callersFunctionfold
src/cuda/flash_attn/cutlass/cute/algorithm/tuple_algorithms.hpp:387
↓ 15 callersFunctionlaunch_dependent_grids
Issuing the launch_dependents instruction hints a dependent kernel to launch earlier launch_dependents doesn't impact the functionality but the perfor
src/cuda/flash_attn/cutlass/cutlass/arch/grid_dependency_control.h:78
↓ 15 callersFunctionload_with_byte_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/scale_bias_tile_iterator.h:258
↓ 15 callersFunctionload_with_pointer_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator.h:419
↓ 15 callersFunctionmax_common_vector
src/cuda/flash_attn/cutlass/cute/layout.hpp:1409
↓ 15 callersMethodnum_moe_layers
Number of MoE layers loaded.
src/moe.rs:2067
↓ 15 callersMethodpartition_fragment_B
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:560
↓ 15 callersFunctionpartition_shape_A
src/cuda/flash_attn/cutlass/cute/atom/mma_atom.hpp:634
↓ 15 callersFunctionpipeline_check_is_producer
src/cuda/flash_attn/cutlass/cutlass/pipeline/sm90_pipeline.hpp:64
↓ 15 callersFunctionprepend
src/cuda/flash_attn/cutlass/cute/algorithm/tuple_algorithms.hpp:816
↓ 15 callersMethodstore_tail
src/cuda/flash_attn/cutlass/cutlass/epilogue/collective/detail.hpp:476
↓ 15 callersFunctionstore_with_pointer_offset
Stores a fragment to memory with additional pointer offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator.h:4636
← previousnext →301–400 of 12,380, ranked by callers