MCPcopy Create free account

hub / github.com/brontoguana/krasis / functions

Functions12,380 in github.com/brontoguana/krasis

↓ 8 callersFunctionfilter_tuple
src/cuda/flash_attn/cutlass/cute/algorithm/tuple_algorithms.hpp:345
↓ 8 callersFunctionflatten_to_tuple
src/cuda/flash_attn/cutlass/cute/algorithm/tuple_algorithms.hpp:544
↓ 8 callersMethodfree
src/cuda/flash_attn/cutlass/cute/arch/tmem_allocator_sm100.hpp:88
↓ 8 callersMethodget
Returns a pointer
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_access_iterator_tensor_op_sm80.h:540
↓ 8 callersMethodget_k
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/ell_predicated_tile_access_iterator.h:888
↓ 8 callersMethodget_mask
Gets the mask
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_access_iterator.h:286
↓ 8 callersFunctionget_run_dir
(default_type: str = "run")
python/krasis/run_paths.py:44
↓ 8 callersMethodget_sk_block_idx
Obtains calling linear threadblock index of the first block to work on the given tile
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/threadblock_swizzle_streamk.h:732
↓ 8 callersFunctionget_tiled_cta_shape_mnl
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/sm100_tile_scheduler_stream_k.hpp:193
↓ 8 callersFunctionis_garbage
Heuristic check for garbage output.
tests/test_network.py:118
↓ 8 callersMethodis_source_needed
src/cuda/flash_attn/cutlass/cutlass/epilogue/collective/default_epilogue.hpp:148
↓ 8 callersFunctionispow2
Returns true if the argument is a power of 2
src/cuda/flash_attn/cutlass/cutlass/array.h:81
↓ 8 callersFunctionload_with_byte_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm80.h:270
↓ 8 callersMethodload_with_byte_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_complex_tensor_op_tile_iterator_sm80.h:1006
↓ 8 callersMethodload_with_byte_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/tile_iterator_planar_complex.h:164
↓ 8 callersMethodload_with_byte_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator_sm70.h:1358
↓ 8 callersMethodload_with_byte_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_iterator.h:3386
↓ 8 callersMethodload_with_pointer_offset
Loads a fragment from memory with additional logical offset
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/tile_iterator_planar_complex.h:176
↓ 8 callersFunctionmake_int8_exception_candidate
( *, layer: int, tensor: str, group: int, benefit: float, cost: int, ready: bool =
tests/test_hqq_self_calibrate.py:157
↓ 8 callersFunctionmake_utccp_copy
src/cuda/flash_attn/cutlass/cute/atom/copy_traits_sm100.hpp:3747
↓ 8 callersFunctionmarlin_repack_int8
Repack a QuantizedInt8 into Marlin GPU INT8 format. Follows the same structure as INT4 marlin_repack but with: - 4 values packed per u32 (not 8) - IN
src/weights/marlin.rs:739
↓ 8 callersFunctionmma
src/cuda/flash_attn/cutlass/cutlass/conv/collective/sm100_implicit_gemm_umma_warpspecialized.hpp:838
↓ 8 callersMethodmma_init
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized.hpp:548
↓ 8 callersFunctionouter_partition
src/cuda/flash_attn/cutlass/cute/tensor_impl.hpp:1003
↓ 8 callersFunctionpipeline_init_arrive_relaxed
Used to guarantee that the Pipeline init is visible to all producers and consumer threadblocks in the cluster
src/cuda/flash_attn/cutlass/cutlass/pipeline/sm90_pipeline.hpp:1368
↓ 8 callersFunctionpipeline_init_wait
Synchronization call. Blocks until barriers are initialized in shared memory.
src/cuda/flash_attn/cutlass/cutlass/pipeline/sm90_pipeline.hpp:1355
↓ 8 callersFunctionquantize_int4_expert_calibrated
( weight_bf16: &[u16], rows: usize, cols: usize, group_size: usize, mode: ExpertInt4CalibM
src/weights/mod.rs:6629
↓ 8 callersMethodreduction_subtile_idx
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/static_tile_scheduler.hpp:79
↓ 8 callersFunctionround_up
src/cuda/flash_attn/cutlass/cute/int_tuple.hpp:356
↓ 8 callersMethodrun
Run all prompts and report results.
python/krasis/stress_test.py:372
↓ 8 callersFunctionsend_chat_request
Send a chat completion request. Returns (status_code, response_body).
tests/test_network.py:52
↓ 8 callersMethodset_iteration_index
Overrides the internal iteration index
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_access_iterator_tensor_op_sm80.h:525
↓ 8 callersMethodset_mask
Sets the predicate mask, overriding value stored in predicate iterator
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_access_iterator.h:276
↓ 8 callersMethodset_tmem_offsets
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_warpspecialized.hpp:480
↓ 8 callersFunctionstore
Stores a fragment to memory at the location pointed to by the iterator
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_simt_tile_iterator.h:525
↓ 8 callersFunctiontensormaps_replace_global_address
Replace address for the global tensor (to be done by single thread)
src/cuda/flash_attn/cutlass/cutlass/gemm/collective/sm100_mma_array_warpspecialized.hpp:769
↓ 8 callersMethodtransposed_problem
Returns arguments for the transposed problem
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/gemm_with_absmax.h:217
↓ 7 callersFunctionLayoutAwareConvert
src/cuda/flash_attn/cutlass/cutlass/detail/collective/mixed_input_utils.hpp:456
↓ 7 callersFunction_base_config
()
tests/test_launcher_matrix.py:91
↓ 7 callersFunction_component_weight_bytes
Weight bytes for per-component quant ("int4", "int8", "awq", or "bf16").
python/krasis/vram_budget.py:291
↓ 7 callersFunction_format_token_count
Format token count: 175000 -> '175K', 1200000 -> '1.2M'.
python/krasis/chat.py:598
↓ 7 callersFunction_load_source_tensor
(model_path: str, entry: Dict[str, Any])
python/krasis/hqq_self_calibrate.py:833
↓ 7 callersFunction_manifest_tensor_entries
(manifest: Dict[str, Any])
python/krasis/hqq_self_calibrate.py:651
↓ 7 callersMethod_moe_forward
MoE forward: routing + shared expert (GPU) + routed experts (CPU). For LatentMoE (Nemotron): wraps expert compute with latent projections.
python/krasis/layer.py:662
↓ 7 callersFunction_packet_brief
(packet: Optional[Dict[str, Any]])
tests/trace_diff.py:1046
↓ 7 callersFunction_parse_csv_ints
(value: str)
tests/test_hqq_rust_quantizer.py:1804
↓ 7 callersFunction_pkg_flag
Return ['-y'] if auto-approve, else [] so the package manager prompts.
python/krasis/setup.py:42
↓ 7 callersFunction_prefix_probe
( method: str, model: Any, tokenizer: Any, full_input_ids: Any, prefix_len: int, *,
tests/reference_backfill_topk.py:519
↓ 7 callersMethod_register_attn_weight
(self, shape: Tuple[int, int])
tests/decode_harness.py:151
↓ 7 callersFunction_section
Format a highlighted section header.
python/krasis/benchmark.py:37
↓ 7 callersFunction_source_contract_summary
(validations: Iterable[Dict[str, Any]])
python/krasis/hqq_self_calibrate.py:998
↓ 7 callersMethod_stream_attn_load
Copy one layer's attention weights from CPU pinned to GPU buffers. Sets the layer's attention attributes to point to the pre-allocated
python/krasis/model.py:4284
↓ 7 callersFunction_validate_source_contract_entry
( *, model_path: str, cache_dir: str, entry: Dict[str, Any], )
python/krasis/hqq_self_calibrate.py:945
↓ 7 callersFunction_warn
Print a warning line (yellow, indented).
python/krasis/server.py:126
↓ 7 callersFunctionabort_if_cuda_context_poisoned
(context: &str, err: &str)
src/server.rs:62
↓ 7 callersFunctionabs_for_integer
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:159
↓ 7 callersFunctionadd_candidate
(recipe: str, names: list[str], row_column_layout: str, current_resolver_status: str)
tests/hqq_attention_diff.py:2490
↓ 7 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator.h:1406
↓ 7 callersMethodarrive
Signal completion of Stage and move to the next stage (group_id) signals to (group_id+1)
src/cuda/flash_attn/cutlass/cutlass/pipeline/sm90_pipeline.hpp:1332
↓ 7 callersMethodbegin_loop
Start of subtile store iteration
src/cuda/flash_attn/cutlass/cutlass/epilogue/fusion/sm90_visitor_tma_warpspecialized.hpp:195
↓ 7 callersMethodbegin_sync_needed
Is a thread sync needed after begin(). Allows chaining async copies across multiple nodes
src/cuda/flash_attn/cutlass/cutlass/epilogue/fusion/sm90_visitor_tma_warpspecialized.hpp:185
↓ 7 callersFunctionbit_width
src/cuda/flash_attn/cutlass/cute/numeric/math.hpp:145
↓ 7 callersFunctionbuild_prompt
Build a prompt of approximately target_tokens length using Gutenberg text.
tests/archive/test_v2lite_profile.py:36
↓ 7 callersMethodclear_mask
Clears the predicate set efficiently
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator.h:1435
↓ 7 callersMethodcompute_epilogue
Returns whether the block assigned this work should compute the epilogue for the corresponding output tile. For the basic tile scheduler, this is alwa
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/sm100_tile_scheduler.hpp:517
↓ 7 callersFunctionconditional_return
src/cuda/flash_attn/cutlass/cute/numeric/integral_constant.hpp:432
↓ 7 callersFunctioncrd2idx
src/cuda/flash_attn/cutlass/cute/layout.hpp:677
↓ 7 callersFunctioncrd2idx
src/cuda/flash_attn/cutlass/cute/stride.hpp:100
↓ 7 callersMethoddata
Returns a pointer to the shared memory buffer
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/epilogue_base.h:175
↓ 7 callersFunctiondequantize_marlin_int8
Dequantize Marlin-repacked INT8 weights back to f32 for verification. Reverses the INT8 permutation and packing to recover the original [N, K] f32 va
src/weights/marlin.rs:1391
↓ 7 callersMethoddeterminant
Computes the determinant of a 2-by-2 matrix
src/cuda/flash_attn/cutlass/cutlass/matrix.h:3231
↓ 7 callersMethoddiv_cluster_size
Divides dividend by the cluster size
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/tile_scheduler_params.h:476
↓ 7 callersFunctionemit_contract_failure_trace
( context: str, reason: str, **fields: Any, )
tests/reference_contract.py:1205
↓ 7 callersMethodenable_mask
Clears the predicate set efficiently
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator.h:1441
↓ 7 callersMethodend_loop
End of subtile store iteration
src/cuda/flash_attn/cutlass/cutlass/epilogue/fusion/sm90_visitor_tma_warpspecialized.hpp:265
↓ 7 callersFunctionexplode_tuple
src/cuda/flash_attn/cutlass/cute/arch/util.hpp:284
↓ 7 callersFunctionfast_min
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:764
↓ 7 callersFunctionfast_tanh
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:922
↓ 7 callersFunctionfp4_shift_A
src/cuda/flash_attn/cutlass/cute/atom/mma_traits_sm120.hpp:203
↓ 7 callersFunctionfp4_shift_B
src/cuda/flash_attn/cutlass/cute/atom/mma_traits_sm120.hpp:207
↓ 7 callersMethodgather
Copy from GPUs to pinned CPU memory (asynchronous).
python/krasis/model.py:614
↓ 7 callersFunctionget
Returns a pointer
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/predicated_scale_bias_vector_access_iterator.h:240
↓ 7 callersMethodget_log_tile
Calculates optimal swizzle width
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/threadblock_swizzle.h:116
↓ 7 callersMethodget_mask
Gets the mask
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator.h:1453
↓ 7 callersFunctionhas_single_bit
src/cuda/flash_attn/cutlass/cute/numeric/math.hpp:127
↓ 7 callersFunctionhqq_auto_promotion_policy
(attention_quant: str)
python/krasis/attention_backend.py:315
↓ 7 callersMethodis_fp8
True if this is an FP8 type that needs dequantization.
src/weights/safetensors_io.rs:68
↓ 7 callersMethodlayout
Returns the layout object
src/cuda/flash_attn/cutlass/cutlass/tensor_ref.h:285
↓ 7 callersFunctionmake_basis_like
src/cuda/flash_attn/cutlass/cute/numeric/arithmetic_tuple.hpp:333
↓ 7 callersFunctionmanifest_matches
(manifest: dict[str, object], contracts: dict[str, dict[str, object]])
scripts/build_sidecars.py:393
↓ 7 callersFunctionnormalize_hqq_attention_group_size
(group_size: Optional[int])
python/krasis/attention_backend.py:385
↓ 7 callersFunctionparse_hf_repo_id
Parse a Hugging Face model URL or repo id into `namespace/name`.
python/krasis/hf_downloader.py:220
↓ 7 callersFunctionprepare_lens
(cu_seqlens: torch.LongTensor)
src/cuda/fla/ops/utils/index.py:37
↓ 7 callersMethodpure
src/cuda/flash_attn/cutlass/cutlass/quaternion.h:201
↓ 7 callersFunctionread_manifest
(path: Path = MANIFEST_PATH)
scripts/build_sidecars.py:347
↓ 7 callersMethodread_u32
(&mut self)
src/gguf.rs:208
↓ 7 callersMethodread_u64
(&mut self)
src/gguf.rs:226
↓ 7 callersFunctionrepack_tiled_int4_packed
Repack INT4 packed weights from [K/8, N] to [N/TILE, K/8, TILE]. Last tile is zero-padded if N is not a multiple of TILE_N.
src/kernel/avx2.rs:1318
↓ 7 callersFunctionrepack_tiled_int8_packed
Repack INT8 data from [K, N] (i8 in u32 container) to [N/TILE, K, TILE] tiled.
src/kernel/avx2.rs:1690
↓ 7 callersFunctionrepack_tiled_scales
Repack INT4/INT8 scales from [K/gs, N] to [N/TILE, K/gs, TILE].
src/kernel/avx2.rs:1340
← previousnext →601–700 of 12,380, ranked by callers