MCPcopy Create free account

hub / github.com/brontoguana/krasis / functions

Functions12,380 in github.com/brontoguana/krasis

↓ 2 callersFunction_validate_sidecar_correction_scale
(value: float)
python/krasis/hqq_self_calibrate.py:3198
↓ 2 callersFunction_variant_case_effects
(variant: Dict[str, Any])
python/krasis/hqq_self_calibrate.py:4653
↓ 2 callersFunction_verify_hcs_vram_floor
Gate readiness on measured free VRAM, not just planned HCS budget.
python/krasis/server.py:2378
↓ 2 callersFunction_weight_bytes
(params: int, quantization: str)
python/krasis/vram_budget.py:119
↓ 2 callersFunction_weight_dtype_code
Map torch dtype to Rust GpuWeight dtype code: 0=BF16, 1=FP32, 2=FP16.
python/krasis/model.py:125
↓ 2 callersFunction_witness_summary_metrics
(summary_path: str)
python/krasis/hqq_self_calibrate.py:4437
↓ 2 callersFunction_write_hqq_attention_artifact_instrumented
( case: str, variant: str, *, model_path: str, layer_idx: int, layer_type: str, te
tests/test_hqq_rust_quantizer.py:489
↓ 2 callersFunction_write_real_artifact_layer_instrumented
( case: str, variant: str, model: KrasisModel, real_case: dict, *, use_rust_shadow: bo
tests/test_hqq_rust_quantizer.py:923
↓ 2 callersFunctionabs
src/cuda/flash_attn/cutlass/cutlass/quaternion.h:391
↓ 2 callersFunctionabsolute_value
src/cuda/flash_attn/cutlass/cutlass/fast_math.h:1073
↓ 2 callersFunctionadd_byte_offset_
Adds a pointer offset in units of element
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/conv3d_dgrad_output_gradient_tile_access_iterator_optimized.h:286
↓ 2 callersFunctionadd_byte_offset_
Adds a pointer offset in units of element
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/conv3d_fprop_activation_tile_access_iterator_optimized.h:281
↓ 2 callersFunctionadd_byte_offset_
Adds a pointer offset in units of element
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_optimized.h:264
↓ 2 callersFunctionadd_pointer_offset
Adds a pointer offset in units of Element
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_access_iterator_pitch_linear_direct_conv.h:155
↓ 2 callersFunctionadd_pointer_offset
Adds a pointer offset in units of Element
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_access_iterator_tensor_op.h:168
↓ 2 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator_triangular_matrix.h:723
↓ 2 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator_2dthreadtile.h:704
↓ 2 callersMethodadd_pointer_offset
Adds a pointer offset in units of Element
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_iterator_pitch_linear_2dthreadtile.h:243
↓ 2 callersFunctionadd_ranked
(bucket: List[Dict[str, Any]], bucket_budget: int)
python/krasis/hqq_self_calibrate.py:2045
↓ 2 callersFunctionadd_samples
(item: dict[str, Any])
tests/hqq_attention_diff.py:746
↓ 2 callersMethodadd_tile_offset
Adds a tile offset
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/regular_tile_iterator_pitch_linear_2dthreadtile.h:249
↓ 2 callersFunctionadvance
Advances the iterator along the advance dimension
src/cuda/flash_attn/cutlass/cutlass/gemm/warp/mma_tensor_op_tile_access_iterator.h:242
↓ 2 callersFunctionadvance_smem_write_stage
Advance global memory read-iterators and shared memory write-iterators to the stage
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/mma_multistage.h:263
↓ 2 callersFunctionadvance_to_next_work
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:323
↓ 2 callersFunctionaffine_candidate_score
(source: torch.Tensor, input_group: torch.Tensor, scale: float, zero: float, q: torch.Tensor)
tests/hqq_attention_diff.py:864
↓ 2 callersMethodapply
src/cuda/flash_attn/cutlass/cute/swizzle.hpp:74
↓ 2 callersFunctionapply_capture_template
(tokenizer: Any, messages: List[Dict[str, str]], capture_settings: Dict[str, Any])
tests/reference_contract.py:512
↓ 2 callersMethodapply_dropout
src/cuda/flash_attn/fa2/dropout.h:27
↓ 2 callersMethodapply_with_tools
Apply with optional tools array for accurate token estimation.
src/chat_template.rs:108
↓ 2 callersMethodapply_with_tools_inner
( &self, messages_json: &str, tools_json: &str, add_generation_prompt: bool,
src/chat_template.rs:141
↓ 2 callersFunctionarrive
src/cuda/flash_attn/cutlass/cutlass/arch/barrier.h:218
↓ 2 callersFunctionas_position_independent_swizzle_layout
src/cuda/flash_attn/cutlass/cute/pointer_flagged.hpp:91
↓ 2 callersMethodas_slice
Get a typed slice view of the allocation. # Safety Caller must ensure the data has been properly initialized as type T.
src/numa.rs:72
↓ 2 callersMethodat
Returns a reference to the element at a given Coord
src/cuda/flash_attn/cutlass/cutlass/tensor_ref_planar_complex.h:286
↓ 2 callersFunctionawq_template_exists
Check if an AWQ calibration template exists for this model. Checks both the bundled templates directory and user cache.
tests/release_test.py:398
↓ 2 callersMethodbegin_epilogue
Called at the start of the epilogue just before iterating over accumulator slices
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/fusion/visitor_2x.hpp:63
↓ 2 callersFunctionbench_decode
Benchmark pure decode speed (timing OFF).
tests/archive/investigate_decode_regression.py:45
↓ 2 callersFunctionbenchmark_decode
Run NUM_RUNS decode benchmarks and return results.
tests/archive/test_qwen3_next_gpu_decode.py:44
↓ 2 callersFunctionbenchmark_decode
Benchmark decode: measure tok/s after prefill.
tests/archive/run_hcs_benchmark.py:63
↓ 2 callersFunctionbenchmark_decode
Benchmark: prefill + decode, measuring TTFT and decode speed separately.
tests/archive/run_hcs_hybrid_bench.py:65
↓ 2 callersFunctionbenchmark_short_decode
Benchmark decode with short prompt to isolate decode speed.
tests/archive/run_hcs_benchmark.py:100
↓ 2 callersFunctionbenchmark_short_decode
Short prompt to isolate decode speed.
tests/archive/run_hcs_hybrid_bench.py:102
↓ 2 callersFunctionbest_fit_affine_params
(source: torch.Tensor)
tests/hqq_attention_diff.py:831
↓ 2 callersFunctionbf16_to_f32
(x: u16)
src/decode.rs:1989
↓ 2 callersFunctionbf16_word_to_f32
(word: int)
tests/hqq_attention_diff.py:151
↓ 2 callersFunctionbuild_base_store
( *, include_fused_config: bool, )
tests/test_hqq_fused_branch_runtime.py:176
↓ 2 callersFunctionbuild_invocation_metadata
(model_name: str, profile_id: str, max_new_tokens: int)
tests/generate_reference.py:184
↓ 2 callersFunctionbuild_prompt
Build a prompt of approximately target_tokens length using Gutenberg text.
tests/archive/run_hcs_benchmark.py:57
↓ 2 callersFunctionbuild_prompt
Build a prompt of approximately target_tokens length using Gutenberg text.
tests/archive/run_hcs_hybrid_bench.py:59
↓ 2 callersFunctionbuild_prompt
Build a prompt of approximately target_tokens length using Gutenberg text.
tests/archive/token_scaling_bench.py:29
↓ 2 callersFunctionbuild_qwen35_source
(model_dir: Path, layer: int, tensor_name: str)
tests/hqq_attention_diff.py:2203
↓ 2 callersFunctionbuild_reference_sanity_report
( reference: Dict[str, Any], *, expected_conversations: int, expected_turns: int, )
tests/reference_contract.py:683
↓ 2 callersFunctionbuild_runtime_contract
( conf_path: str, tokenizer: Any, *, max_new_tokens: int, effective_profile_id: Optional[s
tests/reference_contract.py:929
↓ 2 callersFunctionbuild_ssh_tunnel_command
Build the ssh command for a reverse tunnel without invoking a shell.
python/krasis/ssh_tunnel.py:49
↓ 2 callersFunctionbuild_token_audit_turns
(reference: Dict[str, Any])
tests/generate_reference.py:659
↓ 2 callersFunctionbundle_filename
(bundle_hash: str)
scripts/build_sidecars.py:433
↓ 2 callersFunctioncache_path_cpu
Cache file path for CPU-optimized transposed format (INT4 or INT8).
src/weights/mod.rs:1553
↓ 2 callersFunctioncache_path_gguf_avx2
Cache file path for GGUF-sourced AVX2 transposed CPU cache.
src/weights/mod.rs:1559
↓ 2 callersFunctioncalculate_umma_peer_mask
src/cuda/flash_attn/cutlass/cutlass/pipeline/sm100_pipeline.hpp:95
↓ 2 callersFunctioncall_prefill_logits
Call the /v1/internal/prefill_logits endpoint with raw token IDs.
tests/reference_test.py:442
↓ 2 callersFunctioncall_reference_test
Call the /v1/internal/reference_test endpoint.
tests/reference_test.py:360
↓ 2 callersMethodcan_use
(cls)
src/cuda/fla/ops/backends/__init__.py:40
↓ 2 callersFunctioncandidate_regression_record
( *, scope: dict[str, Any], aggregate: dict[str, Any], candidate_name: str, )
tests/hqq_attention_diff.py:1811
↓ 2 callersFunctioncheck_health
Check if server is healthy.
tests/test_network.py:149
↓ 2 callersFunctionchunk_gated_delta_rule_fwd_h
( k: torch.Tensor, w: torch.Tensor, u: torch.Tensor, g: torch.Tensor | None = None, gk: to
src/cuda/fla/ops/common/chunk_delta_h.py:655
↓ 2 callersFunctionchunk_local_cumsum
( g: torch.Tensor, chunk_size: int, reverse: bool = False, scale: float = None, cu_seqlens
src/cuda/fla/ops/utils/cumsum.py:429
↓ 2 callersFunctioncleanup_marlin_cache_before_build
( model_dir: &Path, config: &ModelConfig, total_moe_layers: usize, config_hash: u64, gpu_b
src/weights/mod.rs:1355
↓ 2 callersMethodclear
< Efficiently disables all accesses guarded by mask
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/predicated_tile_iterator_affine.h:227
↓ 2 callersMethodclear
< Efficiently disables all accesses guarded by mask
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/predicated_tile_iterator_direct_conv.h:154
↓ 2 callersMethodclear
< Efficiently disables all accesses guarded by mask
src/cuda/flash_attn/cutlass/cutlass/epilogue/threadblock/predicated_tile_iterator_strided_dgrad.h:155
↓ 2 callersFunctionclear_mask
Clears the predicates
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/conv2d_fprop_filter_tile_access_iterator_optimized.h:243
↓ 2 callersFunctionclear_mask
Clears the predicates
src/cuda/flash_attn/cutlass/cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_optimized.h:324
↓ 2 callersMethodclear_mask
Clears the predicate set efficiently
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator_triangular_matrix.h:752
↓ 2 callersMethodclear_mask
Clears the predicate set efficiently
src/cuda/flash_attn/cutlass/cutlass/transform/threadblock/predicated_tile_iterator_2dthreadtile.h:733
↓ 2 callersFunctioncluster_arrive
src/cuda/flash_attn/cutlass/cute/arch/cluster_sm90.hpp:57
↓ 2 callersFunctioncoalesce_x
src/cuda/flash_attn/cutlass/cute/layout.hpp:824
↓ 2 callersFunctioncollect_failure_case
(*, path_name: str, use_artifacts: bool)
tests/test_hqq_fused_branch_runtime.py:777
↓ 2 callersFunctioncollect_real_model_success_case
( *, real_case: dict, path_name: str, include_fused_desc: bool, include_fused_config: bool
tests/test_hqq_fused_branch_runtime.py:870
↓ 2 callersFunctioncompact_order
src/cuda/flash_attn/cutlass/cute/stride.hpp:410
↓ 2 callersFunctioncompare_real_artifact_layer
(name: str, real_case: dict)
tests/test_hqq_rust_quantizer.py:1213
↓ 2 callersFunctioncompare_real_case_tensors
( name: str, real_case: dict, tensor_names: list[str], group_indices: list[int], )
tests/test_hqq_rust_quantizer.py:152
↓ 2 callersFunctioncompare_rmse
(name: str, chunk: torch.Tensor, q: torch.Tensor, scale: torch.Tensor, zero: torch.Tensor)
tests/test_hqq_rust_quantizer.py:77
↓ 2 callersFunctioncompare_solve
(name: str, chunk: torch.Tensor, zero: torch.Tensor, scale_seed: torch.Tensor)
tests/test_hqq_rust_quantizer.py:112
↓ 2 callersFunctioncomposition_impl
src/cuda/flash_attn/cutlass/cute/layout.hpp:1031
↓ 2 callersFunctioncompute_metrics
Compute all comparison metrics for one turn.
tests/reference_test.py:506
↓ 2 callersFunctioncompute_prefill_metrics
Compare our prefill logits against BF16 reference at sampled positions.
tests/reference_test.py:461
↓ 2 callersFunctioncompute_rmse
(chunk: &[f32], q: &[u8], scale: f32, zero: f32)
src/hqq.rs:114
↓ 2 callersMethodconsumer_release
src/cuda/flash_attn/cutlass/cutlass/pipeline/sm100_pipeline.hpp:1097
↓ 2 callersMethodcontext_for
( &self, layer_idx: usize, expert_idx: usize, proj_name: &str, row_idx
src/weights/mod.rs:1870
↓ 2 callersFunctioncontinue_current_work_for_linear_idx
src/cuda/flash_attn/cutlass/cutlass/gemm/kernel/sm90_tile_scheduler_stream_k.hpp:305
↓ 2 callersFunctionconvert_from_float
src/cuda/flash_attn/cutlass/cutlass/exmy_base.h:1004
↓ 2 callersMethodconvert_to_float
src/cuda/flash_attn/cutlass/cutlass/float8.h:1083
↓ 2 callersFunctioncooperative_gemm_no_predication
src/cuda/flash_attn/cutlass/cute/algorithm/cooperative_gemm.hpp:279
↓ 2 callersFunctioncoprofile
src/cuda/flash_attn/cutlass/cute/layout.hpp:634
↓ 2 callersFunctioncopy_tiles_and_advance
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/mma_with_reduction_multistage.h:90
↓ 2 callersFunctioncopy_tiles_and_advance
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/mma_planar_complex_multistage.h:94
↓ 2 callersFunctioncopy_tiles_and_advance
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/mma_sparse_multistage.h:97
↓ 2 callersFunctioncopy_tiles_and_advance
src/cuda/flash_attn/cutlass/cutlass/gemm/threadblock/mma_blas3_multistage.h:94
↓ 2 callersFunctioncopy_to_package
(path: Path)
scripts/build_sidecars.py:386
↓ 2 callersFunctioncpu_expert_byte_sizes
Compute per-expert byte sizes for CPU transposed format. INT4 has same sizes as Marlin (same u32 packing, different layout). INT8 has larger packed da
src/weights/mod.rs:1603
← previousnext →1,601–1,700 of 12,380, ranked by callers