MCPcopy Create free account

hub / github.com/Luce-Org/lucebox-hub / functions

Functions2,878 in github.com/Luce-Org/lucebox-hub

↓ 1 callersFunctiongenerate
(name, cfg, tasks, samples_path)
server/scripts/quality_humaneval_plus.py:173
↓ 1 callersMethodgenerate
Send prompt + n_gen request, return generated token ids.
optimizations/pflash/pflash/dflash_client.py:257
↓ 1 callersMethodgenerate
(self, laguna_ids: list[int], n_gen: int)
server/scripts/laguna_pflash_niah.py:278
↓ 1 callersFunctiongenerate_call_id
server/src/server/tool_parser.cpp:32
↓ 1 callersFunctiongenerate_prompt
Generate a prompt that's approximately n_tokens long.
server/scripts/bench_moe_prefill_streaming.py:26
↓ 1 callersFunctionget_bool_or
server/src/laguna/laguna_target_loader.cpp:80
↓ 1 callersFunctionget_f32
server/src/qwen3/qwen3_loader.cpp:105
↓ 1 callersMethodget_feature_range
server/src/common/dflash_draft_ipc.cpp:289
↓ 1 callersFunctionget_or_parse
server/src/server/chat_template.cpp:397
↓ 1 callersFunctionget_prefill_op
(decoder)
optimizations/megakernel/final_bench_nvfp4.py:87
↓ 1 callersMethodget_routing_stats
Get the current routing stats (if tracked). Returns nullptr if the backend does not support routing stats or they are not enabled.
server/src/common/model_backend.h:348
↓ 1 callersFunctionget_u32_arr
server/src/deepseek4/deepseek4_loader.cpp:102
↓ 1 callersFunctionget_u32_arr
Read a u32 array key into a vector (empty if not found or not an array).
server/src/gemma4/gemma4_loader.cpp:117
↓ 1 callersFunctionget_u32_or
server/src/draft/draft_gguf_loader.cpp:53
↓ 1 callersFunctionget_u64_or
server/src/deepseek4/deepseek4_loader.cpp:82
↓ 1 callersFunctiongit_commit
()
server/scripts/profile.py:561
↓ 1 callersFunctiongpt2_unicode_to_byte
server/src/server/tokenizer.cpp:694
↓ 1 callersMethodgpu_embd_table
server/src/laguna/laguna_dflash_target.cpp:531
↓ 1 callersFunctiongpu_empirical_l1
L1 distance between an empirical histogram of GPU draws and the analytic distribution. Pulls the uniform from the same RNG family the CPU path uses.
server/test/test_gpu_sampler_cuda.cpp:151
↓ 1 callersFunctiongpu_info
()
server/scripts/profile.py:550
↓ 1 callersFunctiongpu_sampler_microbench
Per-call latency microbench (gated by env DFLASH_SAMPLER_BENCH=1). Isolates the three regimes that explain the end-to-end numbers: the CPU chain, the
server/test/test_gpu_sampler_cuda.cpp:292
↓ 1 callersFunctiongrade_config
(name, tasks, samples_path, scores_path)
server/scripts/quality_humaneval_plus.py:265
↓ 1 callersFunctiongrade_one
Run completion + task['test'] + check_call in a subprocess with a SIGALRM-based timeout. Returns dict with ok/err. If completion starts with
server/scripts/quality_humaneval_plus.py:237
↓ 1 callersMethodhandle_anthropic
(self, req: dict)
harness/clients/llamacpp_compat_proxy.py:456
↓ 1 callersMethodhandle_compress
server/src/common/layer_split_backend.cpp:193
↓ 1 callersMethodhandle_responses
(self, req: dict)
harness/clients/llamacpp_compat_proxy.py:431
↓ 1 callersMethodhas
server/src/common/kvflash_qk.h:108
↓ 1 callersFunctionhas_request_tools
server/src/server/sse_emitter.cpp:22
↓ 1 callersMethodhas_shared_expert
server/src/common/moe_hybrid_types.h:79
↓ 1 callersFunctionhas_single_request_tool
server/src/server/sse_emitter.cpp:26
↓ 1 callersFunctionhave_nsys
()
server/scripts/profile.py:145
↓ 1 callersFunctionhc_output_batch
server/src/deepseek4/deepseek4_graph.cpp:2132
↓ 1 callersFunctionhc_pre_auto_into
server/src/deepseek4/deepseek4_graph.cpp:2026
↓ 1 callersMethodhealthy
server/src/common/moe_expert_compute.h:52
↓ 1 callersMethodhidden_size
server/src/common/dflash_draft_ipc.h:62
↓ 1 callersMethodhot_experts
server/src/common/moe_hybrid_routing_stats.cpp:113
↓ 1 callersFunctionhttp_sse
( url: str, payload: dict[str, Any], *, headers: dict[str, str] | None = None, expect_subs
harness/client_test_runner.py:420
↓ 1 callersFunctionhybrid_spec_min_accept_rate
server/src/qwen35moe/qwen35moe_backend.cpp:55
↓ 1 callersFunctionhybrid_spec_min_steps_before_ar
server/src/qwen35moe/qwen35moe_backend.cpp:61
↓ 1 callersFunctioninclude_preceding_tool_call_open
server/src/server/tool_parser.cpp:128
↓ 1 callersMethodindex_of
server/src/common/moe_hybrid_routing_stats.cpp:12
↓ 1 callersFunctioninfer_draft_dims_from_safetensors
server/src/draft/draft_safetensors_loader.cpp:230
↓ 1 callersMethodinit
server/src/qwen3/qwen3_backend.cpp:85
↓ 1 callersMethodinit
server/src/server/disk_prefix_cache.cpp:194
↓ 1 callersMethodinit
server/src/qwen35/qwen35_backend.cpp:176
↓ 1 callersMethodinit
server/src/gemma4/gemma4_backend.cpp:33
↓ 1 callersMethodinit_full_cache
server/src/server/prefix_cache.cpp:356
↓ 1 callersFunctioninit_snapshot_test_shard
server/tests/test_deepseek4_unit.cpp:541
↓ 1 callersFunctioninvert_moe_hybrid_placement
server/src/common/moe_expert_compute_ipc.cpp:298
↓ 1 callersFunctionis_billing_header_block
Returns true if `s`, after skipping leading whitespace, starts with kBillingHeader.
server/src/server/prompt_normalize.cpp:11
↓ 1 callersFunctionis_expert_tensor_name
server/src/qwen35/gguf_target_loader.cpp:142
↓ 1 callersFunctionis_laguna_expert_tensor
server/src/laguna/laguna_target_loader.cpp:779
↓ 1 callersFunctionis_norm_tensor
(name: str)
server/scripts/quantize_gemma_dflash_q8.py:57
↓ 1 callersFunctionis_norm_tensor
(gguf_name: str)
server/scripts/quantize_draft_q8.py:80
↓ 1 callersFunctionis_strict_prefix
server/src/server/prefix_cache.cpp:145
↓ 1 callersMethodis_target_parked
server/src/common/layer_split_backend.h:94
↓ 1 callersFunctionkvflash_policy_is_lru
Residency policy from DFLASH_KVFLASH_POLICY (--kvflash-policy): "lru" forces recency-only paging (no drafter probe, no scorer); anything else (default
server/src/common/kvflash_pager.h:670
↓ 1 callersFunctionkvflash_policy_is_qk
"qk": target-QK residency scoring (kvflash_qk.h) — pooled post-RoPE keys vs the current decode query, no drafter involved.
server/src/common/kvflash_pager.h:677
↓ 1 callersMethodkvflash_sync_identity
server/src/common/target_shard_ipc.cpp:308
↓ 1 callersFunctionlaguna_argmax_row
server/src/laguna/laguna_dflash_target.cpp:21
↓ 1 callersFunctionlaguna_chunked_prefill
Chunked prefill loop on top of the shared laguna_step() helper. Reports total prefill time and the argmax / logit at the LAST chunk.
server/test/bench_laguna_pflash.cpp:34
↓ 1 callersFunctionlaguna_layer_step_graph_destroy
server/src/laguna/laguna_target_graph.cpp:947
↓ 1 callersFunctionlaguna_project_hidden
server/src/laguna/laguna_target_graph.cpp:1805
↓ 1 callersFunctionlaguna_snapshot_alloc
---- Cache snapshot helpers (prefix-cache slots) ------------------------ laguna_snapshot_alloc: build a parallel set of K/V tensors RIGHT-SIZED to s
server/src/laguna/laguna_target_graph.cpp:144
↓ 1 callersFunctionlaguna_step_hybrid
---- Single-graph hybrid decode forward -----------------------------------
server/src/laguna/laguna_target_graph.cpp:1854
↓ 1 callersFunctionlaguna_verify_batch
server/src/laguna/laguna_target_graph.cpp:1559
↓ 1 callersFunctionlayer_expert_bytes
server/src/deepseek4/deepseek4_backend.cpp:113
↓ 1 callersFunctionlcp_len
server/src/server/disk_prefix_cache.cpp:140
↓ 1 callersFunctionllama_ftype_name
Map LLAMA_FTYPE_* int → operator-friendly tag (Q4_K_M, IQ4_XS, BF16, …). Kept inline so we don't pull in llama.h here — those enum values are part of
server/src/common/gguf_inspect.cpp:194
↓ 1 callersFunctionload_arch
Resolve the draft's architecture scalars. config.json (next to the safetensors) is authoritative for the transformer hparams; the tensor shape
server/scripts/convert_dflash_to_gguf.py:68
↓ 1 callersFunctionload_cases
(path: Path)
harness/benchmarks/generation_benchmark.py:28
↓ 1 callersFunctionload_hf_tokenizer
()
server/tests/test_tokenizer.py:22
↓ 1 callersFunctionload_multiple_routing_files
Load multiple routing files and return token-shuffled concatenation. Each file is sliced into per-token blocks of ``num_layers`` records. Tok
scripts/train_predictor.py:135
↓ 1 callersFunctionload_remote_moe_runtime
server/src/common/moe_expert_compute_ipc.cpp:647
↓ 1 callersFunctionload_routing_data
Load a single binary routing file into numpy arrays.
scripts/train_predictor.py:86
↓ 1 callersFunctionload_safetensors_header
(path: Path)
server/scripts/convert_dflash_to_gguf.py:190
↓ 1 callersFunctionload_safetensors_header
(path: Path)
server/scripts/quantize_gemma_dflash_q8.py:63
↓ 1 callersFunctionload_safetensors_header
(path: Path)
server/scripts/quantize_draft_q8.py:92
↓ 1 callersFunctionload_sidecar
server/src/server/model_card.cpp:157
↓ 1 callersFunctionload_target_tokenizer
(args)
server/scripts/phase_split_dual_gpu.py:358
↓ 1 callersFunctionload_token_stream
Real-content filler: raw little-endian i32 token stream. Ids are folded below 100000 (drafter-scoreable, same fold as KvFlashDrafterScorer / make_prom
server/test/test_kvflash.cpp:384
↓ 1 callersFunctionload_weights
Load Qwen3.5-0.8B weights and optional GB10 NVFP4 decode weights.
optimizations/megakernel/model_nvfp4.py:340
↓ 1 callersFunctionload_weights
Load Qwen3.5-0.8B weights (bf16 on SM>=80, fp16 on SM<80).
optimizations/megakernel/model.py:66
↓ 1 callersFunctionlog_ds4_expert_memory_info
server/src/deepseek4/deepseek4_backend.cpp:176
↓ 1 callersFunctionlog_staged_cross_gpu_once
server/src/common/peer_access.cpp:41
↓ 1 callersFunctionlog_target_shard_timing
server/src/deepseek4/deepseek4_target_shard_ipc_daemon.cpp:88
↓ 1 callersFunctionlooks_like_path
server/src/common/daemon_loop.cpp:136
↓ 1 callersFunctionloss_listnet
Top-1 ListNet: CE between softmax(scores) and the normalized relevance.
optimizations/spark/spark/train_pregate.py:54
↓ 1 callersFunctionloss_ranknet
Pairwise RankNet logistic loss over (selected, unselected) score pairs. `sel` is the padded (N, 8) selected-expert index matrix (-1 = empty slot)
optimizations/spark/spark/train_pregate.py:61
↓ 1 callersFunctionmain
()
optimizations/pflash/tests/niah_gen.py:98
↓ 1 callersFunctionmain
()
optimizations/pflash/tests/bench_niah_cpp.py:22
↓ 1 callersFunctionmain
()
optimizations/megakernel/final_bench_nvfp4.py:307
↓ 1 callersFunctionmain
()
optimizations/megakernel/bench_pp_tg_nvfp4.py:375
↓ 1 callersFunctionmain
()
optimizations/spark/spark/validate.py:41
↓ 1 callersFunctionmain
()
optimizations/spark/spark/tokenizer.py:47
↓ 1 callersFunctionmain
()
optimizations/spark/spark/extract_sessions.py:113
↓ 1 callersFunctionmain
()
optimizations/spark/spark/calibrate.py:40
↓ 1 callersFunctionmain
()
optimizations/spark/spark/train_pregate.py:127
↓ 1 callersFunctionmain
()
optimizations/spark/spark/bench.py:69
↓ 1 callersFunctionmain
(argv: list[str] | None = None)
harness/client_test_runner.py:1891
← previousnext →1,201–1,300 of 2,878, ranked by callers