MCPcopy Create free account

hub / github.com/Luce-Org/lucebox-hub / functions

Functions2,878 in github.com/Luce-Org/lucebox-hub

↓ 2 callersMethodinit
server/src/deepseek4/deepseek4_backend.cpp:312
↓ 2 callersMethodinit
server/src/laguna/laguna_backend.cpp:135
↓ 2 callersFunctionis_expert_tensor
server/src/deepseek4/deepseek4_loader.cpp:162
↓ 2 callersFunctionis_ws
server/src/server/tool_parser.cpp:385
↓ 2 callersFunctionlaguna_auto_head_major_enabled
server/src/laguna/laguna_backend.cpp:82
↓ 2 callersFunctionlaguna_dspark_confidence_threshold
server/src/laguna/laguna_backend.cpp:116
↓ 2 callersFunctionlaguna_dspark_enabled
server/src/laguna/laguna_backend.cpp:108
↓ 2 callersFunctionlaguna_gpu_argmax_enabled
server/src/laguna/laguna_backend.cpp:90
↓ 2 callersFunctionlaguna_host_stage
---- Public turnkey forward step ---------------------------------------- Pinned host staging for per-step graph inputs. ggml_backend_tensor_set from
server/src/laguna/laguna_target_graph.cpp:1212
↓ 2 callersFunctionlaguna_layer_step_graph_free
server/src/laguna/laguna_target_graph.cpp:935
↓ 2 callersFunctionlist_str
(fl)
optimizations/spark/spark/tokenizer.py:21
↓ 2 callersFunctionload_gemma4_gguf
server/src/gemma4/gemma4_loader.cpp:177
↓ 2 callersFunctionload_hash_routing_cpu
server/src/deepseek4/deepseek4_graph.cpp:2178
↓ 2 callersFunctionload_report
(path: Path)
harness/benchmarks/generation_benchmark.py:317
↓ 2 callersFunctionload_tokenizer
(args)
server/scripts/phase_split_dual_gpu.py:348
↓ 2 callersFunctionlog_split_tel
server/src/deepseek4/deepseek4_layer_split_adapter.cpp:68
↓ 2 callersFunctionlog_step_tel
server/src/deepseek4/deepseek4_backend.cpp:80
↓ 2 callersMethodlookup_full
server/src/server/prefix_cache.cpp:380
↓ 2 callersFunctionmake_ds4_moe_hybrid_config
server/src/deepseek4/deepseek4_graph.cpp:1500
↓ 2 callersFunctionmake_ds4_moe_hybrid_config
server/src/deepseek4/deepseek4_loader.cpp:679
↓ 2 callersFunctionmake_ds4_moe_layer_desc
server/src/deepseek4/deepseek4_graph.cpp:1513
↓ 2 callersFunctionmake_ds4_moe_layer_desc
server/src/deepseek4/deepseek4_loader.cpp:692
↓ 2 callersFunctionmake_ds4_parent_worker_cfg
server/src/deepseek4/deepseek4_backend.cpp:281
↓ 2 callersFunctionmake_temp_gguf_path
server/tests/test_deepseek4_unit.cpp:84
↓ 2 callersFunctionmmq_safe_sub_batch
Sub-batch size for the reduced-hot-stack routed mul_mat_id. The MMQ path (n_tokens > 8) illegal-accesses on a REDUCED expert stack for sparse/ imbalan
server/src/common/moe_hybrid_ffn_eval.cpp:1171
↓ 2 callersFunctionmoe_expert_compute_profile_enabled
server/src/common/moe_expert_compute_ipc.cpp:58
↓ 2 callersFunctionmoe_expert_convert_input_to_ipc_type
server/src/common/moe_expert_compute_ipc.cpp:424
↓ 2 callersFunctionmoe_expert_ipc_input_row_size
server/src/common/moe_expert_compute_ipc.cpp:417
↓ 2 callersFunctionmoe_hybrid_core_bytes_from_memory
server/src/common/moe_hybrid_placement.h:17
↓ 2 callersMethodn_target_layers
server/src/common/dflash_draft_ipc.h:64
↓ 2 callersFunctionnormalize_anthropic_system
Normalize Anthropic's `system` field (top-level on /v1/messages and /v1/messages/count_tokens) into a leading `{role:"system", content:...}` entry on
server/src/server/http_server.cpp:691
↓ 2 callersFunctionnormalize_anthropic_user_text
(text: str)
harness/clients/llamacpp_compat_proxy.py:60
↓ 2 callersFunctionnormalize_text
(text: str)
harness/benchmarks/generation_benchmark.py:50
↓ 2 callersFunctionnow_ms
()
harness/client_test_runner.py:204
↓ 2 callersFunctionnpm_prefix
(work_dir: Path, client: str)
harness/client_test_runner.py:265
↓ 2 callersFunctionnsys_version
()
server/scripts/profile.py:189
↓ 2 callersFunctionnucleus_cutoff
server/src/common/sampler.cpp:32
↓ 2 callersFunctionopen_dflash_floor_log
server/src/qwen35/qwen35_backend.cpp:95
↓ 2 callersMethodpage_in
Recall a chunk into the pool (used by reselect / tests).
server/src/common/kvflash_pager.h:295
↓ 2 callersMethodpage_out
Force a chunk out of the pool (host backing + zeroed slots).
server/src/common/kvflash_pager.h:267
↓ 2 callersMethodpark
server/src/qwen3/qwen3_backend.cpp:125
↓ 2 callersFunctionparse_inline_snap
Parse optional inline-snap suffix: ` snap=<pos>:<slot>`.
server/src/common/daemon_loop.cpp:186
↓ 2 callersFunctionparse_json_tool_call
Parse {"name": ..., "arguments": ...} or {"function": {"name": ..., "arguments": ...}}
server/src/server/tool_parser.cpp:343
↓ 2 callersFunctionparse_moe_expert_compute_ipc_mode
server/src/common/moe_hybrid_ffn_eval.cpp:63
↓ 2 callersFunctionparse_size_arg
server/src/ipc/backend_ipc_main.cpp:64
↓ 2 callersFunctionparse_xml_params
server/src/server/tool_parser.cpp:315
↓ 2 callersFunctionpflash_keep_ratio
server/src/server/http_server.cpp:113
↓ 2 callersFunctionpip_venv
(work_dir: Path, client: str)
harness/client_test_runner.py:269
↓ 2 callersFunctionplacement_backend_supported
server/src/placement/placement_backend.h:49
↓ 2 callersMethodprefill_tokens
Run prompt prefill and return the first generated token id.
optimizations/megakernel/model_nvfp4.py:771
↓ 2 callersMethodproject_hidden_to_logits
Project draft hidden states through the target lm_head and return full f32 logits with vocab as the fastest-changing dimension. Default false (unsuppo
server/src/common/dflash_target.h:137
↓ 2 callersMethodproject_hidden_to_tokens
server/src/qwen35/qwen35_dflash_target.cpp:649
↓ 2 callersFunctionq4km_bytes
Q4_K_M bytes per row: approximate as 0.5625 bytes/element (4.5 bits)
server/test/bench_moe_stream.cpp:41
↓ 2 callersFunctionqwen35_empty_visible_output
server/src/qwen35/qwen35_backend.cpp:141
↓ 2 callersFunctionraxAddChildNoAlloc
Like raxAddChild() but does not allocate the child node. Instead it * returns in 'parentlink' the address of the new child pointer, so that the * ca
server/src/server/rax.c:307
↓ 2 callersFunctionraxCompressNodeNoAlloc
Turn the node 'n', that must be a node without any children, into a * compressed node representing a set of nodes linked one after the other * and h
server/src/server/rax.c:488
↓ 2 callersFunctionraxDefragAddChars
Append characters at the current key string of the defragmentation * iterator. Like the normal iterator, the key is rebuilt incrementally as * the w
server/src/server/rax.c:2371
↓ 2 callersFunctionraxDefragNodeFlags
Build the flags associated with the current item returned by the * defragmentation iterator. */
server/src/server/rax.c:2400
↓ 2 callersFunctionraxDefragStackPush
Push a new node into the defragmentation iterator stack. The parent_child * argument is the child index of this node in its parent, or -1 for the *
server/src/server/rax.c:2319
↓ 2 callersFunctionraxFindParentLink
Return the memory address where the 'parent' node stores the specified * 'child' pointer, so that the caller can update the pointer with another * o
server/src/server/rax.c:1124
↓ 2 callersFunctionraxGenericInsert
Insert the element 's' of size 'len', setting as auxiliary data * the pointer 'data'. If the element is already present, the associated * data is up
server/src/server/rax.c:629
↓ 2 callersFunctionraxIteratorClearInlineLeaf
server/src/server/rax.c:1561
↓ 2 callersFunctionraxIteratorSetInlineLeaf
server/src/server/rax.c:1565
↓ 2 callersFunctionraxNewValueNode
Create a standalone leaf node storing the specified value. */
server/src/server/rax.c:198
↓ 2 callersFunctionraxNodeFindChildPos
Return the position where the edge 'c' should be inserted in order to * preserve lexicographic ordering. */
server/src/server/rax.c:294
↓ 2 callersFunctionraxRemoveChildAtPtr
Low level child removal from node. 'childptr' must point to the child * pointer stored inside the parent node, and is used directly instead of * sea
server/src/server/rax.c:1141
↓ 2 callersFunctionraxRemoveCleanup
Free the useless node 'h' that was left after a deletion, and keep moving * upward while the parent would also become a non-key single-child node. *
server/src/server/rax.c:1226
↓ 2 callersFunctionraxStackInit
Initialize the stack. */
server/src/server/rax.c:94
↓ 2 callersFunctionraxStackPeek
Return the stack item at the top of the stack without actually consuming * it. */
server/src/server/rax.c:139
↓ 2 callersFunctionread_binary_file_exact
server/src/common/io_utils.h:64
↓ 2 callersMethodread_shared_payload
server/src/common/backend_ipc.cpp:324
↓ 2 callersFunctionread_target_layer_scales
server/src/qwen35/gguf_target_loader.cpp:191
↓ 2 callersFunctionread_uncounted_i32
Read a prompt file: raw int32 stream (file size implies token count).
server/src/common/daemon_loop.cpp:143
↓ 2 callersMethodread_verify_logits
server/src/laguna/laguna_dflash_target.cpp:94
↓ 2 callersMethodremember
server/src/server/tool_memory.cpp:14
↓ 2 callersFunctionrender_md
(doc: dict)
server/scripts/profile.py:404
↓ 2 callersMethodreset_request_state
server/src/common/target_shard_ipc.cpp:294
↓ 2 callersFunctionresolve_laguna_kv_types
Laguna honors only the explicit per-axis --cache-type-k/v overrides. The DFLASH27B_KV_F16/_Q4/_TQ3 shorthands are qwen-family toggles - the server aut
server/src/laguna/laguna_backend.cpp:54
↓ 2 callersMethodrollback_to_tree
server/src/qwen35/qwen35_dflash_target.cpp:339
↓ 2 callersFunctionrun
POST to /v1/chat/completions with stream=true. Return (n_tok, wall_secs, decode_secs) where decode_secs starts at the first streamed token (after
server/scripts/bench_daemon.py:32
↓ 2 callersFunctionrun_cases
(args, cases: list[tuple[str, str, str | None, str | None]])
server/scripts/phase_split_dual_gpu.py:498
↓ 2 callersFunctionrun_config
Spin up server with --prefix-cache-slots=slots, replay turns, return latencies.
server/scripts/bench_agent_loop.py:82
↓ 2 callersFunctionrun_dflash_draft_ipc_daemon
server/src/common/dflash_draft_ipc_daemon.cpp:47
↓ 2 callersFunctionrun_eval
()
scripts/train_predictor.py:565
↓ 2 callersMethodrun_forward
server/src/laguna/laguna_layer_split_adapter.cpp:429
↓ 2 callersFunctionrun_laguna_daemon
server/src/laguna/laguna_daemon.cpp:21
↓ 2 callersFunctionrun_qwen35_layer_split_layers_from_activation
server/src/qwen35/layer_split_forward.cpp:353
↓ 2 callersFunctionscan_dir
server/src/common/spark_corpus.cpp:112
↓ 2 callersFunctionscore_tokens_resilient
Score `ids` with allocation-failure resilience: try the full forward; on failure split into two equal halves, score each with the TRUE query tail (las
server/src/qwen3/qwen3_kvflash_scorer.cpp:64
↓ 2 callersMethodscratch_down_data
server/src/common/moe_hybrid_stream.h:66
↓ 2 callersMethodscratch_gate_data
Get tensor pointers into the GPU scratch buffer after a successful stream. Valid until next stream call. Tensors are transient (not owned by any conte
server/src/common/moe_hybrid_stream.h:64
↓ 2 callersFunctionseed_token_buffer
optimizations/megakernel/torch_bindings.cpp:133
↓ 2 callersMethodsend
(self, cmd)
optimizations/spark/spark/_daemon.py:54
↓ 2 callersMethodsend_sse
(self, data: bytes)
harness/clients/llamacpp_compat_proxy.py:391
↓ 2 callersFunctionseq_at
server/src/server/prefix_cache.cpp:61
↓ 2 callersMethodset_keep_verify_logits
server/src/laguna/laguna_dflash_target.h:35
↓ 2 callersMethodset_kvflash_pager
kvflash: route verify writes through the pool (slots allocated here, slot-space mask inside gemma4_verify_batch). Non-owning.
server/src/gemma4/gemma4_dflash_target.h:38
↓ 2 callersMethodset_query
server/src/common/kvflash_qk.h:176
↓ 2 callersMethodset_routing_collector
── Routing data collection ────────────────────────────────────── Set an external routing collector that the backend will call for each token/layer du
server/src/common/model_backend.h:344
↓ 2 callersFunctionshould_keep_ds4_tensor
server/src/deepseek4/deepseek4_loader.cpp:168
← previousnext →801–900 of 2,878, ranked by callers