MCPcopy Create free account

hub / github.com/Luce-Org/lucebox-hub / functions

Functions2,878 in github.com/Luce-Org/lucebox-hub

↓ 2 callersFunctionshould_load_gemma4_tensor
server/src/gemma4/gemma4_loader.cpp:148
↓ 2 callersFunctionshould_load_laguna_tensor
server/src/laguna/laguna_target_loader.cpp:103
↓ 2 callersFunctionshould_load_target_tensor
server/src/qwen35/gguf_target_loader.cpp:159
↓ 2 callersMethodshutdown
server/test/test_server_unit.cpp:2352
↓ 2 callersMethodsnapshot_cur_pos
server/test/test_server_unit.cpp:2348
↓ 2 callersMethodsnapshot_free
server/src/common/target_shard_ipc.cpp:338
↓ 2 callersMethodsnapshot_restore
server/test/test_server_unit.cpp:2154
↓ 2 callersMethodsnapshot_restore
server/src/common/target_shard_ipc.cpp:352
↓ 2 callersMethodsnapshot_save
server/src/common/target_shard_ipc.cpp:323
↓ 2 callersFunctionsnapshot_target_cache_thin
server/src/qwen35/qwen35_target_graph.cpp:1668
↓ 2 callersFunctionsock_set_nonblock
server/src/server/http_server.cpp:55
↓ 2 callersFunctionspark_budget_split
server/src/common/moe_hybrid_storage.cpp:667
↓ 2 callersFunctionsse_try_send
Send data to an SSE client fd with a short (1s) timeout to avoid stalling the inference worker. Returns false if the send fails or times out.
server/src/server/http_server.cpp:947
↓ 2 callersFunctionst_dtype_to_ggml
Map safetensors dtype string to ggml type
server/src/draft/draft_safetensors_loader.cpp:168
↓ 2 callersMethodstart
(self)
server/scripts/phase_split_dual_gpu.py:206
↓ 2 callersFunctionstart_server
( profile: ServerProfile, *, target: Path, draft: Path, bin_path: Path, prefill_drafte
harness/client_test_runner.py:1034
↓ 2 callersMethodstep
Decode one token. Returns next token id.
optimizations/megakernel/model_nvfp4.py:704
↓ 2 callersMethodstep_many
Decode multiple NVFP4 steps without per-token host/device synchronization.
optimizations/megakernel/model_nvfp4.py:742
↓ 2 callersFunctionstop_proc
(proc: subprocess.Popen)
harness/client_test_runner.py:1023
↓ 2 callersMethodstream_expert_sync
server/src/common/moe_hybrid_stream.cpp:157
↓ 2 callersFunctionstrip_leading_think_closers
Strip leading </think> tags (with optional whitespace) from the start.
server/src/server/reasoning.cpp:13
↓ 2 callersFunctionsummarize_backend
(pair_dir: Path, backend: str)
harness/clients/summarize_backend_pair.py:218
↓ 2 callersFunctionsummarize_devices
(monitor: GpuMonitor, devices: list[int], *, target_phase: bool | None = None)
server/scripts/phase_split_dual_gpu.py:326
↓ 2 callersMethodsupports_fast_rollback
server/src/qwen35/qwen35_dflash_target.cpp:505
↓ 2 callersMethodsupports_kvflash
server/test/test_server_unit.cpp:2121
↓ 2 callersMethodsupports_mixed_backend_layer_split
server/test/test_server_unit.cpp:2122
↓ 2 callersFunctiontest_hc_col_normalize
server/tests/test_deepseek4_unit.cpp:63
↓ 2 callersFunctiontest_hc_row_normalize
server/tests/test_deepseek4_unit.cpp:58
↓ 2 callersFunctiontest_sparse_alpha
Sanity-check sparse attention at alpha < 1.0: The output should not be all zeros (basic liveness check). With alpha < 1.0 outputs will differ from den
server/test/test_flash_attn_sparse.cpp:124
↓ 2 callersMethodto_cli_args
(self)
server/scripts/placement/test_dflash_args.py:24
↓ 2 callersMethodto_sse_event
server/src/server/server_status.h:208
↓ 2 callersFunctiontrain_eval
(model, name, Xtr, Ytr, Str, Xte, Yte, n_expert, epochs, lr, seed)
optimizations/spark/spark/train_pregate.py:95
↓ 2 callersFunctiontrim
server/src/server/reasoning.cpp:36
↓ 2 callersMethodupdate_completion_tokens
server/src/server/server_status.h:99
↓ 2 callersMethodvalid
server/src/common/moe_expert_compute_ipc.cpp:152
↓ 2 callersFunctionvalidate_kv_pair_or_abort
server/src/kv_quant.cpp:160
↓ 2 callersFunctionvalue_matches_type
server/src/server/tool_parser.cpp:417
↓ 2 callersFunctionvalue_matches_type_spec
server/src/server/tool_parser.cpp:428
↓ 2 callersMethodverify_batch
server/src/laguna/laguna_dflash_target.cpp:51
↓ 2 callersMethodverify_batch
server/src/gemma4/gemma4_dflash_target.cpp:36
↓ 2 callersFunctionverify_target_derived_scalars
server/src/qwen35/gguf_target_loader.cpp:255
↓ 2 callersFunctionwait_http
(base_url: str, proc: subprocess.Popen | None = None, timeout: int = 240)
harness/client_test_runner.py:979
↓ 2 callersFunctionwrite_counted_i32
(path: Path, ids: list[int])
server/scripts/laguna_pflash_niah.py:48
↓ 2 callersFunctionwrite_counted_i32
(path: Path, ids: Iterable[int])
server/scripts/phase_split_dual_gpu.py:45
↓ 2 callersMethodwrite_shared_payload_segments
server/src/common/backend_ipc.cpp:295
↓ 2 callersMethod~GgufMmap
server/src/common/gguf_mmap.h:91
↓ 1 callersFunction__bfloat162float
hipcc path: use the type's own constructors / conversions
server/hip_compat/cuda_bf16.h:30
↓ 1 callersFunction__float2bfloat16
server/hip_compat/cuda_bf16.h:33
↓ 1 callersFunction_agent_user_message
Synthesise a Codex/Claude-Code style user turn. Structure mimics what an agent client actually sends after a few tool calls have run (read_fi
server/scripts/bench_agent.py:227
↓ 1 callersFunction_auto_max_ctx
(n_prompt, n_gen: int = N_GEN)
server/scripts/bench_llm.py:133
↓ 1 callersFunction_auto_max_ctx
(n_prompt: int, n_gen: int)
server/scripts/bench_agent.py:122
↓ 1 callersFunction_build_agent_user_msg
Build a Codex-style user message padded to ~target_chars.
server/scripts/bench_server.py:268
↓ 1 callersMethod_build_prefill_graph
(self, prompt_len: int)
optimizations/megakernel/model_nvfp4.py:676
↓ 1 callersFunction_detect_arch
()
optimizations/megakernel/setup.py:7
↓ 1 callersFunction_expand_run_to_word_boundaries
Extend [r0,r1] outward until both ends sit on a whitespace boundary. A token's text "starts on whitespace" if its decoded form begins with a
server/scripts/laguna_pflash_niah.py:141
↓ 1 callersFunction_extract_numeric_answer
Extract a numeric answer from model output for GSM-style problems.
harness/benchmarks/generation_benchmark.py:59
↓ 1 callersMethod_forward_raw
Forward request verbatim (no injection needed).
harness/clients/session_inject_proxy.py:40
↓ 1 callersFunction_group_runs
Group consecutive integers into [start,end] inclusive runs.
server/scripts/laguna_pflash_niah.py:125
↓ 1 callersFunction_load_bench_prompts
Load prompts for a bench suite from JSONL file.
harness/client_test_runner.py:1515
↓ 1 callersFunction_load_gsm8k
(n_sample: int)
server/scripts/bench_server.py:147
↓ 1 callersFunction_load_math500
(n_sample: int)
server/scripts/bench_server.py:175
↓ 1 callersFunction_load_swe_rows
()
server/scripts/bench_agent.py:221
↓ 1 callersFunction_looks_like_header
(line: str)
server/scripts/profile.py:197
↓ 1 callersFunction_merge_overlapping
(runs: list[tuple[int, int]])
server/scripts/laguna_pflash_niah.py:178
↓ 1 callersFunction_overlap_1sigma
(base_mean: float, base_std: float, now_mean: float, now_std: float)
server/scripts/profile.py:354
↓ 1 callersFunction_pack_layer_weights
Pack layer weights into device blob matching LayerWeights struct.
optimizations/megakernel/model_nvfp4.py:436
↓ 1 callersFunction_pack_layer_weights
Pack layer weights into device blob matching LayerWeights struct.
optimizations/megakernel/model.py:160
↓ 1 callersFunction_pack_layer_weights_nvfp4
Pack layer weights into device blob matching LayerWeightsNVFP4 struct.
optimizations/megakernel/model_nvfp4.py:456
↓ 1 callersFunction_pack_prefill_fused_layer_weights
(fused_layer_data)
optimizations/megakernel/model_nvfp4.py:476
↓ 1 callersFunction_parse_ar
Parse test_generate stdout → {decode_tps, n_gen}.
server/scripts/bench_agent.py:176
↓ 1 callersFunction_parse_dflash
Parse test_dflash stdout into {prefill_s, decode_tps, al, stages{}}.
server/scripts/bench_agent.py:134
↓ 1 callersMethod_prefill_graph_state
(self, prompt_len: int)
optimizations/megakernel/model_nvfp4.py:698
↓ 1 callersFunction_print_header
()
server/scripts/bench_server.py:351
↓ 1 callersFunction_print_summary
(label: str, results: list[dict])
server/scripts/bench_server.py:369
↓ 1 callersFunction_quantize_matrix_nvfp4_lm
(weight)
optimizations/megakernel/model_nvfp4.py:143
↓ 1 callersFunction_query_nvidia_vram_mib
()
optimizations/pflash/pflash/dflash_client.py:26
↓ 1 callersMethod_read_body
(self)
harness/clients/session_inject_proxy.py:85
↓ 1 callersFunction_read_int32_stream
(r: int, n: int)
server/scripts/parity_laguna.py:35
↓ 1 callersMethod_read_server_log
Read server stderr log if available.
server/tests/test_server_comprehensive.py:64
↓ 1 callersMethod_read_vram_used_mib
GPU memory allocated since daemon spawn, in MiB. Priority order: 1. hipMemGetInfo delta — works on AMD unified-memory APUs (Strix Hal
optimizations/pflash/pflash/dflash_client.py:155
↓ 1 callersFunction_recover_kept_indices
Greedy subsequence match: returns positions in full_ids that align with kept_ids in order.
server/scripts/laguna_pflash_niah.py:112
↓ 1 callersFunction_resolve_draft
()
server/scripts/bench_llm.py:79
↓ 1 callersFunction_resolve_draft
()
server/scripts/bench_agent.py:82
↓ 1 callersFunction_resolve_draft
()
server/scripts/bench_he.py:196
↓ 1 callersFunction_resolve_prefill_graph
()
optimizations/megakernel/model_nvfp4.py:116
↓ 1 callersFunction_resolve_prefill_mode
()
optimizations/megakernel/model_nvfp4.py:102
↓ 1 callersFunction_run_bench_case
Run a single bench case via streaming to capture TTFT and detailed metrics. Always uses streaming to get: walltime, TTFT, prompt_tokens, completi
harness/client_test_runner.py:1530
↓ 1 callersFunction_run_bench_suite
Run all prompts for a given bench suite.
harness/client_test_runner.py:1634
↓ 1 callersFunction_run_dflash_argmax
Returns the argmax token id at the last prompt position via the daemon.
server/scripts/parity_laguna.py:64
↓ 1 callersFunction_run_hf_argmax
(model_id: str, prompt_ids: list[int])
server/scripts/parity_laguna.py:104
↓ 1 callersFunction_score_gsm_response
Score a GSM8K response. Returns (correct, detail_str).
harness/client_test_runner.py:1400
↓ 1 callersFunction_score_he_response
Score a HumanEval response by executing the generated code against test cases. Extracts code from model output, appends the test harness, and run
harness/client_test_runner.py:1456
↓ 1 callersFunction_score_math_response
Score a Math500 response. Returns (correct, detail_str).
harness/client_test_runner.py:1366
↓ 1 callersFunction_score_math_text
Score a math response text against the gold answer. Extracts \\boxed{} answers (after </think> for thinking models), with fallbacks for **bol
server/scripts/bench_server.py:188
↓ 1 callersMethod_start_stdout_reader
(self)
optimizations/pflash/pflash/dflash_client.py:132
↓ 1 callersFunction_target_version
Extract the model version token (e.g. '3.6') from a target filename.
server/scripts/run.py:34
↓ 1 callersFunction_tokenizer_slug
Filesystem-safe slug for tokenizer cache keying.
server/scripts/bench_he.py:221
↓ 1 callersMethod_wait_until_loaded
Block until the daemon emits its explicit ready banner. The memory telemetry helpers are kept for diagnostics, but readiness is drive
optimizations/pflash/pflash/dflash_client.py:186
↓ 1 callersFunction_wrap_prompt
(raw_prompt: str)
server/scripts/bench_llm.py:307
↓ 1 callersMethodabort_inline_snap
server/src/server/prefix_cache.cpp:332
← previousnext →901–1,000 of 2,878, ranked by callers