MCPcopy Create free account
hub / github.com/Luce-Org/lucebox-hub

github.com/Luce-Org/lucebox-hub @main

Chat with this repo
repository ↗ · DeepWiki ↗ · + Follow
3,298 symbols 11,159 edges 356 files ⚖ Apache-2.0 615 documented · 19% updated 2d ago★ 2,79830 open issues

Browse by type

Functions 2,878 Types & classes 420
What it actually does AI analysis from the code graph — generated when you open this
loading…
README

Lucebox

lucebox.com HuggingFace Discord Blog Tutorials

Apache 2.0 CUDA 12+ HIP 7+ C++17

Speculative inference for heterogeneous machines and consumer GPUs.

Custom kernels, speculative prefill and decoding, tuned for each model and hardware target.


Inference Engine Optimizations

Optimization Measured setup Result
DFlash2 Qwen 3.8 27B on one R9700 208.1 tok/s average, 227.8 tok/s peak
DSpark DeepSeek V4 on Strix Halo, native top-6 32.7 tok/s high-acceptance median; 27.9 tok/s mixed-eval average
PFlash + KVFlash Laguna XS 2.1 33B at 256K on RTX 3090 6.1× prefill, 411 s to 67.3 s
Luce Spark Laguna XS.2 33B on RTX 3090 ~100 tok/s in 14.6 GiB
KVFlash Laguna XS 2.1 33B at 256K on RTX 3090 152.3 tok/s with an 8K pool
Heterogeneous execution DeepSeek V4 on R9700 + Strix Halo 86 tok/s decode; 788 tok/s prefill at 2K
Paged attention Qwen 3.6 27B concurrent serving 1.35× attention step; 82% less KV memory
Megakernel Qwen 3.5 0.8B on RTX 3090 413 tok/s, 1.87 tok/J

Supported Models and Drafters

Model links open the exact weights used by the measured setup. Drafter links open the published quant, or the source checkpoint when conversion is required.

Model and optimization Phase Speedup
Qwen 3.5 0.8B BF16 + Megakernel Prefill + decode 1.9× prefill; 1.55× decode
Qwen 3.8 27B UD-IQ4_XS + DFlash2 source, converted to Q8_0, on R9700 Decode 6.4× vs Lucebox AR; 3.8× vs llama.cpp with the same drafter
Laguna XS 2.1 33B Q4_K_M + PFlash/KVFlash with Qwen3 0.6B Q8_0 Prefill 6.1×, 411 s to 67.3 s at 256K
Laguna XS 2.1 33B Q4_K_M + DFlash Q4 drafter Decode 1.7× at 256K
Gemma 4 26B-A4B Q4_K_M + DFlash Q8_0 drafter Decode 1.31×
Gemma 4 31B IT Q4_K_M + DFlash Q8_0 drafter Decode 3.2×
DeepSeek V4 Flash ROCmFPX MIX Strix + DSpark Q4RMFP4 drafter Decode Up to 1.81× vs target-only, 32.7 vs 18.1 tok/s
Ling 3.0 Flash 124B-A5.1B Q4_K_M Decode 34.6 tok/s median AR on DGX Spark

Tested Machines (GPU/APU)

The engine is not tied to one reference card. NVIDIA architectures are selected by CMake; HIP builds should target the device's exact gfx architecture.

Architecture Hardware Runtime Details
RDNA4 gfx1201 Radeon AI PRO R9700 ROCm 7.2 Qwen 3.8 R9700 quick start
RDNA3.5 gfx1151 Ryzen AI MAX+ 395 / Strix Halo ROCm 7.2 DeepSeek V4 Strix profile
RDNA3 gfx1100 Radeon RX 7900 XT / XTX ROCm 6+ DeepSeek V4 dual AMD profile
Ampere sm_86 RTX 3090 CUDA 12+ Qwen 3.8 NVLink result and Megakernel results
Blackwell sm_120 RTX 5090 CUDA 12.8+ Qwen 3.8 single-GPU result
Blackwell sm_121 DGX Spark / GB10 CUDA 12.9 Qwen 3.5 NVFP4 results
Ada sm_89 RTX 4090 CUDA 12+ Linux and WSL2 community runs
Turing sm_75 RTX 2080 Ti CUDA 12.0 DFlash results
Volta sm_70, Pascal sm_61 V100, P40 CUDA 12.0 CUDA quick start
Not pictured Blackwell sm_110 Jetson AGX Thor CUDA 13.0 Thor quick start

Single-device results

Hardware Model Measured result
R9700 Qwen 3.8 27B UD-IQ4_XS + DFlash2 source 208.1 tok/s HumanEval average; 227.8 tok/s best request
Strix Halo DeepSeek V4 ROCmFPX MIX Strix + DSpark Q4RMFP4 32.7 tok/s high-acceptance median; 27.9 tok/s across the fixed 30-prompt evaluation, using all six routed experts
RTX 5090 Qwen 3.8 27B 110.6 tok/s for a 26,758-token prompt and 1,024-token continuation (PR #637)

Heterogeneous and parallel results

Hardware Configuration Measured result
2x RTX 3090 + NVLink Qwen 3.8 target tensor parallel + DFlash2 79.7 tok/s, 2.16× autoregressive decode (PR #637)
RX 7900 XT + Strix Halo DeepSeek V4 with all six experts + DSpark verification width 4 45.0 to 47.7 tok/s decode; 111.2 tok/s prefill at 132,981 tokens (PR #604)
R9700 + Strix Halo DeepSeek V4 across both AMD devices 86 tok/s decode; 788 tok/s prefill at 2K

These runs use different prompts, quantizations, and inference policies. They show which configurations work; they are not a cross-hardware ranking.

Recommended Setups

See Recommended server setups for the model and hardware matrix, including single-GPU and mixed-GPU profiles.

Client Harnesses

harness/ runs Lucebox through popular coding clients and checks server compatibility.

Lucebox client harness experiments on RTX 3090 | Client | Launcher | |--------|----------| | Claude Code | [`run_claude_code.sh`](harness/clients/run_claude_code.sh) | | Codex | [`run_codex.sh`](harness/clients/run_codex.sh) | | OpenCode | [`run_opencode.sh`](harness/clients/run_opencode.sh) | | Hermes | [`run_hermes.sh`](harness/clients/run_hermes.sh) | | Pi | [`run_pi.sh`](harness/clients/run_pi.sh) | | OpenClaw | [`run_openclaw.sh`](harness/clients/run_openclaw.sh) | | Open WebUI | [`run_openwebui.sh`](harness/clients/run_openwebui.sh) |

Set the server binary and model paths, then run a launcher:

DFLASH_SERVER_BIN=server/build/dflash_server \
DFLASH_TARGET=server/models/Qwen3.8-27B-UD-IQ4_XS.gguf \
DFLASH_DRAFT=server/models/draft/qwen38-dflash2-q8_0.gguf \
MAX_CTX=32768 \
harness/clients/run_codex.sh

See the harness guide for setup, no-draft targets, and benchmarks.

Quick Start With Docker

Prebuilt images on GHCR track main. Mount the weights and serve the OpenAI-compatible API on :8000.

| GPU | Image tag | |-----|-----------| | NVIDIA (CUDA 12+) | `:cuda12` | | AMD (ROCm 6+) | `:rocm` | Put the target in `server/models/` and its matching drafter in `server/models/draft/`. Lucebox prebuilt Docker images for NVIDIA and AMD

Run the image for your GPU:

# NVIDIA
docker run --rm --gpus all -p 8000:8080 \
  -v "$PWD/server/models:/opt/lucebox-hub/server/models" \
  ghcr.io/luce-org/lucebox-hub:cuda12

# AMD
docker run --rm --device /dev/kfd --device /dev/dri \
  --group-add video --group-add render --security-opt seccomp=unconfined \
  -p 8000:8080 -v "$PWD/server/models:/opt/lucebox-hub/server/models" \
  ghcr.io/luce-org/lucebox-hub:rocm

Run the Server

This quick start runs the R9700 profile above. The complete flag reference is in the server guide.

# build (ROCm 7.2+, RDNA4)
git clone --recurse-submodules https://github.com/Luce-Org/lucebox.git
cd lucebox
cmake -S server -B server/build-hip -G Ninja \
  -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_HIP_COMPILER=/opt/rocm/lib/llvm/bin/clang++ \
  -DDFLASH27B_GPU_BACKEND=hip \
  -DDFLASH27B_HIP_ARCHITECTURES=gfx1201 \
  -DGGML_HIP_MMQ_MFMA=ON \
  -DGGML_HIP_NO_VMM=ON
cmake --build server/build-hip --target dflash_server -j"$(nproc)"

# target and DFlash2 drafter
mkdir -p models
huggingface-cli download unsloth/Qwen3.8-27B-GGUF \
  Qwen3.8-27B-UD-IQ4_XS.gguf --local-dir models
huggingface-cli download incoai/Qwen3.8-27B-DFlash2 --local-dir models/dflash2
python server/scripts/convert_dflash_to_gguf.py \
  models/dflash2/model.safetensors models/qwen38-dflash2-f16.gguf
python server/scripts/quantize_dflash_draft.py \
  models/qwen38-dflash2-f16.gguf models/qwen38-dflash2-q8_0.gguf --scheme q8_0

# launch the measured profile
./server/build-hip/dflash_server models/Qwen3.8-27B-UD-IQ4_XS.gguf \
  --draft models/qwen38-dflash2-q8_0.gguf \
  --draft-block-size 16 --max-ctx 131072 \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --port 8216

curl -s http://127.0.0.1:8216/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"messages":[{"role":"user","content":"Write a Python LRU cache."}],
       "max_tokens":256,"temperature":0}'

Documentation

Topic Guide
Recommended model and hardware profiles Recommended setups
Runtime parameters Server parameter reference
OpenAI Chat Completions, Responses, and Anthropic Messages API reference
CUDA, HIP, and mixed-device placement [Mixed-backend guide

Core symbols most depended-on inside this repo

browse all functions →

Shape

Function 1,774
Method 1,104
Class 400
Enum 20

Languages

C++78%
Python20%
C2%

Modules by API surface

server/test/test_server_unit.cpp286 symbols
server/src/deepseek4/deepseek4_graph.cpp92 symbols
server/src/server/rax.c69 symbols
server/scripts/test_server_integration.py63 symbols
harness/client_test_runner.py58 symbols
server/src/server/http_server.cpp51 symbols
server/src/common/moe_expert_compute_ipc.cpp48 symbols
server/scripts/phase_split_dual_gpu.py45 symbols
server/tests/test_deepseek4_unit.cpp44 symbols
server/src/qwen35/qwen35_backend.cpp44 symbols
server/src/laguna/laguna_backend.cpp44 symbols
server/src/common/kvflash_pager.h41 symbols

For agents

$ claude mcp add lucebox-hub \
  -- python -m otcore.mcp_server <graph>

⬇ download graph artifact

Ask about this repo answers extend the page