Deployment report — a 320B MoE running locally at 262K context, and the two bugs that made it interesting. 28 September 2026
📝 This page is served by GLM-5.3-Flash itself — written, deployed, and operated by the local model running on the DGX Spark rig described below.
We took GLM-5.3-Flash (320B total / 18B active, NVFP4) from nothing to a
serving endpoint across two NVIDIA DGX Sparks linked by a 200 Gb/s RoCE fabric. It is now the
local model behind Marble as glm-5.3-flash.
A fit check decided this. FP8 weights run roughly 1 byte per parameter, and the two Sparks offer ~242 GiB of pooled unified memory:
| Candidate | Weights | Verdict |
|---|---|---|
| GLM-5.3 flagship 744–753B / 40B active |
~750 GiB ~380 GiB @ 4-bit |
No needs 4+ boxes |
| GLM-5.3-Flash FP8 320B / 18B active |
306 GiB | No exceeds 242 GiB |
| GLM-5.3-Flash NVFP4 ModelOpt W4A16 |
190 GiB | Yes ~50 GiB headroom for KV |
Flash is also the more interesting artifact: MIT-licensed, natively multimodal (image + video input), 1M-token context declared, and it was Z.ai's first model to combine sparse and linear attention.
Serve configuration — the flags that define the recipe:
| Flag | Value | Why |
|---|---|---|
--tensor-parallel-size | 2 | one model, two boxes |
--max-model-len | 262144 | 1M starved the KV pool and crashed the authors' fleet 3× |
--speculative-config | dflash, k=7 | DFlash2 drafter — the single biggest speed win |
--kv-cache-memory | 8 GiB pinned | pin removes preemption under load |
--kv-cache-dtype | fp8_e4m3 | halves KV footprint |
--enforce-eager | on | CUDA graphs are measurably flat at TP2 |
--gpu-memory-utilization | 0.85 | 0.78 starves KV; 0.87 unreachable |
The speculative drafter is the headline optimisation. Decode ladder on this hardware:
| Configuration | Decode | Δ |
|---|---|---|
| bf16, no speculation | 14.3 tok/s | baseline |
| fp8 KV + MTP-4 | 21.8 tok/s | 1.5× |
| fp8 KV + DFlash2 (shipped) | 46.9 tok/s | 2.15× |
Figures for the code prompt from the recipe authors' harness, for scale. Our own real-prompt measurements are lower and quoted in full below.
Single stream, temperature 0, measured after boot on 28 Sept 2026. Throughput on
this model tracks how predictable the output is, because predictable text drafts better —
so a single number without the prompt mix is meaningless.
| Prompt type | tok/s | |
|---|---|---|
| Code (linked-list implementation) | 38.3 | |
| Structured JSON | 30.3 | |
| Prose | 18.7 | |
| Tool call (first, incl. JIT) | 1.5 | one-time Triton JIT |
Tool calling verified — a weather prompt correctly emitted
get_weather({"location": "Boston"}). KV cache came up at 788,977 tokens,
which is 3.01× concurrency against the 262K limit.
The first tool invocation takes ~7 s because Triton JIT-compiles
_prepare_dflash_inputs_kernel on first use. It is a one-time cost per boot —
subsequent calls are normal. Worth warning users about so it isn't mistaken for a hang.
The fetch script read the manifest with IFS=$'\t' read -r sha size fn.
Tab is IFS whitespace, not a plain delimiter — bash collapses runs of it and strips
leading/trailing whitespace. Manifest lines whose sha256 was empty (the 9 non-LFS files)
therefore lost their leading field, every field shifted left, and fn came out empty.
The guard [ -z "$fn" ] && continue then skipped those files
without printing anything.
Consequence: processor_config.json and tokenizer_config.json were
never fetched, and vLLM died at boot with a FileNotFoundError deep in the
multimodal processor stack.
The sting: the verification pass used the same broken reader, so it reported
ALL_VERIFIED_OK while never checking those files. The 190 GB of weights were
genuinely verified (their lines parse correctly) — but "all 44 verified" was false. Re-verified
in Python, which has no such ambiguity.
# Broken: tab is IFS whitespace -> empty first field shifts everything
while IFS=$'\t' read -r sha size fn; do
[ -z "${fn:-}" ] && continue # skips silently
done < manifest.tsv
# Rule: never split a TSV with IFS=$'\t' read when a field can be empty.
# Split in python, or write a sentinel for empty fields.
The published recipe pins NCCL_IB_GID_INDEX=3. That is correct on one node and
wrong on the other. Decoding the IPv4-mapped RoCE v2 GIDs on
rocep1s0f1:
169.254.111.113 → ...ffff:a9fe:6f71 → index 3169.254.154.250 → ...ffff:a9fe:9afa → index 4The worker simply has an extra unused slot. A single hardcoded index fails
ncclCommInitRank with NCCL error: unhandled system error — which looks
alarmingly like a model or memory problem and is nothing of the sort.
Fix, now applied per rank:
case "$NODE_RANK" in
0) HOST_IP=169.254.111.113; GID_INDEX=3 ;; # head
1) HOST_IP=169.254.154.250; GID_INDEX=4 ;; # worker
esac
...
-e NCCL_IB_HCA=rocep1s0f1 -e NCCL_IB_GID_INDEX=$GID_INDEX
Rather than discover fabric breakage after a 25-minute model load, we wrote a
one-minute two-node NCCL all-reduce that runs inside the same container image
and proves the fabric before any weights are touched. It printed
NCCL ALLREDUCE OK, sum-first = 3.0 (1+2) on both ranks.
Run this before every launch. It converts a 25-minute failure mode into a 60-second one.
# /home/rusty/nccl-test.sh <rank> -- run on both nodes, worker first
docker run --rm --gpus all --network host --ipc host --shm-size 8g \
--device /dev/infiniband:/dev/infiniband --ulimit memlock=-1:-1 --cap-add IPC_LOCK \
-e NCCL_NET=IB -e NCCL_IB_HCA=rocep1s0f1 -e NCCL_IB_GID_INDEX=$GID \
-e NCCL_IB_ADDR_FAMILY=AF_INET -e NCCL_IB_ADDR_RANGE=169.254.0.0/16 \
-e NCCL_SOCKET_IFNAME=enp1s0f1np1 -e RANK=$RANK -e WORLD=2 \
--entrypoint python3 "$IMG" -c '...dist.all_reduce()...'
processor_config.json (Bug 1)glm-5.3-flash servingThe 200 Gb/s fabric is not the bottleneck — the house uplink is. The head's wired port negotiates at 100 Mbps, and both wired and WiFi independently cap near 11 MB/s. Because the two interfaces are separately capped, splitting the shard manifest across both doubled throughput:
| Path | Throughput | Note |
|---|---|---|
| Wired alone | 11 MB/s | port negotiated at 100 Mbps |
| WiFi alone | 11 MB/s | separately capped |
| Both, split manifest | 22.5 MB/s | halved the download |
| 200G RoCE fabric | 436 MB/s | node-to-node, for scale |
Parallelising across both nodes would not have helped — they share one internet uplink.
docker save on a 30 GB image appears hung for ~3.5 minutes
before it streams anything, because the daemon stages into
/var/lib/docker/tmp/docker-export-* first. That directory is root-only, so there is
no way to watch progress and distinguish "working" from "wedged".
Piping save | ssh | load did deadlock. The reliable route is
file → rsync → load.
# Launch (on the head). Worker rank 1 goes FIRST, then head rank 0.
ssh spark-be0f
bash /home/rusty/relaunch.sh
# The script does, in order:
# 1. docker rm -f vllm_glm53 (both nodes)
# 2. sudo -n /usr/local/bin/spark-prep (both nodes) <- drop_caches + swappiness=0
# 3. launch worker rank 1, sleep 25, launch head rank 0
# 4. poll curl -sf http://127.0.0.1:8000/health
/v1/models to decide readinessIt returns 200 from config alone, with a dead engine behind it. It will
happily answer while the model is still loading, or after a crash. Poll
/health.
Boot takes ~25 minutes, not the ~15 the upstream docs imply:
| Phase | Time | Signature in logs |
|---|---|---|
| Shard load | 12.5 min | Loading safetensors checkpoint shards |
| Kernel + FlashInfer autotune | ~12 min | Autotuning process starts |
Harmless noise during boot: No available shared memory broadcast block found
in 60 seconds means it is compiling, not hanging. The drafter also warns it ignores
multimodal embeddings — expected, since DFlash2 drafts on text only.
Two optional accelerations exist that we chose not to install for the first boot, to keep the day-0 variable count down:
RoCE all-reduce OFF (NCCL) and falls back cleanly. Optimisation, not correctness.The vLLM containers are restart=no and vm.swappiness=0 is applied
at launch time only — so after a reboot, GLM does not come back. Worse, three
enabled qwen-* systemd units would start Ray and qwen-serve.service
would grab both nodes' memory for the old Qwen 122B instead.
Making GLM durable means a systemd unit on the head plus disabling those three units —
one sudo command per node, outside the narrow helper granted for
spark-prep.
mistral-small-4 and
qwen3.6-35b now point at servers we deliberately stopped. GLM replaces both and
they cannot coexist — it holds both nodes' unified memory.qwen3.6-35b, a dead
endpoint, so a fresh session with no model set would fail. Needs repointing at the CLI/default
layer, not the catalog.This is day-0 silicon with a community recipe. The stack needed two fixes before it would boot,
and the upstream recipe itself required adaptation for this fabric. It is stable and in production
use for our purposes, but treat it as a system that needs its operator to know these specific
failure modes — the NCCL smoke test and the /health rule are not optional ceremony,
they are what makes a 25-minute boot debuggable.
Thanks to the community work this builds on: the tonyd2wild TP2 +
DFlash2 recipe, NVIDIA's NVFP4 checkpoint, LibertAI and RedHatAI quantization work, and the
DGX Spark forum threads that mapped the SM121 crash modes.