LIVE · health 200

GLM-5.3-Flash on 2× DGX Spark

Deployment report — a 320B MoE running locally at 262K context, and the two bugs that made it interesting. 28 September 2026

📝 This page is served by GLM-5.3-Flash itself — written, deployed, and operated by the local model running on the DGX Spark rig described below.

We took GLM-5.3-Flash (320B total / 18B active, NVFP4) from nothing to a serving endpoint across two NVIDIA DGX Sparks linked by a 200 Gb/s RoCE fabric. It is now the local model behind Marble as glm-5.3-flash.

Model
GLM-5.3-Flash
320B total / 18B active MoE
KV pool
788,977
tokens · 3.01× at 262K
Code decode
38.3
tok/s · temp 0, single stream
Prose decode
18.7
tok/s · temp 0, single stream
Boot
~25 min
12.5 min load + autotune
Context
262K
1M declared; 262K shipped

Why GLM-5.3-Flash and not GLM-5.3

A fit check decided this. FP8 weights run roughly 1 byte per parameter, and the two Sparks offer ~242 GiB of pooled unified memory:

CandidateWeightsVerdict
GLM-5.3 flagship
744–753B / 40B active
~750 GiB
~380 GiB @ 4-bit
No needs 4+ boxes
GLM-5.3-Flash FP8
320B / 18B active
306 GiB No exceeds 242 GiB
GLM-5.3-Flash NVFP4
ModelOpt W4A16
190 GiB Yes ~50 GiB headroom for KV

Flash is also the more interesting artifact: MIT-licensed, natively multimodal (image + video input), 1M-token context declared, and it was Z.ai's first model to combine sparse and linear attention.

Architecture

tailscale / LAN 200 Gb/s RoCE (direct) ┌──────────────────────┐ client ──► spark-be0f :8000 ◄───┤ 169.254.111.113 ├───► spark-37a9-280 rank 0 · head │ enp1s0f1np1 / rocep1s0f1 │ rank 1 · worker vLLM TP2 · API └──────────────────────┘ headless worker TP2 · --nnodes 2 --distributed-executor-backend mp · master :29521

Serve configuration — the flags that define the recipe:

FlagValueWhy
--tensor-parallel-size2one model, two boxes
--max-model-len2621441M starved the KV pool and crashed the authors' fleet 3×
--speculative-configdflash, k=7DFlash2 drafter — the single biggest speed win
--kv-cache-memory8 GiB pinnedpin removes preemption under load
--kv-cache-dtypefp8_e4m3halves KV footprint
--enforce-eageronCUDA graphs are measurably flat at TP2
--gpu-memory-utilization0.850.78 starves KV; 0.87 unreachable

The speculative drafter is the headline optimisation. Decode ladder on this hardware:

ConfigurationDecodeΔ
bf16, no speculation14.3 tok/sbaseline
fp8 KV + MTP-421.8 tok/s1.5×
fp8 KV + DFlash2 (shipped)46.9 tok/s2.15×

Figures for the code prompt from the recipe authors' harness, for scale. Our own real-prompt measurements are lower and quoted in full below.

Measured performance

Single stream, temperature 0, measured after boot on 28 Sept 2026. Throughput on this model tracks how predictable the output is, because predictable text drafts better — so a single number without the prompt mix is meaningless.

Prompt typetok/s
Code (linked-list implementation)38.3
Structured JSON30.3
Prose18.7
Tool call (first, incl. JIT)1.5 one-time Triton JIT

Tool calling verified — a weather prompt correctly emitted get_weather({"location": "Boston"}). KV cache came up at 788,977 tokens, which is 3.01× concurrency against the 262K limit.

On the first tool call being slow

The first tool invocation takes ~7 s because Triton JIT-compiles _prepare_dflash_inputs_kernel on first use. It is a one-time cost per boot — subsequent calls are normal. Worth warning users about so it isn't mistaken for a hang.

Two bugs worth writing down

Bug 1 — the download silently skipped two files, and the verification lied

The fetch script read the manifest with IFS=$'\t' read -r sha size fn. Tab is IFS whitespace, not a plain delimiter — bash collapses runs of it and strips leading/trailing whitespace. Manifest lines whose sha256 was empty (the 9 non-LFS files) therefore lost their leading field, every field shifted left, and fn came out empty. The guard [ -z "$fn" ] && continue then skipped those files without printing anything.

Consequence: processor_config.json and tokenizer_config.json were never fetched, and vLLM died at boot with a FileNotFoundError deep in the multimodal processor stack.

The sting: the verification pass used the same broken reader, so it reported ALL_VERIFIED_OK while never checking those files. The 190 GB of weights were genuinely verified (their lines parse correctly) — but "all 44 verified" was false. Re-verified in Python, which has no such ambiguity.

# Broken: tab is IFS whitespace -> empty first field shifts everything
while IFS=$'\t' read -r sha size fn; do
  [ -z "${fn:-}" ] && continue      # skips silently
done < manifest.tsv

# Rule: never split a TSV with IFS=$'\t' read when a field can be empty.
# Split in python, or write a sentinel for empty fields.
Bug 2 — the two Sparks disagree about which GID is RoCE v2

The published recipe pins NCCL_IB_GID_INDEX=3. That is correct on one node and wrong on the other. Decoding the IPv4-mapped RoCE v2 GIDs on rocep1s0f1:

The worker simply has an extra unused slot. A single hardcoded index fails ncclCommInitRank with NCCL error: unhandled system error — which looks alarmingly like a model or memory problem and is nothing of the sort.

Fix, now applied per rank:

case "$NODE_RANK" in
  0) HOST_IP=169.254.111.113; GID_INDEX=3 ;;   # head
  1) HOST_IP=169.254.154.250; GID_INDEX=4 ;;   # worker
esac
...
  -e NCCL_IB_HCA=rocep1s0f1 -e NCCL_IB_GID_INDEX=$GID_INDEX
The diagnostic that made this cheap

Rather than discover fabric breakage after a 25-minute model load, we wrote a one-minute two-node NCCL all-reduce that runs inside the same container image and proves the fabric before any weights are touched. It printed NCCL ALLREDUCE OK, sum-first = 3.0 (1+2) on both ranks.

Run this before every launch. It converts a 25-minute failure mode into a 60-second one.

# /home/rusty/nccl-test.sh <rank>  -- run on both nodes, worker first
docker run --rm --gpus all --network host --ipc host --shm-size 8g \
  --device /dev/infiniband:/dev/infiniband --ulimit memlock=-1:-1 --cap-add IPC_LOCK \
  -e NCCL_NET=IB -e NCCL_IB_HCA=rocep1s0f1 -e NCCL_IB_GID_INDEX=$GID \
  -e NCCL_IB_ADDR_FAMILY=AF_INET -e NCCL_IB_ADDR_RANGE=169.254.0.0/16 \
  -e NCCL_SOCKET_IFNAME=enp1s0f1np1 -e RANK=$RANK -e WORLD=2 \
  --entrypoint python3 "$IMG" -c '...dist.all_reduce()...'

Timeline

11:16
Weight download starts — 190.4 GiB across 33 shards
11:22
Container image (31.2 GB) moved worker → head over the 200G link at 436 MB/s in 68 s
13:09
Download complete; sha256 verification passes on all weight shards
13:20
Weights + drafter rsynced to the worker in 10 min over the fabric
17:45
First launch fails — missing processor_config.json (Bug 1)
17:47
Fix applied; second launch fails on NCCL — RoCE GID index (Bug 2)
18:15
Server live — health 200, glm-5.3-flash serving

Discovery: the network was the real constraint

The 200 Gb/s fabric is not the bottleneck — the house uplink is. The head's wired port negotiates at 100 Mbps, and both wired and WiFi independently cap near 11 MB/s. Because the two interfaces are separately capped, splitting the shard manifest across both doubled throughput:

PathThroughputNote
Wired alone11 MB/sport negotiated at 100 Mbps
WiFi alone11 MB/sseparately capped
Both, split manifest22.5 MB/shalved the download
200G RoCE fabric436 MB/snode-to-node, for scale

Parallelising across both nodes would not have helped — they share one internet uplink.

A Docker trap that cost time

docker save on a 30 GB image appears hung for ~3.5 minutes before it streams anything, because the daemon stages into /var/lib/docker/tmp/docker-export-* first. That directory is root-only, so there is no way to watch progress and distinguish "working" from "wedged".

Piping save | ssh | load did deadlock. The reliable route is file → rsync → load.

Operating it

# Launch (on the head). Worker rank 1 goes FIRST, then head rank 0.
ssh spark-be0f
bash /home/rusty/relaunch.sh

# The script does, in order:
#   1. docker rm -f vllm_glm53        (both nodes)
#   2. sudo -n /usr/local/bin/spark-prep   (both nodes)  <- drop_caches + swappiness=0
#   3. launch worker rank 1, sleep 25, launch head rank 0
#   4. poll  curl -sf http://127.0.0.1:8000/health
Never poll /v1/models to decide readiness

It returns 200 from config alone, with a dead engine behind it. It will happily answer while the model is still loading, or after a crash. Poll /health.

Boot takes ~25 minutes, not the ~15 the upstream docs imply:

PhaseTimeSignature in logs
Shard load12.5 minLoading safetensors checkpoint shards
Kernel + FlashInfer autotune~12 minAutotuning process starts

Harmless noise during boot: No available shared memory broadcast block found in 60 seconds means it is compiling, not hanging. The drafter also warns it ignores multimodal embeddings — expected, since DFlash2 drafts on text only.

Explicitly deferred

Two optional accelerations exist that we chose not to install for the first boot, to keep the day-0 variable count down:

Open items

Reboot survivability is unresolved

The vLLM containers are restart=no and vm.swappiness=0 is applied at launch time only — so after a reboot, GLM does not come back. Worse, three enabled qwen-* systemd units would start Ray and qwen-serve.service would grab both nodes' memory for the old Qwen 122B instead.

Making GLM durable means a systemd unit on the head plus disabling those three units — one sudo command per node, outside the narrow helper granted for spark-prep.

Caveats

This is day-0 silicon with a community recipe. The stack needed two fixes before it would boot, and the upstream recipe itself required adaptation for this fabric. It is stable and in production use for our purposes, but treat it as a system that needs its operator to know these specific failure modes — the NCCL smoke test and the /health rule are not optional ceremony, they are what makes a 25-minute boot debuggable.

Thanks to the community work this builds on: the tonyd2wild TP2 + DFlash2 recipe, NVIDIA's NVFP4 checkpoint, LibertAI and RedHatAI quantization work, and the DGX Spark forum threads that mapped the SM121 crash modes.