fin1te

Writing/·10 min read

Running a 552B model with 1M context on a box that was already busy

The only machine with enough HBM to hold a 552B-parameter model was already running production control-plane work. What fit, how it was installed without a reboot, and how many developers can share it at once.

I spend most of my working week inside coding agents, and the ones worth using are the ones that send the most tokens. A heavy day across a few of them is a few hundred million tokens through somebody else’s API. At some point that stops being an expense line and becomes an architecture decision, which is how I ended up looking at a GPU node we already own: 8× H200, sitting at 0% utilisation, the sort of thing that bothers you quietly until you do something about it.

The node was not spare. It also ran the single-node Kubernetes control plane, the operator that manages our DPUs, and cluster networking for workloads other people depend on. Whatever happened next, that part had to keep running exactly as it was. No reboot, no kernel upgrade, no dropped packets, and no assumptions: every step was checked against a baseline snapshot taken before I touched anything.

This is how DeepSeek-V4.1-Flash, 552B parameters with 1M-token native context, ended up on that node. What fits in 1.13 TB of HBM, the install, the five traps between the download and the first token, and how many developers can share it at once.

The numbers

Metric Value
Model DeepSeek-V4.1-Flash, 552B total, 8B active in prefill and 16B in decode
Context 1M tokens, native
HBM 8× H200 SXM5, 141 GB each, 1.13 TB total
Weights per GPU 50.74 GiB
KV cache 22,181,788 tokens, about 21 full-length conversations at once
Single-user decode 283-348 tok/s at short context
Needle test 906,689-token prompt answered in 43.7 s
Cold start about 4.5 minutes
Cost per token zero marginal

What fits in 1.13 TB

The first question was not which model is best. It was which ones fit, because the weights and the conversation state come out of the same 1.13 TB of GPU memory. The weights are the easy part to reason about. The part people forget is the conversation state, the KV cache: the model’s working memory of what you have said so far, which grows with every token of context. Fit the weights and run out of room for the cache, and a 1M-token model quietly becomes a 64K-token model.

Model Size Verdict
GLM-5.3 744B MoE, 40B active, about 743 GB FP8 Fits, but claims the whole node, and context tops out around 131-160K on 8× H200, not 1M
GLM-5.3-Flash 321B, 18B active, 306 GiB FP8 A good second model on 4 GPUs
DeepSeek-V4-Pro 1.6T Needs multiple nodes on Hopper
Kimi K3 2.8T, about 1.56 TB of weights Does not fit
Qwen 3.8-2.4T 2.4T Does not fit

The number that decided it was the KV cache. V4.1-Flash stores it in FP4, about 890 bytes per token, which means a full 1M-token session needs under 1 GB of memory. That is why 22.2M tokens of cache fit on the node, and it is the figure that decides whether the 1M context on a model card is real on your hardware.

The second number is how much of the model works at once. It is a mixture of experts: 552B parameters in total, but only 8B are active during prefill and 16B during decode, which is how a model this large runs at all on eight GPUs. It has 1M-token context and image input built in, and an MIT licence. It was released in September 2026, and DeepSeek’s own benchmarks put it at 90.6 on Terminal-Bench 2.1 and 74.2% on DeepSWE v1.1. Those are the vendor’s numbers and I did not rerun them.

The checkpoint is 511 GB on disk, split across 48 shards of expert weights, embedding tables and scaling data.

One warning from the selection process. The first comparison I saw, partly AI-written, had two numbers badly wrong: it said GLM-5.3 would do 1M context on this hardware (it cannot), and that V4.1-Flash needed around 200 GB of KV cache for a full context (it needs under 1 GB). Both errors change the decision, in opposite directions. Read the model card and the official vLLM recipe, not a summary of them.

The install: no reboot, no new kernel, no surprises

borrows all eightOne GPU nodeDay jobcontrol plane, DPU operator, networkingvLLMone container, capped RAM and CPUs8× H2001.13 TB of HBM, NVLink mesh
One node, three tenants. The whole install was adding the third without the first two noticing.Open in topo ↗

Before touching anything I saved a baseline: iptables rules, addresses, routes, loaded kernel modules, Kubernetes nodes and pods, DPU status. Every step below ended with a diff against that baseline, and every diff stayed empty.

  1. Drivers, swapped live. The console runs on the BMC’s basic video, so nouveau’s reference count was 0 and it could be removed without a reboot: modprobe -r nouveau, then modprobe nvidia, and the GPUs were visible within minutes. A blacklist keeps nouveau off on future boots.
  2. DKMS against the running kernel. The packaged prebuilt module wanted kernel 6.8.0-138; the node runs 6.8.0-124. DKMS built the module for the kernel that was actually running, which removed the forced upgrade and the reboot along with it.
  3. Driver and Fabric Manager locked together. Both are on apt hold. Unattended upgrades moving one without the other is how you lose NVLink on a node you were not planning to visit.
  4. Docker boxed in before it existed. A four-line daemon.json was written first, with iptables, ip6tables and ip-forward off and no default bridge; the container runs on the host network. Ubuntu’s docker.io reuses the host’s containerd, and k0s has its own, so the cluster never sees the GPUs. The container is capped at 1.5 TB of RAM and 200 CPUs, leaving the control plane its cores.
  5. Verification, not vibes. DPUs Ready, no new unhealthy pods, iptables, addresses and routes identical. Boring, which was the goal.

Support for V4.1-Flash is not in a stable vLLM release yet, so the server runs the nightly image. The shape of the final launch, sanitised:

docker run --gpus all --network host --shm-size 64g --ulimit memlock=-1 \
  --memory 1500g --cpus 200 \
  -v /models/DeepSeek-V4.1-Flash:/model:ro \
  -e VLLM_HOST_IP=127.0.0.1 -e NCCL_SOCKET_IFNAME=lo -e GLOO_SOCKET_IFNAME=lo \
  vllm/vllm-openai:nightly /model \
  --served-model-name deepseek-v4.1-flash --host $HOST_IP --port 8000 \
  --tensor-parallel-size 8 --max-model-len 1048576 \
  --tokenizer-mode deepseek_v41 --tool-call-parser deepseek_v41 \
  --reasoning-parser deepseek_v41 --enable-auto-tool-choice \
  --enable-prefix-caching --speculative-config '{"method":"dspark","num_speculative_tokens":5}'

The five traps

None of these is exotic. Each one was a single default away from breaking something the node does for a living.

Docker against br_netfilter. The host had br_netfilter loaded with bridge-nf-call-iptables=1, which sends bridged traffic through iptables, and Docker sets the FORWARD policy to DROP. On this node that combination would have silently dropped traffic on the bridge carrying DPU management and pod networking. The four lines from step 4 existed for exactly this reason.

A vendor repo that moved under me. The NVIDIA/Mellanox DOCA latest repository changed its release codename from 3.4.0 to 3.5.0. apt refused to refresh it, and its stale dkms package started returning 404s. Accepting the new codename would have touched the OVS and DOCA networking stack on a production node, so I installed Ubuntu’s own dkms package instead and left the repo alone.

Xet said 401 at 223 GB. The Hugging Face download died partway through when the Xet storage backend started returning 401 Unauthorized. Plain HTTP resumed fine, then refused the last two shards, which are 101.5 GB each: the regular client will not download files that large. The fix was 46 shards over plain HTTP and the two big ones over Xet, in that order.

vLLM picked the wrong network. On first start, the model server put its internal listeners on the cluster network instead of loopback, which meant about 20 unauthenticated ports sharing a wire with production traffic. Three settings, VLLM_HOST_IP, NCCL_SOCKET_IFNAME and GLOO_SOCKET_IFNAME, moved 120 sockets to 127.0.0.1. Only the API port answers on a real interface, and it requires a key.

The 95% GPU memory reading. After startup, every GPU reports 136 of 144 GB in use. That is not pressure. vLLM reserves the KV cache up front, and in this configuration the reservation is the 22.2M tokens from earlier. Nothing to fix, but it is the kind of reading that causes a ticket if nobody documents it.

And one for the blooper reel: the coding agent helping me run the install spent its first minutes trying to SSH into the server it was already running on.

Timings

Stage Time
Driver install to eight GPUs visible a few minutes, no reboot
Model download, 511 GB about an hour, 220-350 MB/s through an outbound proxy
Weights loaded into the GPUs 27.9 s, 50.74 GiB per GPU
Cold start to API ready about 4.5 minutes, including profiling and CUDA graph capture
Warm restart about 3 minutes

Add the download and the day is really about an afternoon, most of it spent waiting.

Benchmarks

Method: streaming chat completions with reasoning on, coding-task prompts (“implement X in Python, Go, Rust or TypeScript, with tests”), max_tokens 1500. All users start at the same moment, and every long-context prompt has a unique prefix, so nothing is served from cache. This is the worst case, and each row is a single run. Per-user speed is the median decode rate after the first token; aggregate is total output divided by wall clock.

Short prompts (about 41 tokens)

Users Per-user, median / min TTFT Aggregate
1 283-348 tok/s 0.0-0.1 s 279-346 tok/s
4 279 / 265 0.1 s 1,055 tok/s
16 131 / 113 0.1 s 1,787 tok/s
32 109 / 93 0.2 s 2,959 tok/s
64 69 / 60 0.3 s 3,794 tok/s

Agent-sized contexts (about 43K tokens, unique)

Users Per-user, median / min TTFT, median / max Aggregate
1 316 tok/s 1.2 s 252 tok/s
4 197 / 149 2.9 / 4.4 s 527 tok/s
8 122 / 97 5.1 / 8.6 s 645 tok/s
16 73 / 56 9.2 / 17.2 s 768 tok/s

Large contexts (about 180K tokens, unique)

Users Per-user, median / min TTFT, median / max Aggregate
1 266 tok/s 5.2 s 138 tok/s
4 119 / 71 12.9 / 20.0 s 217 tok/s

The needle test: a 906,689-token prompt, 36,000 synthetic log lines with one secret line at about 72% depth. Answered correctly in 43.7 seconds end to end, about 20,750 prompt tokens a second. Tool calling works on both APIs; the model correctly ran df -h /models and ls -la /tmp when asked.

So how many people does that add up to? Slide it:

16 people
comfortable300 tok/s100 tok/s
Each person
131tok/s
Everyone together
1,787tok/s
First token
0.1s
Experience
comfortable
As more people share the node, the team's total throughput goes up and each person's speed comes down. 15 to 25 is the comfortable range, and around 40 it starts to feel slow. These are the worst-case numbers with no cache hits, so real agent traffic lands somewhat better.

In practice: comfortable for a team of 15 to 25 developers using agents at the same time, busy but usable up to about 40. Real agent traffic does better than these tests, because each turn resends mostly the same prefix, and prefix caching reuses it.

At idle with the model loaded: 0% GPU utilisation, 30-36 °C, about 113-121 W per GPU against a 700 W limit, and a load average around 3 on 256 threads. The node’s day job never blinked.

Pointing a coding agent at it

vLLM serves the Anthropic API natively, /v1/messages, not just the OpenAI-compatible one. Agent CLIs that speak it work with no Anthropic login and no Claude credits:

# the vLLM endpoint the container listens on
ANTHROPIC_BASE_URL=$SERVER_URL
ANTHROPIC_AUTH_TOKEN=<key>
ANTHROPIC_MODEL=deepseek-v4.1-flash
ANTHROPIC_DEFAULT_OPUS_MODEL=deepseek-v4.1-flash
ANTHROPIC_DEFAULT_SONNET_MODEL=deepseek-v4.1-flash
ANTHROPIC_DEFAULT_HAIKU_MODEL=deepseek-v4.1-flash
CLAUDE_CODE_SUBAGENT_MODEL=deepseek-v4.1-flash
CLAUDE_CODE_MAX_CONTEXT_TOKENS=1000000
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
  • Without CLAUDE_CODE_MAX_CONTEXT_TOKENS, the CLI assumes a 200K window and compacts the conversation early, which wastes the whole point.
  • The model calls itself Claude, because the CLI’s system prompt says so. Worth knowing before you file a bug.
  • It also works with OpenCode, Cline, Aider, OpenHands and Goose through the OpenAI-compatible endpoint.

For reference, DeepSeek’s own API lists V4.1-Flash at $0.15 / $0.60 per million input and output tokens off-peak, double at peak. Treat that as context for what a busy team’s meter looks like, not as a measured saving.

What it does not protect you from

vLLM is configured to log no prompts or responses, and it keeps no history between requests, which I checked in the logs. But this is a shared server, not a confidential one. Administrators with root access could still capture traffic, so keep secrets out of what you send through it.

The hardening list, still open: per-user API keys behind a small gateway, removing the shared login account, and pinning the nightly image by its digest.

What I would tell someone doing this

  • Do the memory maths yourself, from the model card and the official recipe. The KV cache number decides your real context length, and the summaries get it wrong in both directions.
  • On a node that does other work, take a baseline and diff against it after every step. “It looks fine” is not a verification.
  • Check br_netfilter before installing Docker anywhere with bridges that matter. If you do not need Docker’s bridge networking, turn its iptables management off first.
  • An unused nouveau can be removed without a reboot, and DKMS saves you from a forced kernel upgrade.
  • Hold the driver and Fabric Manager packages together. Expect multi-hundred-GB downloads to resume, and know that the biggest shards need a different client.
  • Pin vLLM’s internal sockets to loopback before the first start, not after you have counted where they landed.
  • Speculative decoding is what makes single-user speed feel fast; acceptance hovered around 2.94 tokens per step in the interval I sampled.
  • Treat the node’s day job as the acceptance test. If the control plane cannot tell you were there, the install was a success.