Running a 552B model with 1M context on a box that was already busy
The only machine with enough HBM to hold a 552B-parameter model was already running production control-plane work. What fit, how it was installed without a reboot, and how many developers can share it at once.
I spend most of my working week inside coding agents, and the ones worth using are the ones that send the most tokens. A heavy day across a few of them is a few hundred million tokens through somebody else’s API. At some point that stops being an expense line and becomes an architecture decision, which is how I ended up looking at a GPU node we already own: 8× H200, sitting at 0% utilisation, the sort of thing that bothers you quietly until you do something about it.
The node was not spare. It also ran the single-node Kubernetes control plane, the operator that manages our DPUs, and cluster networking for workloads other people depend on. Whatever happened next, that part had to keep running exactly as it was. No reboot, no kernel upgrade, no dropped packets, and no assumptions: every step was checked against a baseline snapshot taken before I touched anything.
This is how DeepSeek-V4.1-Flash, 552B parameters with 1M-token native context, ended up on that node. What fits in 1.13 TB of HBM, the install, the five traps between the download and the first token, and how many developers can share it at once.
The numbers
| Metric | Value |
|---|---|
| Model | DeepSeek-V4.1-Flash, 552B total, 8B active in prefill and 16B in decode |
| Context | 1M tokens, native |
| HBM | 8× H200 SXM5, 141 GB each, 1.13 TB total |
| Weights per GPU | 50.74 GiB |
| KV cache | 22,181,788 tokens, about 21 full-length conversations at once |
| Single-user decode | 283-348 tok/s at short context |
| Needle test | 906,689-token prompt answered in 43.7 s |
| Cold start | about 4.5 minutes |
| Cost per token | zero marginal |
What fits in 1.13 TB
The first question was not which model is best. It was which ones fit, because the weights and the conversation state come out of the same 1.13 TB of GPU memory. The weights are the easy part to reason about. The part people forget is the conversation state, the KV cache: the model’s working memory of what you have said so far, which grows with every token of context. Fit the weights and run out of room for the cache, and a 1M-token model quietly becomes a 64K-token model.
| Model | Size | Verdict |
|---|---|---|
| GLM-5.3 | 744B MoE, 40B active, about 743 GB FP8 | Fits, but claims the whole node, and context tops out around 131-160K on 8× H200, not 1M |
| GLM-5.3-Flash | 321B, 18B active, 306 GiB FP8 | A good second model on 4 GPUs |
| DeepSeek-V4-Pro | 1.6T | Needs multiple nodes on Hopper |
| Kimi K3 | 2.8T, about 1.56 TB of weights | Does not fit |
| Qwen 3.8-2.4T | 2.4T | Does not fit |
The number that decided it was the KV cache. V4.1-Flash stores it in FP4, about 890 bytes per token, which means a full 1M-token session needs under 1 GB of memory. That is why 22.2M tokens of cache fit on the node, and it is the figure that decides whether the 1M context on a model card is real on your hardware.
The second number is how much of the model works at once. It is a mixture of experts: 552B parameters in total, but only 8B are active during prefill and 16B during decode, which is how a model this large runs at all on eight GPUs. It has 1M-token context and image input built in, and an MIT licence. It was released in September 2026, and DeepSeek’s own benchmarks put it at 90.6 on Terminal-Bench 2.1 and 74.2% on DeepSWE v1.1. Those are the vendor’s numbers and I did not rerun them.
The checkpoint is 511 GB on disk, split across 48 shards of expert weights, embedding tables and scaling data.
One warning from the selection process. The first comparison I saw, partly AI-written, had two numbers badly wrong: it said GLM-5.3 would do 1M context on this hardware (it cannot), and that V4.1-Flash needed around 200 GB of KV cache for a full context (it needs under 1 GB). Both errors change the decision, in opposite directions. Read the model card and the official vLLM recipe, not a summary of them.
The install: no reboot, no new kernel, no surprises
Before touching anything I saved a baseline: iptables rules, addresses, routes, loaded kernel modules, Kubernetes nodes and pods, DPU status. Every step below ended with a diff against that baseline, and every diff stayed empty.
- Drivers, swapped live. The console runs on the BMC’s basic video, so nouveau’s reference count was 0 and it could be removed without a reboot:
modprobe -r nouveau, thenmodprobe nvidia, and the GPUs were visible within minutes. A blacklist keeps nouveau off on future boots. - DKMS against the running kernel. The packaged prebuilt module wanted kernel 6.8.0-138; the node runs 6.8.0-124. DKMS built the module for the kernel that was actually running, which removed the forced upgrade and the reboot along with it.
- Driver and Fabric Manager locked together. Both are on apt hold. Unattended upgrades moving one without the other is how you lose NVLink on a node you were not planning to visit.
- Docker boxed in before it existed. A four-line
daemon.jsonwas written first, withiptables,ip6tablesandip-forwardoff and no default bridge; the container runs on the host network. Ubuntu’sdocker.ioreuses the host’s containerd, and k0s has its own, so the cluster never sees the GPUs. The container is capped at 1.5 TB of RAM and 200 CPUs, leaving the control plane its cores. - Verification, not vibes. DPUs Ready, no new unhealthy pods, iptables, addresses and routes identical. Boring, which was the goal.
Support for V4.1-Flash is not in a stable vLLM release yet, so the server runs the nightly image. The shape of the final launch, sanitised:
docker run --gpus all --network host --shm-size 64g --ulimit memlock=-1 \
--memory 1500g --cpus 200 \
-v /models/DeepSeek-V4.1-Flash:/model:ro \
-e VLLM_HOST_IP=127.0.0.1 -e NCCL_SOCKET_IFNAME=lo -e GLOO_SOCKET_IFNAME=lo \
vllm/vllm-openai:nightly /model \
--served-model-name deepseek-v4.1-flash --host $HOST_IP --port 8000 \
--tensor-parallel-size 8 --max-model-len 1048576 \
--tokenizer-mode deepseek_v41 --tool-call-parser deepseek_v41 \
--reasoning-parser deepseek_v41 --enable-auto-tool-choice \
--enable-prefix-caching --speculative-config '{"method":"dspark","num_speculative_tokens":5}'
The five traps
None of these is exotic. Each one was a single default away from breaking something the node does for a living.
Docker against br_netfilter. The host had br_netfilter loaded with bridge-nf-call-iptables=1, which sends bridged traffic through iptables, and Docker sets the FORWARD policy to DROP. On this node that combination would have silently dropped traffic on the bridge carrying DPU management and pod networking. The four lines from step 4 existed for exactly this reason.
A vendor repo that moved under me. The NVIDIA/Mellanox DOCA latest repository changed its release codename from 3.4.0 to 3.5.0. apt refused to refresh it, and its stale dkms package started returning 404s. Accepting the new codename would have touched the OVS and DOCA networking stack on a production node, so I installed Ubuntu’s own dkms package instead and left the repo alone.
Xet said 401 at 223 GB. The Hugging Face download died partway through when the Xet storage backend started returning 401 Unauthorized. Plain HTTP resumed fine, then refused the last two shards, which are 101.5 GB each: the regular client will not download files that large. The fix was 46 shards over plain HTTP and the two big ones over Xet, in that order.
vLLM picked the wrong network. On first start, the model server put its internal listeners on the cluster network instead of loopback, which meant about 20 unauthenticated ports sharing a wire with production traffic. Three settings, VLLM_HOST_IP, NCCL_SOCKET_IFNAME and GLOO_SOCKET_IFNAME, moved 120 sockets to 127.0.0.1. Only the API port answers on a real interface, and it requires a key.
The 95% GPU memory reading. After startup, every GPU reports 136 of 144 GB in use. That is not pressure. vLLM reserves the KV cache up front, and in this configuration the reservation is the 22.2M tokens from earlier. Nothing to fix, but it is the kind of reading that causes a ticket if nobody documents it.
And one for the blooper reel: the coding agent helping me run the install spent its first minutes trying to SSH into the server it was already running on.
Timings
| Stage | Time |
|---|---|
| Driver install to eight GPUs visible | a few minutes, no reboot |
| Model download, 511 GB | about an hour, 220-350 MB/s through an outbound proxy |
| Weights loaded into the GPUs | 27.9 s, 50.74 GiB per GPU |
| Cold start to API ready | about 4.5 minutes, including profiling and CUDA graph capture |
| Warm restart | about 3 minutes |
Add the download and the day is really about an afternoon, most of it spent waiting.
Benchmarks
Method: streaming chat completions with reasoning on, coding-task prompts (“implement X in Python, Go, Rust or TypeScript, with tests”), max_tokens 1500. All users start at the same moment, and every long-context prompt has a unique prefix, so nothing is served from cache. This is the worst case, and each row is a single run. Per-user speed is the median decode rate after the first token; aggregate is total output divided by wall clock.
Short prompts (about 41 tokens)
| Users | Per-user, median / min | TTFT | Aggregate |
|---|---|---|---|
| 1 | 283-348 tok/s | 0.0-0.1 s | 279-346 tok/s |
| 4 | 279 / 265 | 0.1 s | 1,055 tok/s |
| 16 | 131 / 113 | 0.1 s | 1,787 tok/s |
| 32 | 109 / 93 | 0.2 s | 2,959 tok/s |
| 64 | 69 / 60 | 0.3 s | 3,794 tok/s |
Agent-sized contexts (about 43K tokens, unique)
| Users | Per-user, median / min | TTFT, median / max | Aggregate |
|---|---|---|---|
| 1 | 316 tok/s | 1.2 s | 252 tok/s |
| 4 | 197 / 149 | 2.9 / 4.4 s | 527 tok/s |
| 8 | 122 / 97 | 5.1 / 8.6 s | 645 tok/s |
| 16 | 73 / 56 | 9.2 / 17.2 s | 768 tok/s |
Large contexts (about 180K tokens, unique)
| Users | Per-user, median / min | TTFT, median / max | Aggregate |
|---|---|---|---|
| 1 | 266 tok/s | 5.2 s | 138 tok/s |
| 4 | 119 / 71 | 12.9 / 20.0 s | 217 tok/s |
The needle test: a 906,689-token prompt, 36,000 synthetic log lines with one secret line at about 72% depth. Answered correctly in 43.7 seconds end to end, about 20,750 prompt tokens a second. Tool calling works on both APIs; the model correctly ran df -h /models and ls -la /tmp when asked.
So how many people does that add up to? Slide it:
- Each person
- 131tok/s
- Everyone together
- 1,787tok/s
- First token
- 0.1s
- Experience
- comfortable
In practice: comfortable for a team of 15 to 25 developers using agents at the same time, busy but usable up to about 40. Real agent traffic does better than these tests, because each turn resends mostly the same prefix, and prefix caching reuses it.
At idle with the model loaded: 0% GPU utilisation, 30-36 °C, about 113-121 W per GPU against a 700 W limit, and a load average around 3 on 256 threads. The node’s day job never blinked.
Pointing a coding agent at it
vLLM serves the Anthropic API natively, /v1/messages, not just the OpenAI-compatible one. Agent CLIs that speak it work with no Anthropic login and no Claude credits:
# the vLLM endpoint the container listens on
ANTHROPIC_BASE_URL=$SERVER_URL
ANTHROPIC_AUTH_TOKEN=<key>
ANTHROPIC_MODEL=deepseek-v4.1-flash
ANTHROPIC_DEFAULT_OPUS_MODEL=deepseek-v4.1-flash
ANTHROPIC_DEFAULT_SONNET_MODEL=deepseek-v4.1-flash
ANTHROPIC_DEFAULT_HAIKU_MODEL=deepseek-v4.1-flash
CLAUDE_CODE_SUBAGENT_MODEL=deepseek-v4.1-flash
CLAUDE_CODE_MAX_CONTEXT_TOKENS=1000000
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
- Without
CLAUDE_CODE_MAX_CONTEXT_TOKENS, the CLI assumes a 200K window and compacts the conversation early, which wastes the whole point. - The model calls itself Claude, because the CLI’s system prompt says so. Worth knowing before you file a bug.
- It also works with OpenCode, Cline, Aider, OpenHands and Goose through the OpenAI-compatible endpoint.
For reference, DeepSeek’s own API lists V4.1-Flash at $0.15 / $0.60 per million input and output tokens off-peak, double at peak. Treat that as context for what a busy team’s meter looks like, not as a measured saving.
What it does not protect you from
vLLM is configured to log no prompts or responses, and it keeps no history between requests, which I checked in the logs. But this is a shared server, not a confidential one. Administrators with root access could still capture traffic, so keep secrets out of what you send through it.
The hardening list, still open: per-user API keys behind a small gateway, removing the shared login account, and pinning the nightly image by its digest.
What I would tell someone doing this
- Do the memory maths yourself, from the model card and the official recipe. The KV cache number decides your real context length, and the summaries get it wrong in both directions.
- On a node that does other work, take a baseline and diff against it after every step. “It looks fine” is not a verification.
- Check
br_netfilterbefore installing Docker anywhere with bridges that matter. If you do not need Docker’s bridge networking, turn its iptables management off first. - An unused nouveau can be removed without a reboot, and DKMS saves you from a forced kernel upgrade.
- Hold the driver and Fabric Manager packages together. Expect multi-hundred-GB downloads to resume, and know that the biggest shards need a different client.
- Pin vLLM’s internal sockets to loopback before the first start, not after you have counted where they landed.
- Speculative decoding is what makes single-user speed feel fast; acceptance hovered around 2.94 tokens per step in the interval I sampled.
- Treat the node’s day job as the acceptance test. If the control plane cannot tell you were there, the install was a success.