Quad NVIDIA CMP 170HX Server: 256GB VRAM, 563 tok/s on vLLM | Reclaimed Node
← Back to Blog Homelab

Four CMP 170HX Cards, One Server: 256GB of VRAM

Oct 6, 2026 · 4 min read

One 64GB CMP 170HX was enough to run a 40GB image model at full precision. So I built a server with four of them. That is 256GB of VRAM in a single box, built from mining cards nobody else wanted. This post covers the build, first boot and first benchmarks. More detailed testing is coming next.

The Build

{{ s.label }}
{{ s.value }}
256 GB

Total VRAM across four cards.

33–34°C

GPU temperature at idle, with a model loaded.

41–43 W

Idle draw per card, against a 250 W limit.

First Boot

Installation was very simple. It just required the cmpunlocker repo to be cloned and installed: one sudo ./install.sh on top of the NVIDIA 610 open driver, then a reboot. cmpunlocker patches the driver to restore features NVIDIA locked out of the CMP 170HX, including the full 64GB of memory, full compute throughput and PCIe Gen 2 speeds. After the reboot all four cards came up in Ubuntu 26.04 with 64GB each and a Gen 2 x16 link. The first job was vLLM with tensor parallelism across all four GPUs. It splits one language model across the cards, so each holds a quarter of the weights and they work together on every token. Here is nvidia-smi with the model loaded:

{{ smi }}
One vLLM worker per card (TP0–TP3), each holding about 59GB.

A First Look at Speed

A sweep against vLLM’s OpenAI-compatible endpoint running glm-5.3-flash, from 1 up to 32 requests at once, with a mixed workload and up to 2,048 output tokens per request. More detailed benchmarks will follow.

{{ s.label }}
{{ s.value }}
Concurrent
Requests
TTFT median
TTFT p95
Per-user tok/s
Aggregate tok/s
{{ r.conc }}
{{ r.reqs }}
{{ r.ttft }}
{{ r.p95 }}
{{ r.user }}
{{ r.agg }}

TTFT is time to first token. Per-user tok/s is the median speed each request saw after its first token. Aggregate tok/s is all output tokens divided by wall time. No errors across 65 requests. Reasoning tokens count as output.

A single user gets about 160 tokens per second, with the first token in a tenth of a second. As more users join, each one slows down, but total throughput keeps climbing: 166 tokens per second for one user, 421 at 8 and 563 at 32. The gain from 16 to 32 is small, and the slowest first token at 32 jumps to 3.84 seconds, so the sweet spot for this setup sits around 16 users at once.

Single-User Speed by Workload

Structured
222.5 tok/s
Code
159.9 tok/s
Prose
140.0 tok/s

Decode speed at concurrency 1. Predictable output like structured data gives speculative decoding more correct guesses, so it runs fastest.

Keeping It Cool

The CMP 170HX is a passively cooled card. It has no fans of its own and relies on server airflow to push air through its heatsink. A custom shroud with two 10,000 RPM 80mm server fans forces air straight through all four cards. At idle they sit at 33–36°C.

To see how the cooling holds up, I logged power and temperature once a second with nvidia-smi dmon during the full concurrency sweep above, about four and a half minutes of all four cards at 98–100% utilization.

62°C

Hottest GPU reading under load. Memory peaked at 72°C.

180–250 W

Typical draw per card under load, against the 250 W limit.

~990 W

Peak combined draw of all four GPUs in a single sample.

GPU power as reported by the cards, not measured at the wall. The cards spread evenly; the warmest ran about 3°C above the others for the whole test.

Temperatures climbed steadily for the first couple of minutes, then levelled off at 56–62°C and stayed there. The shroud has plenty of headroom, and the HELA 2500Rz has room to spare even with all four cards near their limit.

Coming Soon

{{ u.num }}

{{ u.title }}

{{ u.body }}

One card showed that a reclaimed mining GPU can do serious local AI work. Four should show how far that idea goes. Check back soon.

Related: Local AI image generation on a single 64GB CMP 170HX and the Mac Studio vs. CMP 170HX head-to-head.

Building Something Like This?

Tell us what you are trying to run — local models, passthrough labs, bulk storage — and we will tell you honestly what it takes and what we have on the bench.

Get Sourcing Help
More Posts Shop Inventory