Measurements · Intel Arc Pro B60
Intel Arc Pro B60 for LLM inference: what four cards actually do
Three MoE models held in VRAM at once on 4× Arc Pro B60 24 GB, served by llama.cpp (SYCL) and vLLM-XPU. Measured on 10 and 11 September 2026 on a machine that also carries a production workload, so these numbers are a floor, not a ceiling. Every figure, script and failure is in our public repository.
Open the repository on GitHub →
The machine
| GPUs | 4× ASRock Arc Pro B60 Creator 24 GB (BMG-G21), each at PCIe x8 Gen3 |
|---|---|
| Host | Intel i9-9980XE, X299 (2018), 64 GB DDR4-3200 |
| OS / driver | Ubuntu 24.04, kernel 7.0 HWE, xe driver, GuC 70.72.1 |
| Runtimes | intel/vllm:0.21.0-xpu and llama.cpp SYCL (server-intel image) |
Three models resident at the same time
| Card | Model | Format / runtime | Weights |
|---|---|---|---|
| 0 | gpt-oss-20b | MXFP4 · vLLM-XPU | 13 GB |
| 1 | Qwen3-Coder-30B-A3B | GGUF Q4_K_XL · llama.cpp SYCL | 17.7 GB |
| 2 + 3 | Qwen3-Next-80B-A3B | GGUF Q3_K_XL · llama.cpp SYCL | 35.6 GB |
130 B parameters combined, about 3 B active per token in each model. 81.7 GB of 96 GB in use, weights plus KV cache.
Single-model throughput, warm
| Model | Decode | Prefill (~4k tokens) |
|---|---|---|
| Qwen3-Coder-30B-A3B | 82.4 tok/s | 4 098 tokens in 4.2 s |
| gpt-oss-20b | 41.2 tok/s | 3 307 tokens in 2.4 s |
| Qwen3-Next-80B-A3B | 36.2 tok/s | 4 098 tokens in 10.3 s |
The first call after loading is 3 to 4 times slower: send a warm-up request before real traffic.
Concurrency: where it scales and where it breaks
With the three models serving 12 simultaneous requests, 3 600 tokens came out in 30.3 s (119 tok/s aggregate), with no error and no model slowing another. Per model, the runtime decides everything.
gpt-oss-20b on vLLM-XPU (continuous batching)
| Concurrent requests | Per request | Aggregate |
|---|---|---|
| 1 | 35.1 tok/s | 35.1 |
| 2 | 35.0 tok/s | 70.0 |
| 4 | 33.6 tok/s | 133.9 |
| 8 | — | 241 |
Qwen3-Next-80B on llama.cpp SYCL (no continuous batching)
| Concurrent requests | Per request | Aggregate |
|---|---|---|
| 1 | 32.0 tok/s | 32.0 |
| 2 | 14.5 tok/s | 28.8 |
| 4 | 7.0 tok/s | 20.7 |
Correction, published in the open: our first concurrency run divided token counts by global wall-clock time. It flattered the fast models and hid the 80B regression. The tables above are the corrected figures.
Design rule we took from it: anything latency-sensitive (voice, interactive completion) goes to the model served by vLLM. A large model on llama.cpp serves about two concurrent users, not four.
Host resources and power
| Idle, 3 models loaded | Under 12 concurrent requests | |
|---|---|---|
| Host RAM available | 43.1 GB of 62 GB | 41.5 GB (trough) |
| VRAM used | 81.7 GB of 96 GB | same |
| Power, 4 cards | 163 W | 211 W peak |
| Temperatures | 52–62 °C | 52–62 °C |
64 GB of host RAM is enough for this configuration. The constraint is at load time, not while serving: models must be loaded one after the other.
Eleven failure modes we hit
The part that is not documented anywhere else. Exact error strings and fixes are in the repository.
- vLLM-XPU only serves quantisations that have a native XPU kernel: MXFP4 works, most quantised MoEs fail on an NVIDIA-only kernel path.
- AWQ requires --dtype float16 explicitly.
- oneCCL cannot exchange IPC handles between workers, which blocks tensor parallel.
- Two processes on the same GPU can wedge the host.
- Models must be loaded sequentially: parallel loads froze the machine three times.
- The first inference after load is 3 to 4 times slower.
- Docker --ipc host silently cancels shm_size.
- A card with its 8-pin connector unplugged is simply absent from the PCI bus.
- Battlemage cards show x1 in lspci: that is cosmetic, read the root port instead.
- GuC firmware 70.44.1 was unstable; 70.72.1 from kernel.org fixed it, without a reboot.
- Vulkan is not an alternative to SYCL for decode.
Which motherboards take four cards
Arc Pro blower cards are two slots thick, which rules out most server boards. Measured from vendor drawings: ASUS Pro WS WRX90E-SAGE SE and W790E-SAGE SE take four on 8 bracket positions; Supermicro H13SSL-N and H12SSL-i take only three. Full table in the repository.
Not measured yet
- Tensor parallel at any size (blocked, see above)
- Arc Pro B70 and Arc Pro B60 Dual: no cards yet
- Long context beyond 64k
Use the data
Code under MIT, documentation and measurement data under CC BY 4.0: use the numbers, cite the repository. If you find a mistake, open an issue.
The same stack, delivered
SOKKAN Anchor is this machine, assembled in Geneva with the three models already running and an OpenAI-compatible API. You can test the 30B coding model on these very cards with a trial key before buying.