inabot

Measurements · Intel Arc Pro B60

Intel Arc Pro B60 for LLM inference: what four cards actually do

Three MoE models held in VRAM at once on 4× Arc Pro B60 24 GB, served by llama.cpp (SYCL) and vLLM-XPU. Measured on 10 and 11 September 2026 on a machine that also carries a production workload, so these numbers are a floor, not a ceiling. Every figure, script and failure is in our public repository.

82.4 tok/sQwen3-Coder-30B-A3B, single request, llama.cpp SYCL
119 tok/saggregate, 12 concurrent requests across 3 models
211 Wpeak for the four cards, 163 W idle, 52–62 °C
≈ CHF 25 / GBof VRAM at Swiss retail price (CHF 614 per 24 GB card)

Open the repository on GitHub →

The machine

GPUs4× ASRock Arc Pro B60 Creator 24 GB (BMG-G21), each at PCIe x8 Gen3
HostIntel i9-9980XE, X299 (2018), 64 GB DDR4-3200
OS / driverUbuntu 24.04, kernel 7.0 HWE, xe driver, GuC 70.72.1
Runtimesintel/vllm:0.21.0-xpu and llama.cpp SYCL (server-intel image)

Three models resident at the same time

CardModelFormat / runtimeWeights
0gpt-oss-20bMXFP4 · vLLM-XPU13 GB
1Qwen3-Coder-30B-A3BGGUF Q4_K_XL · llama.cpp SYCL17.7 GB
2 + 3Qwen3-Next-80B-A3BGGUF Q3_K_XL · llama.cpp SYCL35.6 GB

130 B parameters combined, about 3 B active per token in each model. 81.7 GB of 96 GB in use, weights plus KV cache.

Single-model throughput, warm

ModelDecodePrefill (~4k tokens)
Qwen3-Coder-30B-A3B82.4 tok/s4 098 tokens in 4.2 s
gpt-oss-20b41.2 tok/s3 307 tokens in 2.4 s
Qwen3-Next-80B-A3B36.2 tok/s4 098 tokens in 10.3 s

The first call after loading is 3 to 4 times slower: send a warm-up request before real traffic.

Concurrency: where it scales and where it breaks

With the three models serving 12 simultaneous requests, 3 600 tokens came out in 30.3 s (119 tok/s aggregate), with no error and no model slowing another. Per model, the runtime decides everything.

gpt-oss-20b on vLLM-XPU (continuous batching)

Concurrent requestsPer requestAggregate
135.1 tok/s35.1
235.0 tok/s70.0
433.6 tok/s133.9
8—241

Qwen3-Next-80B on llama.cpp SYCL (no continuous batching)

Concurrent requestsPer requestAggregate
132.0 tok/s32.0
214.5 tok/s28.8
47.0 tok/s20.7

Correction, published in the open: our first concurrency run divided token counts by global wall-clock time. It flattered the fast models and hid the 80B regression. The tables above are the corrected figures.

Design rule we took from it: anything latency-sensitive (voice, interactive completion) goes to the model served by vLLM. A large model on llama.cpp serves about two concurrent users, not four.

Host resources and power

Idle, 3 models loadedUnder 12 concurrent requests
Host RAM available43.1 GB of 62 GB41.5 GB (trough)
VRAM used81.7 GB of 96 GBsame
Power, 4 cards163 W211 W peak
Temperatures52–62 °C52–62 °C

64 GB of host RAM is enough for this configuration. The constraint is at load time, not while serving: models must be loaded one after the other.

Eleven failure modes we hit

The part that is not documented anywhere else. Exact error strings and fixes are in the repository.

  1. vLLM-XPU only serves quantisations that have a native XPU kernel: MXFP4 works, most quantised MoEs fail on an NVIDIA-only kernel path.
  2. AWQ requires --dtype float16 explicitly.
  3. oneCCL cannot exchange IPC handles between workers, which blocks tensor parallel.
  4. Two processes on the same GPU can wedge the host.
  5. Models must be loaded sequentially: parallel loads froze the machine three times.
  6. The first inference after load is 3 to 4 times slower.
  7. Docker --ipc host silently cancels shm_size.
  8. A card with its 8-pin connector unplugged is simply absent from the PCI bus.
  9. Battlemage cards show x1 in lspci: that is cosmetic, read the root port instead.
  10. GuC firmware 70.44.1 was unstable; 70.72.1 from kernel.org fixed it, without a reboot.
  11. Vulkan is not an alternative to SYCL for decode.

Which motherboards take four cards

Arc Pro blower cards are two slots thick, which rules out most server boards. Measured from vendor drawings: ASUS Pro WS WRX90E-SAGE SE and W790E-SAGE SE take four on 8 bracket positions; Supermicro H13SSL-N and H12SSL-i take only three. Full table in the repository.

Not measured yet

  • Tensor parallel at any size (blocked, see above)
  • Arc Pro B70 and Arc Pro B60 Dual: no cards yet
  • Long context beyond 64k

Use the data

Code under MIT, documentation and measurement data under CC BY 4.0: use the numbers, cite the repository. If you find a mistake, open an issue.

Open the repository on GitHub

The same stack, delivered

SOKKAN Anchor is this machine, assembled in Geneva with the three models already running and an OpenAI-compatible API. You can test the 30B coding model on these very cards with a trial key before buying.