SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops
Ranran Haoran Zhang, Aysa Xuemo Fan, David Munhรก Correia, Alex Cheema, Rui Zhang
SiliconBench evaluates local LLM serving through speed, memory, and fidelity. Explore how engines handle concurrent workloads while preserving memory headroom and model quality.
The main audit covers nine Apple Silicon serving engines on chat and agent workloads. A complementary NVIDIA DGX Spark track measures serving performance for three shared engine families.
Speed captures throughput and latency under load; memory tracks the system footprint during serving; and fidelity uses a classification task to check for quality regressions against an NVIDIA reference.
Live results
rebuilt automatically from every merged
benchmark run
machinemodel
Single-node serving. Select a machine and model, then click a column header to sort. Read throughput alongside memory and fidelity. โ = crashed (<5/100 requests), n/100 = partial run, โ = not measured. Trend lines show each engine's change across the tested concurrency levels; each row is scaled independently, with first-token latency on a log scale. Hover an engine name for run details.
Runs 2026-08-22 to 2026-08-26 (splits from different runs): 7 of 9 stacks complete every request at every level; 2 crash at least once.
chat split
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm-metal
136.3
445.7
497.2
96 ms
41.8
ollama
146.7
223.0
493.5
160 ms
50.5
llama.cpp
123.9
284.5
285.5
1.94 s
26.6
omlx
119.8
157.9
153.4
2.49 s
28.2
mlx_lm
102.5
132.1
110.0
1.57 s
60.8
sglang
145.4
110.4
105.7
250 ms
61.2
vllm-mlx
127.1
83.4
69.2
1.44 s
59.4
hf_transformers
35.0
52.8
42.9
5 ms
60.3
mistral.rs
118.7
27.940/100
โ
โ
39.0
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm-metal
92.1
244.7
249.9
219 ms
34.9
ollama
100.3
132.0
189.7
1.15 s
48.6
omlx
81.9
86.6
83.8
6.23 s
41.7
llama.cpp
56.6
77.4
76.1
9.88 s
24.0
sglang
80.6
47.5
45.3
655 ms
61.1
vllm-mlx
70.8
39.2
29.9
7.46 s
60.0
mlx_lm
53.970/100
53.270/100
โ
โ
โ
hf_transformers
โ
โ
โ
โ
โ
โ
โ
mistral.rs
73.6
9.913/100
โ
โ
61.7
fidelity (weighted F1, GMRID, vs. one NVIDIA A100 reference)
Stack
0-shot F1
5-shot F1
vllm-nvidia (ref)
0.4094
0.7364
hf_transformers
0.3995
0.7299
llama.cpp
0.4033
0.7377
mlx_lm
0.3970
0.7346
ollama
0.4173
0.4462
omlx
0.3944
0.7379
sglang
0.3955
0.7337
vllm-metal
0.3956
0.7423
vllm-mlx
0.3592
0.7183
Runs 2026-08-21 to 2026-08-25 (splits from different runs): 4 of 9 stacks complete every request at every level; 1 degrade to partial or skip; 4 never start.
chat split
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm-metal
100.9
338.7
456.1
131 ms
35.7
vllm-mlx
95.3
285.1
318.0
430 ms
29.5
llama.cpp
105.5
304.0
296.0
3.06 s
22.9
omlx
113.4
208.0
212.6
3.33 s
โ
mlx_lm
93.9
139.1
144.4
1.24 s
32.8
hf_transformers
โ
โ
โ
โ
โ
โ
โ
mistral.rs
โ
โ
โ
โ
โ
โ
โ
ollama
โ
โ
โ
โ
โ
โ
โ
sglang
โ
โ
โ
โ
โ
โ
โ
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm-metal
76.2
257.5
297.0
181 ms
38.7
llama.cpp
66.9
126.9
135.3
5.93 s
24.0
vllm-mlx
67.6
126.0
109.4
1.19 s
34.6
omlx
77.0
98.5
92.7
5.88 s
โ
mlx_lm
45.425/100
49.526/100
41.726/100
3.29 s
โ
hf_transformers
โ
โ
โ
โ
โ
โ
โ
mistral.rs
โ
โ
โ
โ
โ
โ
โ
ollama
โ
โ
โ
โ
โ
โ
โ
sglang
โ
โ
โ
โ
โ
โ
โ
fidelity (weighted F1, GMRID, vs. one NVIDIA A100 reference)
Stack
0-shot F1
5-shot F1
vllm-nvidia (ref)
0.6953
0.5949
llama.cpp
0.7003
0.5940
mlx_lm
0.6942
0.5903
omlx
0.6981
0.5929
vllm-metal
0.6947
0.6081
vllm-mlx
0.7006
0.5922
Runs 2026-08-22 to 2026-08-25 (splits from different runs): 6 of 10 stacks complete every request at every level; 1 degrade to partial or skip; 3 never start.
chat split
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm-metal
24.3
93.4
151.8
265 ms
39.5
vllm-mlx
24.1
95.0
110.6
1.37 s
37.1
mlx_lm
22.6
77.1
88.5
2.27 s
36.2
llama.cpp
22.4
77.0
81.5
9.72 s
28.3
omlx
24.1
66.1
68.3
8.88 s
โ
hf_transformers
14.1
14.7
14.7
73.55 s
42.7
mistral.rs
โ
โ
โ
โ
โ
โ
โ
ollama
โ
โ
โ
โ
โ
โ
โ
sglang
โ
โ
โ
โ
โ
โ
โ
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm_metal_mtp
20.5
64.2
71.6
1.58 s
49.7
vllm-metal
21.1
52.3
63.6
1.49 s
39.7
omlx
21.6
48.3
50.4
15.05 s
โ
vllm-mlx
19.4
46.7
49.4
2.74 s
41.8
llama.cpp
18.3
44.2
40.2
30.18 s
35.3
mlx_lm
17.556/100
23.456/100
21.856/100
4.40 s
58.9
hf_transformers
โ
โ
โ
โ
โ
โ
โ
mistral.rs
โ
โ
โ
โ
โ
โ
โ
ollama
โ
โ
โ
โ
โ
โ
โ
sglang
โ
โ
โ
โ
โ
โ
โ
fidelity (weighted F1, GMRID, vs. one NVIDIA A100 reference)
Stack
0-shot F1
5-shot F1
vllm-nvidia (ref)
0.8511
0.9128
llama.cpp
0.8497
0.9107
mlx_lm
0.8493
0.9108
omlx
0.8539
0.9120
vllm-metal
0.8512
0.9106
vllm-mlx
0.8530
0.9107
8-bit weights (GGUF Q8_0 for llama.cpp; the mlx-community affine-8bit conversion for vllm-metal and omlx) โ BF16 does not fit this box. Run against three engines only, at concurrency 1/2/4; the other stacks were not attempted, so their absence is not a failure.
Runs 2026-08-24 to 2026-08-25 (splits from different runs): 3 of 3 stacks complete every request at every level.
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=2
tok/s c=4โผ
trend
TTFT p50 c=4
TTFT trend
peak mem GB
omlx
4.7
5.2
5.9
24.13 s
โ
vllm-metal
3.9
4.7
5.2
11.67 s
54.0
llama.cpp
3.8
4.2
4.7
11.07 s
53.0
4-bit MoE (35B total, ~3B active per token), agent split only, at concurrency 1/2/4. Four engines were attempted plus one oMLX variant; the rest of the roster was not run, so its absence is not a failure. mlx_lm is marked partial at every level โ around six in ten requests came back OK with zero generated tokens, so its rates describe only the ~38 that completed and are not a like-for-like comparison with the engines that finished 99โ100. โomlx (RAM cache)โ is the same build with its prefix cache held in memory and nothing written to disk, which is what every other engine here does; the plain oMLX row keeps its own default, a prefix cache backed by SSD. No peak-memory column: the memory sidecar traced only the last concurrency level of each run, so this model's memory needs a re-measure.
Runs 2026-09-02 to 2026-09-03 (splits from different runs): 4 of 5 stacks complete every request at every level; 1 degrade to partial or skip.
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=2
tok/s c=4โผ
trend
TTFT p50 c=4
TTFT trend
peak mem GB
vllm-metal
28.3
32.6
35.3
1.64 s
โ
omlx (RAM cache)
33.7
33.1
33.9
4.48 s
โ
omlx
30.6
27.4
30.1
6.82 s
โ
llama.cpp
21.1
21.6
25.3
2.28 s
โ
mlx_lm
19.738/100
23.540/100
23.538/100
3.10 s
โ
machine 18 cores ยท 64 GB unified memory ยท macOS 26.6 ยท Metal
run dates 2026-08-21 to 2026-09-03
harness78f885226
Framework versions and update commits for each run are recorded in the maintenance journals; structured version fields appear here once the harness emits them. Main platform in the paper's audit. Run details and model coverage appear in the tables above.
Runs 2026-07-05 to 2026-07-08 (splits from different runs): 6 of 9 stacks complete every request at every level; 3 crash at least once.
chat split
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
ollama
130.7
226.0
448.8
385 ms
47.3
llama.cpp
113.5
236.3
252.0
2.22 s
22.3
vllm-metal
57.8
184.1
193.8
254 ms
33.8
omlx
95.2
135.1
141.3
2.67 s
21.3
sglang
100.1
80.0
81.0
607 ms
58.3
mlx_lm
76.1
70.1
62.7
2.28 s
58.7
hf_transformers
26.4
43.8
33.7
1.09 s
57.3
mistral.rs
84.0
3.866/100
โ
โ
61.4
vllm-mlx
110.9
75.223/100
โ
โ
20.2
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
ollama
96.7
126.3
184.3
2.43 s
47.8
vllm-metal
32.1
106.8
101.7
646 ms
34.1
omlx
54.3
91.5
95.1
5.55 s
45.4
llama.cpp
41.6
52.9
51.5
15.42 s
34.5
sglang
38.7
30.4
30.1
2.02 s
58.2
hf_transformers
โ
โ
โ
โ
โ
โ
โ
mistral.rs
โ
โ
โ
โ
โ
โ
โ
mlx_lm
30.270/100
15.817/100
โ
โ
58.5
vllm-mlx
40.9
22.89/100
โ
โ
34.5
fidelity (weighted F1, GMRID, vs. one NVIDIA A100 reference)
Stack
0-shot F1
5-shot F1
vllm-nvidia (ref)
0.4094
0.7364
hf_transformers
0.3995
0.7299
llama.cpp
0.4033
0.7377
mlx_lm
0.3970
0.7346
ollama
0.4173
0.4462
omlx
0.3944
0.7379
sglang
0.3955
0.7337
vllm-metal
0.3956
0.7423
vllm-mlx
0.3592
0.7183
Runs 2026-07-05 to 2026-07-09 (splits from different runs): 5 of 9 stacks complete every request at every level; 1 degrade to partial or skip; 1 crash at least once; 2 never start.
chat split
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm-mlx
82.6
203.6
272.3
320 ms
20.5
llama.cpp
99.0
235.4
234.2
4.05 s
20.7
vllm-metal
67.5
131.7
159.9
701 ms
44.4
omlx
108.6
150.7
153.8
4.51 s
19.8
mlx_lm
68.6
68.5
68.6
11.91 s
17.5
hf_transformers
28.9
29.3
29.076/100
39.73 s
12.7
mistral.rs
โ
โ
โ
โ
โ
โ
โ
ollama
โ
โ
โ
โ
โ
โ
โ
sglang
โ
โ
โ
โ
โ
โ
โ
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm-mlx
66.2
159.7
164.2
578 ms
37.4
omlx
50.7
102.6
103.5
5.19 s
21.7
llama.cpp
50.8
85.5
92.2
9.55 s
18.0
mlx_lm
38.1
38.0
38.0
25.76 s
14.4
vllm-metal
24.9
32.3
31.4
4.26 s
42.0
hf_transformers
24.4
24.5
24.5
134.25 s
12.6
mistral.rs
โ
โ
โ
โ
โ
โ
โ
ollama
โ
โ
โ
โ
โ
โ
โ
sglang
โ
โ
โ
โ
โ
โ
โ
fidelity (weighted F1, GMRID, vs. one NVIDIA A100 reference)
Stack
0-shot F1
5-shot F1
vllm-nvidia (ref)
0.6953
0.5949
llama.cpp
0.7003
0.5940
mlx_lm
0.6942
0.5903
omlx
0.6981
0.5929
vllm-metal
0.6947
0.6081
vllm-mlx
0.7006
0.5922
Runs 2026-07-05 to 2026-07-08 (splits from different runs): 5 of 9 stacks complete every request at every level; 1 crash at least once; 3 never start.
chat split
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm-metal
18.4
59.2
87.8
568 ms
42.1
omlx
27.8
56.8
57.6
9.75 s
31.2
llama.cpp
20.8
35.0
35.2
26.53 s
31.2
mlx_lm
14.3
14.3
14.3
50.52 s
33.8
hf_transformers
10.5
10.7
10.6
105.16 s
32.5
mistral.rs
โ
โ
โ
โ
โ
โ
โ
ollama
โ
โ
โ
โ
โ
โ
โ
sglang
โ
โ
โ
โ
โ
โ
โ
vllm-mlx
โ
โ
โ
โ
โ
โ
โ
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
omlx
14.7
48.7
50.3
13.01 s
38.9
llama.cpp
13.0
21.2
21.0
56.63 s
36.6
vllm-metal
8.4
11.8
10.9
9.77 s
40.7
mlx_lm
5.3
5.4
5.3
114.44 s
32.2
hf_transformers
โ
โ
โ
โ
โ
โ
โ
mistral.rs
โ
โ
โ
โ
โ
โ
โ
ollama
โ
โ
โ
โ
โ
โ
โ
sglang
โ
โ
โ
โ
โ
โ
โ
vllm-mlx
โ
โ
โ
โ
โ
โ
โ
fidelity (weighted F1, GMRID, vs. one NVIDIA A100 reference)
Stack
0-shot F1
5-shot F1
vllm-nvidia (ref)
0.8511
0.9128
llama.cpp
0.8497
0.9107
mlx_lm
0.8493
0.9108
omlx
0.8539
0.9120
vllm-metal
0.8512
0.9106
vllm-mlx
0.8530
0.9107
No runs for Qwen3.8-27B on Apple M2 Max yet. Earlier benchmark runs and maintenance case studies. The paper's main audit uses the M5 Pro.
No runs for Qwen3.6-35B-A3B-4bit on Apple M2 Max yet. Earlier benchmark runs and maintenance case studies. The paper's main audit uses the M5 Pro.
machine 64 GB unified memory ยท macOS 26 ยท Metal
run dates 2026-07-05 to 2026-07-09
harness78f885226
Framework versions and update commits for each run are recorded in the maintenance journals; structured version fields appear here once the harness emits them. Earlier benchmark runs and maintenance case studies. The paper's main audit uses the M5 Pro.
per-framework provenance
framework
benchmarked
hf_transformers
chat 2026-07-08 ยท agent 2026-07-09
llama.cpp
chat 2026-07-05 ยท agent 2026-07-05/2026-07-06
mistral.rs
chat 2026-07-05
mlx_lm
chat 2026-07-05 ยท agent 2026-07-05/2026-07-06
ollama
chat 2026-07-05 ยท agent 2026-07-05/2026-07-06
omlx
chat 2026-07-05 ยท agent 2026-07-05/2026-07-06
sglang
chat 2026-07-05 ยท agent 2026-07-06
vllm-metal
chat 2026-07-05 ยท agent 2026-07-05/2026-07-06
vllm-mlx
chat 2026-07-05 ยท agent 2026-07-05/2026-07-06
Run 2026-07-05: 3 of 3 stacks complete every request at every level.
chat split
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm
103.3
633.3
796.0
52 ms
โ
sglang
106.8
592.3
734.1
42 ms
โ
llama.cpp
119.4
157.7
235.3
178 ms
โ
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
sglang
82.6
323.5
384.1
84 ms
โ
vllm
80.2
294.1
342.7
134 ms
โ
llama.cpp
55.3
70.2
79.7
1.90 s
โ
Runs 2026-07-05 to 2026-08-27 (splits from different runs): 3 of 3 stacks complete every request at every level.
chat split
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
sglang
95.0
580.2
935.3
50 ms
โ
vllm
94.2
529.8
680.6
113 ms
โ
llama.cpp
106.5
162.1
376.1
169 ms
โ
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm
82.0
520.4
684.6
93 ms
โ
sglang
79.4
396.4
684.1
66 ms
โ
llama.cpp
77.5
108.5
172.7
409 ms
โ
Runs 2026-07-05 to 2026-08-27 (splits from different runs): 3 of 3 stacks complete every request at every level.
chat split
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm
17.2
136.6
238.5
161 ms
โ
sglang
17.2
129.4
211.7
167 ms
โ
llama.cpp
20.1
29.8
72.8
1.19 s
โ
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=8
tok/s c=16โผ
trend
TTFT p50 c=16
TTFT trend
peak mem GB
vllm
15.7
133.9
225.1
193 ms
โ
sglang
16.0
125.8
201.9
189 ms
โ
llama.cpp
18.1
24.6
51.7
2.43 s
โ
8-bit weights (GGUF Q8_0 for llama.cpp; the mlx-community affine-8bit conversion for vllm-metal and omlx) โ BF16 does not fit this box. Run against three engines only, at concurrency 1/2/4; the other stacks were not attempted, so their absence is not a failure.
Run 2026-08-27: 1 of 1 stacks complete every request at every level.
agent split (~4K-token prompts)
Stack
tok/s c=1
tok/s c=2
tok/s c=4โผ
trend
TTFT p50 c=4
TTFT trend
peak mem GB
vllm
6.1
13.3
24.0
601 ms
โ
No runs for Qwen3.6-35B-A3B-4bit on NVIDIA DGX Spark yet. Complementary serving-performance reference for three engine families shared with the Apple Silicon audit. Memory measurements are not included in this track.
machine GB10 Grace-Blackwell ยท 128 GB unified ยท Linux / CUDA
run dates 2026-07-05 to 2026-08-27
harness78f885226
Framework versions and update commits for each run are recorded in the maintenance journals; structured version fields appear here once the harness emits them. Complementary serving-performance reference for three engine families shared with the Apple Silicon audit. Memory measurements are not included in this track.
per-framework provenance
framework
benchmarked
llama.cpp
chat 2026-07-05 ยท agent 2026-07-05
sglang
chat 2026-07-05 ยท agent 2026-07-05
vllm
chat 2026-07-05 ยท agent 2026-07-05/2026-08-27
Paper & code
Citation
@misc{zhang2026siliconbenchspeedmemoryfidelity,
title={SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops},
author={Ranran Haoran Zhang and Aysa Xuemo Fan and David Munhรก Correia and Alex Cheema and Rui Zhang},
year={2026},
eprint={2609.19169},
archivePrefix={arXiv},
primaryClass={cs.AR},
url={https://arxiv.org/abs/2609.19169},
}