SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops

SiliconBench evaluates local LLM serving through speed, memory, and fidelity. Explore how engines handle concurrent workloads while preserving memory headroom and model quality.

The main audit covers nine Apple Silicon serving engines on chat and agent workloads. A complementary NVIDIA DGX Spark track measures serving performance for three shared engine families.

Speed captures throughput and latency under load; memory tracks the system footprint during serving; and fidelity uses a classification task to check for quality regressions against an NVIDIA reference.

Live results rebuilt automatically from every merged benchmark run

machine model

Single-node serving. Select a machine and model, then click a column header to sort. Read throughput alongside memory and fidelity. โœ• = crashed (<5/100 requests), n/100 = partial run, โ€“ = not measured. Trend lines show each engine's change across the tested concurrency levels; each row is scaled independently, with first-token latency on a log scale. Hover an engine name for run details.