Ranran Haoran Zhang
PhD Candidate, Penn State University
I am a Ph.D. candidate in Computer Science and Engineering at Penn State University, advised by Dr. Rui Zhang. My research focuses on LLM inference and AI systems, connecting research and engineering to make models more useful, dependable, and accessible in practice. My current work spans:
- Inference systems and runtimes. As a core maintainer of vLLM-metal, I develop efficient serving on Apple Silicon, working across attention kernels, scheduling, memory management, and multi-machine execution.
- Speculative decoding and draft-model training. I investigate decoding correctness and contribute to Speculators, improving training’s numerical accuracy, memory efficiency, and speed.
- Evaluation and benchmarking. I developed SiliconBench to evaluate inference engines on serving speed, memory use, and output fidelity under practical workload and hardware constraints.
My broader background includes learning with imperfect annotations, transfer across NLP tasks, and production AI deployment at eBay. Across these areas, I follow meaningful problems into the parts of the stack where they can be solved.
Previously, I obtained my M.S. in Information Management from the University of Illinois Urbana-Champaign, advised by Dr. Heng Ji, and my B.S. in Computer Science from Changsha University of Science & Technology, advised by Dr. Daojian Zeng.
news
| Sep 12, 2026 | New preprint: SiliconBench, our benchmark of LLM serving speed, memory use, and output fidelity on unified-memory desktops. |
|---|---|
| Jun 12, 2026 | New blog post: One Launch for Any Batch: The Binary Search Inside vllm-metal’s Varlen Attention. |