Skip to yearly menu bar Skip to main content


Organizers

MLSys 2027

Joseph Gonzalez
General Chair

Joseph Gonzalez

UC Berkeley
Joseph is a Professor in the EECS department at UC Berkeley, a co-director and founding member of the UC Berkeley Sky Computing Lab, RISE Lab, and a member of the Berkeley AI Research (BAIR Group). His research interests span artificial intelligence and data systems and he has a wide range of projects including: large language models (LLMs) that interact with external systems (e.g. tool use, RAG, LLM agents, long context windows) system support for large language model deployment (e.g. inference, serving, RAG, batch) large model finetuning machine learning on edge devices new approaches to cloud computing accelerated deep learning for high-resolution computer vision software platforms for autonomous vehicles In 2020, Joseph launched a new stealth company based on his research to radically simplify the process of deploying advanced analytics and machine learning in modern data driven enterprises. Prior to joining Berkeley, Joseph co-founded Turi Inc. (formerly GraphLab), which was based on his thesis work on the GraphLab and PowerGraph Systems. Turi was eventually acquired by Apple.
Song Han
Program Chair

Song Han

MIT
Song Han is an associate professor with tenure at MIT EECS. He got his PhD from Stanford University advised by Bill Dally. Song pioneered efficient AI computing techniques including “Deep Compression” (pruning, quantization) and the “Efficient Inference Engine,” which first introduced weight sparsity to modern AI chips, making it one of the top-5 most cited papers in the 50-year history of ISCA (1953-2023). His innovations, including TinyML and hardware-aware neural architecture search (Once-for-All Network), have advanced AI model deployment on resource-constrained devices. His recent work on LLM quantization and acceleration (SmoothQuant, AWQ, StreamingLLM) has improved efficiency in LLM inference, adopted by NVIDIA TensorRT-LLM. Song received best paper awards at ICLR'16, FPGA'17, and MLSys'24, the NSF CAREER Award, “35 Innovators Under 35,” IEEE “AI’s 10 to Watch,” and the Sloan Research Fellowship. He developed the open lecture series EfficientML.ai to share advances in efficient ML research.
Rashmi Vinayak
Program Chair

Rashmi Vinayak

Carnegie Mellon University
Rashmi received her Ph.D. from UC Berkeley in 2016 where she worked on resource-efficient fault tolerance for big-data systems, and was a postdoctoral scholar at UC Berkeley's AMPLab/RISELab from 2016-17. Broadly, my research interests lie in the area of computer and networked systems. I take a holistic approach towards solving real-world problems considering both theoretical and systems perspectives. I am interested in designing solutions rooted in fundamental theory and in building systems that employ these solutions and insights to advance the state-of-the-art. In the recent past, my research has focused on fault tolerance, resource efficiency, load balancing and reducing latency in large-scale distributed data storage and caching systems. We designed coding theory based solutions that we proved are theoretically optimal; we also built systems using these solutions and evaluated them on Facebook's data-analytics cluster and on Amazon EC2 showing significant benefits over the state-of-the-art. The solutions that we proposed are now a part of Apache Hadoop 3.0 and under consideration by several companies for use in their storage and analytics products.
Wenming Ye
Sponsor Chair

Wenming Ye

Model AI Corp
Wenming Ye works on product and AI systems at Google, with a focus on GenAI and JAX frameworks. He has been an active organizer and sponsor chair for major machine learning and systems conferences.