Wei Huang
Wei Huang is a PhD student at the University of Hong Kong and a Research Scientist Intern at NVIDIA's Efficient AI Group, where he works under the supervision of Prof. Song Han and Prof. Xiaojuan Qi. He is best known for his research on efficient AI systems, including ultra-low-bit quantization for large language models (BiLLM), real-time long video generation infrastructure (LongLive), and efficient long-context reasoning (Tri-Attention). His work has been integrated into NVIDIA TensorRT-LLM, Isaac Lab, and SGLang Diffusion, and he is a core author of multiple open-source projects with over 10,000 GitHub stars collectively.
“Wei Huang and Bohan Zhang, NVIDIA and MIT (equal contribution), Lead authors of the paper”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.