Songlin Yang
Songlin Yang is a Member of Technical Staff at Thinking Machines Lab, focused on machine learning and language model architectures. She earned her PhD from MIT CSAIL, where she was advised by Professor Yoon Kim and specialized in hardware-aware algorithms for efficient sequence modeling. She is best known for developing Gated Linear Attention transformers, creating the Flash Linear Attention library for hardware-efficient attention kernels, and co-developing DeltaNet, which parallelizes linear transformer training with the delta rule across sequence lengths.
“KDA — from the Kimi Mini Linear paper to the 2.8T mainstream model — took less than one year. So if your genius idea is right, you don't need to wait for a new paradigm. Someone will come along and adopt it.”
Source→“Last year, Mandos' Jichao — that's Peak — also wrote a blog about how to design an infra-friendly harness framework, with the most important part being fully utilizing prefix caching.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.