// PERSON
Sasha Rush
ROLE RESEARCHERAT CURSORMENTIONS 1LAST SEEN JUNE 4, 2026
// BIO
Researcher at Cursor who explained the hint-token RL training technique used in Composer 2.5.
// RECENT MENTIONS
// SIGNALS
1 SIGNAL
01
product·Dwarkesh·JUNE 4, 2026
“After Cursor injects these hint tokens they run another forward pass — the trajectory itself doesn't change but the hint causes the model to assign lower probability to the error tokens. Cursor then trains the original model to match those probabilities, basically teaching it to downweight these specific mistakes.”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.