Meta's AI research division, developer of the Segment Anything Model (SAM/SAM-2).
“Robot points in the reconstruction are identified by text-prompted segmentation (SAM3) with multi-view voting”
Source→“the methodology relies on 'the JEPA principle of predicting representations rather than reconstructing inputs [31]' (Section IV-A)”
Source→“LeCun's approach constructs a new latent space... the problem is: after predicting a latent state, if you show it to a language model, the language model can't read it. Show it to a video model, the video model also can't read it.”
AI-extracted from podcast / newsletter / paper summaries. May contain errors.