Thomas Sohmers
“That forward pass, that inference portion of it is heavily, heavily memory-bound due to the fact that basically for every single token that's generated”
Source→“GPT 5.6... could only do this about 70% of the time. GPT 6 Astra does it like over 95% of the time”
Source→“DeepSeek beginning of 2025 with DeepSeek v3... [got a] massive decrease in KV cache size with multi-head latent attention... none of the major US companies... are definitely not using MLA”
Source→“I believe in Oracle's business model and ability to execute... a whole lot more than United States government”
Source→“NVIDIA is actually decreasing the amount of memory per device... based on the market memory conditions”
Source→“Mercor and Surge — Data-labeling/data-economy companies cited as potentially becoming enormous businesses (~$200B)”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.