182: 对话梁琛奇:抖音、猫箱、创业,「他们都搞生产力,我想用 AI 创造开心」
- 01Entertainment is a Structurally Slower, More Constrained Category Than Productivity Tools
- 02The "House-Building vs. Raising-a-Daughter" Framework for Mobile vs. AI Product Design
- 03Building Your Own Model Is Necessary, Not Optional, for AI-Native Entertainment Apps
- 04Three Structural Variables That Make a New AI-Entertainment "Version Answer" Inevitable
- 05Recommendation Systems Never Actually Satisfy True Individual Desire
- 06Human Creativity Remains Irreplaceable in the Mixed AI-Human Content Experience
1. Key Themes
Entertainment is a Structurally Slower, More Constrained Category Than Productivity Tools
Liang argues entertainment products lagged behind productivity tools in the AI wave because they face three binding constraints that had to be cleared first: inference cost, model modality quality, and user habituation. "Every big technology era has its own resource constraint... in the mobile internet era it was traffic... In the AI era it's actually tokens, inference cost" [00:02:54:10]. He notes entertainment sessions are "ten times, even twenty times" longer than tool usage because "it's a killing-time process" rather than a problem-solving one [00:03:24:09], which made ROI impossible until inference costs collapsed. He personally modeled Maopaw's (猫箱/Catbox) unit economics and found ROI turned positive by the end of last year, "and as far as I know, it really did turn positive" [00:03:52:83].
The "House-Building vs. Raising-a-Daughter" Framework for Mobile vs. AI Product Design
Liang's central metaphor for why AI-native product design differs from mobile internet design: mobile products are finite-state systems you must fully specify ("if you don't define something, it won't happen") [01:10:29:33], while AI products require designing a worldview, principles, and goals, then letting the system generate unpredictable but appropriate responses — "like raising a two-year-old daughter... you can't teach her ten thousand things... you can only teach her a framework" [01:10:59:43].
Building Your Own Model Is Necessary, Not Optional, for AI-Native Entertainment Apps
Maopaw's team trained its own story and character models from day one, deliberately orthogonal to general-purpose foundation models. "If you're orthogonal to the current mainline model optimization, you have a very good advantage... how do you construct that orthogonal experience, what kind of things can create a moat on the model layer" [01:14:26:19]. He lists three reasons entertainment models can't just rely on scaling general intelligence: sparse data (new experiences have no existing 1:1 training data), extreme subjectivity (there's no objective right answer for what makes a character interaction "good," only what makes a user happy), and the ability to encode subjective, intuition-driven product taste directly into the model as hidden features that competitors can't easily copy.
Three Structural Variables That Make a New AI-Entertainment "Version Answer" Inevitable
Liang lays out a first-principles case for why a genuinely new entertainment format must emerge: (1) creation-side supply expansion — AI turns "people as a factor of production" into something scalable for the first time, unlocking huge numbers of untapped creative talents; (2) consumption-side new experience — infinite branching interaction plus "the next chapter of recommendation systems," where AI can fit content to an individual rather than the lowest common denominator of millions; (3) removing the human bottleneck in scaling personal interaction — going from "100,000-to-1" (one streamer, many viewers) to "100,000-to-100,000" (everyone gets individualized interaction with a character).
Recommendation Systems Never Actually Satisfy True Individual Desire — AI Can
Liang makes a sharp, under-appreciated point: pre-AI recommendation systems only approximate what a user wants because "that creator is creating for thirty million people... he's finding the greatest common denominator of thirty million people" [00:52:52:79]. He insists this isn't about whether users consciously know what they want: "It's not that he doesn't have that need, it's simple — whether the recommendation is accurate or not, this is something you can see in the data" [00:52:52:79]. AI, by contrast, can build unique context on an individual and create content that fits their specific shape, not just the nearest ellipse to a circle.
Human Creativity Remains Irreplaceable in the Mixed AI-Human Content Experience
Despite AI leverage, Liang insists that pure AI-generated content without human authorship "won't happen" as a dominant paradigm because content needs both "familiar" (empathy bridge) and "strange" (surprise, sourced from diverse human life experience) elements simultaneously. He lists five categories of human work AI struggles to replace: data-scarce domains, non-evaluable/subjective domains, the "originating spark" (defining what an agent's goal even is), embodied intelligence tasks, and — most poetically — the emotional premium of things made specifically for you ("your mom made you a birthday cake... it's not objectively the best cake, but it has emotional premium") [01:02:07:79].
The Danger of Linear Extrapolation in Predicting the Next Big Entertainment Format
Liang explicitly warns against assuming future AI entertainment is just "Character AI and Maopaw extended to more modalities, extended to 3D" — calling this reasonable but almost certainly wrong because it fails to account for how many large, unprecedented variables are converging at once: "if you just do this linear derivation and linear extension... then let me predict, wow, is this planet online really this boring?" [00:56:18:17]
Vertical Integration Beats Being a Pure Applications Company — Even in a "Just an App" Narrative Environment
Liang's team decided, against a market consensus that separates "app companies" from "model companies," to build models themselves from day one because that's where defensible moats live. This is a contrarian bet given VC/market framing during 2023-2024 that model layer and app layer should stay separate.
2. Contrarian Perspectives
Byte-Origin Founders Are Not Automatically Advantaged in AI, and Liang Rejects the "Talk to FA First, Then Decide" Founding Pattern
Liang bluntly criticizes a common ByteDance-alumnus startup pattern: "even some operating methods are: come out first, then don't know what to do, chat with the FA [financial advisor] for a while, then say this direction might be suitable to do... I definitely don't do it this way at all. What I do is I decide what to do, and that's what it is" [01:34:03:41]. This is a direct critique of a widely-assumed advantage (Byte pedigree = strong execution) as potentially masking a lack of independent conviction.
ByteDance's Massive Commercialization Advantage Is a Double-Edged Sword That Can't Be Replicated — And Might Not Even Be Necessary Yet
Liang argues Douyin beat Kuaishou not mainly through product superiority but through "a system that is crushing in this dimension" — near-perfect ROI measurement letting ByteDance invest with total confidence while competitors "see a vast whiteness... not knowing if walking two steps means falling off a cliff" [01:37:28:23]. But he contrarian-ly concludes new entrants shouldn't necessarily rush to build this: it's more important to first prove the experience is worth scaling before obsessing over paid acquisition infrastructure.
Xiaohongshu's Survival Against ByteDance Was Not Luck But a Product of Deliberately "Irrational," Taste-Driven Curation That a Data-Maximizing Culture Like ByteDance's Structurally Cannot Replicate
Liang states Xiaohongshu's community tone was shaped by "non-purely-rational decisions" about what content to promote or kill — decisions ByteDance, whose default mode is "this data is good, so it keeps growing," would likely never make: "This is probably not something Bytedance would do... a more common Bytedance way of doing things is: this data is good, so it keeps growing" [02:25:25:61]. He notes cross-hiring between the two ecosystems has a high "water-and-soil incompatibility" rate — a sharp, rarely-stated observation about organizational culture as a genuine moat.
AI Girlfriend/Boyfriend Apps Will Have Real Value Eventually But Are a Trap Right Now — The Popular Framing Get the Market Backwards
Three years ago, Liang recalls, most people assumed "real-life-simulating" AI companionship apps were the more generalizable, larger market opportunity versus fantastical role-play like Character AI. "The answer is completely not that. The answer is the latter [fantasy role-play] became an independent entertainment platform" [02:34:13:31] — because human faces and relationships are too nuanced ("how many muscle groups are in a human face") for current AI to competently replace, whereas fantasy content's appeal never depended on realism to begin with.
Sora's Buzz Was an Emotional Reaction, Not a Signal of Real Product-Market Fit
Liang is skeptical of the industry's Sora enthusiasm: "at that time there was a lot of a certain specific emotion, called 'wow AI is so amazing' — but this emotion is not demand. After 30 days, remove this emotion, you have to think about whether you yourself still care about this thing" [02:41:06:67]. He also states plainly that Sora's remix feature "hasn't solved this problem" of durable engagement and "wasn't really connected to users' long-term motivation" [02:16:33:89] — a direct critique of a widely-hyped OpenAI product from a competing entertainment-AI founder.
3. Companies Identified
Maopaw (猫箱/Catbox) — Liang's AI role-play/interactive-story app built inside ByteDance's Flow department, launched March 2024. Described as the "version answer" (best current implementation) for AI-native text-based interactive entertainment over the past two-to-three years. Mentioned for reaching hundreds of millions in DAU-scale signals from competitors before launch, and for achieving positive ROI (covering inference/operating costs, excluding pretraining) by end of last year through a mix of ads and gameplay-based monetization (membership, single-play cards, "hear the character's inner voice" paid features). "That time our calculation was, by the end of last year ROI would turn positive... and as far as I know, it really did turn positive" [00:03:52:83].
Douyin (TikTok China) — Liang's first product at ByteDance; used as the extended case study of design excellence, recommendation-driven engagement, and successful internal social-graph expansion. "Douyin, this product, it's like it was designed by a designer... the product character is too obvious" [00:20:29:65].
Songguo Shike (松果时刻) — Liang's first product at Defold, launched September last year: users input a photo plus a narrated story, and AI generates a multi-panel narrative comic. Positioned explicitly as a market-testing "camera for imagination" rather than a final product, used to observe what broad user populations do when given a low-friction creative tool.
Cursor — Cited as proof that even tool-category AI apps still have real growth trajectories and can achieve major outcomes: "it was ultimately... acquired for I don't know, 60 billion dollars" [00:02:24:69] [01:22:47:68] — used as evidence that even a "lesser" outcome relative to Claude Code/Codex is still an extremely strong result.
Character AI (Cartel AI/Karati) — Cited repeatedly as the pioneering AI role-play product and a signal generator that convinced Liang and his team the category (interactive narrative entertainment) had legs, despite never fully cracking monetization itself.
Xingye (星野), Talkie — Chinese AI companion/role-play apps cited alongside Character AI as reaching "several hundred million DAU" scale signals that helped justify Maopaw's launch timing.
Duoshan (多闪) — ByteDance's earlier attempt at a Snapchat-like ephemeral social product, mentioned as a partial but imperfect echo of the "expression, not just connection" insight Liang describes as core to social products.
Roblox — Cited as the clearest existing example of the "hangout as UGC game platform" concept Liang was chasing with his earlier real-time social product at ByteDance: "very close to this platform I just described... he really has this experience — today nothing's going on with me and my friend, I'll just come to Roblox to hang out" [00:12:40:61].
Zenless Zone Zero / 单机派对 (Dandan Party?) — Referenced via Pan Luan's anecdote about his son socializing entirely inside a virtual "island" hangout space with friend lists — used as real-world evidence that Gen Z social behavior has already shifted toward always-on virtual co-presence.
DeepSeek — Not a company Liang worked at, but repeatedly cited as the pivotal catalyst: both for accelerating the "three variables" (cost curve) needed for entertainment AI, and for the emotional/ideological jolt of proving small teams outside giants can do frontier-level work: "I feel like DeepSeek coming out was a fairly clear inflection point and signal" [00:03:52:83], and later, discussing Liang Wenfeng's interviews: "you could sense this person is very, very serious about believing in this thing" [01:19:19:65].
Pinduoduo — Cited via Colin Huang's quote ("Pinduoduo is the Disneyland outside the Fifth Ring Road") as an example of how low-friction gamified mechanics (group-buy "cut a knife") unlock genuine joy/participation value for underserved demographics — used as a model for why "light games" have huge untapped headroom in AI entertainment.
Kuaishou — Extended case study of a company that survived competition against a much larger, resource-advantaged ByteDance because of deep, hard-to-migrate social network effects ("Laotie culture") among its creator/user base, despite making strategic errors (not investing in paid acquisition, sticking with double-column feed instead of full-screen).
Xiaohongshu (小红书) — Cited as a second case of successful defense against ByteDance's Chao Song (可颂) competing product, attributed to deep niche community trust networks ("sisters helping each other") and Bytedance being "too late" to notice the opportunity because it underestimated the graphic/text content market's size.
4. People Identified
Zhang Yiming — Referenced indirectly through the anecdote about ByteDance's early "no one knew recommendation systems" era and Xiang Liang's book, illustrating ByteDance's willingness to make a strong, focused bet once a signal was found.
Alex (Zhu Jun/朱骏) — Former Douyin product leader Liang worked under, described as a "detail freak" ("Alex是一个细节变态") who deeply understood how small changes in key user paths affect entire funnels, and who coined the metaphor of Douyin's interface as "a window, a canvas, a bridge." Liang credits him with shaping his own product philosophy: "he simultaneously has creativity, empathy, imagination, and aesthetic sense... very special, because these are often contradictory" [01:30:37:05].
Juan Jun (卷卷) — Liang's early ByteDance interviewer/manager, praised for consistently encouraging exploration of unconventional ideas rather than dismissing them, and for articulating a clear, confident philosophy that Douyin should cover "all people, all creativity" rather than staying a niche music-video app.
Xiang Liang (向亮) — ByteDance engineer who reportedly self-taught recommendation systems from a book he wrote himself, later became the person Zhang Yiming specifically recruited after reading his work, and who now leads ByteDance's large model efforts — cited as an origin story for ByteDance's "find a signal, then bet hard" pattern.
Liang Wenfeng (梁文峰) — DeepSeek's founder, referenced via his interviews as a source of ideological inspiration for Liang's decision to leave ByteDance and start his own company; specifically his stories about being told as a child that "reading is useless" and his belief that China cannot remain a permanent technological follower.
Pan Luan (潘乱) — Tech commentator referenced for an anecdote about observing his son's virtual social behavior inside a game platform, used to illustrate shifting Gen Z social patterns.
Elon Musk — Quoted regarding SpaceX's mission ("to make the science fiction in science fiction novels no longer fiction") as part of Liang's argument for entertainment/art's civilizational value in inspiring future scientists and engineers.
5. Operating Insights
Build a Parallel "UGC Organization" Inside the Company to Multiplex Product Bets at Low Cost
Liang runs a formal internal structure distinct from his "heavy" product teams: small 1-2 person pods rapidly test a running list of product hypotheses on 2-week-to-1-month cycles, killing or continuing based on fast signal read. "If it's a PGC company, I said we're a UGC company — right now this headcount... being the same as PGC obviously isn't right, logically you should be off by an order of magnitude. If a PGC company has 10, shouldn't your UGC company have 100?" [01:58:52:71] The bottleneck is explicitly headcount of a very specific hybrid profile: people with product sense, some autonomous ownership, and enough coding/vibe-coding fluency to move without heavy engineering support.
Use a Two-Persona Discipline for Innovation: Separate "Diverging Creator" Mode from "Critic/Appraiser" Mode, But Never Let the Critic Retroactively Invalidate the Creative Process Itself
Liang describes an explicit internal practice: generate ideas in pure divergent mode without applying logic filters, then switch to a rigorous critic persona that stress-tests scalability and retention potential. Crucially: "when 10 out of 10 ideas fail, at this point your [critic] persona should not turn around and stab your [creator] persona, saying none of these ideas were reliable ever — because new things emerging is inherently fragile" [01:31:36:51]. The critic's job after failure is to extract learnings and feed them back to the creative persona, not to punish the divergent process.
Don't Bring Big-Company Systems (OKRs, Culture Rituals) Into a Startup Just to Prove You're "Native" — Match Process to the Actual Goal
Liang admits he initially avoided ByteDance-style rigorous review systems because he wanted the startup to feel culturally distinct, influenced by narratives about how relaxed AI labs supposedly are. He later reversed this: "you shouldn't do it because of this mentality of 'we should be different' ... your driving motivation shouldn't be 'this should look different,' but should be 'what actually gets this done'" [02:55:17:91]. He reinstated a strict quarterly review "heartbeat" system specifically because he noticed his own review pressure (and thus rigor) decline without external structure — a rare admission that founders must engineer accountability into themselves deliberately.
Test the Reasonableness of an AI Product Idea by Reverse-Assembling It From Off-the-Shelf Models Before You Ever Train Anything Custom
Liang's stated methodology: use existing market models/agents, combinatorially assemble a prototype experience yourself, and use the demo to gut-check quality before deciding whether it's worth investing in bespoke model training. "If you feel this product is really amazing, you have to be careful — it's possible you're engaging in wishful thinking... because if it's really this good, why hasn't it appeared yet?" [02:58:44:37] Only once the assembled prototype shows real signal do you commit resources to closing the "crude parts" gap via post-training/pretraining investment — inverting the common "build the model first, then find the app" sequence that many model-first competitors follow.
6. Overlooked Insights
The "Familiar vs. Strange" Content Duality Quietly Undercuts the Entire "AI Will Personalize Everything" Narrative
Buried in a dense theoretical passage, Liang makes an argument that most AI-content commentary misses: hyper-personalized AI content generation, taken to its logical extreme, would actually destroy content's appeal, because good content requires "strange" (surprising, genuinely other) material sourced from real human diversity, not just familiarity-optimization. "If you look at something and it's exactly identical to your own life, you won't find it interesting" [00:54:20:59]. This directly contradicts a very common industry assumption (repeated by many AI founders) that hyper-personalization is an unqualified endgame — Liang argues the endgame requires a permanent, non-negotiable human-authored input, not diminishing human involvement over time as AI improves.
The Casual Observation That "90% of User Feedback on Maopaw Was About the Model, Not the UI" Is a Quiet Admission That Application-Layer Differentiation May Be Thinner Than Assumed
When explaining why they had to train their own model, Liang mentions almost in passing: "in this kind of product [Maopaw], user feedback is 90% feedback about the model... it's about the UI, but they say, why does this character talk to me like this, it's driving me crazy" [01:12:28:21]. This is a significant, under-emphasized signal buried mid-conversation: it suggests that in AI-native entertainment products specifically, essentially all of the perceived "product quality" that determines retention is actually model quality — meaning traditional product/UX moats (which dominated the mobile era) may matter far less than founders assume, and the real competitive battle has already fully shifted to who can out-train competitors on subjective, taste-encoded reward models — a much higher and more capital-intensive bar than most entertainment-app entrants are currently prepared for.