What Comes After GPT Models Across Industries?

TL;DR
AI development after GPT-style systems will expand through multimodal models that work with text, images, video, voice, programming, and proteins. The panel expects generated media to reduce production time and cost, increase the volume and variety of entertainment, enhance human performances, improve productivity across organizations, and enable researchers to design proteins for targeted biological functions.
Transcript
welcome to a another conversation on AI I don't think we get enough of this conversation going I do want to thank Richard AAS and the FI team for really increasing the conversation this year on AI because I think there is no greater uh topic of import on the financial side on the leadership side education Side Medical side it's transforming everyth... Read More
Key Insights
- Generated visual media is expected to replace much of today’s rendered production within five to ten years, according to Prem Akkaraju. He argues that AI can reduce some workloads from thousands of compute hours for a single frame to minutes, removing major time and cost constraints.
- Human creativity is expected to remain the starting point for stories, even when separate AI agents help execute them. Akkaraju hopes people will continue seeking stories that other people want to tell, rather than relying primarily on fully personalized films generated automatically from individual preferences.
- Actors and directors are expected to remain important because capturing a real performance can be faster and creatively valuable. AI may enhance or manipulate a performance after one recorded take, but the physical interaction among a director, camera, and actor is not expected to disappear soon.
- Entertainment output could increase by five to twenty times over the coming decade, according to Akkaraju. AI-generated production may also create more formats of different lengths, including two-minute or twenty-minute experiences, while expanding both the amount of content and the number of participating artists.
- Natural language processing has influenced many other areas of artificial intelligence, according to Richard Socher. His work pursued a single neural network for multiple language tasks, leading toward models that users can direct through prompts instead of relying on separate algorithms for every question.
- Multimodal AI is the proposed next frontier beyond text-centered GPT systems. Models are expected to support seamless inputs and outputs across text, programming, images, video, voice, and sound, allowing conversations and productive tasks to span several forms of information within a unified system.
- Proteins are a promising modality for generative models because they function as fundamental building blocks of biology. Socher says a model could be asked to create a protein designed to bind only to SARS-CoV-2 or to a specific type of cancer in the brain.
- AI-generated proteins have already been synthesized in a wet laboratory, according to Socher’s account of work at Salesforce Research. The generated antibacterial protein was described as 40 percent different from any naturally occurring protein, illustrating how models can propose biological structures beyond known natural examples.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: What comes after GPT-style language models?
The next stage described by the panel is multimodal AI that works across text, programming, images, video, voice, sound, and proteins. Instead of treating natural language processing as an isolated field, these systems can accept and produce several kinds of information. Their applications could include organizational productivity, generated entertainment, visual production, biological research, and the design of proteins for targeted functions.
Q: How will generative AI change film and television production?
Generative AI is expected to shift much visual production from rendering toward generation, reducing both production time and cost. Prem Akkaraju says certain frames in Avatar 2 required 6,000 to 7,000 hours of compute time, while comparable work could potentially be reduced to minutes. He expects AI to accelerate creative execution while preserving meaningful roles for human storytellers, directors, and performers.
Q: Will AI replace actors and directors in filmmaking?
The panel does not predict the near-term disappearance of actors or directors. Akkaraju argues that shooting real photography and capturing an actor’s performance can remain easier and creatively important. The physical relationship between a director, camera, and actor is part of the creative process. AI is more likely to enhance or manipulate a recorded performance after a director obtains a usable take.
Q: Will AI generate completely personalized movies for individuals?
AI may eventually have enough information to generate entertainment around a person’s preferences, but Akkaraju does not present that outcome as desirable. He argues that the creative process should begin with a human who directs tools and separate agents to realize a story. His preferred future preserves the experience of hearing stories that other people deliberately choose to tell.
Q: How much more entertainment content could AI enable?
Akkaraju predicts that five to twenty times more content could be created over the coming decade. He also expects greater variation in duration, with experiences designed for short windows such as two minutes or twenty minutes. Lower time and cost barriers could produce an explosion in content creation and increase the number of artists able to participate.
Q: What is multimodal AI and which formats can it handle?
Multimodal AI refers to models that can work with more than text alone. Socher describes systems that support conversations and outputs involving images, video, voice, sound, programming, and proteins. This expansion allows one model to address different kinds of questions and tasks, extending the prompt-based approach from natural language processing into visual, computational, audio, and biological domains.
Q: How can language models contribute to protein design?
Language models can treat proteins as a generative modality and propose new protein structures for specified functions. Socher gives examples of requesting a protein that binds only to SARS-CoV-2 or to a particular type of brain cancer. Because proteins govern biological functions and interactions, generating targeted proteins could unlock new possibilities in medicine and biological research.
Q: Has an AI-generated protein been tested in a laboratory?
Socher says his team at Salesforce Research created a completely new protein with a language model in 2020 and synthesized it in a wet laboratory. The protein had antibacterial properties and was described as 40 percent different from any naturally occurring protein. He presents this experiment as evidence that generative models can design novel biological structures that can be physically produced.
Summary & Key Takeaways
-
Prem Akkaraju predicts that visual media production will move from rendering toward generation, sharply reducing the time and money required to create film and television. Although AI may manipulate recorded performances after a take, he expects human storytellers, actors, directors, cameras, and physical performances to remain important parts of the creative process.
-
Richard Socher describes the next stage of AI as multimodal, with models accepting and producing text, images, video, voice, sound, programming, and biological information. He identifies proteins as a particularly important modality because models can generate protein designs intended to perform specific functions, including binding to SARS-CoV-2 or particular brain cancers.
-
Kai-Fu Lee introduces his experience across research, major technology companies, investment, open-source development, and generative AI products. Together, the panel frames the decade ahead as broader than language models alone, with AI affecting entertainment, organizational productivity, medicine, education, finance, leadership, and other industries through increasingly capable models and applications.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Peter H. Diamandis 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator