What Are the Four Main Battles in the AI Stack?

TL;DR
AI competition at the end of 2023 centered on data access and quality, scarce computing resources, multimodal products, and retrieval infrastructure. Publishers pursued licensing deals or litigation, inference providers cut prices, specialized media tools faced broad multimodal models, and developers debated vector databases, LangChain, LlamaIndex, and the operational work required beyond compelling demos.
Transcript
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in Residence at Decibel Partners, and today I'm joined just by my co-host, Swyx, for a new podcast format. Yeah. And it's a bit uncomfortable because we have to just stare into each other's  eyes lovingly. But in our end of year survey last year, a lot of liste... Read More
Key Insights
- The data war is a conflict over access, attribution, licensing, and the quality of training material. Journalists, writers, artists, researchers, startups, and synthetic-data researchers have different interests in determining how published or generated content may enter model-training datasets.
- Low-background tokens are human-written internet content created before model-generated text became widespread. New Common Crawl updates may contain uncertain mixtures of human and machine writing, making data filtering, provenance checks, and basic detection work increasingly important for model builders.
- First-party data creation is a potential response to declining confidence in public datasets. The discussion suggests that employees could produce small numbers of question-and-answer pairs each day, collectively creating a large, distinctive dataset for instruction tuning or other model-development work.
- The GPU and inference war is driven by unequal access to computing resources and intense provider competition. Mixtral output-token pricing reportedly dropped from about $2 to $0.27 per million tokens in one week, alongside benchmark disputes among inference providers.
- Alternative architectures and computing platforms are attempts to extract more value from limited hardware. Mamba and RWKV represent architectural research, while Modular, tinycorp, and Apple MLX reflect efforts to move workloads away from Nvidia or use available computing resources more efficiently.
- The multimodality war pits focused products against broad models that address several media types. Midjourney, AssemblyAI, Replicate, and Suno advanced image, speech, model-serving, and music capabilities while OpenAI and Google continued developing wider models capable of competing across categories.
- The RAG and operations war concerns whether vector databases and orchestration frameworks are necessary for production AI systems. Developers debated vector-database requirements, watched adoption of turbopuffer, and compared LangChain with LlamaIndex and its step-wise agent-execution approach.
- Code models attracted less visible competition than general reasoning, function calling, and multimodality despite code's importance. Workflow fragmentation and integration friction may explain why engineers often remain with VS Code, Cursor, GitHub, or another tool that already produces satisfactory results.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: What were the four main battles in the AI stack?
The four main battles were the data war, the GPU and inference war, the multimodality war, and the RAG and operations war. They covered access to trustworthy training material, competition for computing resources and cheaper inference, rivalry between specialized media tools and broad models, and disputes over vector databases, orchestration frameworks, retrieval, and production infrastructure.
Q: Why was training data becoming harder to trust?
New internet content could no longer be assumed to come entirely from people because models were increasingly producing text that might enter Common Crawl and other datasets. The discussion compares older human-written material to low-background steel. This uncertainty creates more data-engineering work, including provenance checks, filtering, and searches for obvious phrases that reveal model-generated content.
Q: How can companies create distinctive first-party AI data?
Companies can ask people within their organizations to create small quantities of useful training material regularly. The example offered is having each person write five question-and-answer pairs. Aggregated across a workforce, those contributions could form a large, unique dataset rather than relying blindly on datasets supplied by Common Crawl, EleutherAI, or other external sources.
Q: What defined the GPU and inference war in late 2023?
The GPU and inference war involved unequal access to computing resources, rapidly falling inference prices, benchmark disputes, and attempts to use hardware more efficiently. Mixtral output pricing reportedly began around $2 per million tokens and fell to $0.27 within a week. Mamba, RWKV, Modular, tinycorp, and Apple MLX represented additional architectural or platform responses.
Q: How did specialized multimodal companies compete with broad AI models?
Specialized companies improved individual capabilities in images, speech, music, and model serving. Midjourney soft-launched version 6 and a web interface, AssemblyAI raised a $50 million Series C, Replicate raised a $40 million Series B, and Suno emerged from stealth. At the same time, OpenAI and Google worked on broad models that could compete across several modalities.
Q: Why were vector databases still debated in RAG systems?
Developers disagreed about whether retrieval-augmented generation systems required a dedicated vector database. Early in the year, many vector databases appeared, but the hosts expected fewer new entrants afterward. Turbopuffer was highlighted as a notable serverless exception attracting technically sophisticated adopters, while LangChain and LlamaIndex continued competing over orchestration and agent-execution workflows.
Q: Why did code models appear less competitive than other AI categories?
Code models appeared to receive less money, attention, and experimentation than general function calling, reasoning, and multimodal systems, despite the importance of programming. The suggested reason was workflow fragmentation and integration difficulty. Many engineers already used VS Code, Cursor, GitHub, or another satisfactory setup, so trying a newly released model required effort without an obvious benefit.
Q: Why might AI engineers end up doing data engineering?
Production systems expose work that polished demonstrations can hide. Once an AI application depends on reliable proprietary information, engineers must clean data, establish semantic layers, manage retrieval, inspect provenance, and maintain operational pipelines. The hosts therefore suggest that people entering AI engineering may eventually find themselves performing traditional machine-learning or data-engineering tasks, including the unglamorous work swept aside in demos.
Summary & Key Takeaways
-
The data war concerns who may use published material, under what terms, and with what attribution. OpenAI signed publishing partnerships while the New York Times pursued litigation, Apple offered publishers data contracts, and researchers explored synthetic or first-party data as internet datasets became increasingly difficult to trust as purely human-written.
-
The GPU and inference war combines scarce computing access, aggressive pricing, benchmark disputes, alternative model architectures, and attempts to reduce dependence on Nvidia. Mixtral output pricing reportedly fell from about $2 to $0.27 per million tokens within a week, illustrating how quickly inference providers were competing on cost and performance.
-
The multimodality and RAG/Ops wars concern both product breadth and production infrastructure. Specialized image, speech, music, and model-serving companies improved individual capabilities while OpenAI and Google pursued broad models. Meanwhile, engineers debated vector databases, retrieval frameworks, agent execution, and the data engineering hidden behind polished AI demonstrations.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Latent Space 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator