Why AI Needs Models Built for Software, Not Chat

TL;DR
AI automates too little practical work because chat-oriented models are not designed to produce reliable outputs that software can consume directly. Diogo Almeida argues that Jev addresses this gap as a System One Model optimized for software integration, calibrated decisions, and intelligence per dollar, with data and task selection taking priority over brute-force compute scaling.
Transcript
Like, how can AI be so unbelievably smart? How can we, like, solve Millennium Prize problems in math, but still not automate even the most basics of works? Like, like, really basic rote stuff, but we have this, like, supercharged engine of automation that just does not have, like, the right plugs and stuff to plug into all of this economically valu... Read More
Key Insights
- AI can solve extraordinarily difficult problems while still failing to automate basic work because powerful models lack suitable interfaces and operating characteristics for economically valuable software tasks. Almeida sees this mismatch as evidence that AI needs models designed specifically for programmatic consumption rather than conversational interaction.
- Jev is TypeSafe’s first large programmable model and an example of what Almeida calls a System One Model. Its intended consumer is code, so the model is designed around machine-native outputs instead of internet text completion, chatbot responses, or ordinary instruction-following behavior.
- TypeSafe’s core goal is to make AI as powerful as possible through deep software integration. The company therefore optimizes not only the model’s external interface but also its internal design for use inside software systems, consistent with the meaning behind the TypeSafe name.
- Jev is optimized for intelligence per dollar, a goal reflected in its name through the reference to Jevons paradox. Almeida treats reliability, cost, calibration, and speed as competing dimensions, with Jev representing an explicit choice to pursue the frontier of useful intelligence relative to cost.
- Calibration is essential when model outputs feed directly into software because a system must represent uncertainty appropriately. Almeida argues that preference optimization can distort probability distributions, causing models to favor responses that appear correct or desirable rather than faithfully expressing their internal confidence.
- RLHF can cause mode dropping or mode collapse by pushing models toward common, preferred, or highly acceptable outputs. According to Almeida, this makes models extremely conservative during long-form generation and helps explain why simple mathematical arguments about error accumulation do not always match empirical model behavior.
- Public benchmark performance is not TypeSafe’s central optimization target. The company instead emphasizes real software utility, intelligence per dollar, and the combination of reliability, calibration, cost, and speed required when AI functions as infrastructure rather than as a standalone conversational product.
- The right task and data can matter more than simply scaling compute. Almeida says he would not pre-train a model from scratch even with a billion dollars, reflecting TypeSafe’s view that careful problem selection and model specialization are more important than pursuing scale without a software-specific objective.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: What is Jev and how is it different from a chatbot?
Jev is TypeSafe’s first large programmable model, which Diogo Almeida also describes as a System One Model. It is built for code to consume its outputs directly, unlike pre-trained language models focused on internet text completion or RLHF models focused on replying to instructions. Its internal design and external behavior are optimized for integration into software systems.
Q: Why does AI struggle to automate basic real-world work?
Almeida argues that AI can display extraordinary intelligence while lacking the right connections to practical, economically valuable work. Models designed for text completion or chat do not automatically provide the reliability, calibration, cost profile, and structured behavior required by software. The missing piece is therefore not only greater intelligence, but a model architecture and interface designed for automated programmatic use.
Q: What is a System One Model according to TypeSafe?
A System One Model is TypeSafe’s proposed class of machine-native, large programmable models whose outputs are intended for direct consumption by code. The name is not presented as final, and Almeida avoids limiting the category to decision models because he expects its scope to extend further. Jev is the company’s first released example of this software-oriented model class.
Q: What does intelligence per dollar mean for Jev?
Intelligence per dollar is Jev’s primary optimization objective. Rather than claiming that reliability, cost, calibration, and speed can all be maximized independently, Almeida describes machine learning as a field of trade-offs. Models carrying the Jev name are intended to remain at the frontier of useful intelligence relative to their cost, with other design choices organized around that priority.
Q: Why is calibration important for AI used in software?
Calibration matters because software needs model outputs whose confidence meaningfully reflects uncertainty. Almeida argues that models optimized around human preferences may produce the response that seems most acceptable or likely instead of preserving an accurate probability distribution. That distortion becomes especially consequential when code consumes the result automatically and cannot rely on a person to interpret tone, hesitation, or hidden uncertainty.
Q: How can RLHF cause mode collapse in language models?
Almeida says RLHF can push models toward outputs people prefer to hear or regard as most likely, producing mode dropping or mode collapse. Minority possibilities receive less representation while common, apparently correct responses dominate. This conservatism can support long strings that look coherent, but it can also damage calibration by concealing the full distribution of possible outputs and the model’s underlying uncertainty.
Q: Why does TypeSafe avoid optimizing for public benchmarks?
TypeSafe avoids treating public benchmarks as its central objective because the company is building models for practical software consumption. Its stated priorities include intelligence per dollar, reliability, calibration, cost, and speed. The broader argument is that strong benchmark results do not by themselves establish that a model can function effectively as dependable infrastructure inside real applications and automated workflows.
Q: Why would Almeida avoid pre-training a model from scratch?
Almeida says that even with a billion dollars, he would not pre-train a model from scratch. The position reflects his emphasis on choosing the right task and data instead of assuming that additional compute or full-scale pre-training is the best route. TypeSafe’s approach focuses on specialized, software-oriented models and useful trade-offs rather than pursuing scale as an objective by itself.
Summary & Key Takeaways
-
Jev is TypeSafe’s first large programmable model, also described as a System One Model. Unlike language models trained to autocomplete internet text or follow conversational instructions, it is designed for outputs consumed directly by code. TypeSafe optimizes the model’s internals and external behavior around integration with software rather than chat.
-
TypeSafe prioritizes intelligence per dollar while balancing reliability, cost, calibration, and speed. Almeida argues that machine learning development consists of trade-offs, so Jev’s defining objective is not maximum performance on every dimension. The company also rejects optimizing around public benchmarks and emphasizes selecting appropriate tasks and data over merely expanding compute.
-
Almeida connects weaknesses in RLHF to mode dropping, also called mode collapse. Optimization for preferred or likely responses can make models overly conservative and obscure their actual uncertainty. That behavior matters when AI becomes software infrastructure, where outputs must be calibrated, reliable, and suitable for automated consumption rather than merely persuasive to a human reader.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Latent Space 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator