GPT-5 Review: Strong Agents, Mixed Coding Results

16.7K views
•
August 8, 2025
by
IBM Technology
YouTube video player
GPT-5 Review: Strong Agents, Mixed Coding Results

TL;DR

GPT-5 improves reliability, tool calling, structured outputs, planning, and longer task execution, while offering competitively priced core, mini, and nano models. The hosts consider it an impressive release for agentic workloads, but neither expert was ready to replace Claude Code and Opus 4.1 as a daily coding setup after initial testing.

Transcript

With tool calling improving, as Chris was saying, we're not going to save hundreds of tools. We're going to see thousands or tens of thousands, and there's a lot of opportunity for continuous improvement in just the ecosystem alone. And I think we're going to see a lot more substantial improvements in all these areas outside of just the model perfo... Read More

Key Insights

  • GPT-5 is a unified model family consisting of a core model, mini, and nano, with Thinking and Pro modes available under different rate limits. This structure is intended to simplify access while serving workloads that require different levels of capability and cost.
  • The model router is OpenAI’s response to an increasingly complicated ChatGPT model selector. It places routing in front of the consolidated GPT-5 family so users do not need to decide manually among as many separate model choices for each task.
  • GPT-5 shows benchmark improvements, but the hosts did not characterize its intelligence gains as dramatically beyond competing models. They viewed reliability, practical usefulness, and accessibility as more distinctive themes of the release than raw benchmark performance alone.
  • Reduced hallucination is a central reliability improvement highlighted in the discussion. The practical value is that users may be able to place greater trust in model outputs during everyday work, although the hosts’ reactions were based on early access and initial experimentation.
  • GPT-5 is particularly effective at tool invocation, function calling, MCP usage, and structured outputs according to Mihai’s early testing. These improvements make the model attractive for agentic workloads that must select tools correctly and return information in predictable formats.
  • The nano model is presented as unusually capable for its size, particularly for function calling and agentic tasks. Chris said it outperformed many larger models in the market, while the family’s pricing appeared sustainable for workloads that repeatedly invoke models and external tools.
  • Longer task execution is an important sign of progress in GPT-5. Chris reported that a browser agent spent about 20 minutes solving a Murdle logic game successfully, although an earlier attempt required an explicit instruction not to search online for the answer.
  • GPT-5 did not immediately replace Claude Code and Opus 4.1 for the participating developers. Mihai tested it overnight with different tools and MCP, but returned to his previous coding setup for regular work, describing GPT-5 as impressive but not yet his replacement.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What are the main improvements in GPT-5?

GPT-5’s main improvements are stronger reliability, fewer hallucinations, better tool invocation, more dependable function calling, structured outputs, planning, logic, and longer task execution. The release also introduces a unified family with core, mini, and nano models. The hosts emphasized practical utility and accessibility rather than describing the release as an extraordinary jump in raw intelligence.

Q: Is GPT-5 better than Claude Opus 4.1 for coding?

The participating developers did not conclude that GPT-5 had replaced Claude Opus 4.1 for their daily coding work. Chris answered no when asked whether it would become his daily driver. Mihai tested GPT-5 overnight with different tools and MCP, but returned to Claude Code and Opus 4.1 when he began his regular work the next morning.

Q: How does GPT-5 simplify model selection in ChatGPT?

GPT-5 simplifies model selection by consolidating capabilities around one model family and placing a model router in front of it. The router addresses the growing complexity of choosing manually among many models in ChatGPT. The family includes core, mini, and nano variants, while Thinking and Pro modes are available across pricing tiers with different rate limits.

Q: Why is GPT-5 useful for AI agents?

GPT-5 is useful for AI agents because it appears more reliable at selecting and invoking tools, calling functions, using MCP, and producing structured outputs. Mihai attributed these strengths to focused fine-tuning and considered the associated cost sustainable for agentic workloads. Better planning and longer task completion also support agents that must perform multiple connected actions before producing a result.

Q: What is notable about the GPT-5 nano model?

The GPT-5 nano model stood out because Chris considered it exceptionally capable despite its small size. He said it outperformed many large models in the market, especially for agentic work and function calling. Its ability to invoke functions correctly, combined with the family’s competitive API pricing, makes it relevant when workloads do not require the full core model.

Q: Does GPT-5 hallucinate less than earlier models?

Reduced hallucination was presented as one of GPT-5’s most important reliability improvements. The discussion treated this as practically significant because fewer fabricated outputs can make models more trustworthy for everyday work. The transcript does not provide a specific reduction percentage, so the supported conclusion is that hallucinations improved, not that they were eliminated or reduced by a stated amount.

Q: Can GPT-5 complete longer and more complex tasks?

GPT-5 appears better able to complete longer tasks that require planning, logic, and reasoning. Chris tested browser control by asking an agent to solve a Murdle detective game. One successful attempt took about 20 minutes and solved the puzzle, while an earlier attempt searched for the answer online, prompting him to add an instruction that prohibited cheating.

Q: Does GPT-5 mean that AGI has arrived?

The discussion does not conclude that AGI has arrived. Instead, the hosts describe GPT-5 as progress in reliability, planning, tool calling, and longer task execution. They also argue that further advances may come from the surrounding ecosystem, including thousands or tens of thousands of tools, rather than model performance alone. Those combined improvements could move systems closer to an AGI-like moment.

Summary & Key Takeaways

  • OpenAI introduced a GPT-5 family containing core, mini, and nano models, alongside Thinking and Pro modes with different access limits. ChatGPT uses a model router to reduce the complexity of manual model selection. The release emphasizes practical utility, accessibility, reliability, and consolidation more than a dramatic leap in benchmark performance.

  • The experts were particularly impressed by GPT-5’s tool invocation, function calling, structured outputs, and support for agentic workloads involving MCP. They said these capabilities appeared more reliable than earlier GPT models. The smaller models also attracted attention, especially nano, which Chris described as highly capable for function calling at a sustainable cost.

  • Initial coding reactions were positive but qualified. Chris and Mihai did not expect GPT-5 to replace Claude immediately as their daily development tool. Mihai tested GPT-5 with multiple tools and MCP after receiving access, yet returned to Claude Code and Opus 4.1 when beginning his regular work the following morning.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from IBM Technology 📚