How Does AlphaGo Collaborate in Pair Go?

TL;DR
Pair Go places AlphaGo and a professional on each team, with partners alternating moves without communicating. Success depends on recognizing a partner’s intentions, balancing leadership and support, and adapting across different playing styles, making the format a practical demonstration of human and AI collaboration.
Transcript
good morning everybody I'm Andrew Jackson with the American go Association and I'm joined today with Michael Redmond nine down professional from the nihan Ken and T grael from the uh deep mine team um thanks thanks for being here uh so I guess we should start with a question to you here uh could you tell us a little bit about how you came to join t... Read More
Key Insights
- Pair Go is a collaborative format in which two partners alternate moves without discussing their plans. In this match, each team combines AlphaGo with a professional, requiring the human and AI to interpret one another solely through moves placed on the board.
- The central challenge of Pair Go is understanding when to lead and when to follow. Partners may struggle when they pursue different variations, but alignment around the same positional idea can produce coherent and potentially exciting play.
- Pair Go can educate players of different strengths because the weaker partner learns from stronger moves. The stronger partner also learns by adjusting decisions and expressing plans in a way that the weaker player can recognize and support.
- AlphaGo’s early policy network was trained on professional games. Thore Graepel played it during his first day at DeepMind and watched his position decline, despite initially doubting that a neural network could already play Go so effectively.
- AlphaGo can compensate for weak human moves without becoming frustrated. During internal Pair Go testing, amateur DeepMind players repeatedly made mistakes, while AlphaGo attempted to restore coherence before its human partner disrupted another possible long-term plan.
- Gu Li is characterized as a balanced player who is also capable of fighting. Against AlphaGo in the Master Series, his games were comparatively peaceful, and he appeared to be pushed gradually out of contention instead of losing through an immediate tactical collapse.
- Lian Xiao is characterized as an adventurous and strongly fighting-oriented player. His Master Series games against AlphaGo entered complicated battles and ended relatively quickly, yet the commentators suggest his style might align more closely with the newer, more adventurous AlphaGo.
- AlphaGo’s policy and value networks could support professional study by exposing candidate moves and associated winning percentages. Michael Redmond argues that these possibilities and evaluations would help professionals investigate surprising moves and improve their judgment of complex positions.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: How does Pair Go work with AlphaGo and professionals?
Pair Go divides the participants into two teams, with each team containing AlphaGo and one professional player. The human playing black moves first, followed by the human playing white, AlphaGo for black, and AlphaGo for white. That cycle then repeats. Partners cannot communicate, so each participant must infer the team’s developing plan from the moves already played.
Q: Why is communication difficult in AlphaGo Pair Go?
Communication is difficult because AlphaGo and its professional partner cannot explain their intended variations to each other. If the professional chooses a plan AlphaGo does not recognize, or if the professional misunderstands AlphaGo’s move, their strategy may become inconsistent. Effective cooperation therefore depends on reading intentions from the board and recognizing when to lead, follow, or adjust.
Q: What can weaker players learn from Pair Go?
Weaker players can learn by participating directly in positions shaped by a stronger partner. They see stronger moves in context and must determine how to support them on later turns. The format also challenges the stronger player to choose plans that the weaker partner can understand, so learning and adjustment can occur on both sides of the partnership.
Q: How did Thore Graepel first play an early AlphaGo network?
Thore Graepel played an early neural network during his first day at DeepMind after David Silver invited him to an internal exhibition match. Graepel described himself as roughly a weak amateur one-dan player and initially underestimated the network. As the game continued, he watched his position deteriorate. The network had been trained on professional games and later became AlphaGo’s policy network.
Q: How could AlphaGo help professional Go players improve?
AlphaGo could help professionals by revealing moves proposed by its policy network and the values assigned by its value network. Michael Redmond compares the policy network’s output to human intuition and describes the value network as a source of winning percentages. Studying several candidate moves and their evaluations could help professionals explore unexpected possibilities and improve positional judgment.
Q: How do Gu Li and Lian Xiao differ in playing style?
Gu Li is described as a balanced player who can also fight, while Lian Xiao is presented as an adventurous player whose main strength is fighting. Gu Li’s games against AlphaGo were relatively peaceful and lasted longer. Lian Xiao repeatedly entered large, complicated fights, producing more drastic outcomes and games that generally ended comparatively quickly.
Q: Which professional style may cooperate better with AlphaGo?
The commentators do not give a definite answer before the match. Gu Li’s balanced approach had produced games in which he remained competitive longer, but Lian Xiao’s adventurous fighting style was considered potentially closer to the newer AlphaGo seen in the Master Series. The Pair Go match was expected to show which professional could align with AlphaGo more effectively.
Q: What rules and time controls were announced for the Pair Go match?
The match was announced under Chinese Go rules. The referee stated that black would give 7.5 points, each player would have one hour, and there would be one 60-second countdown period. A player could lose on time. If the playing sequence became incorrect, the moves would remain effective and the original order would not be changed.
Summary & Key Takeaways
-
DeepMind’s Thore Graepel recalls playing an early neural network trained on professional games during his first day at the company. Although he was roughly a weak amateur one-dan player, his position steadily declined. That network later became the policy network used in AlphaGo’s development.
-
Pair Go creates two teams, each consisting of AlphaGo and a professional player. Human and AI partners alternate moves in a fixed cycle without communicating. The challenge is to interpret a partner’s plan, decide when to lead or follow, and recover when moves do not fit a shared strategy.
-
The summit pairs AlphaGo with Gu Li and Lian Xiao, two leading Chinese professionals with contrasting styles. Gu Li is described as balanced and capable in fights, while Lian Xiao favors adventurous, complicated battles. Their match explores which style cooperates more effectively with the stronger, more adventurous AlphaGo.
-
The commentators describe AlphaGo as a potential educational tool for professionals. Examining policy-network candidate moves could reveal its move preferences, while value-network percentages could illuminate positional judgment. The match begins under Chinese rules, with Gu Li taking black and opening at a star point before Lian Xiao plays at the three-three point.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Google DeepMind 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator



