Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The game seemed to be going in AlphaGo's favour when it was half way through. Black (AG) had secured a large area on the top that seemed nearly impossible to invade.

It was amazing to see how Lee Sedol found the right moves to make the invasion work.

This makes me think that if the time for match was three hours instead of two, maybe a professional player will have enough time to read the board deeply enough to find the right moves.



The thing is, Lee Sedol didn't "find the right moves"; the wedge at L11 shouldn't have worked. If black's 79th move had been at L10 instead of K10, Sedol would likely have resigned on the spot.


As far as I can tell from watching Redmond's commentary at the time, there were other options for white in the area.


I was on the other stream; Myungwan Kim and Hajin Lee went way deeper than Redmond typically goes (since they have access to an SGF editor instead of a clumsy demo board, and they aren't performing for the camera as much). They seemed pretty confident in their conclusion that L10 killed white.


Just from Redmond's commentary, killing white in the center is just one of the outcomes that Alphago was impelled to achieve. [for just an arbitrary example] If Alphago killed white in the center but lost the bottom and the top right, it could have still lost. (and Sedol had to acheive multiple objectives also, of course).

Sure, maybe Alphago missed a winning move. But the situation was fluid both strategically and tactically, which might have been why that machine chose it's losing move and moreover, why I don't think it was just a matter of the machine failing to find a kill - especially since I think computers have been able to beat humans on pure tesuji for a while now.


Ah! This matches up with Demis' tweet that black's 79th move was later determined to be a mistake. In MCTS you wouldn't normally go back to moves you've already played unless it's an incomplete information game. However I guess for reinforcement learning (even if it's not actually done during these matches), you would go back and update the estimated values of moves already played, which explains how they know that.


Demis clarified what he meant by that in a subsequent tweet – it wasn't that AlphaGo re-evaluated move 79, it's that the winrate only plummeted after move 87, which was beyond the point of no return: https://twitter.com/demishassabis/status/708934687926804482


But he called move 79 a mistake. So did he learn from a human commentator that it was a mistake, or was that AlphaGo's assessment?

Because it is actually trivial in MCTS to reevaluate moves already taken once you have a better assessment of the positions that follow it.


I would love to hear from the team for sure, but I suspect that they have, by now, fed the game state as of move 78 into AlphaGo, had it look exhaustively at possibilities (way more than in the few minutes it took during the match), and determined that in hindsight move 79 was indeed not an optimal response.


Demis tweeted that while the game was still ongoing.

Anyway I'm pretty sure what happened; I have actually implemented MCTS myself and know that you can trivially update evaluations of old nodes in MCTS as the game continues and you get a better estimate of the line of play actually taken, although you wouldn't normally have a reason to do so. Basically, it can see in hindsight that a move it took was bad, but without additional calculation you wouldn't know whether the alternatives were any better, or also are worse than estimated at the time.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: