Re: CGOS source on github
David Wu <[email protected]> Thu, 21 Jan 2021 19:01:27 -0500
| Newsgroups | gmane.games.devel.go |
|---|---|
| Message-ID | <CAGEydYuUsDYE3hX_V1GxDcQ0L5iFk6OJRAgHEnW=wz1WOpw_Jw@mail.gmail.com> |
--===============4865101546340931587== Content-Type: multipart/alternative; boundary="0000000000004fdd3f05b971e4fa" --0000000000004fdd3f05b971e4fa Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable One tricky thing is that there are some major nonlinearities between different bots early in the opening that break Elo model assumptions quite blatantly at these higher levels. The most noticeable case of this is with Mi Yuting's flying dagger joseki. I've noticed for example that in particular matchups between different pairs of bots (e.g. one particular KataGo net as white versus ELF as black, or one version of LZ as black versus some other version as white), maybe as many as 30% of games will enter into this joseki and the preferences for the bots may happen by chance to line up such that consistently they will play down a path where one side hits a blind spot and begins the game with an early disadvantage. Each different bot may have different preferences such that arbitrarily each possible pairing randomly runs into such a trap or not. And, having significant early-game temperature in the bot itself doesn't always help as much as you would think because this particular joseki is so sharp that a particular bot could easily have such a strong preference for one path or another (even when it is ultimately wrong) so as to override any reasonable temperature. Sometimes, adding temperature or extra randomness simply only mildly changes the frequency of the sequence, or just varies the time before the joseki and trap/blunder happens anyways. If games are to begin from the empty board, I'm not sure there's an easy way around this except having a very large variety of opponents. One thing that I'm pretty sure would mostly "fix" the problem (in the sense of producing a smoother metric of general strength in a variety of positions not heavily affected by just a few key lines) would be to semi-arbitrarily take a very large sampling of positions from a wide range of human professional games, from say, move 20, and have bots play starting from these sampled positions, in pairs once with each color. This would still include many AI openings, because of the way human pros in the last 3-4 years have quickly integrated and experimented with them, but would also introduce a lot more variety in general than would occur in any head-to-head matchup. This is almost surely a *smaller *problem than simply having enough games mixing between different long-running bots to anchor the Elo system. And it is not the only way major nontransitivities can show up, (e.g. ladders). But to take a leaf from computer Chess, playing from sampled forced openings seems to be a common practice there and maybe it's worth considering in computer Go as well, even if it only fixes what is currently the smaller of the issues. On Thu, Jan 21, 2021 at 12:01 PM R=C3=A9mi Coulom <[email protected]> w= rote: > Thanks for computing the new rating list. > > I feel it did not fix anything. The old Zen, cronus, etc.have almost no > change at all. > > So it is not a good fix, in my opinion. No need to change anything to the > official ratings. > > The fundamental problem seems that the Elo rating model is too wrong for > this data, and there is no easy fix for that. > > Long ago, I had thought about using a more complex multi-dimensional Elo > model. The CGOS data may be a good opportunity to try it. I will try when= I > have some free time. > > R=C3=A9mi > _______________________________________________ > Computer-go mailing list > [email protected] > http://computer-go.org/mailman/listinfo/computer-go > --0000000000004fdd3f05b971e4fa Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div></div><div>One tricky thing is that there are some ma= jor nonlinearities between different bots early in the opening that break E= lo model assumptions quite blatantly at these higher levels.=C2=A0</div><di= v><br></div><div>The most noticeable case of this is with Mi Yuting's f= lying dagger joseki. I've noticed for example that in particular matchu= ps between different pairs of bots (e.g. one particular KataGo net as white= versus ELF as black, or one version of LZ as black versus some other versi= on as white), maybe as many as 30% of games will enter into this joseki and= the preferences for the bots may happen by chance to line up such that con= sistently they will play down a path where one side hits a blind spot and b= egins the game with an early disadvantage. Each different bot may have diff= erent preferences such that arbitrarily each possible pairing randomly runs= into such a trap or not.</div><div><br></div><div>And, having significant = early-game temperature in the bot itself doesn't always help as much as= you would think because this particular joseki is so sharp that a particul= ar bot could easily have such a strong preference for one path or another (= even when it is ultimately wrong) so as to override any reasonable temperat= ure. Sometimes, adding temperature or extra randomness simply only mildly c= hanges the frequency of the sequence, or just varies the time before the jo= seki and trap/blunder happens anyways.</div><div><br></div><div>If games ar= e to begin from the empty board, I'm not sure there's an easy way a= round this except having a very large variety of opponents.</div><div><br><= /div><div>One thing that I'm pretty sure would mostly "fix" t= he problem (in the sense of producing a smoother metric of general strength= in a variety of positions not heavily affected by just a few key lines) wo= uld be to semi-arbitrarily take a very large sampling of positions from a w= ide range of human professional games, from say, move 20, and have bots pla= y starting from these sampled positions, in pairs once with each color. Thi= s would still include many AI openings, because of the way human pros in th= e last 3-4 years have quickly integrated and experimented with them, but wo= uld also introduce a lot more variety in general than would occur in any he= ad-to-head matchup.</div><div><br></div><div>This is almost surely a <i>sma= ller </i>problem than simply having enough games mixing between different l= ong-running bots to anchor the Elo system. And it is not the only way major= nontransitivities can show up, (e.g. ladders). But to take a leaf from co= mputer Chess, playing from sampled forced openings seems to be a common pra= ctice there and maybe it's worth considering in computer Go as well, ev= en if it only fixes=C2=A0what is currently the smaller of the issues.</div>= <div>=C2=A0</div></div><br><div class=3D"gmail_quote"><div dir=3D"ltr" clas= s=3D"gmail_attr">On Thu, Jan 21, 2021 at 12:01 PM R=C3=A9mi Coulom <<a h= ref=3D"mailto:[email protected]" target=3D"_blank">[email protected]= m</a>> wrote:<br></div><blockquote class=3D"gmail_quote" style=3D"margin= :0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"= ><div dir=3D"ltr"><div>Thanks for computing the new rating list.</div><div>= <br></div><div>I feel it did not fix anything. The old Zen, cronus, etc.hav= e almost no change at all.</div><div><br></div><div>So it is not a good fix= , in my opinion. No need to change anything to the official ratings.</div><= div><br></div><div>The fundamental problem seems that the Elo rating model = is too wrong for this data, and there is no easy fix for that.</div><div><b= r></div><div>Long ago, I had thought about using a more complex multi-dimen= sional Elo model. The CGOS data may be a good opportunity to try it. I will= try when I have some free time.</div><div><br></div><div>R=C3=A9mi<br></di= v></div> _______________________________________________<br> Computer-go mailing list<br> <a href=3D"mailto:[email protected]" target=3D"_blank">Computer-g= [email protected]</a><br> <a href=3D"http://computer-go.org/mailman/listinfo/computer-go" rel=3D"nore= ferrer" target=3D"_blank">http://computer-go.org/mailman/listinfo/computer-= go</a><br> </blockquote></div> --0000000000004fdd3f05b971e4fa-- --===============4865101546340931587== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ Computer-go mailing list [email protected] http://computer-go.org/mailman/listinfo/computer-go --===============4865101546340931587==--