Re: CGOS source on github

David Wu <[email protected]> Thu, 21 Jan 2021 19:01:27 -0500
Newsgroups gmane.games.devel.go
Message-ID <CAGEydYuUsDYE3hX_V1GxDcQ0L5iFk6OJRAgHEnW=wz1WOpw_Jw@mail.gmail.com>
--===============4865101546340931587==
Content-Type: multipart/alternative; boundary="0000000000004fdd3f05b971e4fa"

--0000000000004fdd3f05b971e4fa
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

One tricky thing is that there are some major nonlinearities between
different bots early in the opening that break Elo model assumptions quite
blatantly at these higher levels.

The most noticeable case of this is with Mi Yuting's flying dagger joseki.
I've noticed for example that in particular matchups between different
pairs of bots (e.g. one particular KataGo net as white versus ELF as black,
or one version of LZ as black versus some other version as white), maybe as
many as 30% of games will enter into this joseki and the preferences for
the bots may happen by chance to line up such that consistently they will
play down a path where one side hits a blind spot and begins the game with
an early disadvantage. Each different bot may have different preferences
such that arbitrarily each possible pairing randomly runs into such a trap
or not.

And, having significant early-game temperature in the bot itself doesn't
always help as much as you would think because this particular joseki is so
sharp that a particular bot could easily have such a strong preference for
one path or another (even when it is ultimately wrong) so as to override
any reasonable temperature. Sometimes, adding temperature or extra
randomness simply only mildly changes the frequency of the sequence, or
just varies the time before the joseki and trap/blunder happens anyways.

If games are to begin from the empty board, I'm not sure there's an easy
way around this except having a very large variety of opponents.

One thing that I'm pretty sure would mostly "fix" the problem (in the sense
of producing a smoother metric of general strength in a variety of
positions not heavily affected by just a few key lines) would be to
semi-arbitrarily take a very large sampling of positions from a wide range
of human professional games, from say, move 20, and have bots play starting
from these sampled positions, in pairs once with each color. This would
still include many AI openings, because of the way human pros in the last
3-4 years have quickly integrated and experimented with them, but would
also introduce a lot more variety in general than would occur in any
head-to-head matchup.

This is almost surely a *smaller *problem than simply having enough games
mixing between different long-running bots to anchor the Elo system. And it
is not the only way major nontransitivities can show up, (e.g. ladders).
But to take a leaf from computer Chess, playing from sampled forced
openings seems to be a common practice there and maybe it's worth
considering in computer Go as well, even if it only fixes what is currently
the smaller of the issues.


On Thu, Jan 21, 2021 at 12:01 PM R=C3=A9mi Coulom <[email protected]> w=
rote:

> Thanks for computing the new rating list.
>
> I feel it did not fix anything. The old Zen, cronus, etc.have almost no
> change at all.
>
> So it is not a good fix, in my opinion. No need to change anything to the
> official ratings.
>
> The fundamental problem seems that the Elo rating model is too wrong for
> this data, and there is no easy fix for that.
>
> Long ago, I had thought about using a more complex multi-dimensional Elo
> model. The CGOS data may be a good opportunity to try it. I will try when=
 I
> have some free time.
>
> R=C3=A9mi
> _______________________________________________
> Computer-go mailing list
> [email protected]
> http://computer-go.org/mailman/listinfo/computer-go
>

--0000000000004fdd3f05b971e4fa
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><div></div><div>One tricky thing is that there are some ma=
jor nonlinearities between different bots early in the opening that break E=
lo model assumptions quite blatantly at these higher levels.=C2=A0</div><di=
v><br></div><div>The most noticeable case of this is with Mi Yuting&#39;s f=
lying dagger joseki. I&#39;ve noticed for example that in particular matchu=
ps between different pairs of bots (e.g. one particular KataGo net as white=
 versus ELF as black, or one version of LZ as black versus some other versi=
on as white), maybe as many as 30% of games will enter into this joseki and=
 the preferences for the bots may happen by chance to line up such that con=
sistently they will play down a path where one side hits a blind spot and b=
egins the game with an early disadvantage. Each different bot may have diff=
erent preferences such that arbitrarily each possible pairing randomly runs=
 into such a trap or not.</div><div><br></div><div>And, having significant =
early-game temperature in the bot itself doesn&#39;t always help as much as=
 you would think because this particular joseki is so sharp that a particul=
ar bot could easily have such a strong preference for one path or another (=
even when it is ultimately wrong) so as to override any reasonable temperat=
ure. Sometimes, adding temperature or extra randomness simply only mildly c=
hanges the frequency of the sequence, or just varies the time before the jo=
seki and trap/blunder happens anyways.</div><div><br></div><div>If games ar=
e to begin from the empty board, I&#39;m not sure there&#39;s an easy way a=
round this except having a very large variety of opponents.</div><div><br><=
/div><div>One thing that I&#39;m pretty sure would mostly &quot;fix&quot; t=
he problem (in the sense of producing a smoother metric of general strength=
 in a variety of positions not heavily affected by just a few key lines) wo=
uld be to semi-arbitrarily take a very large sampling of positions from a w=
ide range of human professional games, from say, move 20, and have bots pla=
y starting from these sampled positions, in pairs once with each color. Thi=
s would still include many AI openings, because of the way human pros in th=
e last 3-4 years have quickly integrated and experimented with them, but wo=
uld also introduce a lot more variety in general than would occur in any he=
ad-to-head matchup.</div><div><br></div><div>This is almost surely a <i>sma=
ller </i>problem than simply having enough games mixing between different l=
ong-running bots to anchor the Elo system. And it is not the only way major=
 nontransitivities can show up, (e.g. ladders). But to take a leaf from  co=
mputer Chess, playing from sampled forced openings seems to be a common pra=
ctice there and maybe it&#39;s worth considering in computer Go as well, ev=
en if it only fixes=C2=A0what is currently the smaller of the issues.</div>=
<div>=C2=A0</div></div><br><div class=3D"gmail_quote"><div dir=3D"ltr" clas=
s=3D"gmail_attr">On Thu, Jan 21, 2021 at 12:01 PM R=C3=A9mi Coulom &lt;<a h=
ref=3D"mailto:[email protected]" target=3D"_blank">[email protected]=
m</a>&gt; wrote:<br></div><blockquote class=3D"gmail_quote" style=3D"margin=
:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"=
><div dir=3D"ltr"><div>Thanks for computing the new rating list.</div><div>=
<br></div><div>I feel it did not fix anything. The old Zen, cronus, etc.hav=
e almost no change at all.</div><div><br></div><div>So it is not a good fix=
, in my opinion. No need to change anything to the official ratings.</div><=
div><br></div><div>The fundamental problem seems that the Elo rating model =
is too wrong for this data, and there is no easy fix for that.</div><div><b=
r></div><div>Long ago, I had thought about using a more complex multi-dimen=
sional Elo model. The CGOS data may be a good opportunity to try it. I will=
 try when I have some free time.</div><div><br></div><div>R=C3=A9mi<br></di=
v></div>
_______________________________________________<br>
Computer-go mailing list<br>
<a href=3D"mailto:[email protected]" target=3D"_blank">Computer-g=
[email protected]</a><br>
<a href=3D"http://computer-go.org/mailman/listinfo/computer-go" rel=3D"nore=
ferrer" target=3D"_blank">http://computer-go.org/mailman/listinfo/computer-=
go</a><br>
</blockquote></div>

--0000000000004fdd3f05b971e4fa--

--===============4865101546340931587==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Computer-go mailing list
[email protected]
http://computer-go.org/mailman/listinfo/computer-go

--===============4865101546340931587==--