Re: Monte-Carlo Tree Search as Regularized Policy Optimization

David Wu <[email protected]> Sun, 19 Jul 2020 09:55:34 -0400
Newsgroups gmane.games.devel.go
Message-ID <CAGEydYsDJcTjKSce9eHuU-sTPabU090LTDWnTJ8f4ZC-pQQRjA@mail.gmail.com>
--===============7139065486792297160==
Content-Type: multipart/alternative; boundary="0000000000004627ab05aacbbf40"

--0000000000004627ab05aacbbf40
Content-Type: text/plain; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

I imagine that at low visits at least, "ACT" behaves similarly to Leela
Zero's "LCB" move selection, which also has the effect of sometimes
selecting a move that is not the max-visits move, if its value estimate has
recently been found to be sufficiently larger to balance the fact that it
is lower prior and lower visits (at least, typically, this is why the move
wouldn't have been the max visits move in the first place). It also scales
in an interesting way with empirical observed playout-by-playout variance
of moves, but I think by far the important part is that it can use
sufficiently confident high value to override max-visits.

The gain from "LCB" in match play I recall is on the very very rough order
of 100 Elo, although it could be less or more depending on match conditions
and what neural net is used and other things. So for LZ at least,
"ACT"-like behavior at low visits is not new.


On Sun, Jul 19, 2020 at 5:39 AM Kensuke Matsuzaki <[email protected]>
wrote:

> Hi,
>
> I couldn't improve leela zero's strength by implementing SEARCH and ACT.
> https://github.com/zakki/leela-zero/commits/regularized_policy
>
> 2020=E5=B9=B47=E6=9C=8817=E6=97=A5(=E9=87=91) 2:47 R=C3=A9mi Coulom <remi=
[email protected]>:
> >
> > This looks very interesting.
> >
> > From a quick glance, it seems the improvement is mainly when the number
> of playouts is small. Also they don't test on the game of Go. Has anybody
> tried it?
> >
> > I will take a deeper look later.
> >
> > On Thu, Jul 16, 2020 at 9:49 AM Ray Tayek <[email protected]> wrote:
> >>
> >>
> https://old.reddit.com/r/MachineLearning/comments/hrzooh/r_montecarlo_tre=
e_search_as_regularized_policy/
> >>
> >>
> >> --
> >> Honesty is a very expensive gift. So, don't expect it from cheap peopl=
e
> - Warren Buffett=EF=BB=BF
> >> http://tayek.com/
> >>
> >> _______________________________________________
> >> Computer-go mailing list
> >> [email protected]
> >> http://computer-go.org/mailman/listinfo/computer-go
> >
> > _______________________________________________
> > Computer-go mailing list
> > [email protected]
> > http://computer-go.org/mailman/listinfo/computer-go
>
>
>
> --
> Kensuke Matsuzaki
> _______________________________________________
> Computer-go mailing list
> [email protected]
> http://computer-go.org/mailman/listinfo/computer-go
>

--0000000000004627ab05aacbbf40
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">I imagine that at low visits at least, &quot;ACT&quot; beh=
aves similarly to Leela Zero&#39;s &quot;LCB&quot; move selection, which al=
so has the effect of sometimes selecting a move that is not the max-visits =
move, if its=C2=A0value estimate has recently been found to be sufficiently=
 larger to balance the fact that it is lower prior and lower visits (at lea=
st, typically, this is why the move wouldn&#39;t have been the max visits m=
ove in the first place). It also scales in an interesting way with empirica=
l observed playout-by-playout variance of moves, but I think by far the imp=
ortant part is that it can use sufficiently confident high value to overrid=
e max-visits.<div><br></div><div>The gain from &quot;LCB&quot; in match pla=
y I recall is on the very very rough order of 100 Elo, although it could be=
 less or more depending on match conditions and what neural net is used and=
 other things. So for LZ at least, &quot;ACT&quot;-like behavior at low vis=
its is not new.</div><div><br></div></div><br><div class=3D"gmail_quote"><d=
iv dir=3D"ltr" class=3D"gmail_attr">On Sun, Jul 19, 2020 at 5:39 AM Kensuke=
 Matsuzaki &lt;<a href=3D"mailto:[email protected]">[email protected]</=
a>&gt; wrote:<br></div><blockquote class=3D"gmail_quote" style=3D"margin:0p=
x 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex">Hi=
,<br>
<br>
I couldn&#39;t improve leela zero&#39;s strength by implementing SEARCH and=
 ACT.<br>
<a href=3D"https://github.com/zakki/leela-zero/commits/regularized_policy" =
rel=3D"noreferrer" target=3D"_blank">https://github.com/zakki/leela-zero/co=
mmits/regularized_policy</a><br>
<br>
2020=E5=B9=B47=E6=9C=8817=E6=97=A5(=E9=87=91) 2:47 R=C3=A9mi Coulom &lt;<a =
href=3D"mailto:[email protected]" target=3D"_blank">[email protected]=
om</a>&gt;:<br>
&gt;<br>
&gt; This looks very interesting.<br>
&gt;<br>
&gt; From a quick glance, it seems the improvement is mainly when the numbe=
r of playouts is small. Also they don&#39;t test on the game of Go. Has any=
body tried it?<br>
&gt;<br>
&gt; I will take a deeper look later.<br>
&gt;<br>
&gt; On Thu, Jul 16, 2020 at 9:49 AM Ray Tayek &lt;<a href=3D"mailto:rtayek=
@ca.rr.com" target=3D"_blank">[email protected]</a>&gt; wrote:<br>
&gt;&gt;<br>
&gt;&gt; <a href=3D"https://old.reddit.com/r/MachineLearning/comments/hrzoo=
h/r_montecarlo_tree_search_as_regularized_policy/" rel=3D"noreferrer" targe=
t=3D"_blank">https://old.reddit.com/r/MachineLearning/comments/hrzooh/r_mon=
tecarlo_tree_search_as_regularized_policy/</a><br>
&gt;&gt;<br>
&gt;&gt;<br>
&gt;&gt; --<br>
&gt;&gt; Honesty is a very expensive gift. So, don&#39;t expect it from che=
ap people - Warren Buffett=EF=BB=BF<br>
&gt;&gt; <a href=3D"http://tayek.com/" rel=3D"noreferrer" target=3D"_blank"=
>http://tayek.com/</a><br>
&gt;&gt;<br>
&gt;&gt; _______________________________________________<br>
&gt;&gt; Computer-go mailing list<br>
&gt;&gt; <a href=3D"mailto:[email protected]" target=3D"_blank">C=
[email protected]</a><br>
&gt;&gt; <a href=3D"http://computer-go.org/mailman/listinfo/computer-go" re=
l=3D"noreferrer" target=3D"_blank">http://computer-go.org/mailman/listinfo/=
computer-go</a><br>
&gt;<br>
&gt; _______________________________________________<br>
&gt; Computer-go mailing list<br>
&gt; <a href=3D"mailto:[email protected]" target=3D"_blank">Compu=
[email protected]</a><br>
&gt; <a href=3D"http://computer-go.org/mailman/listinfo/computer-go" rel=3D=
"noreferrer" target=3D"_blank">http://computer-go.org/mailman/listinfo/comp=
uter-go</a><br>
<br>
<br>
<br>
-- <br>
Kensuke Matsuzaki<br>
_______________________________________________<br>
Computer-go mailing list<br>
<a href=3D"mailto:[email protected]" target=3D"_blank">Computer-g=
[email protected]</a><br>
<a href=3D"http://computer-go.org/mailman/listinfo/computer-go" rel=3D"nore=
ferrer" target=3D"_blank">http://computer-go.org/mailman/listinfo/computer-=
go</a><br>
</blockquote></div>

--0000000000004627ab05aacbbf40--

--===============7139065486792297160==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Computer-go mailing list
[email protected]
http://computer-go.org/mailman/listinfo/computer-go

--===============7139065486792297160==--