Re: Do we even know what a good AI Optimization strategy would be for Wikipedia? Re: Re: Google Zero is coming [was Re: Wikipedia at 25: A Wake-Up Call (essay)]

Gerard Meijssen via Wikimedia-l <[email protected]> Sun, 12 Jul 2026 08:34:36 +0200
Newsgroups gmane.org.wikimedia.foundation
Message-ID <CAO53wxXRa4Li7aWaK4kXcSqXDqjX-oxuAZp+w9o5RoC4JRXDtg@mail.gmail.com>
--===============5562710496409123862==
Content-Type: multipart/alternative; boundary="000000000000aba9000656642b6d"

--000000000000aba9000656642b6d
Content-Type: text/plain; charset="UTF-8"

Hoi,
A few questions:

   - how expensive is it to add all references, undoubled, from all our
   projects to Wikidata AND have them link to where they are used
   - how expensive is it to check if these references are still there and
   link to content at archive.org
   - how expensive is it to create a search engine with the subjects and
   the titles what can be searched for
   - consider this an experiment, what would happen if we make it available
   to the public without fanfare?
   - how expensive would it be to allow a public to add links to youtube
   and whatever when they are linked to subjects [1]


   - how expensive is it to flag to Wikipedias and other projects that
   lists are likely incorrect based on information that we have in the wiki
   sphere?

[1] I watched How gut microbes keep us healthy as we age | Tim Spector &
Nicola Segata - YouTube <https://www.youtube.com/watch?v=d_3dfXbVl2k> could
link it to these two scientists but I am adding citations for the key
paper.. Gut micro-organisms associated with health, nutrition and dietary
interventions - Scholia <https://qlever.scholia.wiki/work/Q140511862>

Thanks,
      GerardM

On Sat, 11 Jul 2026 at 22:59, Michael Snow via Wikimedia-l <
[email protected]> wrote:

> On 7/11/2026 10:43 AM, Alex Stinson via Wikimedia-l wrote:
>
> *Interfaces are cheap, informed curators are expensive*
>
> Agentic AI, Coding tools and LLMs are making the cost of new interfaces *extremely
> cheap, *so cheap that I am commissioning a complex knowledge repository
> for a fraction of the cost and time it would take otherwise. Interface
> projects like WikiProject Med's offline medical Wikipedia App, which used
> to take several months of highly specialized software development, can now
> be spun up in a long weekend with a Claude Code Max subscription. Sage's
> tool is a perfect example of this: perfect for a small market of users,
> unlikely to be a "headline" Wikimedia tactic for getting in front of users,
> because Youtube and AI search interfaces already do this exact thing.
>
> What is not cheap for all of the other platforms (and is often paid for by
> ads), but which we have in abundance, is motivated humans who can
> continue curating the knowledge (and more than 25 years of experiments on
> facilitating knowledge equity focused gap filling). The questions we need
> to address are:
>
>    - Do the curators understand the future of distribution to other
>    humans we need to be building for across multiple future internet
>    scenerios?
>    - Do the curation practices serve diverse forms of access (languages,
>    geographies, topics) from that distribution?
>    - Do the curators understand demand and how curation choices affect
>    distribution to that demand?
>
> These three questions all implicitly require some kind of training and
> education for curators, on matters not directly affecting curation
> activity. Our volunteers are motivated, yes, but what motivates is the
> curation they enjoy, not necessarily tangential considerations. People
> contribute because "I see a typo", "I'm collecting bits of trivia in this
> space", "This leaves out a key fact", or "That's just flat wrong", and the
> wiki makes them all (relatively) easy to fix. For us, it's "cheap" as in
> we've lowered the investment required to participate. Some curators will be
> willing to invest additional engagement on big-picture issues, so it may
> nevertheless be helpful to offer such training, but there's still
> investment cost on both sides and I do not see it moving the needle when
> focusing on the competitive landscape.
>
>
>    - Can we recruit the next generation of curators who don't "assume"
>    that the pageview metric is the reason we contribute?
>
> I question whether that assumption is in fact prevalent. Even in the world
> of professional media, those paid to curate or generate content often
> resist the implications, whether out of principled beliefs or because they
> fear impact to their livelihood. This conversation has emphasized such
> metrics because those participating care about reaching a broader audience.
> But in terms of what motivates curation at a basic level, once it is
> established that there is some audience out there, I don't know that the
> order of its magnitude matters all that much. Audience demand results in
> more information to curate and more space in which to work, and I think
> that motivates equally with reach if not more. People may find it as
> satisfying to work on en:Wedding of Taylor Swift and Travis Kelce, even
> though en:Taylor Swift and en:Travis Kelce undoubtedly have higher
> pageviews.
>
> To use the wheat production infographic as another illustration,
> personally I care less about the interface behavior than what I could do to
> help better curate. Suppose I had information about how much wheat Sudan
> produced in 2011 (currently "no data"), or wanted to address the
> anachronistic use of present-day national boundaries going back all the way
> to 1961? Where would I even start? Otherwise, maybe it increases the visual
> appeal of one article a bit, and we could even reproduce that across a
> bunch of articles, but we'd also be throwing up massive barriers to
> maintenance.
>
> What we most need is to simplify, to make it easier for volunteers to
> steer in the direction we would like them to go. Then they can self-select
> according to what meets their interests. I believe the success of Wikidata
> so far - "success" in the sense that its potential aligns more with "early
> Wikipedia" than "early Wikiversity" - is in large part because it provides
> a vast new space for curation, and has been integrated in a way that makes
> it easy for those interested to go and participate, while it is
> simultaneously easy to ignore for those who are less interested.
>
> --Michael Snow
> _______________________________________________
> Wikimedia-l mailing list -- [email protected], guidelines
> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and
> https://meta.wikimedia.org/wiki/Wikimedia-l
> Public archives at
> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/ZOHA3G3WMEDAGX6DUPOE7QCHK2HRYNWU/
> To unsubscribe send an email to wikimedia-l-leave-RusutVdil2icGmH+5r0DM0B+6BGkLq7r@public.gmane.org

--000000000000aba9000656642b6d
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">Hoi,<div>A few questions:</div><div><ul><li>how expensive =
is it to add all references, undoubled, from all our projects to Wikidata A=
ND have them link to where they are used</li><li>how expensive is it to che=
ck if these references are still there and link to content at <a href=3D"ht=
tp://archive.org">archive.org</a></li><li>how expensive is it to create a s=
earch engine with the subjects and the titles what can be searched for</li>=
<li>consider this an experiment, what would happen if we make it available =
to the public without fanfare?</li><li>how expensive would it be to allow a=
 public to add links to youtube and whatever when they are linked to subjec=
ts [1]</li></ul><div><ul><li>how expensive is it to flag=C2=A0to Wikipedias=
 and other projects that lists are likely incorrect based on information th=
at we have in the wiki sphere?</li></ul></div></div><div>[1] I watched=C2=
=A0<a href=3D"https://www.youtube.com/watch?v=3Dd_3dfXbVl2k" style=3D"backg=
round-color:transparent">How gut microbes keep us healthy as we age | Tim S=
pector &amp; Nicola Segata - YouTube</a>=C2=A0could link it to these two sc=
ientists but I am adding citations for the key paper..=C2=A0<a href=3D"http=
s://qlever.scholia.wiki/work/Q140511862" style=3D"background-color:transpar=
ent">Gut micro-organisms associated with health, nutrition and dietary inte=
rventions - Scholia</a></div><div><br></div><div>Thanks,</div><div>=C2=A0 =
=C2=A0 =C2=A0 GerardM</div></div><br><div class=3D"gmail_quote gmail_quote_=
container"><div dir=3D"ltr" class=3D"gmail_attr">On Sat, 11 Jul 2026 at 22:=
59, Michael Snow via Wikimedia-l &lt;<a href=3D"mailto:[email protected]=
kimedia.org">[email protected]</a>&gt; wrote:<br></div><block=
quote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1=
px solid rgb(204,204,204);padding-left:1ex"><u></u>

 =20
   =20
 =20
  <div>
    <div>On 7/11/2026 10:43 AM, Alex Stinson via
      Wikimedia-l wrote:<br>
    </div>
    <blockquote type=3D"cite">
     =20
      <div dir=3D"ltr">
        <div><b>Interfaces are cheap, informed curators are expensive</b><b=
r>
          <br>
          Agentic AI, Coding tools and LLMs are making the cost of new
          interfaces=C2=A0<i>extremely cheap,=C2=A0</i>so cheap that I am
          commissioning a complex knowledge repository for a fraction of
          the cost and time it would take otherwise. Interface projects
          like WikiProject Med&#39;s offline medical Wikipedia App, which
          used to take several months of highly specialized software
          development, can now be spun up in a long weekend with a
          Claude Code Max subscription. Sage&#39;s tool is a perfect exampl=
e
          of this: perfect for a small market of users, unlikely to be a
          &quot;headline&quot; Wikimedia tactic for getting in front of use=
rs,
          because Youtube and AI search interfaces already do this exact
          thing.<br>
          <br>
          What is not cheap for all of the other platforms (and is often
          paid for by ads), but which we have in abundance, is motivated
          humans who can continue=C2=A0curating the knowledge (and more tha=
n
          25 years of experiments on facilitating knowledge equity
          focused gap filling). The questions we need to address are:<br>
          <ul>
            <li style=3D"margin-left:15px"><span style=3D"background-color:=
transparent">Do the curators
                understand the future of distribution to other humans we
                need to be building for across multiple future=C2=A0</span>=
internet
              scenerios?</li>
            <li style=3D"margin-left:15px"><span style=3D"background-color:=
transparent">Do the curation
                practices serve diverse forms of access (languages,
                geographies, topics) from that distribution?</span></li>
            <li style=3D"margin-left:15px"><span style=3D"background-color:=
transparent">Do the curators
                understand demand and=C2=A0</span>how<span style=3D"backgro=
und-color:transparent">=C2=A0curation choices
                affect distribution to that demand?<br>
              </span></li>
          </ul>
        </div>
      </div>
    </blockquote>
    These three questions all implicitly require some kind of training
    and education for curators, on matters not directly affecting
    curation activity. Our volunteers are motivated, yes, but what
    motivates is the curation they enjoy, not necessarily tangential
    considerations. People contribute because &quot;I see a typo&quot;, &qu=
ot;I&#39;m
    collecting bits of trivia in this space&quot;, &quot;This leaves out a =
key
    fact&quot;, or &quot;That&#39;s just flat wrong&quot;, and the wiki mak=
es them all
    (relatively) easy to fix. For us, it&#39;s &quot;cheap&quot; as in we&#=
39;ve lowered
    the investment required to participate. Some curators will be
    willing to invest additional engagement on big-picture issues, so it
    may nevertheless be helpful to offer such training, but there&#39;s
    still investment cost on both sides and I do not see it moving the
    needle when focusing on the competitive landscape.
    <blockquote type=3D"cite">
      <div dir=3D"ltr">
        <div>
          <ul>
            <li style=3D"margin-left:15px"><span style=3D"background-color:=
transparent">Can we recruit the
                next generation of curators who don&#39;t &quot;assume&quot=
; that the
                pageview metric is the reason we contribute?</span></li>
          </ul>
        </div>
      </div>
    </blockquote>
    <p>I question whether that assumption is in fact prevalent. Even in
      the world of professional media, those paid to curate or generate
      content often resist the implications, whether out of principled
      beliefs or because they fear impact to their livelihood. This
      conversation has emphasized such metrics because those
      participating care about reaching a broader audience. But in terms
      of what motivates curation at a basic level, once it is
      established that there is some audience out there, I don&#39;t know
      that the order of its magnitude matters all that much. Audience
      demand results in more information to curate and more space in
      which to work, and I think that motivates equally with reach if
      not more. People may find it as satisfying to work on en:Wedding
      of Taylor Swift and Travis Kelce, even though en:Taylor Swift and
      en:Travis Kelce undoubtedly have higher pageviews.</p>
    <p>To use the wheat production infographic as another illustration,
      personally I care less about the interface behavior than what I
      could do to help better curate. Suppose I had information about
      how much wheat Sudan produced in 2011 (currently &quot;no data&quot;)=
, or
      wanted to address the anachronistic use of present-day national
      boundaries going back all the way to 1961? Where would I even
      start? Otherwise, maybe it increases the visual appeal of one
      article a bit, and we could even reproduce that across a bunch of
      articles, but we&#39;d also be throwing up massive barriers to
      maintenance.</p>
    <p>What we most need is to simplify, to make it easier for
      volunteers to steer in the direction we would like them to go.
      Then they can self-select according to what meets their interests.
      I believe the success of Wikidata so far - &quot;success&quot; in the=
 sense
      that its potential aligns more with &quot;early Wikipedia&quot; than =
&quot;early
      Wikiversity&quot; - is in large part because it provides a vast new
      space for curation, and has been integrated in a way that makes it
      easy for those interested to go and participate, while it is
      simultaneously easy to ignore for those who are less interested.</p>
    <p>--Michael Snow</p>
  </div>

_______________________________________________<br>
Wikimedia-l mailing list -- <a href=3D"mailto:[email protected]=
rg" target=3D"_blank">[email protected]</a>, guidelines at: <=
a href=3D"https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines" rel=3D"=
noreferrer" target=3D"_blank">https://meta.wikimedia.org/wiki/Mailing_lists=
/Guidelines</a> and <a href=3D"https://meta.wikimedia.org/wiki/Wikimedia-l"=
 rel=3D"noreferrer" target=3D"_blank">https://meta.wikimedia.org/wiki/Wikim=
edia-l</a><br>
Public archives at <a href=3D"https://lists.wikimedia.org/hyperkitty/list/w=
[email protected]/message/ZOHA3G3WMEDAGX6DUPOE7QCHK2HRYNWU/" r=
el=3D"noreferrer" target=3D"_blank">https://lists.wikimedia.org/hyperkitty/=
list/[email protected]/message/ZOHA3G3WMEDAGX6DUPOE7QCHK2HRYN=
WU/</a><br>
To unsubscribe send an email to <a href=3D"mailto:[email protected]=
ikimedia.org" target=3D"_blank">wikimedia-l-leave-RusutVdil2icGmH+5r0DM0B+6BGkLq7r@public.gmane.org</a></=
blockquote></div>

--000000000000aba9000656642b6d--

--===============5562710496409123862==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Wikimedia-l mailing list -- [email protected], guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l
Public archives at https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/AHQV65MFIZNTHTYIBRO6GBSBP7WIFKG5/
To unsubscribe send an email to wikimedia-l-leave-RusutVdil2icGmH+5r0DM0B+6BGkLq7r@public.gmane.org
--===============5562710496409123862==--