Re: Do we even know what a good AI Optimization strategy would be for Wikipedia? Re: Re: Google Zero is coming [was Re: Wikipedia at 25: A Wake-Up Call (essay)]

Michael Snow via Wikimedia-l <[email protected]> Sat, 11 Jul 2026 13:58:58 -0700
Newsgroups gmane.org.wikimedia.foundation
Message-ID <[email protected]>
This is a multi-part message in MIME format.
--===============2806989914482549869==
Content-Type: multipart/alternative;
 boundary="------------5qL6yHjqA5vSNiBpOTxnTBUQ"
Content-Language: en-US
Content-Length: 11224

This is a multi-part message in MIME format.
--------------5qL6yHjqA5vSNiBpOTxnTBUQ
Content-Type: text/plain; charset=UTF-8; format=flowed
Content-Transfer-Encoding: 8bit

On 7/11/2026 10:43 AM, Alex Stinson via Wikimedia-l wrote:
> *Interfaces are cheap, informed curators are expensive*
>
> Agentic AI, Coding tools and LLMs are making the cost of new 
> interfaces /extremely cheap, /so cheap that I am commissioning a 
> complex knowledge repository for a fraction of the cost and time it 
> would take otherwise. Interface projects like WikiProject Med's 
> offline medical Wikipedia App, which used to take several months of 
> highly specialized software development, can now be spun up in a long 
> weekend with a Claude Code Max subscription. Sage's tool is a perfect 
> example of this: perfect for a small market of users, unlikely to be a 
> "headline" Wikimedia tactic for getting in front of users, because 
> Youtube and AI search interfaces already do this exact thing.
>
> What is not cheap for all of the other platforms (and is often paid 
> for by ads), but which we have in abundance, is motivated humans who 
> can continue curating the knowledge (and more than 25 years of 
> experiments on facilitating knowledge equity focused gap filling). The 
> questions we need to address are:
>
>   * Do the curators understand the future of distribution to other
>     humans we need to be building for across multiple future internet
>     scenerios?
>   * Do the curation practices serve diverse forms of access
>     (languages, geographies, topics) from that distribution?
>   * Do the curators understand demand and how curation choices affect
>     distribution to that demand?
>
These three questions all implicitly require some kind of training and 
education for curators, on matters not directly affecting curation 
activity. Our volunteers are motivated, yes, but what motivates is the 
curation they enjoy, not necessarily tangential considerations. People 
contribute because "I see a typo", "I'm collecting bits of trivia in 
this space", "This leaves out a key fact", or "That's just flat wrong", 
and the wiki makes them all (relatively) easy to fix. For us, it's 
"cheap" as in we've lowered the investment required to participate. Some 
curators will be willing to invest additional engagement on big-picture 
issues, so it may nevertheless be helpful to offer such training, but 
there's still investment cost on both sides and I do not see it moving 
the needle when focusing on the competitive landscape.
>
>   * Can we recruit the next generation of curators who don't "assume"
>     that the pageview metric is the reason we contribute?
>
I question whether that assumption is in fact prevalent. Even in the 
world of professional media, those paid to curate or generate content 
often resist the implications, whether out of principled beliefs or 
because they fear impact to their livelihood. This conversation has 
emphasized such metrics because those participating care about reaching 
a broader audience. But in terms of what motivates curation at a basic 
level, once it is established that there is some audience out there, I 
don't know that the order of its magnitude matters all that much. 
Audience demand results in more information to curate and more space in 
which to work, and I think that motivates equally with reach if not 
more. People may find it as satisfying to work on en:Wedding of Taylor 
Swift and Travis Kelce, even though en:Taylor Swift and en:Travis Kelce 
undoubtedly have higher pageviews.

To use the wheat production infographic as another illustration, 
personally I care less about the interface behavior than what I could do 
to help better curate. Suppose I had information about how much wheat 
Sudan produced in 2011 (currently "no data"), or wanted to address the 
anachronistic use of present-day national boundaries going back all the 
way to 1961? Where would I even start? Otherwise, maybe it increases the 
visual appeal of one article a bit, and we could even reproduce that 
across a bunch of articles, but we'd also be throwing up massive 
barriers to maintenance.

What we most need is to simplify, to make it easier for volunteers to 
steer in the direction we would like them to go. Then they can 
self-select according to what meets their interests. I believe the 
success of Wikidata so far - "success" in the sense that its potential 
aligns more with "early Wikipedia" than "early Wikiversity" - is in 
large part because it provides a vast new space for curation, and has 
been integrated in a way that makes it easy for those interested to go 
and participate, while it is simultaneously easy to ignore for those who 
are less interested.

--Michael Snow

--------------5qL6yHjqA5vSNiBpOTxnTBUQ
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: 8bit

<!DOCTYPE html>
<html>
  <head>
    <meta http-equiv="Content-Type" content="text/html; charset=UTF-8">
  </head>
  <body>
    <div class="moz-cite-prefix">On 7/11/2026 10:43 AM, Alex Stinson via
      Wikimedia-l wrote:<br>
    </div>
    <blockquote type="cite"
cite="mid:CAMT4PfMZLk5oZZtUpN9RXATvo7+HDAtH+WB-EQfZzt9JD0=05Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org">
      <meta http-equiv="content-type" content="text/html; charset=UTF-8">
      <div dir="ltr">
        <div><b>Interfaces are cheap, informed curators are expensive</b><br>
          <br>
          Agentic AI, Coding tools and LLMs are making the cost of new
          interfaces <i>extremely cheap, </i>so cheap that I am
          commissioning a complex knowledge repository for a fraction of
          the cost and time it would take otherwise. Interface projects
          like WikiProject Med's offline medical Wikipedia App, which
          used to take several months of highly specialized software
          development, can now be spun up in a long weekend with a
          Claude Code Max subscription. Sage's tool is a perfect example
          of this: perfect for a small market of users, unlikely to be a
          "headline" Wikimedia tactic for getting in front of users,
          because Youtube and AI search interfaces already do this exact
          thing.<br>
          <br>
          What is not cheap for all of the other platforms (and is often
          paid for by ads), but which we have in abundance, is motivated
          humans who can continue curating the knowledge (and more than
          25 years of experiments on facilitating knowledge equity
          focused gap filling). The questions we need to address are:<br>
          <ul>
            <li style="margin-left:15px"><span
                style="background-color:transparent">Do the curators
                understand the future of distribution to other humans we
                need to be building for across multiple future </span>internet
              scenerios?</li>
            <li style="margin-left:15px"><span
                style="background-color:transparent">Do the curation
                practices serve diverse forms of access (languages,
                geographies, topics) from that distribution?</span></li>
            <li style="margin-left:15px"><span
                style="background-color:transparent">Do the curators
                understand demand and </span>how<span
                style="background-color:transparent"> curation choices
                affect distribution to that demand?<br>
              </span></li>
          </ul>
        </div>
      </div>
    </blockquote>
    These three questions all implicitly require some kind of training
    and education for curators, on matters not directly affecting
    curation activity. Our volunteers are motivated, yes, but what
    motivates is the curation they enjoy, not necessarily tangential
    considerations. People contribute because "I see a typo", "I'm
    collecting bits of trivia in this space", "This leaves out a key
    fact", or "That's just flat wrong", and the wiki makes them all
    (relatively) easy to fix. For us, it's "cheap" as in we've lowered
    the investment required to participate. Some curators will be
    willing to invest additional engagement on big-picture issues, so it
    may nevertheless be helpful to offer such training, but there's
    still investment cost on both sides and I do not see it moving the
    needle when focusing on the competitive landscape.
    <blockquote type="cite"
cite="mid:CAMT4PfMZLk5oZZtUpN9RXATvo7+HDAtH+WB-EQfZzt9JD0=05Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org">
      <div dir="ltr">
        <div>
          <ul>
            <li style="margin-left:15px"><span
                style="background-color:transparent">Can we recruit the
                next generation of curators who don't "assume" that the
                pageview metric is the reason we contribute?</span></li>
          </ul>
        </div>
      </div>
    </blockquote>
    <p>I question whether that assumption is in fact prevalent. Even in
      the world of professional media, those paid to curate or generate
      content often resist the implications, whether out of principled
      beliefs or because they fear impact to their livelihood. This
      conversation has emphasized such metrics because those
      participating care about reaching a broader audience. But in terms
      of what motivates curation at a basic level, once it is
      established that there is some audience out there, I don't know
      that the order of its magnitude matters all that much. Audience
      demand results in more information to curate and more space in
      which to work, and I think that motivates equally with reach if
      not more. People may find it as satisfying to work on en:Wedding
      of Taylor Swift and Travis Kelce, even though en:Taylor Swift and
      en:Travis Kelce undoubtedly have higher pageviews.</p>
    <p>To use the wheat production infographic as another illustration,
      personally I care less about the interface behavior than what I
      could do to help better curate. Suppose I had information about
      how much wheat Sudan produced in 2011 (currently "no data"), or
      wanted to address the anachronistic use of present-day national
      boundaries going back all the way to 1961? Where would I even
      start? Otherwise, maybe it increases the visual appeal of one
      article a bit, and we could even reproduce that across a bunch of
      articles, but we'd also be throwing up massive barriers to
      maintenance.</p>
    <p>What we most need is to simplify, to make it easier for
      volunteers to steer in the direction we would like them to go.
      Then they can self-select according to what meets their interests.
      I believe the success of Wikidata so far - "success" in the sense
      that its potential aligns more with "early Wikipedia" than "early
      Wikiversity" - is in large part because it provides a vast new
      space for curation, and has been integrated in a way that makes it
      easy for those interested to go and participate, while it is
      simultaneously easy to ignore for those who are less interested.</p>
    <p>--Michael Snow</p>
  </body>
</html>

--------------5qL6yHjqA5vSNiBpOTxnTBUQ--

--===============2806989914482549869==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Wikimedia-l mailing list -- [email protected], guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l
Public archives at https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/ZOHA3G3WMEDAGX6DUPOE7QCHK2HRYNWU/
To unsubscribe send an email to wikimedia-l-leave-RusutVdil2icGmH+5r0DM0B+6BGkLq7r@public.gmane.org
--===============2806989914482549869==--