Re: Do we even know what a good AI Optimization strategy would be for Wikipedia? Re: Re: Google Zero is coming [was Re: Wikipedia at 25: A Wake-Up Call (essay)]
Michael Snow via Wikimedia-l <[email protected]> Sat, 11 Jul 2026 13:58:58 -0700
| Newsgroups | gmane.org.wikimedia.foundation |
|---|---|
| Message-ID | <[email protected]> |
This is a multi-part message in MIME format.
--===============2806989914482549869==
Content-Type: multipart/alternative;
boundary="------------5qL6yHjqA5vSNiBpOTxnTBUQ"
Content-Language: en-US
Content-Length: 11224
This is a multi-part message in MIME format.
--------------5qL6yHjqA5vSNiBpOTxnTBUQ
Content-Type: text/plain; charset=UTF-8; format=flowed
Content-Transfer-Encoding: 8bit
On 7/11/2026 10:43 AM, Alex Stinson via Wikimedia-l wrote:
> *Interfaces are cheap, informed curators are expensive*
>
> Agentic AI, Coding tools and LLMs are making the cost of new
> interfaces /extremely cheap, /so cheap that I am commissioning a
> complex knowledge repository for a fraction of the cost and time it
> would take otherwise. Interface projects like WikiProject Med's
> offline medical Wikipedia App, which used to take several months of
> highly specialized software development, can now be spun up in a long
> weekend with a Claude Code Max subscription. Sage's tool is a perfect
> example of this: perfect for a small market of users, unlikely to be a
> "headline" Wikimedia tactic for getting in front of users, because
> Youtube and AI search interfaces already do this exact thing.
>
> What is not cheap for all of the other platforms (and is often paid
> for by ads), but which we have in abundance, is motivated humans who
> can continue curating the knowledge (and more than 25 years of
> experiments on facilitating knowledge equity focused gap filling). The
> questions we need to address are:
>
> * Do the curators understand the future of distribution to other
> humans we need to be building for across multiple future internet
> scenerios?
> * Do the curation practices serve diverse forms of access
> (languages, geographies, topics) from that distribution?
> * Do the curators understand demand and how curation choices affect
> distribution to that demand?
>
These three questions all implicitly require some kind of training and
education for curators, on matters not directly affecting curation
activity. Our volunteers are motivated, yes, but what motivates is the
curation they enjoy, not necessarily tangential considerations. People
contribute because "I see a typo", "I'm collecting bits of trivia in
this space", "This leaves out a key fact", or "That's just flat wrong",
and the wiki makes them all (relatively) easy to fix. For us, it's
"cheap" as in we've lowered the investment required to participate. Some
curators will be willing to invest additional engagement on big-picture
issues, so it may nevertheless be helpful to offer such training, but
there's still investment cost on both sides and I do not see it moving
the needle when focusing on the competitive landscape.
>
> * Can we recruit the next generation of curators who don't "assume"
> that the pageview metric is the reason we contribute?
>
I question whether that assumption is in fact prevalent. Even in the
world of professional media, those paid to curate or generate content
often resist the implications, whether out of principled beliefs or
because they fear impact to their livelihood. This conversation has
emphasized such metrics because those participating care about reaching
a broader audience. But in terms of what motivates curation at a basic
level, once it is established that there is some audience out there, I
don't know that the order of its magnitude matters all that much.
Audience demand results in more information to curate and more space in
which to work, and I think that motivates equally with reach if not
more. People may find it as satisfying to work on en:Wedding of Taylor
Swift and Travis Kelce, even though en:Taylor Swift and en:Travis Kelce
undoubtedly have higher pageviews.
To use the wheat production infographic as another illustration,
personally I care less about the interface behavior than what I could do
to help better curate. Suppose I had information about how much wheat
Sudan produced in 2011 (currently "no data"), or wanted to address the
anachronistic use of present-day national boundaries going back all the
way to 1961? Where would I even start? Otherwise, maybe it increases the
visual appeal of one article a bit, and we could even reproduce that
across a bunch of articles, but we'd also be throwing up massive
barriers to maintenance.
What we most need is to simplify, to make it easier for volunteers to
steer in the direction we would like them to go. Then they can
self-select according to what meets their interests. I believe the
success of Wikidata so far - "success" in the sense that its potential
aligns more with "early Wikipedia" than "early Wikiversity" - is in
large part because it provides a vast new space for curation, and has
been integrated in a way that makes it easy for those interested to go
and participate, while it is simultaneously easy to ignore for those who
are less interested.
--Michael Snow
--------------5qL6yHjqA5vSNiBpOTxnTBUQ
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: 8bit
<!DOCTYPE html>
<html>
<head>
<meta http-equiv="Content-Type" content="text/html; charset=UTF-8">
</head>
<body>
<div class="moz-cite-prefix">On 7/11/2026 10:43 AM, Alex Stinson via
Wikimedia-l wrote:<br>
</div>
<blockquote type="cite"
cite="mid:CAMT4PfMZLk5oZZtUpN9RXATvo7+HDAtH+WB-EQfZzt9JD0=05Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org">
<meta http-equiv="content-type" content="text/html; charset=UTF-8">
<div dir="ltr">
<div><b>Interfaces are cheap, informed curators are expensive</b><br>
<br>
Agentic AI, Coding tools and LLMs are making the cost of new
interfaces <i>extremely cheap, </i>so cheap that I am
commissioning a complex knowledge repository for a fraction of
the cost and time it would take otherwise. Interface projects
like WikiProject Med's offline medical Wikipedia App, which
used to take several months of highly specialized software
development, can now be spun up in a long weekend with a
Claude Code Max subscription. Sage's tool is a perfect example
of this: perfect for a small market of users, unlikely to be a
"headline" Wikimedia tactic for getting in front of users,
because Youtube and AI search interfaces already do this exact
thing.<br>
<br>
What is not cheap for all of the other platforms (and is often
paid for by ads), but which we have in abundance, is motivated
humans who can continue curating the knowledge (and more than
25 years of experiments on facilitating knowledge equity
focused gap filling). The questions we need to address are:<br>
<ul>
<li style="margin-left:15px"><span
style="background-color:transparent">Do the curators
understand the future of distribution to other humans we
need to be building for across multiple future </span>internet
scenerios?</li>
<li style="margin-left:15px"><span
style="background-color:transparent">Do the curation
practices serve diverse forms of access (languages,
geographies, topics) from that distribution?</span></li>
<li style="margin-left:15px"><span
style="background-color:transparent">Do the curators
understand demand and </span>how<span
style="background-color:transparent"> curation choices
affect distribution to that demand?<br>
</span></li>
</ul>
</div>
</div>
</blockquote>
These three questions all implicitly require some kind of training
and education for curators, on matters not directly affecting
curation activity. Our volunteers are motivated, yes, but what
motivates is the curation they enjoy, not necessarily tangential
considerations. People contribute because "I see a typo", "I'm
collecting bits of trivia in this space", "This leaves out a key
fact", or "That's just flat wrong", and the wiki makes them all
(relatively) easy to fix. For us, it's "cheap" as in we've lowered
the investment required to participate. Some curators will be
willing to invest additional engagement on big-picture issues, so it
may nevertheless be helpful to offer such training, but there's
still investment cost on both sides and I do not see it moving the
needle when focusing on the competitive landscape.
<blockquote type="cite"
cite="mid:CAMT4PfMZLk5oZZtUpN9RXATvo7+HDAtH+WB-EQfZzt9JD0=05Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org">
<div dir="ltr">
<div>
<ul>
<li style="margin-left:15px"><span
style="background-color:transparent">Can we recruit the
next generation of curators who don't "assume" that the
pageview metric is the reason we contribute?</span></li>
</ul>
</div>
</div>
</blockquote>
<p>I question whether that assumption is in fact prevalent. Even in
the world of professional media, those paid to curate or generate
content often resist the implications, whether out of principled
beliefs or because they fear impact to their livelihood. This
conversation has emphasized such metrics because those
participating care about reaching a broader audience. But in terms
of what motivates curation at a basic level, once it is
established that there is some audience out there, I don't know
that the order of its magnitude matters all that much. Audience
demand results in more information to curate and more space in
which to work, and I think that motivates equally with reach if
not more. People may find it as satisfying to work on en:Wedding
of Taylor Swift and Travis Kelce, even though en:Taylor Swift and
en:Travis Kelce undoubtedly have higher pageviews.</p>
<p>To use the wheat production infographic as another illustration,
personally I care less about the interface behavior than what I
could do to help better curate. Suppose I had information about
how much wheat Sudan produced in 2011 (currently "no data"), or
wanted to address the anachronistic use of present-day national
boundaries going back all the way to 1961? Where would I even
start? Otherwise, maybe it increases the visual appeal of one
article a bit, and we could even reproduce that across a bunch of
articles, but we'd also be throwing up massive barriers to
maintenance.</p>
<p>What we most need is to simplify, to make it easier for
volunteers to steer in the direction we would like them to go.
Then they can self-select according to what meets their interests.
I believe the success of Wikidata so far - "success" in the sense
that its potential aligns more with "early Wikipedia" than "early
Wikiversity" - is in large part because it provides a vast new
space for curation, and has been integrated in a way that makes it
easy for those interested to go and participate, while it is
simultaneously easy to ignore for those who are less interested.</p>
<p>--Michael Snow</p>
</body>
</html>
--------------5qL6yHjqA5vSNiBpOTxnTBUQ--
--===============2806989914482549869==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline
_______________________________________________
Wikimedia-l mailing list -- [email protected], guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l
Public archives at https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/ZOHA3G3WMEDAGX6DUPOE7QCHK2HRYNWU/
To unsubscribe send an email to wikimedia-l-leave-RusutVdil2icGmH+5r0DM0B+6BGkLq7r@public.gmane.org
--===============2806989914482549869==--