Re: Use AI but complement it with structured data

Gerard Meijssen via Wikimedia-l <[email protected]>
Newsgroups gmane.org.wikimedia.foundation
Message-ID <CAO53wxWL528St+emoFs0=OuGxiZewQBJEJ4Tp8kA8o6K2SCQNg@mail.gmail.com>
Hoi,
I am saddened to hear that the relation with the OCLC seems to be one where
we no longer have an open line to their VIAF identifiers..Particularly
because we had a good run together.sharing data. Our data is still
available to them, I hope they still use and update our data; that is what
it is there for. Given the relation we had, I expect that we reached out to
them to see if an accommodation is possible.

PHD theses have a purpose; they enable people to graduate. Once they are, a
typical thesis is never heard about again. As the main objective of this
thread is to kick start improvements on the quality of our data it is
disappointing to hear that numerous studies, thesis exist that explain,
quantify the problems we look away from.

The original suggestion in this thread is that OCLC has its own AI
environment and it uses it to find and merge duplicate records in their
database. It will be good to have our own AI environment (excuse me for not
knowing AI's vocabulary) and give our communities a spin to experiment with
AI, with the objective to benefit the quality of our projects. As Strainu
mentions, the tools to do exactly this exist, probably they will need some
further development to have it work in our own AI data environment. As
Samuel indicates there are these thesis that explain how, where and why we
will benefit in a qualitative manner.

Once the obvious improvements from an AI approach we own become manifest,
we are bound to come up with additional ways to benefit the quality of our
projects.
Thanks.
       GerardM

On Tue, 18 Aug 2026 at 02:21, Samuel Klein via Wikimedia-l <
[email protected]> wrote:

> Gerard writes:
>
> > I read the most wonderful blogpost by the OCLC [1]. For those that do
> not know, OCLC is a longtime partner of the WMF,
> > it is the worldwide organisation of library organisations, it publishes
> the VIAF identifiers and they are heavily used in our projects.
>
> It's a fine post. I don't know about "partner of the WMF"' but they
> maintain important resources like VIAF.  However none of these resources
> are freely licensed, and most of them are being incrementally enclosed and
> paywalled.  We should of course work with OCLC where the result advances
> free knowledge, but should also learn a lesson from VIAF's enclosure:
>
> Back in 2011, Tim Spalding wrote about his concern about VIAF being
> promoted as "open
> <https://blog.librarything.com/2011/03/viaf-oclc-and-open-data/>data"
> <https://blog.librarything.com/2011/03/viaf-oclc-and-open-data/>, when it
> was encumbered by licensing restrictions. Two years ago, Tim's concern was
> realized: bulk updates are no longer provided, and API access is limited to
> 10,000 records a month per IP.  (At which point it would take many
> centuries to update a snapshot!)
>
> The value of having Wikidata as a model of true openness, can not be
> overstated here.
> In 2011, Wikidata did not exist at scale.  Today, it points to an
> alternative that avoids enshittification (as has already happened to the Dewey
> Decimal System
> <https://www.reddit.com/r/librarians/comments/1861395/oclc_classify_is_being_discontinued_alternatives/>
>  and to Worldcat itself
> <https://www.reddit.com/r/BookCollecting/comments/1nhrjmz/any_good_alternatives_to_worldcat_it_has_become/>).
> These are some of the ways our ecosystem can be a pillar of support for
> global free knowledge.  Here is a quote from a paper published by VIAF
> members
> <https://ital.corejournals.org/index.php/ital/article/download/17384/11952> last
> year, uncertain about its future:
>
> " * *Will VIAF continue as an open service, or will it become a paid
> service under OCLC?*
> .
> These doubts and questions caused by the current situation seem to have a
> distant origin; they were already raised in some way back in 2016 and seem
> to involve both VIAF and ISNI.   While awaiting answers, libraries may need
> to consider developing an alternative to VIAF: an open service governed in
> a transparent and shared manner, featuring direct manual contributions,
> open code, modern APIs, a SPARQL engine, and more."
>
>
> Gerard writes:
>
>> we are in a wonderful position that we have both structured data in list,
>> categories and info boxes in our Wikipedias but also in Wikidata. In
>> addition there is the texts that will typically confirm what it says in the
>> structured data.
>>
>
> Absolutely.  There are many PhD theses waiting to be drawn from this line
> of research. 🗼
>
> SJ
>
> [resending after mailing list rejection]
> _______________________________________________
> Wikimedia-l mailing list -- [email protected], guidelines
> at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and
> https://meta.wikimedia.org/wiki/Wikimedia-l
> Public archives at
> https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/7G4LAZ2Z7G7TVZN5NMQNABF5SPJSYO6Y/
> To unsubscribe send an email to wikimedia-l-leave-RusutVdil2icGmH+5r0DM0B+6BGkLq7r@public.gmane.org

_______________________________________________
Wikimedia-l mailing list -- [email protected], guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l
Public archives at https://lists.wikimedia.org/hyperkitty/list/[email protected]/message/PH5W6LKVLDIQOXVVW36IKIVY5TFDLUP3/
To unsubscribe send an email to wikimedia-l-leave-RusutVdil2icGmH+5r0DM0B+6BGkLq7r@public.gmane.org
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.