Re: Packages that work with LLMs

Jean Louis <[email protected]>
Newsgroups gmane.emacs.devel
Organization GNU Support
Message-ID <[email protected]>
On 2026-08-13 21:43, Richard Stallman wrote:
>   > However... is there a write-up on the finding that Apertus 
> qualifies as
>   > free software? Ellie and I took a quick look at what they disclose 
> on
>   > their website (thanks for the pointer here) and flagged some issues 
> to
>   > them
> 
>   > https://github.com/swiss-ai/apertus-tech-report/issues/10

Thanks to [email protected] for pointing that out.

Apertus isn't free at all.

I have allowed myself to contribute there and the copy of my message to 
the thread is here below.

Apertus relies on "openly available data." Note how they didn’t actually 
define "openly." It’s just a vibe.

Their logic is delightfully simple: if a website doesn’t explicitly tell 
a bot to go away in its robots.txt, then the data is "openly available."

So, if a site doesn’t decline bots, Apertus considers the data "open." 
Which, by the way, is the default state for 99% of the internet.

So, Apertus didn’t use freely licensed datasets. They used copyrighted 
information. And now they’re calling it "Apertus enough."

Subject: A gentle (but necessary) clarification on "Openness" and 
Copyright

Dear Apertus Team,

I’m writing to you with great enthusiasm for your work, particularly 
your focus on precision and European values. These are noble and 
necessary pillars for the next generation of language models. However, I 
believe there is one critical piece missing from your narrative that 
needs to be addressed in a public statement to ensure true transparency.

The core issue lies in your definition of "openly available data."

On your website, you state:

     "Particular attention has been paid to data integrity and ethical 
standards: the training corpus builds only on data which is publicly 
available. It is filtered to respect machine-readable opt-out requests 
from websites, even retroactively..."

Here is the nuance that often gets lost in translation: "Publicly 
available" is not the same as "Public Domain" or "Freely Licensed."

     The Copyright Contradiction: By your definition, if a website’s 
robots.txt doesn’t explicitly block bots, the data is "open." Yet, the 
vast majority of that content is under standard copyright by default. It 
is "open" in the sense that anyone can read it, not that it is free of 
rights. By using this data without explicit permission, you are arguably 
infringing on copyright, even if you are respecting technical opt-outs.

     "European Values" & Copyright: One of the bedrock principles of 
European data ethics is the respect for intellectual property and 
authorship. If "openly available" means "everything on the public web 
that isn't technically blocked by a crawler," then you are using 
predominantly copyrighted material. This dilutes the term "fully open 
language model," as you are not using freely licensed content, but 
rather scraping the public commons.

     Traceability Questions: You mention traceability, but we cannot 
trace the specific origins of much of this data. How can we verify 
integrity if the "publicly available" label covers millions of 
copyrighted texts without clear provenance?

Why this matters:

You are positioning yourselves as the ethical, European alternative. But 
if "ethical" means "respecting copyright," then using broadly scraped, 
copyrighted web data without explicit licensing is a fundamental flaw in 
the "open" narrative.

A simple public acknowledgment from you—clarifying that your "openness" 
is based on technical accessibility (robots.txt) rather than legal 
permissiveness (copyright)—would go a long way in building trust. It 
would show that you are not just "open enough," but truly transparent 
about how open you are.

Thank you for considering this. I look forward to seeing Apertus define 
what "European Openness" really means and that next "Apertus", or 
Apertus 2G or second generation becomes truly compatible with the free 
software principles applied analogously on LLM creation.

More precisely, free software means users of a program have the four 
essential freedoms:

     The freedom to run the program as you wish, for any purpose (freedom 
0).
     The freedom to study how the program works, and change it so it does 
your computing as you wish (freedom 1). Access to the source code is a 
precondition for this.
     The freedom to redistribute copies so you can help others (freedom 
2).
     The freedom to distribute copies of your modified versions to others 
(freedom 3). By doing this you can give the whole community a chance to 
benefit from your changes. Access to the source code is a precondition 
for this.

Evaluation based on your methodology

     Freedom 0: The freedom to run the program as you wish, for any 
purpose.

Invalidation: While you can technically run the model, the value of the 
model is tied to its training data. Because Apertus scraped "publicly 
available" (copyrighted) data, the output is inherently tied to those 
specific copyrighted sources. If a content owner claims their 
copyrighted material was used in a way that violates their rights, the 
"freedom" to use the output is subject to legal gray areas. You aren't 
running a truly free tool; you are running a tool that holds latent 
legal liabilities inherited from the internet's copyright laws.

     Freedom 1: The freedom to study how the program works, and change it 
so it does your computing as you wish.

Invalidation: This requires access to the source code and the data that 
built it. Apertus claims to use "openly available data," but they do not 
provide a reproducible, precise list of which specific copyrighted texts 
were used. You cannot truly study or change the model's behavior 
regarding specific content (e.g., "remove the section based on 
Copyrighted Article X") because the provenance is fuzzy. You are 
studying a black box where the ingredients are labeled "internet soup" 
rather than a precise recipe.

     Freedom 2: The freedom to redistribute copies so you can help 
others.

Invalidation: When you redistribute the model, you are redistributing a 
compressed version of copyrighted works from the web. If the model 
reproduces substantial portions of copyrighted texts (as they admit it 
can), you are effectively redistributing copyrighted material without a 
license. You cannot freely give away the model to help others if doing 
so might inadvertently expose them to copyright infringement claims from 
the original content owners who didn't consent to this "open" usage.

     Freedom 3: The freedom to distribute copies of your modified 
versions to others.

Invalidation: This is the most critical failure. To modify the model 
(e.g., to fine-tune it for a specific domain or remove biases), you 
often need to understand its weight distribution, which is derived from 
its training data. Since the training data is a vast, uncurated mix of 
copyrighted web content, any modification you make is built on a 
foundation of potentially infringing material. You are distributing a 
modified version of a "copyrighted collage," meaning your freedom to 
distribute your own improvements is hampered by the underlying rights 
issues of the base "open" model. You are not distributing a free 
derivative; you are distributing a derivative of a copyrighted whole.

-- 
Jean Louis
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.