Re: Packages that work with LLMs
Jean Louis <[email protected]>
| Newsgroups | gmane.emacs.devel |
|---|---|
| Organization | GNU Support |
| Message-ID | <[email protected]> |
On 2026-08-13 21:43, Richard Stallman wrote: > > However... is there a write-up on the finding that Apertus > qualifies as > > free software? Ellie and I took a quick look at what they disclose > on > > their website (thanks for the pointer here) and flagged some issues > to > > them > > > https://github.com/swiss-ai/apertus-tech-report/issues/10 Thanks to [email protected] for pointing that out. Apertus isn't free at all. I have allowed myself to contribute there and the copy of my message to the thread is here below. Apertus relies on "openly available data." Note how they didn’t actually define "openly." It’s just a vibe. Their logic is delightfully simple: if a website doesn’t explicitly tell a bot to go away in its robots.txt, then the data is "openly available." So, if a site doesn’t decline bots, Apertus considers the data "open." Which, by the way, is the default state for 99% of the internet. So, Apertus didn’t use freely licensed datasets. They used copyrighted information. And now they’re calling it "Apertus enough." Subject: A gentle (but necessary) clarification on "Openness" and Copyright Dear Apertus Team, I’m writing to you with great enthusiasm for your work, particularly your focus on precision and European values. These are noble and necessary pillars for the next generation of language models. However, I believe there is one critical piece missing from your narrative that needs to be addressed in a public statement to ensure true transparency. The core issue lies in your definition of "openly available data." On your website, you state: "Particular attention has been paid to data integrity and ethical standards: the training corpus builds only on data which is publicly available. It is filtered to respect machine-readable opt-out requests from websites, even retroactively..." Here is the nuance that often gets lost in translation: "Publicly available" is not the same as "Public Domain" or "Freely Licensed." The Copyright Contradiction: By your definition, if a website’s robots.txt doesn’t explicitly block bots, the data is "open." Yet, the vast majority of that content is under standard copyright by default. It is "open" in the sense that anyone can read it, not that it is free of rights. By using this data without explicit permission, you are arguably infringing on copyright, even if you are respecting technical opt-outs. "European Values" & Copyright: One of the bedrock principles of European data ethics is the respect for intellectual property and authorship. If "openly available" means "everything on the public web that isn't technically blocked by a crawler," then you are using predominantly copyrighted material. This dilutes the term "fully open language model," as you are not using freely licensed content, but rather scraping the public commons. Traceability Questions: You mention traceability, but we cannot trace the specific origins of much of this data. How can we verify integrity if the "publicly available" label covers millions of copyrighted texts without clear provenance? Why this matters: You are positioning yourselves as the ethical, European alternative. But if "ethical" means "respecting copyright," then using broadly scraped, copyrighted web data without explicit licensing is a fundamental flaw in the "open" narrative. A simple public acknowledgment from you—clarifying that your "openness" is based on technical accessibility (robots.txt) rather than legal permissiveness (copyright)—would go a long way in building trust. It would show that you are not just "open enough," but truly transparent about how open you are. Thank you for considering this. I look forward to seeing Apertus define what "European Openness" really means and that next "Apertus", or Apertus 2G or second generation becomes truly compatible with the free software principles applied analogously on LLM creation. More precisely, free software means users of a program have the four essential freedoms: The freedom to run the program as you wish, for any purpose (freedom 0). The freedom to study how the program works, and change it so it does your computing as you wish (freedom 1). Access to the source code is a precondition for this. The freedom to redistribute copies so you can help others (freedom 2). The freedom to distribute copies of your modified versions to others (freedom 3). By doing this you can give the whole community a chance to benefit from your changes. Access to the source code is a precondition for this. Evaluation based on your methodology Freedom 0: The freedom to run the program as you wish, for any purpose. Invalidation: While you can technically run the model, the value of the model is tied to its training data. Because Apertus scraped "publicly available" (copyrighted) data, the output is inherently tied to those specific copyrighted sources. If a content owner claims their copyrighted material was used in a way that violates their rights, the "freedom" to use the output is subject to legal gray areas. You aren't running a truly free tool; you are running a tool that holds latent legal liabilities inherited from the internet's copyright laws. Freedom 1: The freedom to study how the program works, and change it so it does your computing as you wish. Invalidation: This requires access to the source code and the data that built it. Apertus claims to use "openly available data," but they do not provide a reproducible, precise list of which specific copyrighted texts were used. You cannot truly study or change the model's behavior regarding specific content (e.g., "remove the section based on Copyrighted Article X") because the provenance is fuzzy. You are studying a black box where the ingredients are labeled "internet soup" rather than a precise recipe. Freedom 2: The freedom to redistribute copies so you can help others. Invalidation: When you redistribute the model, you are redistributing a compressed version of copyrighted works from the web. If the model reproduces substantial portions of copyrighted texts (as they admit it can), you are effectively redistributing copyrighted material without a license. You cannot freely give away the model to help others if doing so might inadvertently expose them to copyright infringement claims from the original content owners who didn't consent to this "open" usage. Freedom 3: The freedom to distribute copies of your modified versions to others. Invalidation: This is the most critical failure. To modify the model (e.g., to fine-tune it for a specific domain or remove biases), you often need to understand its weight distribution, which is derived from its training data. Since the training data is a vast, uncurated mix of copyrighted web content, any modification you make is built on a foundation of potentially infringing material. You are distributing a modified version of a "copyrighted collage," meaning your freedom to distribute your own improvements is hampered by the underlying rights issues of the base "open" model. You are not distributing a free derivative; you are distributing a derivative of a copyrighted whole. -- Jean Louis