Re: reading LLM code may rule out white room re-implementation; was: On keybindings and the slow erosion of help's utility

Jean Louis <[email protected]>
Newsgroups gmane.emacs.devel
Organization GNU Support
Message-ID <[email protected]>
On 2026-07-18 07:48, Dr. Arne Babenhauserheide wrote:
> Jean Louis <[email protected]> writes:
> 
>> On 2026-07-17 18:03, Dr. Arne Babenhauserheide wrote:
>> Yet one thing is seriously important and forgotten, that such legal
>> concerns are valid only if underlying datasets used to train LLM are
>> non-free.
> 
> This is only guaranteed if by "free" you mean that they are in the
> public domain.
> 
>> Today there are millions of LLMs and many of some important models
>> have been trained on fully free datasets, thus using such represents
>> one of legal solutions.
> 
> Please show an example of a practically usable LLM that is only trained
> on public domain works.

So the trap is that you set the rules, and then you ask me to show you 
if your rule applies.

You are unilaterally redefining that "free" means "public domain" and 
then you demand that I prove the existence of the LLM trained only on 
public domain works, on the standard you invented.

All of the organizations fostering free software such as FSF, clearly 
distinguish between the public domain and free licenses.

Free as public domain only is simply inaccurate representation.

Even if an LLM were trained entirely on non-free datasets, the output is 
not necessarily non-free. The legal question is whether the output 
contains substantial copied expression from the training data. If it 
doesn't, then the output is a new work. The license of the training data 
does not automatically "infect" the output.

You are quite capable of finding references for that.

-- 
Jean Louis
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.