Re: Packages that work with LLMs

Jean Louis <[email protected]>
Newsgroups gmane.emacs.devel
Organization GNU Support
Message-ID <[email protected]>
Hello Dr. Arne,

On 2026-08-14 12:27, Dr. Arne Babenhauserheide wrote:
> Jean Louis <[email protected]> writes:
> 
>> Their logic is delightfully simple: if a website doesn’t explicitly
>> tell a bot to go away in its robots.txt, then the data is "openly
>> available."
> 
> This complies with the EU copyright directive article 4, so it is
> legally sound in the EU (but might not be elsewhere in the world):
> https://en.wikipedia.org/wiki/Directive_on_Copyright_in_the_Digital_Single_Market#Article_3_and_4

Thanks, I got you. Though that copyright exception is for purposes of 
training LLM (text and data mining). The LLM can be trained on it, no 
issues.

But that LLM is trained on copyrighted datasets.

This creates issue for the user using the LLM to generate text. When 
text is generated it may potentially infringe on someone's copyrights, 
especially if text is given verbatim.

So it is generally known that having LLMs doesn't infringe on someone's 
copyrights. Yet having outputs generated can infringe on someone's 
copyright, by the user starting to distribute such copyrighted 
materials.

The legal responsibility for distributing infringing generated text 
still rests entirely on the user or provider of the LLM service. This 
directive does not absolve LLM outputs from copyright infringement 
claims.

It is remarkably simple to generate verbatim copyrighted texts from LLM.

That means we have to find LLM which is trained on fully free and known 
list of datasets. Apertus isn't "aperto" enough.

> It may be the reason why Apertus is Text only: the exception granted 
> for
> model training by Article 4 only applies to text an data mining.

I rather believe they do not have skilled people to do it.

-- 
Jean Louis
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.