Re: Packages that work with LLMs
Jean Louis <[email protected]>
| Newsgroups | gmane.emacs.devel |
|---|---|
| Organization | GNU Support |
| Message-ID | <[email protected]> |
Hello Dr. Arne, On 2026-08-14 12:27, Dr. Arne Babenhauserheide wrote: > Jean Louis <[email protected]> writes: > >> Their logic is delightfully simple: if a website doesn’t explicitly >> tell a bot to go away in its robots.txt, then the data is "openly >> available." > > This complies with the EU copyright directive article 4, so it is > legally sound in the EU (but might not be elsewhere in the world): > https://en.wikipedia.org/wiki/Directive_on_Copyright_in_the_Digital_Single_Market#Article_3_and_4 Thanks, I got you. Though that copyright exception is for purposes of training LLM (text and data mining). The LLM can be trained on it, no issues. But that LLM is trained on copyrighted datasets. This creates issue for the user using the LLM to generate text. When text is generated it may potentially infringe on someone's copyrights, especially if text is given verbatim. So it is generally known that having LLMs doesn't infringe on someone's copyrights. Yet having outputs generated can infringe on someone's copyright, by the user starting to distribute such copyrighted materials. The legal responsibility for distributing infringing generated text still rests entirely on the user or provider of the LLM service. This directive does not absolve LLM outputs from copyright infringement claims. It is remarkably simple to generate verbatim copyrighted texts from LLM. That means we have to find LLM which is trained on fully free and known list of datasets. Apertus isn't "aperto" enough. > It may be the reason why Apertus is Text only: the exception granted > for > model training by Article 4 only applies to text an data mining. I rather believe they do not have skilled people to do it. -- Jean Louis