Re: Ballot option: Allow AI-Assisted Contributions

Adrian Bunk <[email protected]> Wed, 5 Aug 2026 15:51:30 +0300
Newsgroups gmane.linux.debian.devel.vote
Message-ID <anMx0jhibgJanOvr@localhost>
On Wed, Aug 05, 2026 at 01:17:44PM +0200, Aigars Mahinovs wrote:
> On Wed, 5 Aug 2026 at 11:15, Gerardo Ballabio <[email protected]>
> wrote:
> 
> > Aigars Mahinovs wrote:
> > > Lawyers that I have spoken with are of the opinion that AI providers
> > (either model weight providers or service providers) can not *really*
> > legally claim any kind of copyright on the outputs of the AI models, so
> > they are very explicitly NOT doing that.
> >
> > As I understand it, the problem with copyright isn't that AI providers
> > might claim copyright. It's that *someone else* might claim copyright
> > because the AI scraped and regurgitated their code. That's still an
> > open legal question AFAIK.
> 
> 
> If that were ruled to be true, then the *entire* AI landscape would
> collapse as there is no way to gather enough data with *compatible*
> licenses to produce a coherent AI model. Even if the pure-legal
> interpretation were in favor of that outcome, there are strong social and
> commercial incentives against the law moving that way.
>...
> So the current consensus seems to be to operate on the presumtion that
> various legal loopholes (like fair-use in the USA and data mining exception
> in EU) are sufficient to decouple the copyright of the model from the
> copyrights of the training data (if the actual acquiring of the training
> data is done without violating the copyrights).
>...

Not only the acquiring, the use for training also has to be legal.

The law (17 U.S. Code ยง 107) says that a factor when determining whether 
a particular case is a fair use is the effect on the market and the 
value of the copyrighted work.

Judges in the US have already questioned fair use for model training 
based on that.[1]

Regarding model output, a judge in New York wrote last year:

  Turning to the merits, the Court must determine whether the Consolidated Class
  Action Complaint adequately pleads an output-based infringement claim. 
  It does.[2]

This is not a decision on the merits, but it is an ongoing lawsuit where 
the judge did compare training data and model output in the decision not 
to dismiss.

cu
Adrian

[1] https://lists.debian.org/debian-vote/2026/07/msg00157.html
[2] https://assets.law360news.com/2404000/2404371/https-ecf-nysd-uscourts-gov-doc1-127138452540.pdf