Re: We should require AI disclosure
Dave Cridland <[email protected]> Thu, 9 Jul 2026 14:12:50 +0100
| Newsgroups | gmane.network.jabber.standards-jig |
|---|---|
| Message-ID | <CAKHUCzyGXbDFJwMgTUCHAMmOwpH0WKFeVQkM_kFPGXTePYBfTQ@mail.gmail.com> |
--===============3975533669392672897== Content-Type: multipart/alternative; boundary="000000000000317e4e06562d62d3" --000000000000317e4e06562d62d3 Content-Type: text/plain; charset="UTF-8" On Wed, 8 Jul 2026 at 18:08, MSavoritias via Standards <[email protected]> wrote: > 1. It seems to be a consensus in the list that XEPs can be rejected if > they are "bad" in some way. The issue here is that its not clear what > "bad" means. This would help the editor (or council) a lot for example > to be able to point at "something" that is public and clear. This can of > course also help for the Code of Conduct and for the Experimental > process among other things. > > I think your last suggestion actually handles the "definition of bad", or rather, defines a "definition of good enough" that I think - with maybe some tweaks - is ideal. > IMHO a good baseline is: > > - The author of said XEP *CAN* give the copyright to XSF > > Right. I think (hope!) we have this already. I refer to this as the copyright warranty. > - The author of said XEP *CAN* explain why and how in a XEP > > I'm less bothered about this, but I see the logic, and think it's covered by Guus's PR (of Ralph's words). If you want clarification, I mean that obviously it's sensible to understand the spec you're submitted in all its details, but I'm not sure we need a warranty to that effect. It equally obviously does no harm, though, so I've no objections either! > - The author of said XEP *HAS* communicated with the XSF (in one of our > rooms, mailing list, summit, etc.) or is vouched by one of the members > of XSF before making the XEP that their approach is at least desired and > may make some sense for initial experiments. > > I think this needs tweaking but the essential concept - that other people support the submission - is really important. This does two things: - It means that the kind of mindless AI slop that I think we're all rightly concerned about never gets traction - not because it's AI, but because it's mindless slop. - It also means that the early stage workload is offloaded from Council. Council only need to take notice of submissions that actually gain any kind of traction in the community. I would s/one of the members of XSF/participants in the Standards SIG/ because we've not selected our membership on the basis of technical scrutiny, and it just feels like if a ProtoXEP is getting positive discussion and engagement on the standards list it's probably good for Experimental. Council's effort is then a judgement call on that, rather than having to scrutinise the specification itself as much. > > This solves: XEP "dumps" where we don't know why it is like this, who > wrote this, is this in good faith, is this wanted/needed etc., solves > the work of somebody submitting XEPs without knowing what it is written > in there and also covers liability. Note that the first is already > written (and imo that makes llms unable to be used already as openjdk > and other projects have stated) > > Your parenthetical opinion is far from universal. First, the situation with code is radically different to that for text, and second, the copyrightability of AI output varies heavily by jurisidiction and human effort involved. Here in the UK, for instance, we have existing primary legislation (from 1986!) that says LLM output is copyrightable. Other jurisdictions have case law concerning the extremes of AI output, so we know that a XEP coming from a US citizen solely generated by a single prompt of something like "Create a new XEP for something" would likely not be copyrightable. But, this doesn't matter, because by requiring the author to warrant they can assign copyright, we push that liability onto the author. Hoorah! > 2. Regarding what approach we can have to actually write this down some > examples are that *DO NOT* ban LLMs: > All these are code based, and our concern here is prose. We're protected in any case because of the warranty we demand from submitters. You can stop reading here if you want, the rest is just saying "I understand the argument but reject it". There's an argument that because the training data for LLMs was (possibly misused) copyright code, the output is a derived work (in copyright terms) of the original training data. This is supported because at least in some cases in the early days, if you asked for particularly niche code, you'll end up with output looking very similar to pre-existing projects. I have not duplicated these results, and I absolutely tried hard - I have the only open source code based on certain specifications which are not public, so niche in some interesting ways, but I've seen other claims and have no reason to doubt it can (or could) happen. The problem with this is that few people are making this argument with prose - people actually complain that AI generated prose doesn't look human enough - and also nobody makes this argument with human generated anything. I can fully assure you that this email is not considered a derived work of Asimov or Bujold or Scalzi, despite the fact I have read all three an almost unhealthy amount. Similarly, the code I write has never been considered a derived work even though I certainly read other people's copyright code to learn how to write it. Whether or not the LLMs were originally trained with copyright works in contravention of their licences is a whole other matter, but from a purely legal perspective I think that's between the copyright owners and the LLM developers. You are welcome to take an ethical stance on that one on a personal basis. Dave. --000000000000317e4e06562d62d3 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div dir=3D"ltr"><br></div><br><div class=3D"gmail_quote g= mail_quote_container"><div dir=3D"ltr" class=3D"gmail_attr">On Wed, 8 Jul 2= 026 at 18:08, MSavoritias via Standards <<a href=3D"mailto:standards@xmp= p.org">[email protected]</a>> wrote:<br></div><blockquote class=3D"gmai= l_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,20= 4,204);padding-left:1ex">1. It seems to be a consensus in the list that XEP= s can be rejected if <br> they are "bad" in some way. The issue here is that its not clear = what <br> "bad" means. This would help the editor (or council) a lot for ex= ample <br> to be able to point at "something" that is public and clear. This= can of <br> course also help for the Code of Conduct and for the Experimental <br> process among other things.<br> <br></blockquote><div><br></div><div>I think your last suggestion actually = handles the "definition of bad", or rather, defines a "defin= ition of good enough" that I think - with maybe some tweaks - is ideal= .</div><div>=C2=A0</div><blockquote class=3D"gmail_quote" style=3D"margin:0= px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"> IMHO a good baseline is:<br> <br> - The author of said XEP *CAN* give the copyright to XSF<br> <br></blockquote><div><br></div><div>Right. I think (hope!) we have this al= ready. I refer to this as the copyright warranty.</div><div>=C2=A0</div><bl= ockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-lef= t:1px solid rgb(204,204,204);padding-left:1ex"> - The author of said XEP *CAN* explain why and how in a XEP<br> <br></blockquote><div><br></div><div>I'm less bothered about this, but = I see the logic, and think it's covered by Guus's PR (of Ralph'= s words).</div><div><br></div><div>If you want clarification, I mean that o= bviously it's sensible to understand the spec you're submitted in a= ll its details, but I'm not sure we need a warranty to that effect. It = equally obviously does no harm, though, so I've no objections either!</= div><div>=C2=A0</div><blockquote class=3D"gmail_quote" style=3D"margin:0px = 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"> - The author of said XEP *HAS* communicated with the XSF (in one of our <br= > rooms, mailing list, summit, etc.) or is vouched by one of the members <br> of XSF before making the XEP that their approach is at least desired and <b= r> may make some sense for initial experiments.<br> <br></blockquote><div><br></div><div>I think this needs tweaking but the es= sential concept - that other people support the submission - is really impo= rtant.</div><div><br></div><div>This does two things:</div><div>- It means = that the kind of mindless AI slop that I think we're all rightly concer= ned about never gets traction - not because it's AI, but because it'= ;s mindless slop.</div><div>- It also means that the early stage workload i= s offloaded from Council. Council only need to take notice of submissions t= hat actually gain any kind of traction in the community.</div><div><br></di= v><div>I would s/one of the members of XSF/participants in the Standards SI= G/ because we've not selected our membership on the basis of technical = scrutiny, and it just feels like if a ProtoXEP is getting positive discussi= on and engagement on the standards list it's probably good for Experime= ntal. Council's effort is then a judgement call on that, rather than ha= ving to scrutinise the specification itself as much.</div><div>=C2=A0</div>= <blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-= left:1px solid rgb(204,204,204);padding-left:1ex"> <br> This solves: XEP "dumps" where we don't know why it is like t= his, who <br> wrote this, is this in good faith, is this wanted/needed etc., solves <br> the work of somebody submitting XEPs without knowing what it is written <br= > in there and also covers liability. Note that the first is already <br> written (and imo that makes llms unable to be used already as openjdk <br> and other projects have stated)<br> <br></blockquote><div><br></div><div>Your parenthetical opinion is far from= universal.</div><div><br></div><div>First, the situation with code is radi= cally different to that for text, and second, the copyrightability of AI ou= tput varies heavily by jurisidiction and human effort involved. Here in the= UK, for instance, we have existing primary legislation (from 1986!) that s= ays LLM output is copyrightable. Other jurisdictions have case law concerni= ng the extremes of AI output, so we know that a XEP coming from a US citize= n solely generated by a single prompt of something like "Create a new = XEP for something" would likely not be copyrightable.</div><div><br></= div><div>But, this doesn't matter, because by requiring the author to w= arrant they can assign copyright, we push that liability onto the author. H= oorah!</div><div>=C2=A0</div><blockquote class=3D"gmail_quote" style=3D"mar= gin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1= ex"> 2. Regarding what approach we can have to actually write this down some <br= > examples are that *DO NOT* ban LLMs:<br></blockquote><div><br></div><div>Al= l these are code based, and our concern here is prose. We're protected = in any case because of the warranty we demand from submitters.</div><div><b= r></div><div>You can stop reading here if you want, the rest is just saying= "I understand the argument but reject it".</div><div><br></div><= div>There's an argument that because the training data for LLMs was (po= ssibly misused) copyright code,=C2=A0the output is a derived work (in copyr= ight terms) of the original training data. This is supported because at lea= st in some cases in the early days, if you asked for particularly niche cod= e, you'll end up with output looking very similar to pre-existing proje= cts. I have not duplicated these results, and I absolutely tried hard - I h= ave the only open source code based on certain specifications which are not= public, so niche in some interesting ways, but I've seen other claims = and have no reason to doubt it can (or could) happen.</div><div><br></div><= div>The problem with this is that few people are making this argument with = prose - people actually complain that AI generated prose doesn't look h= uman enough - and also nobody makes this argument with human generated anyt= hing. I can fully assure you that this email is not considered a derived wo= rk of Asimov or Bujold or Scalzi, despite the fact I have read all three an= almost unhealthy amount. Similarly, the code I write has never been consid= ered a derived work even though I certainly read other people's copyrig= ht code to learn how to write it.</div><div><br></div><div>Whether or not t= he LLMs were originally trained with copyright works in contravention of th= eir licences is a whole other matter, but from a purely legal perspective I= think that's between the copyright owners and the LLM developers. You = are welcome to take an ethical stance on that one on a personal basis.</div= ><div><br></div><div>Dave.</div></div></div> --000000000000317e4e06562d62d3-- --===============3975533669392672897== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ Standards mailing list -- [email protected] To unsubscribe send an email to [email protected] --===============3975533669392672897==--