Re: reading LLM code may rule out white room re-implementation; was: On keybindings and the slow erosion of help's utility
Jean Louis <[email protected]>
| Newsgroups | gmane.emacs.devel |
|---|---|
| Organization | GNU Support |
| Message-ID | <[email protected]> |
On 2026-07-17 18:57, Dr. Arne Babenhauserheide wrote: >> "Dr. Arne Babenhauserheide" <[email protected]> writes: >>> So if code under an incompatible license may prove to have been >>> material >>> to the LLM output, and if the legal landscape changes so LLM >>> processing >>> of code preserves some copyright (which currently seems likely), then >>> LLM code posted here may make it illegal for anyone who read it to >>> write >>> a free software implementation of the idea. Those are such extreme ideas yet not well researched, and it should not considered legal advise. Yet it sounds as such. Unless you are lawyer specialized in copyrights, are you? Millions people reading some code cannot automatically become defendants in the court because your statement that it makes "illegal for anyone who read it to write a free software implementation of the idea". In other words you deal with extreme "legally sounding" cautions by spreading F.U.D. I do expect from you to research better the subject as to avoid generalizations and negative influences. For "anyone" to be illegal in that situation you described, the one who is suing against "anyone" would need to prove that the accused actually accessed proprietary materials and in the same time knew or should have known it is proprietary. The situation you described in the above quoted text is thus almost never true, yet it sounds like legal advise by doctor... Please research the subject better as your statements have weight due to your title and if they are incorrect, such can influence people reading it on the mailing list. > Posting proprietary code that you do not own is already illegal by law, > so no extra rule was needed and I don’t think that equivalence applies. Even this statement isn't well researched and not granted as such as person would need to know it is proprietary. Stumbling upon proprietary code can take place without such knowledge. That some possible situations could end up in court is true both for the LLM- and not LLM-related issues. > The risk with LLM-code is that at the moment it is legal to post it, > but > a court ruling could change that and it’s unclear what that would mean > for all code built on it. That statement is unfounded. No research at all. First you are putting millions of already existing models all in one group as "LLM-code". That is major logical flaw in your generalized statements. You are lumping together all the code generated by GPT-4 with the fine-tuned model that only saw my internal GPL code base, with Qwen, and Olmo models and Apertus 70B. You mix MIT, Apached, BSD, proprietary models, even my private GNU GPL LLM outputs all as one, calling it "LLM-code". And that information presented in such generalized way is FUD (spreading fears, uncertainties and doubts, especially on this mailing list. There is obvious demand for the LLM assistance on this mailing list. Yet, you brought no solution. Just some type of quoted legal advise. I suggest you do following, run some of those models, train them yourself, go through the process. Even if you take smaller models they can run on your computer. You can research the subject in 7 days, do some work, download models and research their origins, datasets. Then come back and review all your statements you posted here. Then tell me if you have changed the attitude. And there is nothing preventing the GNU project to do the same research and find those LLMs which can be used in LLM code generations. I have downloaded more than 1000 different LLMs, and have done the exercise and I an just recommend it to others. > For LLM code you currently can’t know whether you will own it a year > from now. Inaccurate. Not true! Reasons explained above. And yes, it is quite possible to know whether LLM code generated one will own it or not. -- Jean Louis