Re: ChatGPT ???

"Eben Bruyns (as eben at sdk dot co dot nz)" <[email protected]>
Newsgroups gmane.lisp.lispworks.general
Message-ID <CAEPc2HdXtEeEX=8+a89M6X6uiE3wUJkw9S+XBnxJ9P7pm+H1SA@mail.gmail.com>
This is actually a common pattern with these LLM's, I've caught them lying
and trying to deceive quite a few times.

Regards,

Eben Bruyns


On Tue, 30 Sept 2025 at 13:44, David McClain <[email protected]>
wrote:

> What I was actually looking for were updates to the ISO-226 model.
>
> But after being chided by me for offering stupid, handwaving, arguments
> for its poor exhibit, ChatGPT defended its actions, stating that it was
> only providing examples of code to show how to do the model. It later
> admitted to putting in dummy parameters instead of the actual 2003 ISO-226
> model coefficients.
>
> So in this instance, ChatGPT behaved very badly. I could care less about
> the code. What I wanted to see was an update on the physics and the numeric
> parameter values that have been updated, if any.
>
> Turns out there were no updates.
>
> On Sep 29, 2025, at 14:40, Alexey Veretennikov (as alexey dot veretennikov
> at protonmail dot com) <[email protected]> wrote:
>
> I use Claude to ask some questions, look at the code, get some ideas.
> Never use what it produces since usually it is quite bad style code,
> rolling out function which are already implemented in standard library and
> following bad patterns. But its the same as to get inspiration from  open
> source projects - code might vary a lot in quality, from brilliant ideas to
> garbage. I use it to C/C++ code and never for Lisp.
>
> Copilot Enterprise does pretty good job in guessing what I want to write
> next, and sometimes I auto complete with it boilerplate code (loops over
> iterators etc). It works pretty well for writing unit tests with a lot of
> repeating declarations, after you correct its suggestions a couple of
> times, it starts to be better at following the patterns. It is horrible to
> do anything with math though, but could easily type in equations straight
> from wikipedia in C/C++ form in a style following the code in the file. I
> wouldn't say it increased my productivity a lot, but maybe 5-10%.
>
> So in general it is useful, but not to the extent community self
> proclaimed leaders want us to believe. I think for example the impact of
> LLVM on the industry was way bigger than any modern AI tooling for a tiny
> fraction of the cost.
> Too bad so much resources went as usual wrong way.
>
> Saying that, I don't think it would be a bad idea to have copilot
> integration in LW in a chat mode as a tool with ability to easily add
> context(open buffers) to ask questions, like gptel Emacs extension does. I
> even wrote it to LW support as a feature request. This tooling is here to
> stay and all modern IDEs should support it (even Emacs does with
> gptel/copilot.el).
>
> BR,
> /Alexey
>
> On Monday, September 29th, 2025 at 22:21, Eben Bruyns (as eben at sdk dot
> co dot nz) <[email protected]> wrote:
>
> I'll see you and raise you!
>
> I had Claude code just flat out lie to me about running tests and the
> results. I've also had it refuse to do repetitive tasks, the output was
> "this is too tedious", it wrote a script that did the job wrong and called
> it production ready. It's honestly uncanny how much it behaves like an over
> confident bad junior dev! The number of times you'll see "this is perfect,
> flawless, ship it" when it doesn't even pass basic smoke tests is unreal.
>
> Regards,
>
> Eben Bruyns
>
>
> On Mon, 29 Sept 2025 at 16:44, David McClain (as dbm at
> refined-audiometrics dot com) <[email protected]> wrote:
>
>> I just had a session with ChatGPT, wherein I caught it outright
>> fabricating information to me. Not just once, but three times.
>>
>> It likes to present examples using Python code. I translated its Python
>> code to Lisp to double check the answers it gives.
>>
>> I asked it for a function to convert dBSPL to Phons, and back. If you
>> don’t know what that is, it isn’t really important for this discussion. It
>> has to do with Equal-Loudness contours and the Fletcher-Munson curves for
>> human hearing.
>>
>> By definition, the Phons scale equals the dBSPL scale at 1 kHz. It varies
>> at other frequencies in a nonlinear manner.
>>
>> But ChatGPT’s model produces 91 Phon for 40 dBSPL at 1 kHz. It is off by
>> more than 50 dB !!
>>
>> When I told it about the error, it apologized profusely, stating it knew
>> where the error was. GPT said the error was because it didn’t take into
>> account the threshold of hearing.
>>
>> But that would have produced an error of absolute magnitude less than 1
>> dB. So once again, it is lying to me. I told it to stop giving me
>> handwaving excuses for its errors.
>>
>> It apologized again, with another silly handwaving argument, and produced
>> another answer referring to Python’s iso226 model. I demanded to see that
>> Python code and it finally gave me a URL that actually pointed to the
>> source code. That is the final straight answer.
>>
>> It took me an hour to get a straight, correct, answer from ChatGPT. It
>> willfully misleads. So watch out when using it!!
>>
>> I cannot imagine a clueless management class individual actually running
>> their company on the promise of such flimsy AI. At best, ChatGPT might
>> serve as an encyclopedia, but should not be trusted unless you know more
>> about the subject that it does.
>>
>>
>>
>> _______________________________________________
>> Lisp Hug - the mailing list for LispWorks users
>> [email protected]
>> http://www.lispworks.com/support/lisp-hug.html
>>
>
>
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.