Newer AI Coding Assistants Are Failing in Insidious Ways

https://spectrum.ieee.org/ai-coding-degrades

134 points · 11 comments · view on lemmy.world

11 Comments

unexposedhazard@discuss.tchncs.de · 39 pts · 207d (1 reply)

I use LLM-generated code extensively in my role as CEO of Carrington Labs, a provider of predictive-analytics risk models for lenders.

Well you know where not to buy now...

panda_abyss@lemmy.ca · 2 pts · 206d

Oh yeah, I read that and thought “this has all the problems of the garden of forking paths”

Based on the description, the whole company is just p-hacking. There’s a reason why nobody uses stepwise feature selection, you just replaced the evaluation/step mechanism with AI.

Feyd@programming.dev · 29 pts · 207d (3 replies)

Unexpected???

TaviRider@reddthat.com · 20 pts · 207d (1 reply)

Unexpected to AI true believers.

WanderingThoughts@europe.pub · 8 pts · 207d

I've lived long enough to go through multiple AI winters. This is just business failure as usual.

luciferofastora@feddit.org · 2 pts · 207d

Unexpected in the same way as you don't expect leopards to eat your face...

nightlily@leminal.space · 25 pts · 207d

Paint huffer surprised when other paint huffers are happy to accept any old solvent.

riskable@programming.dev · 21 pts · 207d (1 reply)

Correction: Newer versions of ChatGPT (GPT-5.x) are failing in insidious ways. The article has no mention of the other popular services or the dozens of open source coding assist AI models (e.g. Qwen, gpt-oss, etc).

The open source stuff is amazing and gets better just as quickly as the big AI options. Yet they're boring so they don't make the news.

count_dongulus@lemmy.world · 17 pts · 207d

That's because OpenAI is in panic mode. They're now spending their resources on making the LLM cheaper to operate and capable of injecting paid results.

atrielienz@lemmy.world · 7 pts · 207d

GiGo.

1Fuji2Taka3Nasubi@piefed.zip · 1 pts · 207d

Not failing, just Skynet doing its work.