Major benchmarks measure what AI can do. None measure whether it does what you mean: the distance between what you ask an AI to do and the unspoken assumptions about how you want the AI to do it. We propose a new metric: the Genie coefficient.
Apparently Schneier never heard about alignment, which is a whole field with benchmarks and techniques.
Much of AI is a scam. Probably the large majority.
I'm not sure you've quite captured how overfitting functions in current language models. Nor how baselines and benchmarks are used for machine learning products. For serious models of the current year, overfitting is something that happens when you train too long on the same data. It happens less if you add new data, and has nothing to do with new benchmarks and evaluations. You can read about overfitting more generally (mostly the classical kind) on wikipedia.
I think Bruce Schneier has a reasonable understanding of how benchmarks, public policy, large language models, and overfitting are at work here.
OK buddy. It's been running your auto correct, search, a bunch of medical scans, and map routing for awhile now with no complaints. But now that we've learned some LLM vocabulary...
11 Comments
keepthepace@tarte.nuage-libre.fr · 1 pts · 23d
Apparently Schneier never heard about alignment, which is a whole field with benchmarks and techniques.
FiniteBanjo@feddit.online · 0 pts · 24d
AI is a scam and the more constraints you add the less it functions, called "overfitting"/
Artisian@lemmy.world · 1 pts · 24d
Much of AI is a scam. Probably the large majority.
I'm not sure you've quite captured how overfitting functions in current language models. Nor how baselines and benchmarks are used for machine learning products. For serious models of the current year, overfitting is something that happens when you train too long on the same data. It happens less if you add new data, and has nothing to do with new benchmarks and evaluations. You can read about overfitting more generally (mostly the classical kind) on wikipedia.
I think Bruce Schneier has a reasonable understanding of how benchmarks, public policy, large language models, and overfitting are at work here.
FiniteBanjo@feddit.online · -3 pts · 24d
Everything LLM is a scam and a large portion of other neural network tech has also proven unreliable.
venusaur@lemmy.world · 3 pts · 24d
That’s one way for everything you say to lose credibility.
FiniteBanjo@feddit.online · 0 pts · 23d
OK, Slopper.
venusaur@lemmy.world · 1 pts · 23d
Ah, confirmed bias
hohoho@lemmy.world · 2 pts · 24d
Would you please cite your sources? I’d like to read more about this topic.
FiniteBanjo@feddit.online · 0 pts · 23d
While I'm at it I'll show you the proof that god definitively does not exist. /s
Artisian@lemmy.world · 0 pts · 23d
OK buddy. It's been running your auto correct, search, a bunch of medical scans, and map routing for awhile now with no complaints. But now that we've learned some LLM vocabulary...
FiniteBanjo@feddit.online · 0 pts · 23d
lmfao
so far off the mark with that one.