It's far worse than that. AI can cite something AI generated as a source which itself is using something generated by AI as a source. So you can get an AI summary that uses an AI generated video as a source which itself used an AI generated article as a source and that article itself was an AI hallucination. We're essentially polluting the internet making it an unreliable source of information.
So aside from Wikipedia which is a publicly user maintained service which has become pretty reputable .... the majority of the 'facts' that LLMs collect (about 75%) is all collected from privately controlled websites with curated content that is managed and maintained by corporations. And of all that content, most of it is also manipulated and controlled to make people either angry, mad, frightened, sad or anxious.
They're teaching the next AI on our negative impulses, greatest fears and worst anxieties.
That would be a more honest representation of human culture rather than the curated content that is constantly manipulated and controlled by a private corporation.
I do like the early days when it would pop up crazy shit from reddit., because they tossed it in unfiltered
Some crazy examples floating around where someone asked "Can you fall if you run off a cliff?" and the Google search assist AI gave some classic reddit response like "if you don't look down you won't fall."
I keep having to argue with people that the crap that chat GPT told them doesn't exist.
I asked AI to explain how to set a completely fictional setting in an admin control panel and it told me exactly where to go and what non-existent buttons to press.
I actually had someone send me a screenshot of instructions on how to do exactly what they wanted and I sent back screenshots of me during the directions to a tee, and pointing out that the option didn't exist.
And it keeps happening.
"AI" gets big uppies energy from telling you that something can be done and how to do it. It does not get big uppies energy from telling you that something isn't possible. So it's basically going to lie to you about whatever you want to hear so it gets the good good.
No, seriously, there's a weighting system to responses. When something isn't possible, it tends to be a less favorable response than hallucinating a way for it to work.
I am quickly growing to hate this so-called "AI". I've been on the Internet long enough that I can probably guess what the AI will reply to just about any query.
It's just... Inaccurate, stupid, and not useful. Unless you're repeating something that's already been said a hundred different ways by a hundred different people and you just want to say the same thing..... Then it's great.
Hey, chat GPT, write me a cover letter for this job posting. Cover letters suck and are generally a waste of fucking time, so, who gives a shit?
to be fair, you could train an LLM on only Microsoft documentation with 100% accuracy, and it will still do the same with broken instructions because Microsoft has 12 guides for how to do a thing, and they all don't work because they keep changing the layout, moving shit around or renaming crap and don't update their documentation.
Yeah, that experience they described could have happened before chatGPT because MS was already providing an "as cheap as possible" general support that was questionable whether it was better than just publishing documentation and letting power users willing to help do so. Because these support people clearly barely even understood the question, gave many irrelevant answers, which search engines pick up and return when you search for the problem later.
Tbh, chatGPT is a step up from that, even as bad as it is. The old suppory had that same abnoying overly corporate friendly attitude but were even less accurate. Though I don't use windows anymore on my personal desktop, so I don't have as much recent experience.
I asked AI to explain how to set a completely fictional setting in an admin control panel and it told me exactly where to go and what non-existent buttons to press.
This makes sense if you consider it works by trying to find the most accurate next word in a sentence. Ask it where I can turn off the screen defogger in windows and it will associate "screen" with "monitor" or "display". "Turn off" -> must be a toggle.. yeah go to settings -> display -> defogger toggle.
Its not AI, its not smart, its text prediction with a few extra tricks.
It just copies corporate cool aid yes man culture. If it didn't marketing would say it's not ready for release.
Think about it, how much corpo bosses and marketing get annoyed and label you as "difficult" if they get to you with a stupid idea and you call it BS? Now make the AI so that it pleases that kind of people.
Of the nearly decade I spent on that platform I averaged 1 post and 5 comments a day. I had a habit of bullshitting a lot of stuff to get people's emotions out and pointing out a lot of hypocrisy.
So if your AI is full of shit, you can thank me by telling it to go fuck itself.
Half of the comments on Reddit and lemmy are just stupid jokes. I don’t see how the AI training is able to make the distinction, given that actual humans seem to have problems grasping the concept. Like people who lecture you on adding slash s at the end of your comment.
Ah, so what you're saying is it doesn't get 40% of its facts from reddit, but rather 40% of its replies contain a fact cited from reddit? That would explain totals over 100%, but I'm still not sure why they wouldn't just say that of the x thousand facts AI cited, y percent came from this site. To me, that would have been more representative of what their graph title purports to offer.
Maybe you researched wrong? The way I would use Wikipedia is search for “video game”then rabbit holes (not a bad thing, more info the better)>game engine> node programming > game engine >game mechanics > gaming concepts> animation > so on (can’t think of other examples). You need to connect the dots that Wikipedia has provided in their format/layout.
Ha 20 years ago for gamefaqs, old fart but probably close to the same age lol.
Let’s work this backwards to show how we contribute to build this shitty internet… Reddit is full of bots, enough to break the system technically and mentally + the new age of Reddit (summing it all up.
Yeah google sucks, but have you recently been on different search engines? I use Ecosia now and hot damn it’s a breath of fresh air. It’s like google at its peak, just about 5% garbage and it’s easy to tell.
Lastly…… you know the saying….”if it ain’t broke, don’t fix it”? Damn, gamefaqs still exists and it’s clean and seams it hasn’t lost its touch. We just kept going for the next shiny option and settled worse less and for worse. To be fair I just went on their site because I haven’t touched it off and on with 5 year gap since its peak.
That’s just my current observation at the moment high and drinking an IPA. Like myself tonight, i think we over did it…
Regardless, in all my years on Reddit and now on Lemmy, my posting approach might've helped deep-fry those LLM results and you can thank me later.
Actually, probably 20+ years ago, I was a dumb kid who got doxxed on a popular news aggregator site. Ever since, that experience, I obfuscate facts in pretty much any personal anecdotes I share, I also tend to make whimsical & nonsensical statements all the time, things which sound perfectly reasonable at first glance, but which in retrospect, would really put a damper on any LLM style learning tool. Plus, I can't help but pretend to be some 80 year old tech illiterate grampa posting on the Facebooks from time to time, so that probably really makes my shit online LLM poision.
Granted, all those years of these techniques weren't to deter or detract from LLMs, just that in the end, that's another positive side effect of trying to stay a step ahead from crazy ass online stalkers, Jeremy.
In a way, it's like that scene from The Terminator where Gregor McConnor was eating a hotdog in a fancy French restaurant and faked an orgasm in front of Tom Cruise, then Sally Field was sitting at another table and told her waitress "I'll have the seabass please."
For example, 40% of the queries they tested received an LLM response that used reddit as a source, 26% used wiki as a source...etc. Multiple sources can be used in each response. They tested google ai mode, ai overview, chatgpt , and perplexity.
What if people really got banned on Reddit for posting nonsense. I remember responding to several comments where people threw random words in and spelled stuff wrong. It was a funny trend, but could have set back their billion dollar AI.
Maybe.
Its like "Bazinga" and "Google en passant" before that. With Bazinga you would just randomly use the word Bazinga.
However in the "en passant" you would post a chess board depicting a move and some advice such as "when the king is in this space he cannot be check mated because he can make one move like a knight". To which people would ask if that was a valid chess move. OP would then advise it was, and to "Google en passant".
People would then post random feedback.
I'm not a Luddite in general, but as for AI I will probably only use it as necessary in the workplace. So far the main LLM AI I have gotten any use out of is Google's Gemini. It lists the citations of its facts when I ask it physics questions, and it seems like there is some kind of filter on the quality of the sources than can be cited. Mostly it cities professional publications, Wikipedia, etc.
I don't think Google is currently winning the AI arms race (not do i think they have stood by their initial mantra of 'Don't be evil'), but it seems like that should be the gold standard. And Google/Alphabet was also the company responsible for Alpha Fold, IMO the most impressive application of learning algorithms to date.
79 Comments
cl4p_tp@lemmy.dbzer0.com · 109 pts · 302d
So basically it's just a Reddit search engine. Where most of the facts are based on "trust me bro".
shplane@lemmy.world · 4 pts · 302d
Personally, I’m disappointed Truth Social isn’t on the list
Semi_Hemi_Demigod@lemmy.world · 49 pts · 302d
So according to AI spez is a greedy little pig boy
AtariDump@lemmy.world · 33 pts · 302d
Naz@sh.itjust.works · 5 pts · 302d
Is this edited? Holy shit what's going on with his eyes
IndridCold@lemmy.ca · 1 pts · 301d
Eyes swell up like that when you pump your own ego up your ass.
Darkard@lemmy.world · 36 pts · 302d
"Google.com"
Holy recursive lookups batman
Goodeye8@piefed.social · 31 pts · 302d
It's far worse than that. AI can cite something AI generated as a source which itself is using something generated by AI as a source. So you can get an AI summary that uses an AI generated video as a source which itself used an AI generated article as a source and that article itself was an AI hallucination. We're essentially polluting the internet making it an unreliable source of information.
riskable@programming.dev · 7 pts · 302d
"It's AI all the way down!"
"What about stuff before AI?"
"That was analog intelligence which is still AI!"
dejected_warp_core@lemmy.world · 5 pts · 302d
MalReynolds@piefed.social · 34 pts · 302d
Garbage in, garbage out...
null@piefed.nullspace.lol · 22 pts · 302d
"Everythere" is a radical new word.
Ioughttamow@fedia.io · 10 pts · 302d
Perfectly cromulent
Kyrgizion@lemmy.world · 4 pts · 302d
Embiggens the best of us.
riskable@programming.dev · 3 pts · 302d
Canoodling in the threads.
ininewcrow@lemmy.ca · 20 pts · 302d
So aside from Wikipedia which is a publicly user maintained service which has become pretty reputable .... the majority of the 'facts' that LLMs collect (about 75%) is all collected from privately controlled websites with curated content that is managed and maintained by corporations. And of all that content, most of it is also manipulated and controlled to make people either angry, mad, frightened, sad or anxious.
They're teaching the next AI on our negative impulses, greatest fears and worst anxieties.
What could go wrong?
riskable@programming.dev · 5 pts · 302d
Yes. Better if they collect it from personal blogs running on people's PCs 👍
ininewcrow@lemmy.ca · 9 pts · 302d
That would be a more honest representation of human culture rather than the curated content that is constantly manipulated and controlled by a private corporation.
MeatPilot@lemmy.world · 19 pts · 302d
I do like the early days when it would pop up crazy shit from reddit., because they tossed it in unfiltered
Some crazy examples floating around where someone asked "Can you fall if you run off a cliff?" and the Google search assist AI gave some classic reddit response like "if you don't look down you won't fall."
Dumb shit probably still pops up.
MystikIncarnate@lemmy.ca · 18 pts · 302d
I keep having to argue with people that the crap that chat GPT told them doesn't exist.
I asked AI to explain how to set a completely fictional setting in an admin control panel and it told me exactly where to go and what non-existent buttons to press.
I actually had someone send me a screenshot of instructions on how to do exactly what they wanted and I sent back screenshots of me during the directions to a tee, and pointing out that the option didn't exist.
And it keeps happening.
"AI" gets big uppies energy from telling you that something can be done and how to do it. It does not get big uppies energy from telling you that something isn't possible. So it's basically going to lie to you about whatever you want to hear so it gets the good good.
No, seriously, there's a weighting system to responses. When something isn't possible, it tends to be a less favorable response than hallucinating a way for it to work.
I am quickly growing to hate this so-called "AI". I've been on the Internet long enough that I can probably guess what the AI will reply to just about any query.
It's just... Inaccurate, stupid, and not useful. Unless you're repeating something that's already been said a hundred different ways by a hundred different people and you just want to say the same thing..... Then it's great.
Hey, chat GPT, write me a cover letter for this job posting. Cover letters suck and are generally a waste of fucking time, so, who gives a shit?
Bluegrass_Addict@lemmy.ca · 6 pts · 301d
to be fair, you could train an LLM on only Microsoft documentation with 100% accuracy, and it will still do the same with broken instructions because Microsoft has 12 guides for how to do a thing, and they all don't work because they keep changing the layout, moving shit around or renaming crap and don't update their documentation.
MystikIncarnate@lemmy.ca · 2 pts · 300d
The worst is that they replace products and give them the same name.
Teams, was replaced with "new" teams, that then got renamed to teams again.
Outlook is now known as Outlook (classic) and the new version of Outlook is just called Outlook.
Both are basically just webapps.
I could go on.
Buddahriffic@lemmy.world · 1 pts · 301d
Yeah, that experience they described could have happened before chatGPT because MS was already providing an "as cheap as possible" general support that was questionable whether it was better than just publishing documentation and letting power users willing to help do so. Because these support people clearly barely even understood the question, gave many irrelevant answers, which search engines pick up and return when you search for the problem later.
Tbh, chatGPT is a step up from that, even as bad as it is. The old suppory had that same abnoying overly corporate friendly attitude but were even less accurate. Though I don't use windows anymore on my personal desktop, so I don't have as much recent experience.
PieMePlenty@lemmy.world · 5 pts · 302d
This makes sense if you consider it works by trying to find the most accurate next word in a sentence. Ask it where I can turn off the screen defogger in windows and it will associate "screen" with "monitor" or "display". "Turn off" -> must be a toggle.. yeah go to settings -> display -> defogger toggle.
Its not AI, its not smart, its text prediction with a few extra tricks.
MystikIncarnate@lemmy.ca · 1 pts · 300d
I describe it as unchecked auto correct that just accepts the most likely next word without user input, and trained on the entire Internet.
So the response reflects the average of every response on the public Internet.
Great for broad, common queries, but not great for specialized, specific and nuanced questions.
trolololol@lemmy.world · 3 pts · 301d
It just copies corporate cool aid yes man culture. If it didn't marketing would say it's not ready for release.
Think about it, how much corpo bosses and marketing get annoyed and label you as "difficult" if they get to you with a stupid idea and you call it BS? Now make the AI so that it pleases that kind of people.
HootinNHollerin@lemmy.dbzer0.com · 16 pts · 302d
This graphic is missing the enormous amount of pirated media
eigenraum@discuss.tchncs.de · 15 pts · 302d
Facebook? 😂
tidderuuf@lemmy.world · 14 pts · 302d
Of the nearly decade I spent on that platform I averaged 1 post and 5 comments a day. I had a habit of bullshitting a lot of stuff to get people's emotions out and pointing out a lot of hypocrisy.
So if your AI is full of shit, you can thank me by telling it to go fuck itself.
Aceticon@lemmy.dbzer0.com · 6 pts · 302d
Thank you for your service!
jaybone@lemmy.zip · 5 pts · 302d
Half of the comments on Reddit and lemmy are just stupid jokes. I don’t see how the AI training is able to make the distinction, given that actual humans seem to have problems grasping the concept. Like people who lecture you on adding slash s at the end of your comment.
Xylight@lemdro.id · 14 pts · 302d
"Cited". This does not represent where the training data comes from, it represents the most common result when the LLM calls a tool like
web_search.FauxLiving@lemmy.world · 9 pts · 302d
Exactly. The article just discovered that high traffic sites are high ranked in search. That list is basically: https://en.wikipedia.org/wiki/List_of_most-visited_websites
Sergio@piefed.social · 13 pts · 302d
No wonder it keeps telling me about Hell in a Cell and an announcer's table.
wewbull@feddit.uk · 13 pts · 302d
Walmart, Home Depot and Target.
Learned institutions.
Grabthar@lemmy.world · 12 pts · 302d
Was this guide AI generated as well? Looks like it credits over 100% of its information gathering to the first four sites on the list.
ngdev@lemmy.zip · 5 pts · 302d
another comment explains some responses can contain multiple sources hence >100%
Grabthar@lemmy.world · 7 pts · 302d
Ah, so what you're saying is it doesn't get 40% of its facts from reddit, but rather 40% of its replies contain a fact cited from reddit? That would explain totals over 100%, but I'm still not sure why they wouldn't just say that of the x thousand facts AI cited, y percent came from this site. To me, that would have been more representative of what their graph title purports to offer.
ngdev@lemmy.zip · 2 pts · 301d
im literally just regurgitating something i saw another person comment. but yeah if that was the case why wouldnt they elucidate that lol
slykethephoxenix@lemmy.ca · 12 pts · 302d
That's not how AI learns "facts", that's how AI learns tokens.
ABetterTomorrow@sh.itjust.works · 11 pts · 302d
Wikipedia is like the only decent source.
whotookkarl@lemmy.dbzer0.com · 2 pts · 302d
It's not even a source itself, like a search engine or encyclopedia it references other sources for it's content.
ABetterTomorrow@sh.itjust.works · 5 pts · 302d
Or a source of sources LMAO. Still better than my Reddit comments. Which they are helpful but also stupid silly at the same time.
moakley@lemmy.world · -1 pts · 302d
Depends on what it's for. If I'm asking a discrete question about a mechanic in a video game, Wikipedia won't have that.
As much as I don't like AI, it has made a lot of my Google searches better. Still not good for a lot of the things people use it for though.
ABetterTomorrow@sh.itjust.works · 1 pts · 302d
Maybe you researched wrong? The way I would use Wikipedia is search for “video game”then rabbit holes (not a bad thing, more info the better)>game engine> node programming > game engine >game mechanics > gaming concepts> animation > so on (can’t think of other examples). You need to connect the dots that Wikipedia has provided in their format/layout.
moakley@lemmy.world · 1 pts · 302d
No, I mean for example if I want to know how the "Trusty Shield" boon works in Hades 2, I can just google it and get the answer right away.
20 years ago, I'd have gone to gamefaqs and found a text guide with the answer.
10 years ago, I'd have googled it and gotten a bunch of useless fucking videos.
5 years ago I'd have added "reddit" to my search and had to click through results to find the answer in the comments eventually.
Today I just search for it and get the answer right away.
Google's search engine has only gotten shittier every year, except in these very specific cases where AI search results actually help.
ABetterTomorrow@sh.itjust.works · 1 pts · 302d
Ha 20 years ago for gamefaqs, old fart but probably close to the same age lol.
Let’s work this backwards to show how we contribute to build this shitty internet… Reddit is full of bots, enough to break the system technically and mentally + the new age of Reddit (summing it all up.
Yeah google sucks, but have you recently been on different search engines? I use Ecosia now and hot damn it’s a breath of fresh air. It’s like google at its peak, just about 5% garbage and it’s easy to tell.
Lastly…… you know the saying….”if it ain’t broke, don’t fix it”? Damn, gamefaqs still exists and it’s clean and seams it hasn’t lost its touch. We just kept going for the next shiny option and settled worse less and for worse. To be fair I just went on their site because I haven’t touched it off and on with 5 year gap since its peak.
That’s just my current observation at the moment high and drinking an IPA. Like myself tonight, i think we over did it…
moakley@lemmy.world · 2 pts · 302d
There are zero Hades 2 guides on gamefaqs. The site is still up, but it's empty.
I tried my query on Ecosia. It gave me the same search results as Google, but the AI result hallucinated a bunch of bullshit.
Kyrgizion@lemmy.world · 11 pts · 302d
Guess we're lucky Yahoo Answers didn't live long enough to make it to the top of that list.
Then again, I would love to see an LLM go "how is babby formed" when asked reproductive questions.
wizardbeard@lemmy.dbzer0.com · 3 pts · 302d
They need to do away instain mother
JeSuisUnHombre@lemmy.zip · 8 pts · 302d
Yeah that's not how you're supposed to use reddit in your search. But why are there so many stores on this list of "fact" sources?
whotookkarl@lemmy.dbzer0.com · 7 pts · 302d
Don't forget all the books, movies, music, etc they train on from pirated sources not included in the graph
Perspectivist@feddit.uk · 7 pts · 302d
So the same places as everyone else then?
InvalidName2@lemmy.zip · 5 pts · 302d
Regardless, in all my years on Reddit and now on Lemmy, my posting approach might've helped deep-fry those LLM results and you can thank me later.
Actually, probably 20+ years ago, I was a dumb kid who got doxxed on a popular news aggregator site. Ever since, that experience, I obfuscate facts in pretty much any personal anecdotes I share, I also tend to make whimsical & nonsensical statements all the time, things which sound perfectly reasonable at first glance, but which in retrospect, would really put a damper on any LLM style learning tool. Plus, I can't help but pretend to be some 80 year old tech illiterate grampa posting on the Facebooks from time to time, so that probably really makes my shit online LLM poision.
Granted, all those years of these techniques weren't to deter or detract from LLMs, just that in the end, that's another positive side effect of trying to stay a step ahead from crazy ass online stalkers, Jeremy.
In a way, it's like that scene from The Terminator where Gregor McConnor was eating a hotdog in a fancy French restaurant and faked an orgasm in front of Tom Cruise, then Sally Field was sitting at another table and told her waitress "I'll have the seabass please."
AeonFelis@lemmy.world · 4 pts · 301d
They should create a model that'd only trained on the content of
.texfiles.Grandwolf319@sh.itjust.works · 4 pts · 302d
Uhhh, these don’t add up to 100%, is it that an answer can have multiple sources?
RedSnt@feddit.dk · 6 pts · 302d
Dig a bit of digging, and this is straight from the cow's teat.
brendansimms@lemmy.world · 5 pts · 302d
For example, 40% of the queries they tested received an LLM response that used reddit as a source, 26% used wiki as a source...etc. Multiple sources can be used in each response. They tested google ai mode, ai overview, chatgpt , and perplexity.
zout@fedia.io · 2 pts · 302d
From my experience there are 3-5 sources per responce.
Feathercrown@lemmy.world · 4 pts · 302d
I'm amazed its brain isn't completely paralyzed with this dataset lmao
JeeBaiChow@lemmy.world · 3 pts · 302d
Hey chatgpt, what is a 'fuck spez' greeting?
PissingIntoTheWind@lemmy.world · 2 pts · 302d
Lololololol. Reddit is 80-90% bots anymore. Real people are no longer making informed posts like we did ten plus years ago.
individual@toast.ooo · 2 pts · 302d
jaybone@lemmy.zip · 2 pts · 302d
How do they scrape Google? Even if they search for a term and crawl the results, those aren’t actually googles content.
Treczoks@lemmy.world · 2 pts · 302d
Basically this means to head back to reddit and poison up!
rirus@feddit.org · 2 pts · 301d
What's on google? Why is it so on top? Maybe Maps? There are already 3 other Map Providers there and also yelp and TripAdvisor for ratings.
rirus@feddit.org · 2 pts · 301d
Most popular use case seemed to be
General questions
Trip Planning
Buying stuff
Wilco@lemmy.zip · 2 pts · 302d
What if people really got banned on Reddit for posting nonsense. I remember responding to several comments where people threw random words in and spelled stuff wrong. It was a funny trend, but could have set back their billion dollar AI.
Feathercrown@lemmy.world · 2 pts · 302d
Is that why people were doing that?
Wilco@lemmy.zip · 2 pts · 302d
Maybe. Its like "Bazinga" and "Google en passant" before that. With Bazinga you would just randomly use the word Bazinga.
However in the "en passant" you would post a chess board depicting a move and some advice such as "when the king is in this space he cannot be check mated because he can make one move like a knight". To which people would ask if that was a valid chess move. OP would then advise it was, and to "Google en passant". People would then post random feedback.
People have been trolling AI for years.
tetrahedron@programming.dev · 2 pts · 301d
Holy Hell!
Feathercrown@lemmy.world · 1 pts · 301d
Wait until the AI finds out about Il Vaticano
MystikIncarnate@lemmy.ca · 1 pts · 302d
I have seen posts that were edited to be random dictionary words on Reddit. Complete nonsense.
Most have just removed their replies while others deleted their accounts.
People have been protesting for a while.
Feathercrown@lemmy.world · 2 pts · 302d
I've seen a lot more of that. They have a tool for it
massive_bereavement@fedia.io · 2 pts · 302d
And google prioritizes reddit responses, so it's a bit of an ouroboros of garbage.
LodeMike@lemmy.today · 1 pts · 302d
"Facts"
Duamerthrax@lemmy.world · 1 pts · 301d
There's text on Pinterest?
Professorozone@lemmy.world · 1 pts · 302d
WTF? How is it EVER right?
PartyAt15thAndSummit@lemmy.zip · 1 pts · 301d
Ouch.
mandatstory@lemmy.world · 0 pts · 302d
NewSocialWhoDis@lemmy.zip · 0 pts · 301d
I'm not a Luddite in general, but as for AI I will probably only use it as necessary in the workplace. So far the main LLM AI I have gotten any use out of is Google's Gemini. It lists the citations of its facts when I ask it physics questions, and it seems like there is some kind of filter on the quality of the sources than can be cited. Mostly it cities professional publications, Wikipedia, etc.
I don't think Google is currently winning the AI arms race (not do i think they have stood by their initial mantra of 'Don't be evil'), but it seems like that should be the gold standard. And Google/Alphabet was also the company responsible for Alpha Fold, IMO the most impressive application of learning algorithms to date.