New compression technique

1.2k points · 160 comments · view on lemmy.world

160 Comments

MalReynolds@slrpnk.net · 314 pts · 2d (3 replies)

Trying to find it funny, but in 2026 it's way too close to the bone and just makes me sad.

UnderpantsWeevil@lemmy.world · 106 pts · 2d

Had Sam Altman personally vocalized this article in full sincerity, I would not have been surprised.

jafra@slrpnk.net · 19 pts · 2d

Yeah. I first thought it's about jpeg, then i read ai and i wasnt sure if its worrying stupidity reporting or bad satire. Edit: i thought '92 i mean

AmyAye@nord.pub · 6 pts · 1d

Yeah, like how they are destroying rare books.

I mean, you don't get that back just by asking the content theft machine.

DaddleDew@lemmy.world · 180 pts · 2d (4 replies)

Amateur. I can compress entire seasons of a TV series to a few bytes. All I have to do is type its title in Netflix and then BAM, gigabytes of video come out.

rockSlayer@lemmy.blahaj.zone · 63 pts · 2d (2 replies)

Sure, but Netflix has a lot of cache misses for my favorite shows. I found a random website that has compressed matches for thousands more shows and movies than Netflix

3abas@lemmy.world · 3 pts · 2d (1 reply)

I pirate my shit.

msage@programming.dev · 10 pts · 1d

TTJ.jpg

bampop@lemmy.world · 13 pts · 1d

Back in the day we had this neat trick to do 100% compression on photos. Just print them out on pieces of paper, for completely byte-free storage!

Sanctus@anarchist.nexus · 83 pts · 2d (1 reply)

Lmfao "tears of joy"

toy_boat_toy_boat@lemmy.world · 6 pts · 2d

Precisely! We human people all know that tears definitely do not express joy.

Correct?

OldGrayDog@fedinsfw.app · 56 pts · 2d (1 reply)

I've read that the Trump administration is hiring him to archive all of the Epstein files using that format.

some_designer_dude@lemmy.world · 14 pts · 2d

Compressing the Epstein files is easy: “Donald Trump raped children.” What’s my Weissman score?

dylanTheDeveloper@lemmy.world · 49 pts · 1d

This is so sad, Chat GPT generate me an email responding to this article

ShellMonkey@piefed.socdojo.com · 39 pts · 2d (7 replies)

On a similar note, I saw a story a bit back of someone saving input tokens by feeding the bot an image of a wall of text rather than the text itself and having it read the image via OCR.

Satire and reality are too hard too distinguish these days.

skisnow@lemmy.ca · 8 pts · 2d

I see this one all the time on LinkedIn, the management rubes love it

SorryQuick@lemmy.ca · 3 pts · 1d (2 replies)

Saves input tokens and uses more output because of OCR, and output is priced way higher.

Rooster326@programming.dev · 3 pts · 1d

It's an accounting trick. It uses a different budget....

boonhet@sopuli.xyz · 1 pts · 1d

Also if you're doing anything multi-prompt, that output is re-used as (luckily cached) input tokens for the next prompt anyway lol

topperharlie@lemmy.world · 3 pts · 1d (1 reply)

I mean, back in my days (👴) people sent a photo pasted into a word document, so things haven't change that much...

Natanael@infosec.pub · 2 pts · 1d

They still do

brian@programming.dev · 1 pts · 1d

this is the default compaction method in omp now. afaik it's exploiting the pricing structure of the big providers more than anything

kamen@lemmy.world · 36 pts · 1d (3 replies)

Lossless -> lossy -> mindless.

dellish@lemmy.world · 10 pts · 1d (2 replies)

-> mindy?

hancock@retrolemmy.com · 4 pts · 1d

mindly?

kamen@lemmy.world · 3 pts · 1d

Probably not very.

pressanykeynow@lemmy.world · 28 pts · 1d (5 replies)

Could have just stored those in PiFS https://github.com/philipl/pifs

sukhmel@programming.dev · 18 pts · 1d

That's nice, albeit I want to point out for anyone wondering that this is only conjectured and not guaranteed:

One of the properties that π is conjectured to have is that it is normal, which is to say that its digits are all distributed evenly, with the implication that it is a disjunctive sequence, meaning that all possible finite sequences of digits will be present somewhere in it.

There is no guarantee for any specific sequence to appear in π, but for short chunks chances are better (it's not really a probability, but it's simpler to say and I can't explain in details anyway). That's because (from wiki):

It is widely believed that the (computable) numbers √2, π, and e are normal, but a proof remains elusive.

tetris11@feddit.uk · 1 pts · 1d (3 replies)

It's pronounced "Piff Es" for anyone wondering

Axolotl_cpp@feddit.it · 4 pts · 1d (1 reply)

Idk if it's irony, but the real pronounce is probably /paɪ ɛf-ɛs/ (english)
or /'pi fs/ in every other language that reads letters as they are written normally

sukhmel@programming.dev · 1 pts · 1d

I think it's akin to ‘GIF is pronounced ziph’

beejboytyson@lemmy.world · 1 pts · 1d

Nice datpiff.com

gera@feddit.nu · 27 pts · 1d

TrickDacy@lemmy.world · 24 pts · 2d (7 replies)

a typical jpeg of 20 Mb

Uh what. Literally should be the highest fucking possible quality from a $10K camera if it's that big. My raw images aren't even that big usually.

PieMePlenty@lemmy.world · 18 pts · 1d (1 reply)

Its not typical, but you can get 20 Mb+ jpegs out of an entry level 18MP dlsr. Especially if there's lots of color and at like 5500x3300 resolutions and created with 100% quality preset.
I checked my immich and I have some (and larger), but yeah, not exactly typical.

TrickDacy@lemmy.world · 2 pts · 1d

I am not sure I've ever used the 100 quality setting on jpegs. Many years ago I experimented with that a lot and decided that anything over 90 was not different to my eye but the file size was much bigger, relatively speaking. So yeah I suppose if you did use 100, a 20 MB jpeg image is not hard to reach.

bstix@feddit.dk · 15 pts · 1d (2 replies)

Raw images is where you go wrong. You need at least 10 mb of metadata tags to achieve professional levels of file sizes. How can you even look at a picture without having a full description of all your childhood memories that lead you to take this beautiful picture of yesterdays mac''n'cheese dinner. This is why we need more data centers.

TrickDacy@lemmy.world · 1 pts · 1d (1 reply)

Lol I'm so lost what the joke is here

HerbGrower@slrpnk.net · 2 pts · 1d

Highest quality jpg?

Lucidlethargy@sh.itjust.works · 1 pts · 1d

Next you'll be questioning the tears of joy!

whereitsat@lemmy.zip · 23 pts · 1d

brilliant satire that critiques all of magazine journalism.

i love going to [insert publication here] and reading another article about 'so and so is ready for their next chapter in life.' the so and so always an uninteresting, overly wealthy fuckwit that hasn't accomplished anything other than going to college and having a wealthy parent.

einkorn@feddit.org · 21 pts · 2d

Throw some angel investment cash at this guy already!

M1k3y@discuss.tchncs.de · 18 pts · 1d (1 reply)

The sad thing is that this has been possible for decades using convolutional autoencoders, but with LLMs we forgot that AI architectures other than transformers still exist.

CanadaPlus@lemmy.sdf.org · 6 pts · 1d

Yeah, it's really awful. With any luck, AI winter will follow AI summer, like usual, and the serious people can come out again. Although, aren't CNNs more of a this century thing? I guess two decades is still decades...

IIRC autoencoders actually produce the same image to within our ability to notice, as well.

klar@lemmy.world · 18 pts · 1d (2 replies)

git-llmfs for JPEG, I kind of wish this was real just to laugh at users

Axolotl_cpp@feddit.it · 7 pts · 1d (1 reply)

Before you consider using the project, I highly recommend you read and understand the last paragraph of the LICENSE file. If you are seriously considering using it, it will become important.

This is so funny

For whoever is wondering, the license is MIT

boonhet@sopuli.xyz · 5 pts · 1d

And also

Ironically, aside from the obvious use of LLMs for the actuall project, everything was written by hand. I wanted to use this as a learning opportunity about git filters, so everything you see is my fault.

JasonDJ@lemmy.zip · 15 pts · 2d (7 replies)

All fun and games until you post a story about how clumsy your dad was and now all your photos are updated to show him as just a giant thumb...with thumbs for arms, and legs, and hands and feet, and fingers and toes. Just all thumbs, all the way down. Like tripping on acid with Rita Repulsa.

OR3X@lemmy.world · 21 pts · 2d (4 replies)

Hamartia@lemmy.world · 9 pts · 2d

Dd...ddaadd?

spankysalmon@fedinsfw.app · 4 pts · 2d

All I see is a police man wearing red, I don't get it.

some_kind_of_guy@lemmy.world · 1 pts · 2d

sknaht

Natanael@infosec.pub · 1 pts · 1d

Are you spying on me?

boonhet@sopuli.xyz · 1 pts · 1d (1 reply)

For some reason all my photos now show me having two left hands.

JasonDJ@lemmy.zip · 1 pts · 1d

Sounds like every picture of me dancing.

Shanmugha@lemmy.world · 15 pts · 2d

Here is another one, no compression:

  • read image
  • send a prompt "regenerate image, here is pixel-by-pixel description"
  • enjoy the result

(sarcasm)

Anchorxiety@reddthat.com · 14 pts · 2d (1 reply)

I compress my meals by putting them into a blender to liquify them before consumption!

psx_crab@lemmy.zip · 1 pts · 2d

No no no, you have to freeze it first

bingus@lemmy.zip · 14 pts · 1d

I suppose they weren't worth a 1000 words after all

tal@lemmy.today · 13 pts · 1d (3 replies)

It's humorous, but last I checked, the best general-purpose compressors with the highest levels of compression---even lossless, which is probably not what most people think of when they think of neural nets---are neural net based.

Neural net-based compressors are computationally expensive, which is why we don't normally use them for most day-to-day tasks, but they really can produce really small outputs.

I'm going to take the text of the US Constitution and stick it in a text file.

$ wget https://www.gutenberg.org/cache/epub/5/pg5.txt
$ stat -c %s pg5.txt 
48326

Okay, so 48326 bytes.

Let's do lzo. You'd expect a limited amount of compression --- LZO is "fast" compression, usually only used where compression speed is really important, like where you want to be compressing stuff that's going to be decompressed once and your bottleneck is throughput to disk:

$ lzop <pg5.txt >pg5.txt.lzo 
$ stat -c %s pg5.txt.lzo
24843

Okay, how about gzip? That's Deflate, an older, but pretty-widely-used general-purpose compression algorithm.

$ gzip <pg5.txt >pg5.txt.gz
$ stat -c %s pg5.txt.gz
16660

Okay, what about LZMA? That's a newer, more-CPU-intensive thing that's probably a good general-purpose choice that'll generally give better compression ratios. It's the kind of thing that I'd probably use in a lot of cases. (Personally, these days, I tend to use pixz, which provides both indexed access for tarballs and parallel compression and decompression, which is important for modern processors.)

$ xz <pg5.txt >pg5.txt.xz
$ stat -c %s pg5.txt.xz 
15488

Okay, now PAQ, a neural-net-based compressor:

$ zpaq a pg5.txt.zpaq a pg5.txt -method 5
$ stat -c %s pg5.txt.zpaq 
13063
BradleyUffner@lemmy.world · 12 pts · 1d (1 reply)

"Neutral net based" compression isn't even in the same universe as "compressed to a prompt" via LLM

tal@lemmy.today · 5 pts · 1d

It actually is. I mean, it's building a dictionary off of a variety of content ahead-of-time, rather than training it on the specific item in question, but that's not uncommon for non-general-purpose compressors.

I mean, doing so to a (probably short) prompt is (a) lossy (and I gave a lossless example) and (b) lossy to an extreme degree, to where it's probably not incredibly useful option for the kinds of systems that exist today.

But...existing diffusion models aren't actually intended for this, either. I'd bet that you could train a model to do image compression along these lines, with a large dictionary, that could do usable compression along the lines of what is (jokingly) described in the article. Probably have a larger compressed form than what they're thinking of.

EDIT: At one point in time, about over a quarter-century ago now, I went out and banged on a neural net post-processor for JPEG artifacts. The idea here is that JPEG very probably isn't optimally representing the final image, as a human, using their knowledge of what the world looks like, can manually (if time-consumingly) clean these up. I didn't meet with a lot of success; I only wanted to put a small amount of time into it, and I was working with much weaker hardware than people are running neural nets on today. But that generated a pre-existing dictionary, a pre-trained neural net, off a training corpus of uncompressed images. It didn't try to reconstruct the image from scratch, the way something like this would, just clean up artifacts, but it has that same pre-generated neural net approach.

mrmanager@lemmy.today · 8 pts · 1d

I did not know about PAQ actually, thats cool. :)

ryannathans@aussie.zone · 12 pts · 2d (6 replies)

Well actually....

https://www.nature.com/articles/s42256-025-01033-7

Lossless compression by LLM

Mikina@programming.dev · 14 pts · 2d (1 reply)

Is it a compression when you need presumably gigabites of a model to reconstruct the data?

It's basically the same as saying I can shatter compressions records by hashing the thing and using a rainbow table to decompress it.

ryannathans@aussie.zone · 8 pts · 2d

A roughly 340MB model at fp16 is all it takes to beat xz and zstd by 2x compression on text

urushitan@kakera.kintsugi.moe · 11 pts · 2d
WesternInfidels@feddit.online · 3 pts · 2d (1 reply)

They're using domain-specific LLMs to compress narrow-domain data. Their text compression LLM was trained on, and then tested on, legal text and medical text.

There's no reason one couldn't apply the domain-specific-compressor idea to a conventional lossless text compressor, essentially moving much of the dictionary from the compressed file to the program itself. I don't know if anyone's tried that. I'd like to know how that compares.

ryannathans@aussie.zone · 3 pts · 2d

There's a number of compression projects working on this, this is the first that comes to mind

https://bellard.org/nncp/

https://bellard.org/ts_zip/

BlackLaZoR@lemmy.world · 2 pts · 2d

I'd be more interested in lossy compression tho. Few % of image quality loss for multiplying compression ratio is mega useful.

green_red_black@slrpnk.net · 12 pts · 2d (7 replies)

I know this is a joke but honestly if an engineer figured out how to make a Jpeg 600 bytes it would be impressive (obviously don't destroy the original even if all worked right.)

UnderpantsWeevil@lemmy.world · 45 pts · 2d (1 reply)

if an engineer figured out how to make a Jpeg 600 bytes it would be impressive

Why use many pixel when few pixel do trick?

wonderingwanderer@sopuli.xyz · 6 pts · 2d

If you zoom out really far, you can kinda see it...

NeatNit@discuss.tchncs.de · 21 pts · 2d

something something information theory.

wonderingwanderer@sopuli.xyz · 3 pts · 2d (3 replies)

That seems fundamentally impossible.

If you're just doing plain black and white, no grey scale, even ignoring headers the most you can get out of that is 4800 pixels. Say maybe 120x40 or 75x64.

Add in headers to tell the machine where each pixel actually goes and you're looking at way less.

Want greys or color in your image? Forget about it.

glibg10b@lemmy.zip · 8 pts · 2d (2 replies)

If you're still thinking in pixels, then you're already behind JPEG. There are other ways to compress visual data, such as in the form of its frequency spectrum, or with vector embeddings

wonderingwanderer@sopuli.xyz · 5 pts · 2d (1 reply)

Or in a tensor field consisting of billions of weighted parameters across several matrices which then get multiplied along specific embedded vectors, apparently...

CanadaPlus@lemmy.sdf.org · 2 pts · 1d

To explain a bit more, most of the information in an image is actually stuff we would never notice. The very specific way a few of a pixels slightly deviate from a perfect gradient, for example.

The basic idea of a jpeg is to convert an image to a kind of frequency space, where each pixel corresponds to a specific wave-like pattern. It's reversable, preserving the information. and fast conversion to do, thanks to FFT. Even without, it would only be quadratic time. Then, the algorithm crops it in that new space. (IIRC jpeg actually has a few extra features, as well)

vegafjord@slrpnk.net · 10 pts · 1d

He discovered alt text.

BuckFutter@timeyak.com · 9 pts · 2d (2 replies)

This is satire right? Right? Lmao.

ranzispa@mander.xyz · 8 pts · 2d (1 reply)

Yes.

BuckFutter@timeyak.com · 4 pts · 2d

Sweet, glad we cleared that up 😉

AI_toothbrush@lemmy.zip · 9 pts · 2d (7 replies)

The fuuuucking annoying part is that these weight trained models would be perfect for translation models, compression, etc. An llm is already kind of a really efficient lossy compressor but you could actually make it lossless and an actual compressor if used properly. But instead people are literally telling llms to translate instead of training models that are for translating. The technology isnt the problem itself, its the industry and capitalism.

spankysalmon@fedinsfw.app · 8 pts · 2d (4 replies)

lol no you cannot make it lossless.

TrickDacy@lemmy.world · 4 pts · 2d

Enhance!

sukhmel@programming.dev · 1 pts · 1d

I think, they don't mean lossless compression with LLM, but with a neural network in general. I think it might be possible, but I'm not sure

AI_toothbrush@lemmy.zip · -2 pts · 1d (1 reply)

Yeah you can? A neural network that only relies on wheights and doesnt use random numbers will always give the same output for the same input.

spankysalmon@fedinsfw.app · 2 pts · 1d

So does lossy compression... same output for input of an algorithm is deterministic, not lossless...

ChaoticNeutralCzech@feddit.org · 5 pts · 2d (1 reply)

There are not many places where 95% compression of UTF-8 plain text could outweigh needing a 30GB model in memory and a significant fraction of current LLM inference cost to decompress it − and good luck convincing librarians to adopt it.

Natanael@infosec.pub · 2 pts · 1d

There's a 300 MB library for audio compression using it. If you have large audio libraries it could eventually become worth the tradeoff.

https://huggingface.co/facebook/encodec_32khz

The image versions are probably more useful though.

The important part for these codes is that it has pre-LLM functions for quality metrics to judge if the output is close enough to indistinguishable (preserves detail, doesn't add any).

Although there is also variants deriving a neural net from the media to recreate it from the smaller model.

PattyMcB@lemmy.world · 6 pts · 1d (1 reply)
fruitycoder@sh.itjust.works · 8 pts · 21h

No no IPoAC is far more viable

BilSabab@lemmy.world · 6 pts · 1d (16 replies)

Remember that Sloot thing from late 90s?

Slashme@lemmy.world · 4 pts · 1d (15 replies)

No, what was it?

BilSabab@lemmy.world · 9 pts · 1d (5 replies)

found it - here's a detailed account https://corecursive.com/sloot-digital-coding-system/

IsoKiero@sopuli.xyz · 2 pts · 1d (1 reply)

Very interesting piece of history. Thank you!

BilSabab@lemmy.world · 2 pts · 1d

you're welcome

AlboTheGuy@feddit.nl · 2 pts · 1d (2 replies)

Wow that was quite the read, thanks!

PartyAt15thAndSummit@lemmy.zip · 1 pts · 1d

Yeah, that had me hooked to the screen.

BilSabab@lemmy.world · 1 pts · 1d

you're welcome

BilSabab@lemmy.world · 6 pts · 1d (7 replies)

so basically there was a dutch tech bro in the late 90s who claimed he had developed a compression technique that could turn a movie into just a couple of kilobytes and right when he was going to sign with some investors he mysteriously passed away leaving no documents explaining his invention.

khanh@lemmy.zip · 2 pts · 1d (5 replies)

why would someone kill him over it? what do you think?

BilSabab@lemmy.world · 3 pts · 1d (2 replies)

most likely he just overstressed because he sold folks a bill of goods

CanadaPlus@lemmy.sdf.org · 2 pts · 1d (1 reply)

Yup. It seems a lot more likely than solving a problem that we still can't solve, and that might not even be possible, in the 90's.

BilSabab@lemmy.world · 1 pts · 1d

and now we have AI and even this wondrous wonder of technology can't figure it out.

Nalivai@lemmy.world · 1 pts · 1d (1 reply)

Not saying that's what happened, but if I ever find myself in a similar situation, I'm faking my death and moving halfway across the world in heartbeat. After I get the money of course

BilSabab@lemmy.world · 1 pts · 1d

the whole story screams that kind of thing.

BCsven@lemmy.ca · 1 pts · 1d

And his investors family were stabbed, and another guy relates to it had is car blow up. Random or coordinated?

GodSpeeD808@feddit.nl · 4 pts · 1d
richardisaguy@lemmy.world · 6 pts · 1d (4 replies)

This is not funny, you don't understand how long and how many nights i spend overengineering compression pipelines for family photos and videos...

grrgyle@slrpnk.net · 7 pts · 1d (2 replies)

I've got mine down to 32 byte string to describe the location, datetime, and people in the photograph, with flags for who's smiling and/or blinking.

Knock_Knock_Lemmy_In@lemmy.world · 7 pts · 1d (1 reply)

I want to see your porn metadata.

qarbone@lemmy.world · 4 pts · 1d

Absolutely debauched. Especially when you haven't even bought them coffee yet.

CanadaPlus@lemmy.sdf.org · 1 pts · 1d

Then I bet your stuff is actually good.

Bluewing@lemmy.world · 5 pts · 1d

New Math and vibe coding. Am I right? (insert canned laughter here).

anon_8675309@lemmy.world · 5 pts · 1d

That’s … redefining the word I would think.

Jolteon@lemmy.zip · 4 pts · 1d

It's basically just tricking the AI model to be your cloud storage provider.

bitjunkie@lemmy.world · 4 pts · 1d

🤢

Phoenix3875@lemmy.world · 3 pts · 1d

laughs when reading this

cries when paying thousands of dollars for DLSS

cloudshouter@lemmy.zip · 3 pts · 2d (1 reply)

Two thoughts. First, there is a way to high chance that some people are basically doing this already or are taking inspiration from this. Second, the tell that it’s a joke article is that they bothered to destroy the prints, which implies that they bothered to scan them as well.

sukhmel@programming.dev · 1 pts · 1d

It seems to mean the prints of digital photos they were compressing, as it wouldn't involve scanning

bandwidthcrisis@lemmy.world · 2 pts · 2d (4 replies)

Whatever happened to fractal image compression, which compressed an image by finding parts similar to larger regions of the picture, so that is could reconstruct the image from itself.

BlackLaZoR@lemmy.world · 4 pts · 2d (3 replies)

It only works if the image is actually fractal-like. Real world is a bitch

ChickenLadyLovesLife@lemmy.world · 4 pts · 2d

Sometimes the world is a Mandel Bro and sometimes it isn't.

phonics@lemmy.world · 2 pts · 2d (1 reply)

Were in a fractal just too zoomed in to notice

Wiz@midwest.social · 1 pts · 1d

Actually, a simulation of a fractal

Jankatarch@lemmy.world · 2 pts · 1d

Lossy compression, don't mind.

goatinspace@feddit.org · 1 pts · 13h

Snapz@lemmy.world · 1 pts · 1d

♪ Hey, hey, billy, can you

compress the Buick ? ♪

Well, all right, but

he'll probably Pu-ick.

wallwood@lemmy.ml · 1 pts · 1d

I suspected everything the moment I read 13-year-old-boy

Tenderizer@aussie.zone · 1 pts · 5h

I'd be game for that. I hate the fact that photos of me exist in this world.

PhoenixDog@lemmy.world · -7 pts · 1d (3 replies)

"First, it uses AI"

Nope. Stopped reading. Don't care.

HereIAm@lemmy.world · 35 pts · 1d (2 replies)

Looks like someone got Onioned.

PhoenixDog@lemmy.world · -3 pts · 1d (1 reply)

Shrug. Still don't care.

boonhet@sopuli.xyz · 5 pts · 1d

Perhaps you're not in the right community then, this one is for humor not serious discussion.

BlackLaZoR@lemmy.world · -8 pts · 2d (1 reply)

I imagine you could actually make an insane compression algo if you used AI as support for regular jpg compression. Not this way tho.

spankysalmon@fedinsfw.app · 9 pts · 2d

Your imagination is not how computer science works

chunes@lemmy.world · -12 pts · 2d (22 replies)

There is a future where AI can simply look up the DNA for any individual you specify and it can perfectly reconstruct what they look like, at any age.

BrainBow65@lemmy.world · 15 pts · 1d (18 replies)

You're joking, right? There's a reason genotype and phenotype aren't synonyms.

chunes@lemmy.world · -1 pts · 1d (15 replies)

There's also a reason identical twins look alike.

HerbGrower@slrpnk.net · 6 pts · 1d

Wow, Jack and Ross look so alike! How did he get his arm back, isn't that after the accident?

hirihit640@sh.itjust.works · -3 pts · 1d (13 replies)

Don't know why you're being downvoted, it's a good point that shows that genetics is the main factor for your appearance

onlinepersona@programming.dev · 8 pts · 1d (8 replies)

Maybe but there is more to the outcome than just genes. Certain ones can be activated of deactivated and that can be done by epigentics for example. Gene expression is a thing.

The AI would need way more information that just DNA to generate an image of the person.

hirihit640@sh.itjust.works · 2 pts · 1d (7 replies)

But they could probably get 90% the way there with just DNA. I think that was the point.

onlinepersona@programming.dev · 6 pts · 1d (2 replies)

No, they couldn't, that's what's being said. If you have a huge state machine with billions of parameters, toggling a different subset can yield wildly different results.

hirihit640@sh.itjust.works · 0 pts · 1d

But there are examples of twins who grew up apart and still look almost identical

Susaga@sh.itjust.works · 5 pts · 1d (1 reply)

They said "perfectly", not "90%". You're trying to move the goalposts.

hirihit640@sh.itjust.works · 0 pts · 1d

I just re-read their post and yes you're right, in which case I disagree with them. "Perfect" is not achievable by any means. Maybe accurate enough to be mistaken as a twin for the real person, but not perfect

sukhmel@programming.dev · 1 pts · 1d (1 reply)

Maybe if you mean things like two eyes, two hands, two feet, one nose, but not in any details

hirihit640@sh.itjust.works · 1 pts · 1d
[ removed ]
Slashme@lemmy.world · 7 pts · 1d (3 replies)

Identical twins also share the same womb, so it's not 100% genes.

T156@lemmy.world · 7 pts · 1d (2 replies)

Same womb, similar growing environment, etc.

But even then, they're not perfectly identical. Very similar, but not the exact same.

hirihit640@sh.itjust.works · -4 pts · 1d (1 reply)

I don't think root comment was saying the AI would generate perfectly accurate images. I think 90% accuracy is doable, just from DNA

BrainBow65@lemmy.world · 6 pts · 1d

If you're just talking about a strand of DNA with no epigenetic info I think 90% is being very optimistic.

Slashme@lemmy.world · -1 pts · 1d (1 reply)

Forensic DNA phenotyping is actually being used. Of course it's not perfect, but it can give a direction.

BrainBow65@lemmy.world · 3 pts · 1d

Hadn't heard of this before but I'm more than a little skeptical, especially if its being used for policing

"More recently, companies such as Parabon NanoLabs and Identitas have begun offering forensic DNA phenotyping services for U.S. and international law enforcement. However, the science behind the commercial services offered by Parabon NanoLabs has been criticized as it has not been subjected to scrutiny in peer-reviewed scientific publications. It has been suggested that it is not known "whether their ability to estimate a face’s appearance is better than chance, or if it’s an approximation based on what we know about ancestry” -DNA phenotyping wiki

T156@lemmy.world · 6 pts · 1d

Not unless it can predict all the contributing factors.

DNA isn't the full story.

Maybe they got hit by a car, were a hobbyist boxer, or had surgery, for example.

Slashme@lemmy.world · 1 pts · 1d (1 reply)

You're kind of describing forensic DNA phenotyping.

sukhmel@programming.dev · 1 pts · 1d

Whic is very much not precise, as explained in introduction to this article, for instance