So in the european union thanks to GDPR we have all these new requirements for encryption of personal data in transit. Which are reasonable on their face. The problem is that nobody knows how to do encryption properly. They'll put a password on the zip or PDF file during export and the password will be (and I swear I've actually encountered this one in the wild) my date of birth in the form mmyy. Fun fact, a decent computer can brute force a 4 digit numeric password in less than a second.
This is called security theater. Its there to make you feel safe. Any incompetent hacker could breeze through it with a chatGPT answer. And ironically we have better systems. See that padlock next to the URL on your web browser? That means your connection to this website is secured by https. That means bad guys can't see your password just because they're on the same wifi network as you. That uses asymetric encryption (where the encryption and decryption keys are different and actually unrelated). That's real security. But I guess that's only for technical users not the lady from the doctors office.
Real security would be asymetric encryption where only the end users have the encryption and decryption keys, not any of the systems in between like google gmail. Sending via gmail will actually encrypt your data in transit because email can also have SSL protection. But google can still scan your emails, realise you have cancer and push you ads for miracle mineral solution. But use a tool like PGP or GPG to encrypt before sending through gmail… Nobody is seeing your diagnosis without stealing your secret key. And nobody needs to see your secret key to encrypt in the first place.
This one will be controversial but maybe some people will find this cool
Local AI has gotten really good. I love checking out ways to integrate smaller AIs in my software and workflows (whisper for example, captioning or image semantics).
I am the singer of a band and I already made a speech to text thing so that I can hold a button and say a thing and it changes my Autotune key as soon as it recognizes the song name. It's awesome. I am currently looking into switching to a faster model so that maybe I can do this live with more effects and use it to reassign foot triggers and pedals etc.
I also love the idea that open source models are incredibly good already. That thing where you say "AI please turn the lights on", you could build that reliably already with a good agent I model and a formatted tool calls. It's incredible. And you don't have to subscribe or anything, it all runs at home.
I have this hope that we are getting so good with AI that paid models become redundant and everyone can just use local AI models willy nilly.
I'm curious about local models. How much storage does it take to give them a good library of data? And are you able to curate that or do they just come with a dataset that's community made?
Sry for not answering for a while but I am not sure myself on a lot of this.
I know for a while the big companies scraped petabytes of data into datasets for training. It was quantitative, the more the better.
About 2 years ago people found out removing bad data is really helpful and more is not necessarily better. Datasets got smaller and more focused on beneficial data, which of course looks very different depending on the purpose of the model.
Basically what happens is someone trains a model, then tests the model, then trains then tests, all the while they are slowly tweaking stuff and rewriting basic parts of how it's trained or how the model does inference "using the model". And they try to see how to make it so the model performs the best. It's like 3 fields of science combined and a lot of the time they release a whole scientific paper together with the model. Especially if there's some new technical discovery they stumbled upon in training which might advance the field.
But I can't do that, I am smart and experienced in that field yet still I will happily admit I could not keep up with those brilliant scientists under any circumstance.
When I wanna run a local AI, I download a model. For consumers you expect it to be a maximum of 30gb for the model and maybe another 5gb for libraries and stuff. Then using those libraries you run the model, feed it your data (the microphone input) and let it have at it. You rarely do any more training, because you will almost never do it better than a scientist who dedicates most of their time to that specialty.
Think of the model itself as a chaotic big blob. Like a zip archive but inside is a bunch of unreadable data, usually called "weights". Those weights give the AI a way to calculate an output from an input. It's not a dataset, but it's everything that the model could learn from the dataset when it was trained. And then I myself never have to deal with terabytes of data from a dataset, it's all neatly packed into a 30gb model. And I can just run that.
I can explain it in more detail but even I can't grasp the full technical details anymore because every step has tons of optimizations and transformations baked in, but the very basic model still functions as we are used to from deep neural networks. If you are interested in learning, that's a good way to start.
Anyway, you sly dog caught me monologuing. Thank you for letting me share all of this stuff :)
No, not a problem at all :) thanks for getting back to me.
I didn't realize you couldn't just feed it raw data, but that makes sense. That's too bad, I would have gone all in on local ai if it was something end users could specifically curate their individual models with. Definitely still sounds useful though.
30gb is still way smaller than what I was expecting. That's pretty actionable, just like uninstall a single videogame.
Thank you, and I hope you have a great day as well
Yeah you can use your own data but it's extremely unlikely that you get comparable results and it takes much more time and more trial and error to train such a model yourself.
Even just using a model that's trained on royalty free works like the dolphin dataset is more feasible than using your own data I assume.
But yes, legally and ethically the boundaries of personal identity and copyright will be part of the discussion for at least a few more years.
That is extremely cool. Maybe in the future instead of Serial Experiments Lain style building computers to be highly customized it's primarily in the software. Less aesthetically interesting maybe but still amazing DIY.
I think it might be helpful to know this is a modification of a meme where the different % autistic creature is seen as annoying.
I'm sure the original has more than one meaning but I read it as highlighting that folks with higher support needs are sometimes looked down on by those with lower support needs..
I'm late diagnosed and it's really forcing me to confront my own ableism, which I wasn't even conscious of.
Adding more for all readers, not a direct reply: It's not a flaw to have certain situations be difficult to manage. For example, my brother (who is not diagnosed) has a stim/tic that drives me FUCKING NUTS. It would be ableist to expect him to stop or that he can control it, or to say I am better because I don't have an obvious constant stim (or haven't unmasked enough to unleash it yet lol). I can remove myself from the situation when it gets too much without blaming him for it. My siblings have a lot of resentment towards me because from their perspective I was seen as the "good" one that they should be more like. From my perspective I was never good and spent all my energy trying to suppress the "bad", so I didn't even realize this was happening. The way the adults framed it isn't my responsibility, but what I can do now is speak up when neurodivergence behaviours are framed as character flaws instead of differences we should work together to find mutually beneficial accommodations for.
I don't think the version I posted is calling people out for being annoyed at others. I think it's calling out when people compare themselves favourably to people with higher support needs, or allow others to use them as examples of "good" or "better". I believe it's alluding to the autistic/Asperger's distinction and how that can be used to frame some people with autism as better than others. Again, I'm late diagnosed so I'm not going to dive into that label and how people who were given it should relate to it. It's not my experience, I don't get a say! People with higher support needs may have less capacity to advocate for themselves. Those who are able to advocate for the needs of autistic people should strive to include higher support needs people too.
(I hope it's clear that this advocacy I am talking about is a wider activity than seeking personal accommodations. I'm not saying you need to include others in that, but awesome if you have the ability!
I don't know if you understand the interaction that's being described in the meme. It's depicting of a thing that's famously annoying to most people, autistic or not. We kind of get a bad rep for it
Kudos to you if you're also one of the people who never felt that way. I've encountered one or two
14 Comments
TheGingerNut@piefed.blahaj.zone · 18 pts · 17d
So in the european union thanks to GDPR we have all these new requirements for encryption of personal data in transit. Which are reasonable on their face. The problem is that nobody knows how to do encryption properly. They'll put a password on the zip or PDF file during export and the password will be (and I swear I've actually encountered this one in the wild) my date of birth in the form mmyy. Fun fact, a decent computer can brute force a 4 digit numeric password in less than a second.
This is called security theater. Its there to make you feel safe. Any incompetent hacker could breeze through it with a chatGPT answer. And ironically we have better systems. See that padlock next to the URL on your web browser? That means your connection to this website is secured by https. That means bad guys can't see your password just because they're on the same wifi network as you. That uses asymetric encryption (where the encryption and decryption keys are different and actually unrelated). That's real security. But I guess that's only for technical users not the lady from the doctors office.
Real security would be asymetric encryption where only the end users have the encryption and decryption keys, not any of the systems in between like google gmail. Sending via gmail will actually encrypt your data in transit because email can also have SSL protection. But google can still scan your emails, realise you have cancer and push you ads for miracle mineral solution. But use a tool like PGP or GPG to encrypt before sending through gmail… Nobody is seeing your diagnosis without stealing your secret key. And nobody needs to see your secret key to encrypt in the first place.
Arcanepotato@crazypeople.online · 8 pts · 17d
But how will profit be made if we don't allow Gmail to read our emails?
(Very cool info, thanks for sharing)
Grass@sh.itjust.works · 4 pts · 17d
I wish more people cared about stuff like this
hoshikarakitaridia@lemmy.world · 10 pts · 17d
This one will be controversial but maybe some people will find this cool
Local AI has gotten really good. I love checking out ways to integrate smaller AIs in my software and workflows (whisper for example, captioning or image semantics).
I am the singer of a band and I already made a speech to text thing so that I can hold a button and say a thing and it changes my Autotune key as soon as it recognizes the song name. It's awesome. I am currently looking into switching to a faster model so that maybe I can do this live with more effects and use it to reassign foot triggers and pedals etc.
I also love the idea that open source models are incredibly good already. That thing where you say "AI please turn the lights on", you could build that reliably already with a good agent I model and a formatted tool calls. It's incredible. And you don't have to subscribe or anything, it all runs at home.
I have this hope that we are getting so good with AI that paid models become redundant and everyone can just use local AI models willy nilly.
sad_detective_man@sopuli.xyz · 5 pts · 17d
I'm curious about local models. How much storage does it take to give them a good library of data? And are you able to curate that or do they just come with a dataset that's community made?
hoshikarakitaridia@lemmy.world · 2 pts · 13d
Sry for not answering for a while but I am not sure myself on a lot of this.
I know for a while the big companies scraped petabytes of data into datasets for training. It was quantitative, the more the better.
About 2 years ago people found out removing bad data is really helpful and more is not necessarily better. Datasets got smaller and more focused on beneficial data, which of course looks very different depending on the purpose of the model.
Basically what happens is someone trains a model, then tests the model, then trains then tests, all the while they are slowly tweaking stuff and rewriting basic parts of how it's trained or how the model does inference "using the model". And they try to see how to make it so the model performs the best. It's like 3 fields of science combined and a lot of the time they release a whole scientific paper together with the model. Especially if there's some new technical discovery they stumbled upon in training which might advance the field.
But I can't do that, I am smart and experienced in that field yet still I will happily admit I could not keep up with those brilliant scientists under any circumstance.
When I wanna run a local AI, I download a model. For consumers you expect it to be a maximum of 30gb for the model and maybe another 5gb for libraries and stuff. Then using those libraries you run the model, feed it your data (the microphone input) and let it have at it. You rarely do any more training, because you will almost never do it better than a scientist who dedicates most of their time to that specialty.
Think of the model itself as a chaotic big blob. Like a zip archive but inside is a bunch of unreadable data, usually called "weights". Those weights give the AI a way to calculate an output from an input. It's not a dataset, but it's everything that the model could learn from the dataset when it was trained. And then I myself never have to deal with terabytes of data from a dataset, it's all neatly packed into a 30gb model. And I can just run that.
I can explain it in more detail but even I can't grasp the full technical details anymore because every step has tons of optimizations and transformations baked in, but the very basic model still functions as we are used to from deep neural networks. If you are interested in learning, that's a good way to start.
Anyway, you sly dog caught me monologuing. Thank you for letting me share all of this stuff :)
Hope you have great day ^^
sad_detective_man@sopuli.xyz · 2 pts · 13d
No, not a problem at all :) thanks for getting back to me.
I didn't realize you couldn't just feed it raw data, but that makes sense. That's too bad, I would have gone all in on local ai if it was something end users could specifically curate their individual models with. Definitely still sounds useful though.
30gb is still way smaller than what I was expecting. That's pretty actionable, just like uninstall a single videogame.
Thank you, and I hope you have a great day as well
hoshikarakitaridia@lemmy.world · 2 pts · 12d
Yeah you can use your own data but it's extremely unlikely that you get comparable results and it takes much more time and more trial and error to train such a model yourself.
Even just using a model that's trained on royalty free works like the dolphin dataset is more feasible than using your own data I assume.
But yes, legally and ethically the boundaries of personal identity and copyright will be part of the discussion for at least a few more years.
Arcanepotato@crazypeople.online · 5 pts · 17d
That is extremely cool. Maybe in the future instead of Serial Experiments Lain style building computers to be highly customized it's primarily in the software. Less aesthetically interesting maybe but still amazing DIY.
87Six@lemmy.zip · 5 pts · 17d
I feel like human interaction is more a good human trait than autism but ig it's down to autists now :-/
Arcanepotato@crazypeople.online · 20 pts · 17d
I think it might be helpful to know this is a modification of a meme where the different % autistic creature is seen as annoying.
I'm sure the original has more than one meaning but I read it as highlighting that folks with higher support needs are sometimes looked down on by those with lower support needs..
sad_detective_man@sopuli.xyz · 4 pts · 17d
That's actually more insightful than I had interpreted it. Thank you
Arcanepotato@crazypeople.online · 5 pts · 17d
I'm late diagnosed and it's really forcing me to confront my own ableism, which I wasn't even conscious of.
Adding more for all readers, not a direct reply: It's not a flaw to have certain situations be difficult to manage. For example, my brother (who is not diagnosed) has a stim/tic that drives me FUCKING NUTS. It would be ableist to expect him to stop or that he can control it, or to say I am better because I don't have an obvious constant stim (or haven't unmasked enough to unleash it yet lol). I can remove myself from the situation when it gets too much without blaming him for it. My siblings have a lot of resentment towards me because from their perspective I was seen as the "good" one that they should be more like. From my perspective I was never good and spent all my energy trying to suppress the "bad", so I didn't even realize this was happening. The way the adults framed it isn't my responsibility, but what I can do now is speak up when neurodivergence behaviours are framed as character flaws instead of differences we should work together to find mutually beneficial accommodations for.
I don't think the version I posted is calling people out for being annoyed at others. I think it's calling out when people compare themselves favourably to people with higher support needs, or allow others to use them as examples of "good" or "better". I believe it's alluding to the autistic/Asperger's distinction and how that can be used to frame some people with autism as better than others. Again, I'm late diagnosed so I'm not going to dive into that label and how people who were given it should relate to it. It's not my experience, I don't get a say! People with higher support needs may have less capacity to advocate for themselves. Those who are able to advocate for the needs of autistic people should strive to include higher support needs people too.
(I hope it's clear that this advocacy I am talking about is a wider activity than seeking personal accommodations. I'm not saying you need to include others in that, but awesome if you have the ability!
sad_detective_man@sopuli.xyz · 2 pts · 17d
I don't know if you understand the interaction that's being described in the meme. It's depicting of a thing that's famously annoying to most people, autistic or not. We kind of get a bad rep for it
Kudos to you if you're also one of the people who never felt that way. I've encountered one or two