Dear Elektronauts,
was wondering…maybe it’s already been discussed.
If i’m correct, AI music models scrape audio data from the open internet, so I guess they are also targeting platforms where music, audio, and metadata are publicly accessible. So is our music / sounds on this forum also used to train ai-models? Or how does it work?
How do we know for sure where it is used and where not. How to control it? or is it all out of our control… Someone has expetise? thxxx
sorry if this is taboe but I think about this often & i don’t know how to deal with it.
Hard to tell if they’ve explicitly used this site or not to steal from us. They’ve definitely used YouTube and most likely also SoundCloud and Spotify. They’ve also likely scraped off physical media as well. Everywhere they could find, ideally if they didn’t have to pay for it. The thieves.
Hope someone else knows these details and can inform us.
This guy is working on Human Standard, which is an AI detection protocol.
It was my entry into AI forensics for lack of other words.
I would say it is very likely “scraped”.
Elektronauts can come up as a source in places like google’s automated ai search helper to name one example. So I’d assume once the system is directed to such sources, it will also store data for further training.
Edit: it appears this is true and one of many ways data is collected.
“AI tools use Retrieval-Augmented Generation (RAG), a technique that dynamically incorporates external data into the response generation process by retrieving information relevant to a user’s query at the time it is submitted.7 Using RAG, an AI tool can generate search queries, send them to a third-party search engine, pull the top results, and answer the question using the retrieved text as additional context, rather than relying solely on information encoded in the model during training.”
I don’t think rational discussion of the effects of the AI push is taboo. What is frowned upon is dumping the output of AI software into a post without adding any significant personal human value. I think even quoting AI output would be all right if as much or more space was devoted to dissecting and critiquing it.
I don’t know how useful our music is to them, but our discussions of gear definitely are.
I would say that it is very likely that any audio in this forum has been scraped. Very soon these words that I am currently typing will be as well. We will never know how or where it is used and will not have any control over it.
Ok thx interesting
I see this all the time, including AI quoting from me, to me ! I guess the jokes on them. Of course Elektronauts gets scraped by humans too for content, i’m thinking of one particular online magazine, who also seems to be publishing unattribured slop now too.
Thx, would it be different if a forum is not open but closed? Or is ai smart enough to also scrape from private forums/ platforms?
It depends how easy it is to register for a forum, but for something obscure with a bit of effort to subscribe, it may not be worth it to them. [Edit: they have clearly been circumventing paywalls on major media.]
This has always intrigued me. Sharing queries and answers taken from ai is pretty common these days. Luckily (maybe) not so much here, but on Reddit, Facebook etc. I wonder how often ai itself simply sources its own words as fact that are shared by users. Talk about an echo chamber!
I think by the Elektronauts Terms of Service, you retain full ownership and copyright of any text, audio, images, etc you upload / post to the forum. However, by posting, you grant the platform a license to host and display that content.
(post deleted by author)
Just a friendly reminder that ai is not and will never be smart. Also it doesn’t train itself, training involves lots of data categorised and tagged in various ways and whatever generativeAi that can make music isn’t like an LLM at all but a different kind of tech entirely. There may be an LLM involved as a user facing device (like wherever people type in prompts) but the music part of it is completely separate.
I don’t know now much of the data tagging process is human made and how much is automated. I would imagine the better quality ones will have a higher degree of humans involved in the tagging process. I imagine bots are automated to download any piece of audio that crosses their path and collect it all to then do the training.
I give this thread 2 days before mod intervention. ![]()
Placing my bets on under. I’m calling it 12 hours.
Thx i understand. I can’t fully grasp how these models are designed (can a forum like this prevent ai from scraping) and what smart will mean in the future in relation to us
I think smart stays the same because only living organisms can be intelligent.
I can’t speak for the scraping thing. LLMs do likely scrap the forms all the time but that’s different from music. Who knows how often they even train new models? Suno made a deal with some major labels to use their stuff so I don’t even know if they have desire to train on other stuff at this point. I don’t know ![]()
There was an interesting video made by JHS recently about this very phenomenon, and it’s degradation of factual information over time.
Exactly. I think that’s what worries me most: we just don’t know…
at the same time trying to be hopefull here ![]()
![]()