OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models.
There is just not enough non-copyrighted data out there.
You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models
without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.
I'm pretty sure we would know if they did that. And we don't.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
You mean very likely the Anna's archive torrent dump because it's MUCH better quality than the general internet and beyond a certain amount of input data (which is a lot, but much less than the internet) the only thing that matters in training is the quality of the data, to the point that now many labs have thousands of people just making and improving essentially school exercises full time?
Hell, I know that for one "lab" (kindof AI lab) since 2020 or so has determined wikipedia quality is dropping fast. It was already dropping slowly before that, but now it's getting bad.
No, I mean the WebText corpus whose construction from 45 million Reddit post with at least 3 karma is described in section 2.1 of the PDF I linked. They did remove all Wikipedia documents.
I wonder if anyone has run the numbers on what the actual cost, both in cash and logistical headache, contacting so many copyright holders would be. That seems like quite the feat to calculate.
There's an established network of intermediaries that can supply a large variety of books for a few dollars apiece, so no need to contact copyright holders directly.
This is very true. As someone with quite the experience with materials published under Penguin, Scholastic, etc. you effectively have a "dictionary attack" on the matter, rather than true "brute force," but that still leaves quite a list to compile to send to each and is easier for larger titles than smaller ones. I wonder how that leads to a bias in what materials get used for training. You are not getting many local self-published books this way.
It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
Pretty easy to assertain that they don't acquire them legally due to the plethora of evidence and court cases against them. No copyright holder would be sueing them if they knew they sold the works in the first place.
I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.
I'm not aware of any successful accusations against OpenAI for illegally obtaining copyrighted material, in contrast to Anthropic, who settled for $3000 per work and then still had to buy legal copies to keep using them (likely for much less).
Instead, the ongoing lawsuits focus on the idea that AI training involves making additional copies, for which they would need a copyright license instead of just one legal copy.
You are kind of right, but you also did not look very hard. They deleted huge datasets in anticipation of lawsuits, at least that much is known.
Of course plaintiffs were unable to depose their internal lawyers (who apparently know why they were frantically deleted) due to 'attourney client priviledge' further refusing to provide any kind of transparency. But yeah, I guess they are they are better at covering their tracks and destroying evidence.
Also, as one more example, I find it hard to believe that their models could generate 'Studio Ghibli' style images without training on the movies. There is no licensing deal between them.
I think the real issues here are two-fold:
Firstly, Copyright is very ill equipped to handle these cases. Just because the model is tuned not to output the exact training data does not mean that compressing mostly-copyrighted datasets into a proprietary model is ethical, fair or /should/ be allowed, simply because they might destroy entire livelihoods. If you take those copyrighted works away you are left with, in OpenAIs own words, a cute little experiment.
Secondly, there is absolutely no transparency. Datasets are easily deleted and its impossible to tell what the models have been trained on, especially after fine tuning. Moreover, only the biggest most successfull works would be easily identifiable without the fine tuned model. Once again, sticking it to the little man.
Hmm... I think this could be a case of the organic brain soaking up LLM-isms by the way as it is using these tools intensively. As much as I don't like robot writing, this all seems overly paranoid to me.
Pretty sure you misunderstood this person. I don't think they are advocating for the smashing of looms, rather fear that the movement that could redistribute the benefits of this new technology will go down that dark road again.
I also think the reckoning is overdue, maybe it will never come to be.
What reckoning? The reckoning for large scale IP theft, scaling up what they call "ghost work", the usage of natural resources and land which power data centers with no return to the people that depend on those and the already existing predatory and disgusting business practices that they have always partaken in. (I assume)
I don't think that is right. There is always a commercial and non commercial duality to art. Artists need money. Many artists hone their craft and make their money by producing commercial products, corporate music or web ads or whatever. I think your prediction is naive; this will take a huge, irreversible toll on art.
Word. But here we have about 300 sweaty nerds claiming 'AI training is like learning' and similarly misguided hogwash. Not sure how such a culture war can be won against people with no regard for intellectual property, the beauty of art and a fleet of propaganda chatbots (yes, I'm dying on that conspiracy theory hill any day).
Uhm, properly compensate passengers that have concrete harm to point to instead of offering a laughable alibi payment? What I'll never understand is why some neoliberal schools of thought defend conglomerates like they're their children... This company has enough money to compensate for fuck-ups, at least it should have. Every insurance on the planet can do the simple math involved for this risk assesment.
Why the nazi-esque "Frakturschrift" Logo? I get the font was not originally associated like that, but your name and that logo make me question what's behind it.
I'm Austrian.
Such fonts were in use many centuries before NS time in many countries in north Europe. There were eventually replaced by modern ones, but in Germany they lasted long enough, until WWII. It's a common misconception, that only nazis used it. Even more, they forced replacement of fraktur fonts with modern ones. Nowadays such fonts are in use in places where an old-looking font is needed, like in signs. Also modern nazis use it sometimes, but only because they don't know history well.
I selected fraktur for a logo of my language because it looks "cool" and is already pretty stylized compared to a boring logo using some modern font. I don't insist to use it forever, maybe I will change it later with some better one, but for this someone else need to draw it properly, I don't have necessary design skills for this.
My nick-name is easy to explain. Many years ago I have enjoyed playing Wolfenstein and Call of Duty and these games have Panzerschreck as one of weapons. I choose this as my nick-name, but I unfortunately misspelled it.