← Back to Gadget Pulse US

Apple Quietly Trained Its AI on a Controversial Dataset Nobody Noticed

Persona #4 · Vol: 0

Apple spent two years telling anyone who would listen that its AI would be different.

Then researchers started pulling at a thread, and what they found has the tech world asking uncomfortable questions about what "Apple Intelligence" actually learned from.

The thread leads to a dataset called Books3 — a sprawling, unauthorized pile of roughly 180,000 pirated books that has been floating through the AI training ecosystem for years.

Now a growing stack of evidence suggests Apple's open-source model, released quietly under the name OpenELM, was trained on a dataset called RedPajama, which itself contains Books3.

Apple has never confirmed or denied the connection directly.

Why should you care if you own an iPhone 15 Pro or any device running Apple Intelligence features?

Because the company's entire marketing identity rests on trust.

Apple isn't selling you the smartest assistant.

It's selling you the assistant you can actually trust with your messages, your photos, your calendar.

That pitch collapses fast if the foundation was built on books authors never agreed to feed into a machine.

Here is the part that gets genuinely strange.

Apple has publicly bragged about licensing real news archives from publishers like NBC and Condé Nast to train its models — paying for content, doing it the "right" way.

So the company clearly understands consent matters.

Yet the Books3 pipeline ran through an intermediary dataset that bundled pirated material alongside legitimate sources.

Whether Apple knew or not is the question nobody inside Cupertino wants to answer on the record.

For everyday users, the practical fallout is murky.

Siri won't suddenly start reciting Stephen King novels.

But the reputational damage compounds a bigger problem: Apple Intelligence has already been criticized for shipping late, arriving incomplete, and underwhelming compared to what Google and OpenAI are putting out.

Add a training-data scandal to a product that's already playing catch-up, and the "trust" advantage starts looking less like a moat and more like a liability.

Watch what happens next with the lawsuits.

Authors have already sued Meta and others over Books3.

If discovery pulls Apple into that pile, the company that built its brand on privacy could find itself defending the exact behavior it spent years mocking its rivals for.

Apple announced Apple Intelligence with enormous fanfare, then delivered features in slow drips across multiple iOS updates.

That's not how you ship something you're confident in.

That's how you ship something you're still patching together while the clock runs out.

My take: Apple built its AI reputation on a promise it may not have fully earned.

The company that lectures everyone about privacy should be the loudest voice explaining exactly where its training data came from — and so far, it's been the quietest.

Final Thoughts

Until that changes, every "private and secure" tagline deserves a healthy dose of skepticism.

Continue Reading