• 1.75K Posts
  • 776 Comments
Joined 3 years ago
cake
Cake day: June 12th, 2023

help-circle




























  • The distinction of using pre-existing CSAM is notable because law enforcement and child protection organizations frequently give those illegal materials a kind of digital fingerprint – known as a hash – which allows them to track images and videos when they appear online. In this case, attorneys for the plaintiff stated that the Canadian Centre for Child Protection used images’ fingerprints to identify AI-generated CSAM on X that depicted their client.

    How does that work? Are there hashes of child sex abuse images somehow reproduced by LLM? Surely the new images don’t have the same hashes. Guardian is not really being clear here.

    edit: Ars’ is less badly written: “This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.””