Google is not just scanning for known examples of kiddie porn, they are scanning for images of naked children, and reporting parents to the police.
for instance, they tried to get a parent arrested for using their Android phone to get medical assistance for their child during lockdown.
>With help from the photos, the doctor diagnosed the issue and prescribed antibiotics, which quickly cleared it up. But the episode left Mark with a much larger problem, one that would cost him more than a decade of contacts, emails and photos, and make him the target of a police investigation.
“I knew that these companies were watching and that privacy is not what we would hope it to be,” Mark said. “But I haven’t done anything wrong.”
I wouldn't be surprised if Microsoft wasn't doing the same thing with photos viewed in their default Photo app (The names of the images viewed in that app are sent as telemetry already along with the times they were viewed and how long each image was left open).
Windows defender may even be checking the files it scans for whatever content MS considers objectionable. We likely wouldn't won't know unless it becomes a news story or a whistleblower tells us.
> I wouldn't be surprised if Microsoft wasn't doing the same thing with photos viewed in their default Photo app
Google has apparently shared their algorithm with others, but only Facebook was mentioned by name.
>in 2018, Google developed an artificially intelligent tool that could recognize never-before-seen exploitative images of children. That meant finding not just known images of abused children but images of unknown victims who could potentially be rescued by the authorities. Google made its technology available to other companies, including Facebook.
I had no idea the photos app was doing this. This seems beyond excessive and I’m glad I happen to use the old photo viewer. The new one always takes a second to load the photo, and I suppose that’s because it’s doing all this other stuff in the background.
I often hear the argument that the average user does not care about privacy, but I don't think that this is true.
Instead I think the average user is just not aware of the insane amount of data that is exfiltrated and what can be done with that data, because the software does not make it clear what it actually is doing.
Honestly, why on Earth does Microsoft want to know the filenames of the photos being opened? And, how often each one is viewed? What possible enhancement to the app can be justified because of this information?
Cynically, I'm sure it could be "justified" in all kinds of ways like "knowing more about the photos viewed most often with our Photos app helps us to determine what new features would best enhance the user experience. Only metadata is being collected, so no employee is ever able to view your pictures. We value your privacy blah blah blah"
Even more cynically, it just gives them more data about you and your interests (with viewing frequency/duration being a measure of interest) which lets them push more narrowly targeted ads at you, which gives them more money, which lets them invest more in improving Windows as a whole or whatever else they tell themselves so they can face their own reflections.
It stores that information in a local index, and also stores telemetry data which counts interactions with the app, but there is no evidence it is uploading indexes to MS.
These documents were massive and filled with lots of information buried in places that seemed odd and scattered across different sections so that you couldn't get all the information about what an app collected in one section/table of the document.
Another thing I remember from those pages was that in some circumstances short clips (less than 1 minute) of the videos you played using window's default player would be sent to MS as well, but they said that would only happen in the event of crashes.
Yes, pretty easy. Keyboard-warrioring in general is.
But that's not the point. Victim blaming doesn't do any good for today's victim, but maybe it will help show the next person that their fate is under their control.
If the photos are hosted on Google's servers, they have every right to scan whatever they want. If you are privacy conscious, you can upload files you encrypt yourself at the cost of not being able to share easily. I recognize edge cases like the one in the article you linked, but I don't see an alternative. Not scanning for CSAM on your own servers isn't a realistic expectation.
> If the photos are hosted on Google's servers, they have every right to scan whatever they want.
Google is trying to get people arrested because Android backed up their personal photo libraries to Google Photos.
> The day after Mark’s troubles started, the same scenario was playing out in Texas. A toddler in Houston had an infection in his “intimal parts,” wrote his father in an online post that I stumbled upon while reporting out Mark’s story. At the pediatrician’s request, Cassio, who also asked to be identified only by his first name, used an Android to take photos, which were backed up automatically to Google Photos. He then sent them to his wife via Google’s chat service.
> Not scanning for CSAM on your own servers isn't a realistic expectation.
Google isn't just scanning for known examples of kiddie porn, they have developed an algorithm that scans for any image of naked children and, as usual, there doesn't seem to be a human in the loop when their algorithm makes a bad call.
It's Google's servers. I imagine every content hosting company has a heuristics based approach to detecting CSAM. Google has a habit of tuning moderation algorithms for false positives rather than missed positives.
Not having a human in the loop is terrible and a problem frequently when it comes to getting support from Google companies. That's a separate issue from Google or Dropbox being able to scan unencrypted files for CSAM. Google's policy of automating as much as it can has tons of downsides. But it's understandable when you look at the scale Google functions.
It's important to separate the policy of scanning from Google's terrible appeal process and the algorithm false positives.
I would feel differently if the story was that Android scanned an outgoing SMS or a photo saved locally. I am not sure where the balance point is to identify and report CSAM while also respecting user's rights to privacy.
> It's Google's servers. I imagine every content hosting company has a heuristics based approach to detecting CSAM.
The slippery slope claim made in the past was that they were only looking for files that match known examples of kiddie porn.
For instance, this tech from Microsoft:
>PhotoDNA creates a unique digital signature (known as a “hash”) of an image which is then compared against signatures (hashes) of other photos to find copies of the same image. When matched with a database containing hashes of previously identified illegal images, PhotoDNA is an incredible tool to help detect, disrupt and report the distribution of child exploitation material.
Reporting parents to police because they took pictures of their child's first bath and those photos were automatically backed up to Google servers is much, much worse.
Please stop using 'child's first bath' as an indicator of overreaction by the algorithm. It is not helpful to argue something which will make you look hyperbolic or histrionic once your argument is looked into.
1. The picture(s) in question mentioned above which got an account shut down and police notified was not a 'child in a bath', it was close ups of a child's genitals
2. You are probably not a parent since a child's 'first bath' is not a thing. They are bathed as infants starting immediately. If you are referring to a child's first independent bath, then no one should be taking pictures since private bathing is not something that should be intruded upon for picture taking
> 2. You are probably not a parent since a child's 'first bath' is not a thing. They are bathed as infants starting immediately. If you are referring to a child's first independent bath, then no one should be taking pictures since private bathing is not something that should be intruded upon for picture taking
Are you a parent? A prudish parent perhaps?
I have pictures of me as a kid in a bath (genitals obscured underwater). It's a common parental thing among parents who AREN'T DRIED UP PRUDES AND FRIGHTED OF NUDITY.
The bottom line is - another case of our ability to use our technology as we desire being interrupted and interfered with. I hate to even remotely invoke RMS, but... He wasn't wrong!
Please re-read what I wrote. Was it really unclear? I said that if a child is bathing independently that it is not appropriate to barge in and take pictures of them. They would be at least 5 or 6 years old at that point. Is that prudish? Why is everyone so emotional about this? All I said was that it was a terrible and inaccurate argument to talk about pictures of a child in a bath because that is not what flagged the algorithm and there is no such thing as a child's first bath.
Please try to parse a response before yelling at someone for something they didn't state.
> 1. The picture(s) in question mentioned above which got an account shut down and police notified was not a 'child in a bath', it was close ups of a child's genitals
No, the example of Google calling the police on a parent given in the NYT article was because the parent sought medical assistance for their child during lockdown and the doctor requested pictures of the issue the child was having.
This is no more an acceptable reason for Google to report parents to the police for child abuse than taking pictures of your child's first bath, and taking pictures of your child's milestones has long been something that people have done.
>> 1. The picture(s) in question mentioned above which got an account shut down and police notified was not a 'child in a bath', it was close ups of a child's genitals
> No, the example of Google calling the police on a parent given in the NYT article was because the parent sought medical assistance for their child during lockdown and the doctor requested pictures of the issue the child was having.
These things are not mutually exclusive.
Stop being defensive. It doesn't help you. I am trying to help you by telling you that you are making a shitty argument.
> I am trying to help you by telling you that you are making a shitty argument.
OK. Allow me to tell you that, in my opinion, you are the person making the shitty argument
Your argument that "no one should be taking pictures" of their child's first bath is just cringe-worthy.
Google is to blame for the situation when they call the police over a single false positive from their faulty algorithm, not parents who take pictures of their own children.
Are you even reading what I am writing? I never said that no one should take pictures of their children's first bath -- I said that there is no such thing as a child's first bath, just like there is no such thing as a child's first meal, or a child's first nap, or a child's first bowel movement. They are given baths as infants.
You are obviously either unable or unwilling to parse what I am trying to tell you, and fixated on this notion that you have to be 'right' no matter what even though I am not telling you that you are wrong!
I am at a loss as to what to reply to you with besides I urge you to read things before continuing to respond because otherwise you end up talking around people instead of with them.
How do you take a photo on android to only store it locally?
The whole photos experience on Android has been simplified to break that distinction and there is no way to even really know or realize what is happening - BY DESIGN.
The real answer here is, again, Android is not privacy first by default.
BTW parents taking medical pictures of their children != CSAM which is an acronym for "Child Sexual Abuse Material"... Taking care of your child's medical problem isn't "sexual abuse" despite the conflation of the two people are glossing over in this thread.
I installed an offline gallery and disabled Google Photos. For years and years Samsung android phones didn't have the option of cloud backups.
There's a very clear distinction between local and cloud, every single photo taken using any camera app is local only. Galleries, like Google Photos, and backup services like Dropbox have an explicit setting to enable cloud backup. Google Photo backup is very distinctly different from Google Drive and phone backups in general.
I use FolderSync to sync it with a self hosted NextCloud instance. If I am travelling, I prefer using Syncthing, a wifi file server running off the phone, or a physical wire to transfer photos. I know exactly what pictures are sent where and backed up with what method. I am not sure why you believe Android treats photo and file management like a black box, I don't believe Apple does either. Apple iOS devices are much more difficult to manually backup as they only allow background photo sync to iCloud, if the screen is turned off syncing to a third party will most likely fail. But there is a straightforward toggle to disable Photos backup to iCloud.
Android file management is superior to any other mobile OS. Android is file privacy first by design, as long as you are comfortable managing backup and syncing with self-hosted and self-managed solutions. I have gone the additional step of rooting my phone as well to get around any Android limitations on what folders certain apps have access to.
No one is arguing against the last point you made. No one was conflating the two.
> There's a very clear distinction between local and cloud, every single photo taken using any camera app is local only.
>Android is file privacy first by design
Yet we have the NYTimes article stating that the only ways Google could have access to the photos they tried to have a parent arrested for would be to scan content on the device, scan automatic backups, or spy on the text messages sent from the device.
The examples in the story were parents sharing images directly with their doctor, not the with the general public.
There are many examples of Google's automated systems making egregious mistakes while scanning user data with no human in the loop to review the decision.
>Ed Francis studies the evolution of military technology over at his YouTube channel, Armoured Archives. But this week, Francis says five years’ worth of research stored on Google Drive has become inaccessible thanks to Google’s automated error.
Francis says the file in question was simply a collection of data on various tanks for a coming video on how military vehicles have evolved across historical conflicts. But Google’s automated systems deemed the file a terrorist threat, resulting in a complete lockdown of his YouTube, GMail, and Google Drive accounts.
I remember that story, I came across a post asking people to upvote the support ticket on the r/datahoarder subreddit by a fan on the 17th. I came across it and cross posted a link on hacker news on the 22nd in the evening and by the following morning the account was restored and taken care of. I can't say for sure how much posting it here helped, but based on the timeline I think it made a difference. I believe it was also one of the highest upvoted stories of that day.
What stood out to me what one of the first couple comments in the support thread was the contact information for the Museum Director of the Swedish Tank Museum.
My unconfirmed theory is that Google's OCR pdf service flagged specific text and pages in the pdf of tank plans and repair manuals as they are considered classified and shouldn't be in the public unredacted. It's historically significant and absolutely worth preserving, but I can see why it might get flagged.
This was also interesting in that it got the entire account taken down. Usually Google flags a file and disables sharing completely. Google disabled sharing backups of a recent Kanye interview that was filled with hate speech called "Kanye champs removed video.mp4". I have not watched it but it was my first time seeing something removed because it "may violate Google Drive's Hate Speech Policy. Some features related to this file were restricted."
I am not sure if it was the post to HN or various videos other made that helped, but it made me realize that the one of the best ways to get in touch with a human at Google is posting here on HN and that Google will continue to become more aggressive with scanning files.
Section 230 was modified by FOSTA-SESTA when related to hosting sexual solicitation, and EARN IT is back under review by congress which would effectively make E2EE illegal and make hosts liable for illegal content uploaded by its users.
Would scanning be okay if there is a government entity with a human you can appeal to that would override any flags made by automated systems? No corporation wants more government oversight, and the only way to avoid it is to do good faith self policing. EARN IT would receive a lot more support if the public sentiment changes to believe Google and Apple and other hosts intentionally choose to ignore problematic materials. Reddit is criticized often because it allowed problematic subreddits to grow. I am not sure where things will end up.
I'm not defending this at all, but one of the reasons why there are no (or few) humans that can be contacted is that they* said that it was tried before and it caused a lot more issues with mistakes/takeovers due to social engineering.
* Can't remember who said it but it was at a town hall this year
> one of the reasons why there are no (or few) humans that can be contacted is that they* said that it was tried before
This just sounds like yet another excuse for holding payroll down as much as possible.
If I am a customer of Amazon, Apple, Netflix, Walmart or any number of the other companies with a similar market cap, I can get access to real live human beings who provide customer support.
> If the photos are hosted on Google's servers, they have every right to scan whatever they want.
I don't get why that should be the case. When you rent an apartment the owner is generally no longer able to just walk in and do whatever he likes, including installing cameras in the shower. So I would expect that when Google tries to sign you up for its online storage (repeatedly) you should get the same protections, especially with smart phones being such an important part of modern life.
It's more about shareable media. Relative to your analogy, I feel storing files in encrypted containers is renting an apartment and feeling confident about the landlord not walking in. On the other hand, if you hang up an obscene image in your window of your apartment, I imagine it's understandable that a landlord would use a key and take it down. If hanging up obscene images becomes a pattern, I think the landlord would kick you out.
From one perspective, I understand that cloud storage should include certain rights to privacy. If I buy some compute time on AWS, I manage and control all flow of data. It would be a complete violation if AWS policed the kind of content I could store. I think self-managed vs company managed is what dictates my expectations. Google photos and drive, especially since most people don't pay for it, fall into the company managed category in my opinion.
In the last few years Google Drive became more strict about sharing copyright work and started flagging files, which limits the ability the "share to anyone with a link". For people who use it as a backup, the advice I started hearing was to encrypt the data before uploading. The opsec while considering nation-state level threats to privacy already recommended only using encrypted containers since the beginnings of cloud storage.
The Communications Decency Act is an interesting law that dictates the liability of online content and service providers. There is an ongoing battle to increase ISP and host liability as a way for the entertainment industry to combat piracy.
FOSTA-SESTA and the ruling on Backpage [1] made things more difficult for hosts like Google and changed the liability surrounding sexual solicitation. If a host like Google knows what's being hosted and does nothing, they get in trouble. They don't have the protections they did in the past.
> It's more about shareable media... if you hang up an obscene image in your window of your apartment
Again, the parents Google is reporting to the police are not trying to share photos with anyone except their own doctor.
This isn't scanning information because you made it publicly available. This is scanning information on your device because it was automatically backed up by Android.
Maybe monitor companies should put chips inside that scan the current screen every second to check for CSAM matches or whatever else they don’t want the customer using their monitors for and diligently report violations to the police?
They were wanting to implement end-to-end encryption for Photos and iCloud backups (both announced today). Scanning uploaded data in the cloud for CSAM wouldn’t work then. Hence on device scanning
> Doing it on device is probably preferable to on cloud and at worst no different.
I own my device, I don't own their cloud. That's a big difference. Don't co-opt my property to do work you want done. Data stored on your servers is your business, so doing the checking there is fine, as long using them isn't mandatory.
They can't do the check there because clearly they had already been working on encrypting the data end to end making that impossible. So the middle ground was end to end encryption with on device scanning which is a step up from no encryption. Somehow we ended up with the best option of no scanning at all which is nice.
Are you trolling? What Federighi proposed before was scanning "for CSAM" on device [1]. Same angle.
> Doing it on device is probably preferable to on cloud and at worst no different.
Please elaborate. How is it better to force users to run software they don't want than to let them decide whether or not to have their photos scanned when they choose to upload them to the cloud?
Anyway it's a false dichotomy. Apple isn't doing on-device scanning, and now they've announced they won't do it in the cloud either.
Well the scanning was allegedly supposed to take place only when uploading. If a user chooses not to opt-in for cloud library the device scanning was allegedly supposed to be turned off.
So yeah allegedly no worst.
You might notice i used the word "allegedly" a lot, it's because we are speculating about a feature that was never actually rolled out and that nobody audited externally to my knowledge. If you don't trust Apple then this argument don't apply and you are probably better of not using an iPhone.
Nonetheless it's still not worst than actually rolled out CSAM scanning feature of Google Cloud that already had major user adversial effect. So you should trust Google even less and you definitely shouldn't use a stock Android device.
iOS is closed source. Literally every part is unverifiable and "forced". I have no way of proving that my iphone isn't and hasn't always been scanning my photos. But I don't have the time or energy to care about that, I've decided that using an Apple product is safer than an Android I didn't audit (which is essentially impossible). By scanning on device it enables the possibility of end to end encryption which reduces the risk of a hack or bug exposing my photos.
Yes, this whole thing is just another public relation exercise of "Apple cares about your privacy" bullshit when they are actually saying that they still plan to scan your device for CSAM. "End to end" encryption of backup on iCloud is also a joke when they are going to store the encryption keys on the iDevices on which you can run no other system software apart from the closed source ones provide by Apple.
My understanding was: it was only ever scanning things on the client that were uploaded to iCloud Photos, in the same exact way they are scanned server-side today but arguably a step more “transparently” for the user (at least if the user is a reverse engineer, heh).
What was crazy about this? The media outrage around the PR was incoherent to me at the time, and it seemed no one took any time to understand the present realities or details.
Your apartment complex decides they don’t want to be party to anything illegal. Just in case, they set up a police precinct in the lobby. They set up hidden cameras in every room of your apartment, and if their AI model detects anything suspicious, they send the video to the detective. Because you aren’t doing anything illegal, you have nothing to worry about, right?
What’s crazy about this?
Another way of expressing the concern: does your iPhone work for you, with the help of Apple’s services, or does your iPhone work for Apple? Working for me means not having software designed to report me to the police for how I use my device. The hash database for now includes CSAM hashes; there’s no reason it couldn’t be extended to include similarly heinous material, like Winnie the Pooh memes, Hong Kong freedom posters, rainbow flags, hormone therapy instructions, or anything else offensive to the regime.
People that are comfortable with their device working for a company with the assistance of the user choose Android.
Your apartment complex decides they don’t want to be party to anything illegal. Just in case, they set up a police precinct in the lobby. They set up hidden cameras in the lobby, the hallway leading to your apartment, and the balcony of your apartment. "All public spaces of course" the crowd consents. And if their AI model detects anything suspicious, they send the video to the detective. Because you aren’t doing anything illegal, you have nothing to worry about, right?
You’ve just described every photo upload service in America, although my understanding was Apple would use a list of hashes of known bad content, not an “AI” as google does.
Everyone scans for CSAM. I am not conjecturing on the ethics of scanning photos here, I am suggesting that moving from server- to client-side scanning had no effect on any of the things you are ranting about. Hence why I do not understand the outrage.
> Working for me means not having software designed to report me to the police for how I use my device.
Let me fix it: “how I use iCloud Photos, a hosted service on Apple’s servers”.
Isn’t this literally done in most office buildings in the US and probably a lot of apartment buildings already? There are CCTV cameras everywhere in the US (and other countries).
You agree _in your rental contract_ not to dump explosives in the apartment complex trash compactor, and that the apartment complex has the right to process your trash.
The apartment complex stops looking inside everyone's bags that land in the compactor (a bit late to prevent it from getting into the trash), instead sniffing bags as you put them into the trash chute on your floor, and tagging the bag with a red sticker if it the chemical composition matches a particular pre-registered explosive from a list.
At the other end of the chute, a counter checks if you've dumped a dozen bags with red stickers, and if so, a human may open one to see WTF you're up to. If you're indeed dumping explosives in bulk, they let the bomb squad know.
EITHER WAY, that bag would be headed to the compactor, you are affirmatively putting it there, and any explosives are a breach of agreement. In no other case besides you opening the chute to dump a bag is it sniffing anything else anywhere, except when you open the chute to send these bags to the compactor, bags you decide the contents of.
You'd have a hard time verifying that said on-device scanning would only have been run only on iCloud-uploaded content.
Once the feature exists your local data might become accessible to a government warrant, which would make the iPhone the opposite of a privacy oriented device.
If it's only for iCloud uploaded data they can simply do the scanning there. There's no reason to use customer's CPU/battery against them.
This necessitates a workflow where both the photos and decryption keys are accessible by the same server and that there is a security workflow to request the users decryption keys without the user involved.
This is specifically what Apple is trying to avoid - they are intentionally pushing an environment where the user must be involved to get the keys, by way of their account password and/or other enforcement mechanisms designed to ensure only the real user can access such keys.
All of the facial recognition, object identification, etc is all done on-device for the same reason. By contrast Google can and does do this in the cloud - and there are Google servers that intentionally have access to decrypt your photos.
iCloud backup was previously a vector that bypassed this, however they also today announced they are fixing that: https://news.ycombinator.com/item?id=33897793 "Apple introduces end-to-end encryption for backups"
Sure, from a technical perspective, it is a nice solution to a set of problems.
But there are many serious problems with it. I'll isolate one: an on-device content surveillance mechanism is a slippery slope to a bad, bad place.
As the saying from the 90s goes, "child porn [CSAM] is the root password to the US Constitution." It is bad enough that people are willing to suspend their better judgement to do something about it.
But after you already have the mechanism accepted an in place, it is far easier to add to it. Pick your boogeyman: you're giving all of them a lovely tool to address their desires. Dissent suppression and Winnie-the-Poo detection for Xi, Erdogan gets to sniff out Golem memes, and choose-your-own horror for the coming dictator of the US.
It is far harder to tell a sovereign, "we could easily do that, but will not" than "we don't have a mechanism to do that." And pretending that it won't happen doesn't pass the laugh test - we have seen this show many, many times. But if you want to argue, start by explaining how Apple's jumping to implement the 10-minute-max sharing limit shows how they'd stand up to China about this.
I agree there is a slippery slope concern, and Apple has themselves made related arguments such as the FBI case and not wanting to create a software update to decrypt the contents of a phone.
However that is contrasted with a very real and much more practical concern of malicious parties getting access to your cloud stored photos.
It would help to note Apple also today announced "Advanced Data Protection" in iCloud which closes the hole where iCloud Backups, iCloud Photos and various other bits of data were technically decryptable by Apple. They've closed that (but it's opt-in, to balance the average user losing all their photos against other users desire to be secure even if it means losing their data). Details: https://support.apple.com/en-us/HT202303#advanced
However even without the "Advanced Data Protection" what I said about not having a workflow for any Apple server to "normally" request both the keys and photo data is also still good security.
To detect Winnie-the-poo, it required a code push of a new database to all clients. If that’s the bar, than a corrupt apple could also push a software update tomorrow that enabled such scanning, whether this scheme was implemented or not.
> If it's only for iCloud uploaded data they can simply do the scanning there.
This is what Apple was trying to avoid. Scanning on iCloud also requires that Apple can see your photos.
If the scanning is done on device, Apple could encrypt the photos in the cloud so that they can't decrypt them at all. Neither could the authorities.
> There's no reason to use customer's CPU/battery against them.
The amount of processing ALL phones do for phones is crazy, adding a simple checksum calculation in the workflow does fuck-all to your battery life. Item detection alone is orders of magnitude more CPU/battery heavy.
> Once the feature exists your local data might become accessible to a government warrant, which would make the iPhone the opposite of a privacy oriented device
Why does that dystopia require on device scanning? Why couldn't they just do it with an OS update today? It's not a reasonably slippery slope, given the actual mechanics of how the CSAM system was designed (perceptual image hashes, not access to arbitrary files)
> There's no reason to use customer's CPU/battery against them.
That's the better argument, but still not super strong. On-device scanning means you know and can verify what hashes are being scanned for, who is providing them, and when they change. Cloud scanning is a complete black box. None of us would know if Google was doing ad-hoc scans of particular users' photos at the behest of random LEOs.
> None of us would know if Google was doing ad-hoc scans of particular users' photos at the behest of random LEOs.
Not your device, not your software. You should assume anything you upload unencrypted is scanned. This distinction was clearly voiced by the majority during the debacle of Apple's on-device scanning proposal. They basically said, "Scanning in the cloud is [choose one: fine, skeezy], but we draw the line at doing on-device scans. I don't want that software on my device."
The point of the feature is to use the data in court cases, which are public record. So word would get out there, via journalist, a whistleblower, etc. They had to make the proposal public before implementing it.
The crazy part is forcing a user to run software that undermines their own interests, on a device they purportedly own that is likely a huge part of their life. Software running on a computer should represent that computer owner's interest, period.
Apple can run software that represents Apple's interests on their own servers.
First, the huge possibility of false positives. They're not "checksums", but rather "perceptual hashes", which are nowhere near as solid as a cryptographic hash. When you multiply a small chance of a false positive across hundreds of millions of deployed devices, you end up with many false positives - false positives that result in people being characterized as "child abusers".
Second, the overall aim of the technology also targets legitimate or inadvertent activities, that the law has also unjustly criminalized - say taking family pictures of your kids, 17 year olds sexting, or someone trolling you. Once again unjustly branding people as "child abusers", this time with "evidence" to back up the narrative.
But even in the case of a proper match to the NCMEC list for its bona fide purpose, that still is undermining the phone's owner's interests. That's the philosophical contour, even if you personally wish to brush it aside with the desire to catch people looking at evil images. Our society respects similar longstanding privacy boundaries, even if it does end up helping some "bad people".
It's really not in the NCMEC's interest to respect any of this, as their foundational dynamic is that there is a horrible thing happening in the world, and they must do everything possible to stop it. That kind of advocacy is certainly needed, but we shouldn't just accept their desires as if they're an unbiased neutral party.
Yes, the internet went on a "generate false perceptual hash positives" binge.
Everyone just forgot that in the Apple system you don't get automatically banned and reported to the authorities if you get a match (or five). The matches get manually checked by an actual human, who will see in 0.2 seconds that it's not actual CP, but a highly distorted picture of a cat or some gibberish. Zero action will be taken.
Also: the scanning was only done if iCloud is enabled. If iCloud is enabled, the EXACT SAME scan could be done in the cloud. Why are people not objecting to this with the same fervour?
A lot of the concern is that others may not trust the NCMEC, or that they don't trust that other images won't be added in whatever ends up on your device.
It took me a while of ruminating but I think I like this answer the best. Fundamentally here some people are upset that something they own would betray them.
Which is weird, because it implies anything you used but didn’t buy can and should betray you, and it’s your own fault. You should have been smarter. But it makes a bit of sense, for some odd reason I can’t quite make tangible. When the samsung TV I bought starts showing me ads, when my pricey cable or Hulu package interrupts for commercials, I feel a tinge of this.
So now I think I somewhat get the feeling. Thank you.
> because it implies anything you used but didn’t buy can and should betray you, and it’s your own fault
The contrapositive isn't automatically true. You can be upset when a cloud host or ISP violates your expectations, even though you don't own them. You can expect Apple to basically work in your interests even on their own servers, since you're paying them.
The ownership thing is more about the structure of society. If computing and communications technology is going to move us forward as a distributed democratic society, then those capabilities must remain distributed throughout the population rather than being controlled by a handful of centralized gatekeepers.
And yes, we've strayed quite far from this ideal. Most people have little control over what their mobile phone does. Still, there is a distinction between a device not having your desired functionality because the manufacturer didn't create it (or even disallowed others from creating it), and the manufacturer deliberately creating functionality that harms users.
Meh. Running it on device means that it is narrowly scoped and transparent. There is no way to secretly add a new hash or to vary what you're searching for depending on a user's location or identity.
Running it on their servers means nobody knows how expansive the scans are, and there is nothing preventing ad-hoc scans for arbitrary content for individual users.
Truthfully generating that fingerprint is not in the user's interest. User representing software would skip doing that. If the storage service required it, it would generate plausible dummy values.
Pretty much every "file" you create on your iphone is uploaded to iCloud by default. For example if you take a screenshot, it saves to your photos library which defaults to uploading new photos to iCloud immediately.
It's my device, working against me; and it's also not going to catch anyone remotely intelligent, since everyone and their mother will know about this feature.
What’s always so interesting about this then is what’s the point? Apple catches the least sophisticated of criminals who don’t know how to turn off iCloud (likely the least harmful nodes in this network of horror)?
And to accomplish that all we have to put up with is a huge lack of privacy and tons of false positives?
The point was to encrypt all iCloud data so that Apple could not give access to it to authorities.
Now the authorities can go "there might be CP in there, think of the children!" and a judge will give them the warrant.
On-device check agains KNOWN CP checksums + E2EE encrypted iCloud -> authorities need to get the actual phone, they can't just go looking around on people's iCloud accounts to see if they might find something.
Either commit and say “scanning for CSAM is the top priority” and scan on device or say “privacy is the top priority” and don’t scan at all.
I’m literally saying this halfway path is the worst of both worlds, since people still lose privacy, but very little meaningful progress against csam will be made.
> Additionally, the core of the protection is Communication Safety for Messages, which caregivers can set up to provide a warning and resources to children if they receive or attempt to send photos that contain nudity.
That was announced at the same time, but is different from what the parent is talking about.
One program scans photos that were about to be messaged and if they look like nudity asks "are you sure you want to do that" (and also optionally notified the parents, although it looks like they aren't going through with that part). The scanning happens on device, and the results of the scanning are never sent to Apple or any other third party.
The other program scanned all images that were being uploaded to iCloud and if it found any that matched a government maintained image fingerprint database would notify Apple who would forward that information to the government. Big difference in scope and impact.
They haven't really - I logged in to specifically post about that. Like you pointed out, the real outrage was never about scanning for CSAM in the "cloud" but on your device. And this clever fluff of public relation exercise by Apple is just to cover up the fact that device scanning for CSAM is still in the picture.
Communication Safety for Messages is opt-in and analyzes image attachments users send and receive on their devices to determine whether a photo contains nudity ... The company told WIRED that while it is not ready to announce a specific timeline for expanding its Communication Safety features, the company is working on adding the ability to detect nudity in videos sent through Messages when the protection is enabled.
.... “Additionally, because the minor is typically sending newly or recently created images, it is unlikely that such images would be detected by other technology, such as Photo DNA. While the vast majority of online CSAM is created by someone in the victim’s circle of trust, which may not be captured by the type of scanning mentioned, combatting the online sexual abuse and exploitation of children requires technology companies to innovate and create new tools. Scanning for CSAM before the material is sent by a child’s device is one of these such tools and can help limit the scope of the problem.”
As for the newly announced "end-to-end encryption" on iCloud, note that the keys will be stored on the iDevices, and as such, available to Apple through its software anytime. This is exactly why the US government and BigTech have been pushing for "passwordless" authentication so strongly. It's no coincidence that Microsoft suddenly started asking for TPM to run Windows. With "passwordless" authentication, we will even lose control over the digital keys that allow us to encrypt and access various services, and BigTech becomes in charge.
Yes, I am being very cynical here for 2 major reasons - (1) Governments around the world want control and access to our device, and BigTech can deliver. This is not an Apple or Google problem, this is a social and democratic problem that we need to fight politically with the government by demanding stringent data protection and privacy laws. (2) Apple is a corporate who needs to be profitable. Collecting data and selling it (either to the government through programs like PRISM or to advertisers through a platform is a very lucrative source of revenue. Anybody who things Apple will let go of billions of dollars from that is delusional (and that is exactly why Apple is pivoting to become a service company). Apple's history when it comes to invading its user's privacy is just as bad as Google's or Facebook.
Arguably, scanning on iCloud is precisely the right answer, but in specific circumstances. I figure that cloud storage accessible only to yourself should be considered as private as the storage on your own physical device. But when you share content with other people, especially if you share via a publicly reachable URL (not sure if iCloud can even do that, though), then scanning for illegal content is fair game.
It wasn’t device scanning. A fingerprint of every photo was uploaded to the cloud as metadata for each photo. It’d be like being spooked if the title of the photo is uploaded to the cloud and examined for known csam titles.
And to make things more secure, the only way apple could read those fingerprints was if you uploaded 5 csam photos.
How is that worse than Google where engineers can look at every photo you upload regardless of the content as part of their server side scanning?
I took their argument to be that scanning at all was the problem. I may have misunderstood. I've made way too many comments today on HN, and I should most probably step away from the keyboard. In fact, I'm going to do that for real, my daughter just got home from school and I need to remember priorities.
Scanning on iCloud means that Apple can see the content of all your scannable data in iCloud. Scanning on-device is compatible with Apple never having access to your data in an unencrypted format. If Apple has a legal obligation to ensure that iCloud does not store CSAM/etc. then either you have to scan on device before upload _or_ you have to store iCloud data without E2E encryption. From a privacy perspective, on-device scanning before upload is obviously better.
> Scanning on-device is compatible with Apple never having access to your data in an unencrypted format.
Only if you exclude the following network transmission, which is the easy part that needs no special code. The privacy concern comes in with those two things together. So yeah if you take away one half of a bad thing such that the bad thing no longer becomes possible, it's not bad anymore. The concern is the whole process, on-device scanning being the key not-yet-implemented component.
> If Apple has a legal obligation to ensure that iCloud does not store CSAM/etc. then either you have to scan on device before upload _or_ you have to store iCloud data without E2E encryption.
Apple does not have that legal obligation. If they can't decrypt the content on their servers, then their only response to a government-issued warrant would be to hand over encrypted data.
Also, CSAM is not the concern. The concern is this would be used against dissidents in authoritarian countries. On-device scanning takes us a step towards becoming one and further empowering the existing ones.
This doesn't necessarily follow. Law enforcement having near realtime access to everything you ever photographed is a worse situation then them having to know in advance which things you might be storing, so they can add a fingerprint of it to a database and then wait until your device matches and uploads it.
CSAM is a serious problem and you're ignoring Apple's moral or even repetitional obligation to try to address it. However poorly.
We shouldn't be hyperbolic and just flatten levels of badness, or we'll can never find a compromise (which is, I suspect, just how you want things).
" If Apple has a legal obligation to ensure that iCloud does not store CSAM/etc"
My understanding is in the USA companies like Apple cannot be legally obligated to ensure that iCloud does not store CSAM. Something about the US Constitution, but I can't remember what. Apple is legally obligated to report CSAM if they come across it themselves though.
> Something about the US Constitution, but I can't remember what.
The first amendment comes into play. The government cannot compel Apple to write software in a certain way, such as "write your encryption so you have keys that access all of users' data". That would be "compelled speech". So if the government provides Apple with a warrant, Apple can only provide encrypted or whatever meta information they have, not decrypted content.
EDIT: Whether or not Apple should be scanning for CSAM in some way is an entirely separate argument to whether it should be "on device" or "in cloud". That's not obvious, so my explanation below explains why that is the case.
I disagree. On-device scanning is desirable here and scanning it in iCloud though the industry-norm, is not desirable.
Apple specifically stated they would only scan the photos being uploaded to iCloud [1] and not local photos which were not being uploaded.
Most cloud storage providers are scanning in the cloud, this necessitates a workflow where both the photos and decryption keys are accessible by the same server and that there is a security workflow to request the users decryption keys without the user involved (by way of their account password, or other enforcement mechanisms designed to ensure only the real user can access such keys, Apple has previously detailed Hardware HSM enforced details of that for iCloud in their Apple Platform Security guide [2]).
By extension that opens up the potential for a malicious employee or third-party to either intercept those photos while being scanned or otherwise allows a workflow that could be exploited to request the keys and photos. This is a major loss of security and privacy which Apple is specifically and intentionally designing to avoid. Until now iCloud Backups were a major weakness in that story however they also today announced they are fixing that [3].
Hence: the on-device scanning was just an entirely technical solution to scanning "only photos uploaded to icloud" (like everyone else) while also "ensuring some server in icloud can't decrypt your photos".
You could of course argue that it makes a later change of policy to scan local-only photos easier. While this is true, Apple could decide to add code to iOS to do anything they want at any time, so they could just decide to do that anyway, even if they were scanning iCloud photos in the cloud. At the same time, by not scanning them in the cloud, they have a technical solution to ensure it's much harder for a malicious actor to get unauthorised access to iCloud user photos. Because there is no place/server that is "supposed" to be able to access both the photos and the keys.
What are the benefits? It would tell Apple (and hence law enforcement) who had copies of existing, known CSAM. I guess that would help catch such people. But such individuals would also have strong incentive to simply turn of iCloud to completely avoid detection.
What are the risks? Apple employees rummaging through your photos (after 30 NeuraHash matches, a human reviewer would go in and check your photos, and as some analysts pointed out, hash matches can be spurious, that is, completely different photos can have the same hash by coincidence).
Over time, would the program expand beyond CSAM? Consider the intersection of countries in which Apple sells iPhones and countries that criminalise being LGBTQ [1]
All of your concerns are valid for "should Apple do CSAM scanning at all" and I share those.
However my argument is specfically to the OPs point that they should just scan the photos "in the cloud" and not "on your device". But the reality is they were only scanning photos that were being uploaded to the cloud, they were just doing it on your device so that Apple doesn't need to have a server that can decrypt your photos to do it in the cloud. Which is what most others (including Google) do.
> as some analysts pointed out, hash matches can be spurious, that is, completely different photos can have the same hash by coincidence
I think you might have misinterpreted whatever those analysts wrote. The chances of thirty images all coincidentally matching both perceptual hashes is essentially nil.
I can't agree. Normalizing scanning of files on people's devices is a very very bad precendent. It turns Apple into a low level government policeman rifling through your device which would spread to your PC and android devices. Normalizing this behavior would be truly awful, CSAM would come first, then tax audits, anything that looks "suspicious", etc. That is a very very bad precedent, police should only be able scan your device if they have a warrant.
The normalisation aspect of this would have been the biggest horror. It would mean owning a device that doesn't do this becomes associated with criminality. "Don't you know that only pedophiles run Linux?"
It wasn’t scanning photos looking for csam, it was computing a fingerprint of every photo that could only be read on the server if 5 of those fingerprints matched known csam.
I can’t see how that is worse than unencrypted photos in the cloud that can be scanned at any time server side.
If Apple were to calculate the perceptual hash client-side but only do the actual hash comparison server-side following an upload, it would preserve the 'no keys on iCloud' policy while also preserving the 'no local scanning' policy most of us want.
That the server may know the hash is IMHO much less a problem (it shouldn't be reversible anyway) than the legitimization of on-device scanning.
I would think calculating the perceptual hash client side is precisely the "local scanning" part most are concerned about. Pushing a perceptual hash of every photo to Apple would actually decrease your privacy as now Apple can find any known photo in your library - not just CSAM but on any topic - memes, political protests, porn, whatever.
IMHO, it's the comparison to a stored hash value which is the scanning part. In my scenario, the hash calculation is only done during upload and the hash is not saved on device.
Let's say tomorrow China tells Apple to do client-side scanning, and let's say Apple decides to comply.
Scenario A (Apple already did on-device scanning):
Apple flips a few bits. iOS no longer cares whether images where uploaded on not. The code change can easily be hidden assuming no reverse engineering. There's little measurable difference unless the user has a matching image in which case the user is already done for [EDIT: in retrospect, the device should be running a scan on the local images every blacklist fetch. I still think that scanning more images is much less of a difference than the other scenario, where the device has to do stuff it just did not do before]
Scenario B (Hash Comparison is server-side):
The code change can be hidden of course. But otherwise there are significant changes:
The device needs to calculate hashes to every image upon creation which it did not before.
The device needs to fetch the blacklist every X days which it did not before.
The device needs to run the local scan every blacklist fetch which can be slow if many images.
All these things are far more likely to be noticeable and far more likely to generate network traffic.
>Pushing a perceptual hash of every photo to Apple would actually decrease your privacy as now Apple can find any known photo in your library - not just CSAM but on any topic - memes, political protests, porn, whatever.
The Hash shouldn't be reversible, so the only way for Apple to do that is to have an existing dataset. Which is possible, but preferable to the no E2E situation and (IMHO) the on-device scanning situation.
That’s not how their encryption worked. The only way for apple to know a csam photo matched was if you had 5 photos that had a csam fingerprint from their database uploaded to iCloud. This matching only happened server side for photos uploaded to the cloud.
This seems far better than having every photo scanned server side and have those photos sit there unencrypted.
I think people objected to having the flagged hash database and comparison locally. I suppose the split of perceptual hash client side, comparison server-side might be objectionable on the "it's draining my battery" angle, but I suspect those who prefer cloud-side scanning would generally like hash creation client-side if it were the price of E2EE.
Of course that's a lot less secure and leaks more privacy info, but the objections seem more about the morality of having a device you own doing the scanning, not so much about any privacy concerns.
The comparison didn’t happen locally in Apple’s scheme. Along with each photo a encrypted fingerprint was uploaded and all scanning happened sever side.
>If the user image hash matches the entry in the known CSAM hash list, then the NeuralHash of the user image exactly transforms to the blinded hash if it went through the series of transformations done at database setup time. Based on this property, the server will be able to use the cryptographic header (derived from the NeuralHash) and using the server-side secret, can compute the derived encryption key and successfully decrypt the associated payload data
So my understanding is that the server won't even know the hash of the image unless the client-side operations determine that it matches one of the hashes in the CSAM database, which is stored on the client.
I’m seeing the same. There are a lot of people saying apple’s scheme was somehow worse than Google’s. I was arguing that apples was at least attempting to be privacy preserving.
I think the state of the art has progressed past that? A non-trivial reversal of a perceptual hash would mean that every cloud provider maintaining a CSAM scan list violates CSAM laws - if they could reverse the hash to get too close to the original image, the hash is just a lossy storage format.
My own devices should not snitch on me to the government. Obviously this mandatory snooping of all private data is not going to stop at the single most politically salable application; that's just the starting point.
You are right, but by local-only photos I meant "not using iCloud photos" - e.g. "none".
The point is, I think if Apple decided to scan your photos for CSAM even if you weren't using iCloud, that this would be a distinctly different situation to scanning only the ones uploaded to iCloud.