#25 – Privacy definitions for engineers
The Privacy world has done a great job of over complicating things for engineers with conflicting definitions in the GDPR, CCPA, ISO and elsewhere. In this episode get to the heart of what you need to know for some of the main privacy terminology out there.
Audio
Transcript
Hello and welcome to the GDPR Guy. I’m Carl Gottlieb, your host and resident privacy advisor.
A quick warning for this podcast. I’m in a bit of a Candid Carl mood so there will be some swearing. If you don’t like that kind of thing then please don’t feel the need to listen any further.
Working with lots of tech clients, and therefore lots of different engineers, means I hear the same questions a lot. And I mean a LOT.
And the most common ones centre around what some privacy terms mean as they relate to engineers, specifically:
What is PII?
What does processing mean? and
What does anonymous really mean?
In the future I’ll probably point my client engineers to this podcast, but in the meantime this is one for you listeners.
So let’s start with PII.
Now first of all, privacy people often act like complete dickheads when they hear the term PII, as in Personally Identifiable Information. This is mainly because they’re annoying pricks that have spent all of their time studying the GDPR and zero time in the real world helping people. And I’m here to help you, so I’m happy to use this term and I’ll explain why.
There’s a long and mixed history behind these three terms, of Personal Information, Personal Data and PII that spans continents and decades of legislation, standards and general principles. It also spans both privacy and security domains as well as other areas of IT. And in each area, you’ve historically found differences in exact definitions, but the key thing to say is that all three terms are legally becoming synonymous. They’re all starting to mean the same thing. And more crucially, all of these terms become the same when you’re looking at contracts, which for tech companies especially in B2B is most of what you need to care about.
Let me elaborate on that. If you’re in B2B, your primary set of rules you need to comply with is your customer contract. And likely that contract will stipulate what further laws you need to comply with, such as some local law of your client. And of course on top of that, you’ll need to comply with your own local laws, but for B2B it’s the contract that will stipulate how you’re to process your client’s data.
And in that contract, or maybe its Data Processing Addendum, the DPA, it will refer specifically by definition to personal data, or PII, or whatever term it wants, to mean the same thing in almost all cases. The document will only refer to one definition, and it can define it how it likes.
Let’s quickly explain that for normal people that don’t spend their days in contracts. In a contract you have all the boring sentences, that sporadically, will point to randomly capitalised words mid-sentence. These capitalised words are known as Defined Terms and they’re important because they mean a very specific thing. And they can mean anything. The document will define them somewhere – hopefully at the start of the document in a dedicated section or sometimes throughout the text in brackets somewhere.
Defined Terms are important because they clarify and precisely define the exact scope of a word and mean you can be as creative as you like with them in your contract. Want to define a Security Incident as an event that requires at least a billion records stolen? Go for it. Want to define Applicable Laws as meaning only US laws? Why not?
And crucially, here’s where you define Personal Data, or Personal Information or whatever term you want to use to refer to PII throughout the DPA and contract. This means that you can remove any ambiguity of what the term means. You can state that it excludes publicly available information, you could remove any scope of data for which the processor already owns, and you could limit the term to only apply to a subset of the data the client gives you. A common example of this is where half the data you receive is for you to process as a data processor on behalf of the customer, and the other half is for you to share out in the public domain as a data controller like a media company would.
Being precise with the term allows you to be extremely narrow on what contract terms and what protections apply to what data.
And for engineers, this means that in the wider world, the term itself doesn’t really matter – it’s how we define it on paper and what rules we contractually apply to it that matters. Yep, you’re going to need to look at the contract, or more likely, just ask your friendly privacy or legal person what it says.
And in reality, what you’ll find in most DPAs is that PII is just any data that the client gives you that could identify one of their people, even indirectly. If the client thinks it’s PII then it’s PII to you. I’ll say that again because it’s ultra important. If the client thinks it’s PII then it’s PII to you.
And back to my earlier point – I hope you now see why for people on the ground it’s stupid to argue over whether PII or Personal Information is the right term when talking about the this stuff.
And lastly, if you want to have a final word on this, just remind people that the International Standards Organisation, as in ISO, you know who does ISO 27001 and all that. Yes, The Standards organisation, uses PII for its definition.
The other two questions I wanted to address are “what is processing” and “what is anonymous”. Again, both will often be addressed in contracts as Defined Terms, but it’s worth quickly talking about why engineers will ask about this.
Engineering is all about data, and so is building a product these days. AI relies on good data and so do high tech valuations so everyone wants lots of data to play with. And this unfortunately has led over the last few decades to people building huge warehouses of data or “data lakes” where they can store all this data and supposedly find some analytical value from it one day.
It does remind me of people getting their bodies frozen when they die so that one day science will work out a way to cure their diseases and bring them back to life. It ain’t gonna happen.
And so with all this data around them, engineers, and product people, will always get excited about what they can do with it, hence the questions of what contract terms and regulations might limit them with their rules on PII, data processing and anonymisation.
We know that PII is usually anything that helps identify people, but again, you’ll be amazed how that term gets adjusted in contracts so please check.
Processing is almost always what it seems to us privacy people, as in any viewing, editing, touching, storing, deleting – basically anything at all that could involve some visibility of the data, even if it just passes through your wires or through your system for a microsecond. Just because you don’t technically use it for anything or don’t want to use it, if you have the ability to see or interrupt that data, such as by pulling out a physical or virtual wire, then you’re processing it.
Whether you’re processing PII is a very different question, as often you’ll be processing non-PII data. But any PII you come into contact with is you processing it. End of story.
Anonymisation is a both simple and super complicated area. And engineers will always look to analyse this at a technical level, happy to debate for hours over how anonymous something actually is. But in almost all cases, that debate ends up with the same conclusion, that PII only becomes anonymous when we can’t stand up in court and prove that we did everything we could to remove all the identifiable parts of it. And specifically, all within our bubble of what we’re allowed to do legally. It’s fair to say something is anonymous if you don’t have a legal way of deanonymizing it. A phone number might be anonymous data to me, but I could make it PII by breaking into Vodafone and stealing their customer records to deanonymize it. Just as a large multinational could illegally share data between its businesses to deanonymize pockets of data. It could technically do it, but it wouldn’t be legal.
But simply, anonymous data isn’t PII because it can’t help identify someone
The anonymisation test is really about people and not data. Are you genuinely happy you stripped all identifiers from the data? Could you explain that to a judge, to a regulator, to a child’s parent or to the press? Mathematical measures of anonymisation don’t matter when it’s your job or your company’s reputation on the line. Ordinary people, just like you engineers, want to answer the binary question. Was it anonymous or not?
And for B2B, again it’s your contract that will help clarify the waters here by stipulating what it means by anonymous. Now these contracts also have audit rights, so be ready to explain to a customer, or even an ex-customer what anonymous data of theirs you have and process and how. Just remember that if a customer gives you PII, then it is still PII. It doesn’t matter that it might appear anonymous to you. It’s still PII. If the customer strips out all identifiers and tells you it’s anonymous then it’s now anonymous.
Generally you’ll have a clause in a DPA that says you can anonymise their PII for your own unrestricted use. Customers often get concerned about these clauses, but if it’s properly anonymised then it shouldn’t really matter, and there’s no need for them to worry.
But people do.
Remember when Spotify put up advertising billboards mocking a few anonymous users for how they used the app? Back in 2017 Spotify had an ad in London that said, “Be as loving as the person who put 48 Ed Sheeran songs on their “I Love Gingers” playlist.”
Do you know who this person is? No. The ad is anonymous. Maybe the person doesn’t even exist. And if they do, then Spotify knows who they are, but they created an anonymous advert.
The problem is that people freaked out over these ads saying that Spotfy was spying on them and all this bollocks.
There was no invasion of privacy here.
But the key lesson is that people get weird about how you use their data, often irrationally so. So if you play in the margins of definitions of anonymous data, or even go head first into the completely anonymous data zone, you might personally get burned. So protect your data and yourself as much as you can.
So let’s do a quick recap:
PII will likely mean the same to you as Personal Information and Personal Data. Just go and read the contracts that apply to you.
Processing means literally any viewing or relaying of data, and
Anonymous is where you’ve stripped all the bits of PII from a data set such that you can’t reidentify someone.
We Privacy people do a bad job at simplifying all this, so please try and ignore all the technical privacy wordage out there. Ignore the GDPR, ignore the CCPA. Just come and talk to us in the Legal or Privacy team, and we can run through what it says in the contract or the privacy policy, so you know what you’re dealing with.
Thanks for listening.
More Episodes of The GDPR Guy
