“What started to grow was the notion of a privacy officer or privacy manager as someone who could run a program that could pull together the technical and the legal piece, and I think everyone in the profession at the time thought that was a really good thing,” Kosa said. “But as the discipline grew, as the domain evolved, a lot more people got interested in it, but a lot of those people got interested not for the same reasons the people who grew the field were interested in it.”
In other words, it turned into a compliance-based exercise.
That shift didn’t sit well with her. What irked her was her sense that the field was losing its strong base of privacy advocates, replaced by professionals who were saying to companies, “I can knock out a privacy impact assessment for you for $50,000, no problem.”
Showing posts with label Ethics. Show all posts
Showing posts with label Ethics. Show all posts
Sunday, 5 March 2017
Should the privacy profession adopt a code of ethics?
Sunday, 10 April 2016
Information wants to be free...?
I both like and dislike the quote "information wants to be free", mainly because it opens up a very nice philosophical discussion on what 'free' means but also because - and this is part I hate - it is some damned meaningless without any grounding in any form of semantics; and we've seen this before!
For the first part, this statement treats information in an anthropomorphic manner. Is it really information itself that has the need to be free? Let's assume that it does, though in a very fairy tale like way, it seems to me.
So let's then look at the word 'free', which I assume does not mean 'free' as in 'without cost' in the sense that someone has to pay for it. Though this is a curious idea in that information is somehow prostituting itself and despite all attempts someone (the information's pimp perhaps?) insists on controlling things. I guess this is the idea that information is going through some pre-1960's sexual revolution...
Let's stick with the word 'freedom' and its naive or common-sense meaning. What does it mean to be free? We can turn further to the the UN Declaration of Human Rights and the EU Charter of Fundamental Rights for further clarification, though here I'm sure we go into more of a legal-political debate more than anything. Evidently freedom means either the right to do or be something or the right to be protected from something.
So the question is, if information want to be free:
Under the first question, the freedom to be 'free' as in 'without cost' certainly falls. What about the freedom to be private, or the freedom not to be abused - as in excessive privacy violations? Do we further need a notion of agency - does information have an owner or provenance?
Without answering those - I don't think I can give a definitive answer anyway - here's another thought. Given that matter = energy, isn't the use of the term information quite literally another way of saying 'humans' (or 'men' as in the Declaration of Independence). In which case the question 'information wants to be free' is just an expression of man's desire to define what freedom is - ostensibly in terms of freedom to do/be and freedom from.
And here comes the practical part, which freedoms to do/be or from do we allow or deny in order to be "free"?
For the first part, this statement treats information in an anthropomorphic manner. Is it really information itself that has the need to be free? Let's assume that it does, though in a very fairy tale like way, it seems to me.
So let's then look at the word 'free', which I assume does not mean 'free' as in 'without cost' in the sense that someone has to pay for it. Though this is a curious idea in that information is somehow prostituting itself and despite all attempts someone (the information's pimp perhaps?) insists on controlling things. I guess this is the idea that information is going through some pre-1960's sexual revolution...
Rather I think the word 'free' refers to 'freedom' albeit in a Western sense of the word. Think of the use of the concept freedom as used in the US Declaration of Independence.
"We hold these truths to be self-evident, that all information is created equal, that it is endowed by its Creator with certain unalienable Rights, that among these are Life, Liberty and the pursuit of Happiness"...or should that be Existence, Communication and Semantics perhaps....?
Let's stick with the word 'freedom' and its naive or common-sense meaning. What does it mean to be free? We can turn further to the the UN Declaration of Human Rights and the EU Charter of Fundamental Rights for further clarification, though here I'm sure we go into more of a legal-political debate more than anything. Evidently freedom means either the right to do or be something or the right to be protected from something.
So the question is, if information want to be free:
- What does information want the freedom to do/be?
- What does information want the freedom from?
Under the first question, the freedom to be 'free' as in 'without cost' certainly falls. What about the freedom to be private, or the freedom not to be abused - as in excessive privacy violations? Do we further need a notion of agency - does information have an owner or provenance?
Without answering those - I don't think I can give a definitive answer anyway - here's another thought. Given that matter = energy, isn't the use of the term information quite literally another way of saying 'humans' (or 'men' as in the Declaration of Independence). In which case the question 'information wants to be free' is just an expression of man's desire to define what freedom is - ostensibly in terms of freedom to do/be and freedom from.
And here comes the practical part, which freedoms to do/be or from do we allow or deny in order to be "free"?
Friday, 23 May 2014
Surgical privacy: Information Handling in an Infectious Environment
What has privacy engineering, data flow modelling and analysis got to do with how infectious materials and the sterile field are handled in medical situations? Are there things we can learn by exploiting by drawing an analogy between these seemingly different fields?
We've discussed this subject earlier and a few links can be found here. Indeed privacy engineering has a lot to learn from analogous environments such as aviation, medicine, anaesthesia, chemical engineering and so on; the commonality here is that those environments understood they had to take a whole systems approach rather than relying upon a top-down driven approach or relying upon embedding the semantics of the area in one selected discipline.
We've discussed this subject earlier and a few links can be found here. Indeed privacy engineering has a lot to learn from analogous environments such as aviation, medicine, anaesthesia, chemical engineering and so on; the commonality here is that those environments understood they had to take a whole systems approach rather than relying upon a top-down driven approach or relying upon embedding the semantics of the area in one selected discipline.
Monday, 19 May 2014
Foundations of Privacy - Another Idea
This got triggered by a post on LinkedIn about what a degree in privacy might contain. I've certainly thought about this before, at least in terms of software engineering, and even have a whole course that could be taken over a semester ready to go.
Aside: CMU has the "World's First Privacy Engineering Course": a Master of Science in Information Technology—Privacy Engineering (MSIT-PE) degree. So, close, but a major university here in Finland turned down the chance to create something similar a few years back...
Aside: CMU has the "World's First Privacy Engineering Course": a Master of Science in Information Technology—Privacy Engineering (MSIT-PE) degree. So, close, but a major university here in Finland turned down the chance to create something similar a few years back...
That aside, I've been wondering about how to present they various levels of things we need to consider to properly define privacy and put it on strong foundations. Though in the guise of information theory we already have this, though admittedly Shannon's seminal work from the 1930's is maybe a little too deep. On the other hand understanding concepts such as channels, entropy are fundamental building blocks, so maybe they should be there along with privacy law - now that would make some course!
Even just sketching out areas to present and what might be contained therein...how about this, even if a linear map from morality to mathematics is too constraining?
There are missing bits - we still have a semantic gap between the "legal world" and the "engineering world"; parts that I'm hoping that things such as the many conferences, academic works and books such as the excellent Privacy Engineer's Manifesto and Privacy Engineering will play a role in defining. Maybe the semantic gap goes away once we start looking at this...is there even a semantic gap?
However, imagine for a moment starting anywhere in this stack and working up and down and keeping everything linked together in the context of privacy and information security. Imagine seeing the link between EU privacy laws and type theory, or between the construction of policies and entropy, the algebra of HIPAA, a side course in homotopy type theory and privacy...maybe with that last one I'm getting carried away, but, this is exactly what we need to have in place.
Each layer provides the semantics to the layer above - what do our morals and ethics means in terms of formalised laws, what do laws mean in terms of policies, what do policies mean in terms of software engineering structures, and down to the core mathematics and algebras of information.
Privacy and privacy engineering in particular almost has everything: law, algebra, morals, ethics, semantics, policy, software, entropy, information, data, BigData, Semantic Web etc etc etc. Furthermore, we have links to areas such as security, cryptography, economic theory etc!
Aren't these the very things any practitioner of privacy (engineering) should know, or at least have knowledge of? Imagine if lawyers understood information theory and semantics, and, software engineers understood law?
OK, so there might be various ways of putting this stack together, competing theories of privacy etc, but that would be the real beauty here - a complete theory of privacy from the core mathematics through physics, computation, type theory, software engineering, policies, law and even ethics and morals.
But again, no more naivety, no more terminological or ontological confusions, policies and laws being traceable right down to the computation structures and code. Quite a tall order, but such a course bringing all these together really would be wonderful...
And wouldn't that be something!
An Access Control Paradox
The canonical case for data flow and privacy is some data collection from a set of identifiable individuals and generate insights (formerly called reports) about these. In order to protect privacy we will apply the necessary security and access controls and anonymisation of log files as necessary.
Let's consider the case where where generate a number of reports, and we'll order them according to some metric of their information content and specifically how easy or possible it is to re-identify the original sources.
Consider the system below, we collect from a user their user ID, device ID and location - this is some kind of tracking application, or for that matter, any kind of application we typically have on our mobile devices, eg: something for social media, photo sharing etc...
We've taken necessary precautions for privacy - we'll assume there's notice and consent given - in that the user's data is passed using a secure channel into our system. Process of this data takes place and we generate two reports:
In many cases this is considered sufficient - we've the notice and consent and all necessary access controls and channel security. Protecting the report or file with the sensitive data in it is a given. But now the less sensitive data is often forgotten in all of this:
Let's consider the case where where generate a number of reports, and we'll order them according to some metric of their information content and specifically how easy or possible it is to re-identify the original sources.
Consider the system below, we collect from a user their user ID, device ID and location - this is some kind of tracking application, or for that matter, any kind of application we typically have on our mobile devices, eg: something for social media, photo sharing etc...
We've taken necessary precautions for privacy - we'll assume there's notice and consent given - in that the user's data is passed using a secure channel into our system. Process of this data takes place and we generate two reports:
- The first containing specific data about the user
- The second using some anonymous ID associated with certain event data for logging purposes only. This report is very obviously anonymous!
For additional security purposes we'll even restrict access to the former because it contains PII - but the second which is anonymous doesn't need such protection.
In many cases this is considered sufficient - we've the notice and consent and all necessary access controls and channel security. Protecting the report or file with the sensitive data in it is a given. But now the less sensitive data is often forgotten in all of this:
- How is the identifier generated?
- How granular is the time stamp?
- What does the "event" actually contain?
- Who has access?
- How is this all secured?
Is the identifier some compound of data, hashed and salted, for example:
salt = "thesystem";id = sha256( deviceId + userid + salt);
This would at least allow analysis over unique user+device combinations and the salt, if specific to this logfile or system, then restricts matching to this log file only. Assuming of course the salt isn't know outside of here.
The timestamp is of less importance but if of very high granularity would prevent the sequencing of events.
The contents of the event are always interesting - what data is stored there? What needs to be and how? If this is some debug log then there's probably just as much here as there is in the report containing the PII. Often it might just be stack traces (with or without parameters), or memory dumps - both of which contain interesting data, even if it is just a pointer to where a weakness in the system might exist.
Now come the questions of who has access and how is this secured? Given that such a report has interesting content shouldn't this be as secure as the report containing specific and identifiable user data? If there's some shared common knowledge could rainbow tables of hashes etc be constructed?
Consider this situation:
Where two separate systems exist, but there exists a common path between these systems which can be exploited because access control wasn't considered necessary for such "low grade", non-personal data.
Any common path is the precursor to de-anonymisation of data.
This might seem to be a rather trivial situation, except that such shared access and common knowledge of things such as salts, keys etc exist in most companies, large and small. In the latter it is often hard to avoid. Mechanisms such as employee contracts and awareness training actually do very little to solve this problem as they aren't designed to address or even understand this problem.
And here lies the paradox of access control: while we guard reports, files, datasets containing PII, we fail to address the same when working with anonymous data - whatever anonymous means.
Tuesday, 11 June 2013
Privacy, Data Collection and Surveillance
The privacy debate about the collection
of data by the NSA continues with many asking questions about the moral and
ethical issues surrounding this. The phrase "the death of privacy" is
abound.
This is true I'm afraid, we lost our
privacy, but not when the NSA starting collecting data but when we starting communicating
using technologies that were readily and easily available - that probably dates
back to the birth of written communication.
Data collection concerns me certainly,
but here I want to focus on one of the maxims of privacy: "if you don'tuse it, don't collect it" and the fact that privacy is much more about the
usage of data, not its collection (viz. the above maxim).
One can argue that merely using Google,
Facebook and all the rest of the social media services one has already lost
one's privacy, but interaction with these services is voluntary - no-one forced
you to post those party pictures to the entire World and dog (complete with
EXIF and location information).
We
admittedly do have a problem with other more hidden aspects of data collection
and processing, for example with infrastructure and derived data.
In the above respects we have not lost
privacy but moved the bounds of what personally and socially we call privacy –
obviously people are not placing emphasis on the moral and ethical issues but
rather on the economic benefit of using such data consuming services. In
writing this blog I am losing my privacy, but with the economic gain of brand
building and knowledge sharing.
Using this data consumers and users can
be profiled and classified; typically for the serving of the perfect advertisement.
However this is not unlike what an "old style shopkeeper" did through
personally knowing his customers. The major difference is that today this is
done automatically and impersonally by computer. We lost the link with that
corner shop keeper who knew us and our families personally. Ever try contacting
the customer service departments of practically any company these days?
This also touches on the point that users
start or have started to feel that they are not in control of their data.
Most advertising and profiling companies are
using classification structures that are fairly coarse grained but then further
refined those with additional [coarse] grained data such as location and social
network. This for the most part is nothing more than could be understood by
reflecting on one's own life, place of abode and neighbourhood. For the most
part this is just reasserting what is already derivable from a person’s
postcode.
Much of the data collected by the NSA in
the current revelations is somewhat innocuous; primarily this seems to be just
telephone record meta-data like the kind you see on an itemized bill. But such
innocuous data can easily be cross-referenced and fingerprinted.
The trouble here is that government
authorities can have a more insidious effect upon a person's life than a
supermarket or credit card provider can. Indeed there are safe guards and
protections through the rule of law - though as we have seen these can be
constructed so that under some circumstances the law can allow whatever is
necessary to get a/the job done.
Before however we dismiss the above,
consider two points:
- automatic guilt, or, guilty until proven innocent
- scope creep
The first derives from the fact that all
your actions may be used against you in the future. If you think you have nothing to hide then consider all the crimes you committed
today? Did you drive over the speed limit, run a red light, have you ever stolen something/anything etc?
The second derives from the first that
once you have this information then it could be used for purposes well beyond
its original intent. Worse are the twin possibilities of false positives and
false negatives. Consider councils in the UK using CCTV cameras originally
intended to catch terrorists and prevent crime (in general) for catching dog owners not cleaning up after
their dogs.
From the above the moral and ethical
arguments are easily fashioned, the economic arguments are much more difficult
and vary depending upon the context and our view of what society should be:
- Is personal freedom, privacy and liberty greater than that of society's?
- Is mass surveillance better than letting one "terrorist" commit an act of atrocity?
These questions however go right to the
heart of the definitions of freedom, liberty, privacy, security, society and our
own control over our own data. I don't think any of us even remotely comprehend
the repercussions and difficulties of even trying to address, let alone answer
such questions.
But until we start having this debate in an impartial, focused and formal manner with the terms and definitions clearly stated, judging and/or condemning any form of data collection and any form of processing and usage of data is not going to be possible in any meaningful, lasting manner.
In another way we're back to a question
posted by a group of mathematicians regarding the esoteric nature of things as
we move away from the fundamental building blocks, and losing sight of what
those building blocks [of society and humanity] actually mean.
Whether the NSA and everyone else's collection of data is right or wrong I can't answer, but the debate about what privacy actually is and our relationship personally and as a society with the concepts of privacy, security and trust is going to be an extremely interesting debate with wide repercussions.
Subscribe to:
Posts (Atom)


