One of the main aspects of personal [information] privacy is that much of the topic is that other parties would not collect nor perform any analysis of your data. The trouble is that this argument is often made in isolation, in that it somewhat assumes that the acts we perform by computer exist in a place where we can hide. For example, what someone does behind closed doors usually remains private. But, if that act is made in a public place, say, in the middle of the street by default whatever is done is not private - even if we hoped no-one saw.
Anything and everything we do on the internet is in public by default. When we perform things in public, then other people may or will see, find out and perform their own analysis to form a profile of you.
Many privacy enhancing technologies are akin to standing in the middle of a busy street and shouting "don't look!". Even if everyone looks away, more often than not there is a whole raft of other evidence to show what you've been doing.
Admittedly most of the time nobody really cares nor are actually looking in the first place. Though as it has been found out recently (and this really isn't a surprise) that some such as the NSA and GCHQ are continually watching. Even the advertisers don't really care that much; their main interest is trying to categorise you to ship a generic advertisement - and advertisers are often really easy to game...
If we really do want privacy on the internet then rather than concentrating on how to be private (or pretending that we are), we need to concentrate on how to reduce the evidence trail that we leave. Such evidence is in the form of web logs, search queries, location traces from your navigator, tweets, Facebook postings etc.
Once we have understood what crumbs of evidence is being left, we can start exploring all the side avenues where data flows (leaks) and the points where data can be extracted surreptitiously. We can also examine what data we do want released, or have no choice about.
At this moment, I don't really see a good debate about this, at least not at a technical level though there are some great tools such as Ghostery that assist in this. Certainly there is little discussion at a fundamental level which would really help us define what privacy really is.
I personally tend to take the view at the moment that privacy might even be the wrong term, or at best, somewhat a misleading term.
On the internet every detail of what we do is potentially public and can be used for good as well as evil (whatever those terms actually mean), our job as privacy professionals is to make that journey as safe as possible, hence the use of the term "information safety" to better describe what we do.
Showing posts with label Rhetoric. Show all posts
Showing posts with label Rhetoric. Show all posts
Sunday, 24 November 2013
Friday, 19 July 2013
Systems Safety - Defining Moments
As I've been concentrating on "safety improvements", or at least techniques for the improvement of system I've tended to concentrate on four areas:
Above Diagram Key: Y-Axis: relative degree of safety embedded into that discipline, X-Axis, year of time since seminal incident.
Aviation safety's seminal moment was the 1935 crash of a Boeing Model 299 aircraft during a presentation flight. Instead of blaming the pilots, effort was made to understand the causes of the accident and develop techniques to help prevent similar accidents in the future.
For industrial safety the seminal moment was the 1974 Flixborough Disaster in the UK. This resulted in work on the design of industrial plants and the development of the notion of "inherent safety".
Surgical safety has quite a long tradition especially with the development of anaesthetic safety from the 1960s and the introduction of a proper systems approach. However anesthetists seem not to feature prominently as surgeons and doctors so the fame would probably go to Peter Pronovost et.al. for the Central Line checklist. This was probably one of the major contributors to the WHO Surgical Safety Checklist discussed in detail in Atul Gawande's book The Checklist Manifesto which brings together much of the above incidents.
If you're still in doubt maybe Atul Gawande's article in the New Yorker magazine entitled The Checklist: If something so simple can transform intensive care, what else can it do? (Dec 10, 2007) might help.
Getting back to the crux of this article, what is the incident that will cause the wholesale change in attitudes and techniques to software engineering that instills such a sense of discipline that we can eradicate errors to such a degree that we could compare ourselves favourably with other disciplines?
The increasingly frequent hacking and information leaks? The NSA wiretapping and mass surveillance? Facebook and Google's privacy policies? None of these have had any lasting effect upon the very core of software engineering if any at all. Which either means that we place such low value on the safety of our information or that the economics of software are so badly formulated in society that the catastrophe would have to be so huge that it would have to cause societal change?
Interestingly, in software engineering and computer science we're certainly not short on techniques for improving the quality and reliability of the systems we're developing: formal methods (eg: Alloy, B, Z, VDM etc), proof, simulation, testing, modelling (in general). What we probably lack is the simplicity of a checklist to guide us through the morass of problems we encounter. In this last respect, this is why I think we're more like surgeons that modern day aviators; or, maybe some of us are like the investigators to the 1935 Boeing crash and other aviation heroes learning their trade?
- Aviation
- Industrial
- Medical (specifically surgical)
- Software Engineering (specifically information privacy)
Above Diagram Key: Y-Axis: relative degree of safety embedded into that discipline, X-Axis, year of time since seminal incident.
Aviation safety's seminal moment was the 1935 crash of a Boeing Model 299 aircraft during a presentation flight. Instead of blaming the pilots, effort was made to understand the causes of the accident and develop techniques to help prevent similar accidents in the future.
For industrial safety the seminal moment was the 1974 Flixborough Disaster in the UK. This resulted in work on the design of industrial plants and the development of the notion of "inherent safety".
Surgical safety has quite a long tradition especially with the development of anaesthetic safety from the 1960s and the introduction of a proper systems approach. However anesthetists seem not to feature prominently as surgeons and doctors so the fame would probably go to Peter Pronovost et.al. for the Central Line checklist. This was probably one of the major contributors to the WHO Surgical Safety Checklist discussed in detail in Atul Gawande's book The Checklist Manifesto which brings together much of the above incidents.
If you're still in doubt maybe Atul Gawande's article in the New Yorker magazine entitled The Checklist: If something so simple can transform intensive care, what else can it do? (Dec 10, 2007) might help.
Getting back to the crux of this article, what is the incident that will cause the wholesale change in attitudes and techniques to software engineering that instills such a sense of discipline that we can eradicate errors to such a degree that we could compare ourselves favourably with other disciplines?
The increasingly frequent hacking and information leaks? The NSA wiretapping and mass surveillance? Facebook and Google's privacy policies? None of these have had any lasting effect upon the very core of software engineering if any at all. Which either means that we place such low value on the safety of our information or that the economics of software are so badly formulated in society that the catastrophe would have to be so huge that it would have to cause societal change?
Interestingly, in software engineering and computer science we're certainly not short on techniques for improving the quality and reliability of the systems we're developing: formal methods (eg: Alloy, B, Z, VDM etc), proof, simulation, testing, modelling (in general). What we probably lack is the simplicity of a checklist to guide us through the morass of problems we encounter. In this last respect, this is why I think we're more like surgeons that modern day aviators; or, maybe some of us are like the investigators to the 1935 Boeing crash and other aviation heroes learning their trade?
Tuesday, 11 June 2013
Privacy, Data Collection and Surveillance
The privacy debate about the collection
of data by the NSA continues with many asking questions about the moral and
ethical issues surrounding this. The phrase "the death of privacy" is
abound.
This is true I'm afraid, we lost our
privacy, but not when the NSA starting collecting data but when we starting communicating
using technologies that were readily and easily available - that probably dates
back to the birth of written communication.
Data collection concerns me certainly,
but here I want to focus on one of the maxims of privacy: "if you don'tuse it, don't collect it" and the fact that privacy is much more about the
usage of data, not its collection (viz. the above maxim).
One can argue that merely using Google,
Facebook and all the rest of the social media services one has already lost
one's privacy, but interaction with these services is voluntary - no-one forced
you to post those party pictures to the entire World and dog (complete with
EXIF and location information).
We
admittedly do have a problem with other more hidden aspects of data collection
and processing, for example with infrastructure and derived data.
In the above respects we have not lost
privacy but moved the bounds of what personally and socially we call privacy –
obviously people are not placing emphasis on the moral and ethical issues but
rather on the economic benefit of using such data consuming services. In
writing this blog I am losing my privacy, but with the economic gain of brand
building and knowledge sharing.
Using this data consumers and users can
be profiled and classified; typically for the serving of the perfect advertisement.
However this is not unlike what an "old style shopkeeper" did through
personally knowing his customers. The major difference is that today this is
done automatically and impersonally by computer. We lost the link with that
corner shop keeper who knew us and our families personally. Ever try contacting
the customer service departments of practically any company these days?
This also touches on the point that users
start or have started to feel that they are not in control of their data.
Most advertising and profiling companies are
using classification structures that are fairly coarse grained but then further
refined those with additional [coarse] grained data such as location and social
network. This for the most part is nothing more than could be understood by
reflecting on one's own life, place of abode and neighbourhood. For the most
part this is just reasserting what is already derivable from a person’s
postcode.
Much of the data collected by the NSA in
the current revelations is somewhat innocuous; primarily this seems to be just
telephone record meta-data like the kind you see on an itemized bill. But such
innocuous data can easily be cross-referenced and fingerprinted.
The trouble here is that government
authorities can have a more insidious effect upon a person's life than a
supermarket or credit card provider can. Indeed there are safe guards and
protections through the rule of law - though as we have seen these can be
constructed so that under some circumstances the law can allow whatever is
necessary to get a/the job done.
Before however we dismiss the above,
consider two points:
- automatic guilt, or, guilty until proven innocent
- scope creep
The first derives from the fact that all
your actions may be used against you in the future. If you think you have nothing to hide then consider all the crimes you committed
today? Did you drive over the speed limit, run a red light, have you ever stolen something/anything etc?
The second derives from the first that
once you have this information then it could be used for purposes well beyond
its original intent. Worse are the twin possibilities of false positives and
false negatives. Consider councils in the UK using CCTV cameras originally
intended to catch terrorists and prevent crime (in general) for catching dog owners not cleaning up after
their dogs.
From the above the moral and ethical
arguments are easily fashioned, the economic arguments are much more difficult
and vary depending upon the context and our view of what society should be:
- Is personal freedom, privacy and liberty greater than that of society's?
- Is mass surveillance better than letting one "terrorist" commit an act of atrocity?
These questions however go right to the
heart of the definitions of freedom, liberty, privacy, security, society and our
own control over our own data. I don't think any of us even remotely comprehend
the repercussions and difficulties of even trying to address, let alone answer
such questions.
But until we start having this debate in an impartial, focused and formal manner with the terms and definitions clearly stated, judging and/or condemning any form of data collection and any form of processing and usage of data is not going to be possible in any meaningful, lasting manner.
In another way we're back to a question
posted by a group of mathematicians regarding the esoteric nature of things as
we move away from the fundamental building blocks, and losing sight of what
those building blocks [of society and humanity] actually mean.
Whether the NSA and everyone else's collection of data is right or wrong I can't answer, but the debate about what privacy actually is and our relationship personally and as a society with the concepts of privacy, security and trust is going to be an extremely interesting debate with wide repercussions.
Tuesday, 5 February 2013
Deconstructing Privacy
Some very constructive comments after my previous posting on the naivety of privacy - thanks to all who participated. So to address this problem that we are often talking cross purposes and without any common frame of reference we need to first take a look at in what terms we're framing privacy [of information systems].
Typically we see that privacy is addressed or framed in seven broad areas:
Each of these areas most certainly overlap but we have the difficulty of switching between these frames. For example, it is often the case that if we have great system security, then privacy is of little concern because we've addressed the problem of data leakage; however we haven't addressed the problem of data content because this is largely irrelevant to security. Similarly if we have great access control we don't have to worry about the data getting into the wrong hands? Or possibly that if we've presented the user with the necessary consents then all is fine?
If we firstly deconstruct each area and examine how each views privacy, then attempt a cross-referencing exercise between these, then we might actually have a basis for constructing, at least a framework for a common terminology and semantics.
Typically we see that privacy is addressed or framed in seven broad areas:
Each of these areas most certainly overlap but we have the difficulty of switching between these frames. For example, it is often the case that if we have great system security, then privacy is of little concern because we've addressed the problem of data leakage; however we haven't addressed the problem of data content because this is largely irrelevant to security. Similarly if we have great access control we don't have to worry about the data getting into the wrong hands? Or possibly that if we've presented the user with the necessary consents then all is fine?
If we firstly deconstruct each area and examine how each views privacy, then attempt a cross-referencing exercise between these, then we might actually have a basis for constructing, at least a framework for a common terminology and semantics.
Friday, 25 January 2013
On The Naivety of Privacy
Recent events regarding privacy and the internet have left me wondering if we are being somewhat naïve. We are starting to see a slew of new laws, strategies and technologies for protecting our privacy in what is effectively a public space. The end-user however is not, as far as I can tell, really getting the benefit of this - indeed if anyone is it is the emerging privacy-industrial complex [1] as some have written.
It is utterly naïve to believe that laws, strategies, intentions, grand speeches, certifications, automatic filtering, classification iconography and so on make for better end-user privacy. The more we do this the more confused we become, and simultaneously we lose sight of what we're really trying to achieve. Spare a thought for the poor end-users.
There is a great deal that is misunderstood or not known by privacy advocates about how the internet, computers and information systems work - I fear in a lot of cases either some don't want to understand because it takes them outside of their comfort zone, or the semantic gap between the engineers and the legal/advocacy side is too great and that bridging this gap is extraordinarily difficult for both parties.
I worry about our lack of formality and discipline, possibly in equal quantities. We – the privacy community – lack these aspects to really understand and accept the fundamentals of our area and how to apply this to the information systems we are trying to protect. In some cases we are actively fighting against the need to scientifically ground our chosen area.
We must take a moment to think and understand what problem we are really trying to solve. The more philosophical works by Solove and Nissenbaum address the overall concept of privacy. I'm not sure that the implications of these are really understood. Part of the problem is that general theories of information [2] are very abstract and obtuse when compared with the legal views of the above authors, and we've done very little to tie these areas together to produce the necessary scientific formalisation of privacy we need.
As an example, the Privacy by Design (PbD) manifesto is being waived by many to be the commandments of privacy and following these magically solves everything. This only leads to “technical debt” and greater problems in the future. Often we find the engineers, R&D teams and the scientists excluded from, and outside of, this discussion.
I think we're missing the point what privacy really is and certainly we have little idea at this time how to effectively build information systems with inherent privacy [3] as a property of those systems. I have one initial conclusion:
We have no common definitions, common language, common semantics nor mappings between our individual worlds: legal, advocacy and engineering. Worse, in each of these worlds terminology and semantics are not always so well internally defined.
When an [software] engineer says “data collection”, "log" or "architecture", these do not mean the same to a lawyer or a consumer advocate. Indeed I don't think these terms semantically map even remotely cleanly – if at all - between these groups.
A set of PowerPoint slides with a strategy, a vision, a manifesto, good intentions, project plan or a classifications scheme mean very little and without some form of semantics are wasted, token efforts that only add to the complexity and confusion of a rapidly changing field.
We desperately need to address the problem that we must create a way of communicating amongst ourselves through which all of the internal factions within the privacy community can effectively understand each other's point of view. Only then might we even have a chance of realistically and effectively addressing the moving target of privacy issues facing end- users and businesses that rely so much on the interchange and analysis of information.
The problem with formally (or rigorously) defining anything is that it has the nasty tendency to expose holes and weaknesses in our thinking. Said holes and weaknesses are not entirely appreciated, especially when it challenges an established school of thought or a political or dogmatic balance [4].
The privacy community is constantly developing new laws and legal arguments, new sets of guidelines, manifestos and doom scenarios while the engineers are trying to address these often inconsistent and complex ideas through technical means. From the engineering perspective not only we are internally exposing flaws in database design, information system architecture and user experience but also the mismatch between engineering, legal, the world of the consumer advocate and ultimately a company's information strategy.
An information strategy needs to address everything from how the engineers develop software to how you want your company to be perceived by the consumer. How many information strategies actually address the role that information plays in a modern, global consumer ecosystem where the central concept is the collection and processing of user information? Of those, how many address the engineering and scientific levels of information?
We must take a serious retrenchment [5] step and look back at what we have created. Then we need ruthlessly carve away anything that does not either solve the communication issue within the privacy community or does not immediately serve the end-user. Reemphasizing the latter point, this explicitly means the end-user values, not what we as a privacy community might perceive to be valued by the end-user.
We must fully appreciate the close link between privacy and information, and that privacy is one of the most crosscutting of disciplines. Privacy is going to expose every single flaw in the way we collect, manage, process and use information from the user experience, as well as the application and services eco-system, and even the manner in which we conduct our system and software engineering processes and information governance. The need to get privacy right is critical not just for the existence of privacy as a technical discipline in its own right (alongside security, architecture, etc) but also for the consumer and the business.
The emphasis must be placed on the deep technical knowledge of experts in information systems – these must be the drivers and unifiers between the engineers, the lawyers, the advocates and ultimately the users. Without this deep, holistic, scientific and mathematical foundation we will not be able to sufficiently nor consistency address or govern any issues that arise in the construction of our information systems at any level of abstraction.
If the work we do in privacy does not have a scientific, consistent and formal underpinning [6] that brings together the engineers, lawyers and advocates then privacy is waste of time at best and deeply destructive to the information systems at worst.
Without this we are a disjointed community caring for ourselves and not the business or consumer and privacy becomes just a bureaucratic exercise to fulfill the notions of performing a process and metrics rendered as meaningful as random numbers.
PostScript:
Via Twitter I came across a talk given by Jean Bezivin entitled "Should we Resurrect Software Engineering?" presented at the Choose Forum in December 2012. Many of the things he presented are analogous to what is happening in privacy. He makes the point a number of times that we have never addressed the missing underlying theory of software engineering and how to really unify the various communities, fields and techniques in this area. Two points I particularly liked was that we concentrated on the solution (MDE) but never thought about the problem; the other point is the use Albert Camus' quote
Notes
[1] #pii2012: The Emergent Privacy-Industrial Complex
[2] Jerry Seligmann, Jon Barwise (1997) Information Flow. Cambridge University Press.
[3] I like the idea of privacy being an inherent construct in system design in much the same way that inherent safety emerged from chemical/industrial plant design
[4] A blog article discussing “mathematical catastrophes” – two that come to mind are Russel and Frege and also Russel and Gödel. Both related but the latter’s challenge to the mathematical school of thought was dramatic to say the least.
[5] A formal retrenchment step in that we not just start again but actively record what we’re backtracking on. Poppleton et.al. constructed a theory of retrenchment for software design using formal methods; the same principles apply here.
[6] If you’re still in doubt just remember that whatever decisions are made with respect to privacy, there’s a programmer writing formal, mathematical statements encoding this. Lessig’s Code is Law principle.
It is utterly naïve to believe that laws, strategies, intentions, grand speeches, certifications, automatic filtering, classification iconography and so on make for better end-user privacy. The more we do this the more confused we become, and simultaneously we lose sight of what we're really trying to achieve. Spare a thought for the poor end-users.
There is a great deal that is misunderstood or not known by privacy advocates about how the internet, computers and information systems work - I fear in a lot of cases either some don't want to understand because it takes them outside of their comfort zone, or the semantic gap between the engineers and the legal/advocacy side is too great and that bridging this gap is extraordinarily difficult for both parties.
I worry about our lack of formality and discipline, possibly in equal quantities. We – the privacy community – lack these aspects to really understand and accept the fundamentals of our area and how to apply this to the information systems we are trying to protect. In some cases we are actively fighting against the need to scientifically ground our chosen area.
We must take a moment to think and understand what problem we are really trying to solve. The more philosophical works by Solove and Nissenbaum address the overall concept of privacy. I'm not sure that the implications of these are really understood. Part of the problem is that general theories of information [2] are very abstract and obtuse when compared with the legal views of the above authors, and we've done very little to tie these areas together to produce the necessary scientific formalisation of privacy we need.
As an example, the Privacy by Design (PbD) manifesto is being waived by many to be the commandments of privacy and following these magically solves everything. This only leads to “technical debt” and greater problems in the future. Often we find the engineers, R&D teams and the scientists excluded from, and outside of, this discussion.
I think we're missing the point what privacy really is and certainly we have little idea at this time how to effectively build information systems with inherent privacy [3] as a property of those systems. I have one initial conclusion:
WE HAVE NO UNDERLYING THEORY OF PRIVACY
We have no common definitions, common language, common semantics nor mappings between our individual worlds: legal, advocacy and engineering. Worse, in each of these worlds terminology and semantics are not always so well internally defined.
When an [software] engineer says “data collection”, "log" or "architecture", these do not mean the same to a lawyer or a consumer advocate. Indeed I don't think these terms semantically map even remotely cleanly – if at all - between these groups.
A set of PowerPoint slides with a strategy, a vision, a manifesto, good intentions, project plan or a classifications scheme mean very little and without some form of semantics are wasted, token efforts that only add to the complexity and confusion of a rapidly changing field.
We desperately need to address the problem that we must create a way of communicating amongst ourselves through which all of the internal factions within the privacy community can effectively understand each other's point of view. Only then might we even have a chance of realistically and effectively addressing the moving target of privacy issues facing end- users and businesses that rely so much on the interchange and analysis of information.
The problem with formally (or rigorously) defining anything is that it has the nasty tendency to expose holes and weaknesses in our thinking. Said holes and weaknesses are not entirely appreciated, especially when it challenges an established school of thought or a political or dogmatic balance [4].
The privacy community is constantly developing new laws and legal arguments, new sets of guidelines, manifestos and doom scenarios while the engineers are trying to address these often inconsistent and complex ideas through technical means. From the engineering perspective not only we are internally exposing flaws in database design, information system architecture and user experience but also the mismatch between engineering, legal, the world of the consumer advocate and ultimately a company's information strategy.
An information strategy needs to address everything from how the engineers develop software to how you want your company to be perceived by the consumer. How many information strategies actually address the role that information plays in a modern, global consumer ecosystem where the central concept is the collection and processing of user information? Of those, how many address the engineering and scientific levels of information?
We must take a serious retrenchment [5] step and look back at what we have created. Then we need ruthlessly carve away anything that does not either solve the communication issue within the privacy community or does not immediately serve the end-user. Reemphasizing the latter point, this explicitly means the end-user values, not what we as a privacy community might perceive to be valued by the end-user.
We must fully appreciate the close link between privacy and information, and that privacy is one of the most crosscutting of disciplines. Privacy is going to expose every single flaw in the way we collect, manage, process and use information from the user experience, as well as the application and services eco-system, and even the manner in which we conduct our system and software engineering processes and information governance. The need to get privacy right is critical not just for the existence of privacy as a technical discipline in its own right (alongside security, architecture, etc) but also for the consumer and the business.
The emphasis must be placed on the deep technical knowledge of experts in information systems – these must be the drivers and unifiers between the engineers, the lawyers, the advocates and ultimately the users. Without this deep, holistic, scientific and mathematical foundation we will not be able to sufficiently nor consistency address or govern any issues that arise in the construction of our information systems at any level of abstraction.
If the work we do in privacy does not have a scientific, consistent and formal underpinning [6] that brings together the engineers, lawyers and advocates then privacy is waste of time at best and deeply destructive to the information systems at worst.
Without this we are a disjointed community caring for ourselves and not the business or consumer and privacy becomes just a bureaucratic exercise to fulfill the notions of performing a process and metrics rendered as meaningful as random numbers.
* * *
PostScript:
Via Twitter I came across a talk given by Jean Bezivin entitled "Should we Resurrect Software Engineering?" presented at the Choose Forum in December 2012. Many of the things he presented are analogous to what is happening in privacy. He makes the point a number of times that we have never addressed the missing underlying theory of software engineering and how to really unify the various communities, fields and techniques in this area. Two points I particularly liked was that we concentrated on the solution (MDE) but never thought about the problem; the other point is the use Albert Camus' quote
<< Mal nommer les choses, c'est ajouterau malheur du monde >>
[To misname things is to add misery to the world]A subtle hint to getting the fundamentals right: terminology and semantics!
Notes
[1] #pii2012: The Emergent Privacy-Industrial Complex
[2] Jerry Seligmann, Jon Barwise (1997) Information Flow. Cambridge University Press.
[3] I like the idea of privacy being an inherent construct in system design in much the same way that inherent safety emerged from chemical/industrial plant design
[4] A blog article discussing “mathematical catastrophes” – two that come to mind are Russel and Frege and also Russel and Gödel. Both related but the latter’s challenge to the mathematical school of thought was dramatic to say the least.
[5] A formal retrenchment step in that we not just start again but actively record what we’re backtracking on. Poppleton et.al. constructed a theory of retrenchment for software design using formal methods; the same principles apply here.
[6] If you’re still in doubt just remember that whatever decisions are made with respect to privacy, there’s a programmer writing formal, mathematical statements encoding this. Lessig’s Code is Law principle.
Thursday, 22 November 2012
Information Privacy: Art or Science?
I was handed a powerpoint deck today containing notes for a training course on privacy. One thing that struck me was the statement on one of the slides, in fact it was the only statement on that slide:
This troubles me greatly and the interpretation of this probably goes a long way into explaining some things about the way information privacy is perceived and implemented.
What do we mean by art, and does this mean that privacy is not a science?
Hypothesis 1: Privacy is an art
If you've ever read great code it is artistic in nature. You can appreciate the amount of understanding and knowledge that has gone into writing that code. Not just at the act of writing, or the layout and indentation, but in the design of the algorithms, the separation of concerns, the holistic bigger picture of the architecture. Great code requires less debugging, performs well, stays in scope, and if it ever does require modification, it is easy to do. Great programmers are scientists - they understand the value to the code, they avoid technical debt, they understand the theory (maybe only implicitly) and the science and discipline behind their work and in that respect they are the true artists of their trade.
For example, Microsoft spent a lot of effort in improving the quality of its code with efforts such as those the still excellent book Code Complete by Steve McConnell. This book taught programmers great techniques to improve the quality of their code. McConnell obviously knew what works and what didn't from a highly technical perspective based on a sound, scientific understanding of how code works, how code is written, how design is made and so on.
I don't think information privacy is an art in the above sense.
Hypothesis 2: Privacy is an "art".
In the sense that you're doing privacy well in much the same was as a visitor to an art gallery knows "great art". Everyone has their own interpretation and religious wars spring forth over whether something is art or not.
Indeed here is the problem, and in this respect I do agree that privacy is art. Art can be anything from the formal underpinnings of ballet to the drunken swagger of a Friday night reveler - who is to say that the latter is not art? Compare ballet with forms of modern and contemporary dance: ballet is almost universally considered "art" while some forms of contemporary dance is not - see our drunken reveler at the local disco...this is dance, but is it art?
Indeed sometimes the way we practice privacy is very much like the drunken reveler but telling everyone at the same time that "this is art!"
What elevates ballet, or the great coder, to become art is that they both have formal, scientific underpinnings. Indeed I believe that great software engineering and ballet have many similarities and here we can also see the difference between a professional dancer and a drunken reveler on the dance floor: one has formal training in the principles and science of movement, one does not.
Indeed if we look at the sister to privacy: security, we can be very sure that we do not want to practice security of our information systems in an unstructured, informal, unscientific manner. We want purveyors of the art - artists - of security to look after our systems: those that know and intuitively feel what security is.
There are many efforts to better underpin information privacy, rarely do these come through in the software engineering process in any meaningful manner unless explicitly required or audited for. Even then we are far from a formal, methodical process by which privacy becomes an inherent property of the systems we are building. When we achieve this as a matter of the daily course of our work then, and only then, privacy will become an art practiced by artists.
PRIVACY IS AN ART
This troubles me greatly and the interpretation of this probably goes a long way into explaining some things about the way information privacy is perceived and implemented.
What do we mean by art, and does this mean that privacy is not a science?
Hypothesis 1: Privacy is an art
If you've ever read great code it is artistic in nature. You can appreciate the amount of understanding and knowledge that has gone into writing that code. Not just at the act of writing, or the layout and indentation, but in the design of the algorithms, the separation of concerns, the holistic bigger picture of the architecture. Great code requires less debugging, performs well, stays in scope, and if it ever does require modification, it is easy to do. Great programmers are scientists - they understand the value to the code, they avoid technical debt, they understand the theory (maybe only implicitly) and the science and discipline behind their work and in that respect they are the true artists of their trade.
For example, Microsoft spent a lot of effort in improving the quality of its code with efforts such as those the still excellent book Code Complete by Steve McConnell. This book taught programmers great techniques to improve the quality of their code. McConnell obviously knew what works and what didn't from a highly technical perspective based on a sound, scientific understanding of how code works, how code is written, how design is made and so on.
I don't think information privacy is an art in the above sense.
Hypothesis 2: Privacy is an "art".
In the sense that you're doing privacy well in much the same was as a visitor to an art gallery knows "great art". Everyone has their own interpretation and religious wars spring forth over whether something is art or not.
Indeed here is the problem, and in this respect I do agree that privacy is art. Art can be anything from the formal underpinnings of ballet to the drunken swagger of a Friday night reveler - who is to say that the latter is not art? Compare ballet with forms of modern and contemporary dance: ballet is almost universally considered "art" while some forms of contemporary dance is not - see our drunken reveler at the local disco...this is dance, but is it art?
Indeed sometimes the way we practice privacy is very much like the drunken reveler but telling everyone at the same time that "this is art!"
What elevates ballet, or the great coder, to become art is that they both have formal, scientific underpinnings. Indeed I believe that great software engineering and ballet have many similarities and here we can also see the difference between a professional dancer and a drunken reveler on the dance floor: one has formal training in the principles and science of movement, one does not.
Indeed if we look at the sister to privacy: security, we can be very sure that we do not want to practice security of our information systems in an unstructured, informal, unscientific manner. We want purveyors of the art - artists - of security to look after our systems: those that know and intuitively feel what security is.
There are many efforts to better underpin information privacy, rarely do these come through in the software engineering process in any meaningful manner unless explicitly required or audited for. Even then we are far from a formal, methodical process by which privacy becomes an inherent property of the systems we are building. When we achieve this as a matter of the daily course of our work then, and only then, privacy will become an art practiced by artists.
Tuesday, 13 November 2012
Measuring Privacy against Effort to Break Security
As part of my job I've needed to look at metrics and measurement of privacy. Typically I've focussed on information entropy versus, say, number of records (define "record") or other measurements such as amount of data which do not take into consideration the amount of information, that is, the content of the data being revealed.
So this lead to an interesting discussion* with some of my colleagues where we looked at a graph like this.
The y-axis is a measure of information content (ostensibly information entropy wrt to some model) and the x-axis a measure of the amount of force required to obtain that information. For any given hacking technique we can deliniate a region on the x-axis which corresponds to the amount of sophistication or effort placed into that attack. The use of the terms, effort and force here come from the physics and I think we even have some ideas on how the dimensions of these map to the security world, or actually what these dimensions might be.
So for a given attack 'x', for example an SQL inject attack against some system to reveal some information 'M', we require a certain amount of effort just for the attack to reveal something. If we make a very sophisticated attack then we potentially reveal more. This is expressed as the width of the red bar in the above graph.
One conclusion here is that security people try to push the attack further to the right and even widen it, while privacy people try to lower and flatten the curve, especially through the attack segment.
Now it can be argued that even with a simple attack, over time the amount of information increases, which brings us to a second graph which takes this into consideration:
Ignoring the bad powerpoint+visio 3D rending, we've just added a time scale (z-axis, future towards back), we can now capture or at least visualise the statement above that even an unsophisticated attack over time can reveal a lot of information. Then there's a trade-off between a quick sophisticated attack versus a long, unsophisticated attempt.
Of course a lot of this depends upon having good metrics and good measurement in the first place and that we do have real difficulties with, though there is some pretty interesting literature [1,2] on the subject and in the case of privacy some very interesting calculations that can be performed over the data such as k-anonymity an l-diversity.
I have a suspicion that we should start looking at privacy and security metrics from the dimensional analysis point of view and somewhat reverse engineer what the actual units and thus measurements are going to be. Something to consider here is that the amount of effort or force of an attack is not necessarily related to the amount of computing power, for example, brute forcing an attack on a hash function is not as forcible as a well planned hoax email and a little social engineering.
If anyone has ideas on this please let me know.
References
[1] Michele Bezzi (2010) An information theoretic approach for privacy metrics. Transactions on Data Privacy 3, pp:199-215
[2] Reijo M. Savola (2010) Towards a Risk-Drive Methodology for Priavcy Metrics Development. IEEE International conference on Social Computing/IEEE International Conference on Privacy, Security, Risk and Trust.
*for "discussion" read 'animated and heated arguments, a fury of writing on whiteboards, excursions to dig out academic papers, mathematics, coffee etc' - all great stuff :-)
So this lead to an interesting discussion* with some of my colleagues where we looked at a graph like this.
The y-axis is a measure of information content (ostensibly information entropy wrt to some model) and the x-axis a measure of the amount of force required to obtain that information. For any given hacking technique we can deliniate a region on the x-axis which corresponds to the amount of sophistication or effort placed into that attack. The use of the terms, effort and force here come from the physics and I think we even have some ideas on how the dimensions of these map to the security world, or actually what these dimensions might be.
So for a given attack 'x', for example an SQL inject attack against some system to reveal some information 'M', we require a certain amount of effort just for the attack to reveal something. If we make a very sophisticated attack then we potentially reveal more. This is expressed as the width of the red bar in the above graph.
One conclusion here is that security people try to push the attack further to the right and even widen it, while privacy people try to lower and flatten the curve, especially through the attack segment.
Now it can be argued that even with a simple attack, over time the amount of information increases, which brings us to a second graph which takes this into consideration:
Ignoring the bad powerpoint+visio 3D rending, we've just added a time scale (z-axis, future towards back), we can now capture or at least visualise the statement above that even an unsophisticated attack over time can reveal a lot of information. Then there's a trade-off between a quick sophisticated attack versus a long, unsophisticated attempt.
Of course a lot of this depends upon having good metrics and good measurement in the first place and that we do have real difficulties with, though there is some pretty interesting literature [1,2] on the subject and in the case of privacy some very interesting calculations that can be performed over the data such as k-anonymity an l-diversity.
I have a suspicion that we should start looking at privacy and security metrics from the dimensional analysis point of view and somewhat reverse engineer what the actual units and thus measurements are going to be. Something to consider here is that the amount of effort or force of an attack is not necessarily related to the amount of computing power, for example, brute forcing an attack on a hash function is not as forcible as a well planned hoax email and a little social engineering.
If anyone has ideas on this please let me know.
References
[1] Michele Bezzi (2010) An information theoretic approach for privacy metrics. Transactions on Data Privacy 3, pp:199-215
[2] Reijo M. Savola (2010) Towards a Risk-Drive Methodology for Priavcy Metrics Development. IEEE International conference on Social Computing/IEEE International Conference on Privacy, Security, Risk and Trust.
*for "discussion" read 'animated and heated arguments, a fury of writing on whiteboards, excursions to dig out academic papers, mathematics, coffee etc' - all great stuff :-)
Thursday, 18 October 2012
Two things: Mars and Mathematics
An interlude to my hiatus of posting...trying to write a paper based on the earlier DNT article...
There are couple very nice things I want to post here, mainly for future reference and because, well just because :-)
The first is a panorama of Mars taken by Curiosity found via the Bad Astronomy blog (regular and compulsory reading): what would it look like if you could stand (where Curiosity is) on Mars...
As explained by Phil Plait, these pictures were stitched together by Denny Bauer from a series of pictures from Curiosity's MastCam - amazing work-
This I rate in the same category as the Titan surface picture taken by Huygens, and talking of Huygens I found (via BadAstronomy) a link to a posting about the surface of Titan being somewhat like wet sand which then led to an article about Huygen's landing and then to an ESA page which details the landing with a video reconstruction. I know that Curiosity's landing was pretty spectacular and we have videos, but Huygens did it much, much further from home after a longer journey onto a moon that was a complete mystery - isn't science amazing!
Then there is a posting on n-Category Cafe about set theory and order theory and the dependence of one on the other: The Curious Dependence of Set Theory on Order Theory (Tom Leinster). The posting the question:
and offers two answers: yes and no ....
Rather than go into the mathematics, the discussion of the point is fascinating from two angles: firstly, that isn't it amazing how results in one area of mathematics appear to be inextricably linked with results in seemingly unrelated areas. I guess elliptic curves and modular forms via the Taniyama-Shimura conjecture is a good example. Secondly the discussions open up quite a debate on the philosophy of mathematics and ends up discussing computer programming, data structures and the Axiom of Choice.
Then there's Einstein's letter on religion which is being auctioned on eBay in which he detailed his views and more importantly his understanding of religion; this is a quote from a 1930's essay by Einstein:
:-)
There are couple very nice things I want to post here, mainly for future reference and because, well just because :-)
The first is a panorama of Mars taken by Curiosity found via the Bad Astronomy blog (regular and compulsory reading): what would it look like if you could stand (where Curiosity is) on Mars...
As explained by Phil Plait, these pictures were stitched together by Denny Bauer from a series of pictures from Curiosity's MastCam - amazing work-
This I rate in the same category as the Titan surface picture taken by Huygens, and talking of Huygens I found (via BadAstronomy) a link to a posting about the surface of Titan being somewhat like wet sand which then led to an article about Huygen's landing and then to an ESA page which details the landing with a video reconstruction. I know that Curiosity's landing was pretty spectacular and we have videos, but Huygens did it much, much further from home after a longer journey onto a moon that was a complete mystery - isn't science amazing!
Then there is a posting on n-Category Cafe about set theory and order theory and the dependence of one on the other: The Curious Dependence of Set Theory on Order Theory (Tom Leinster). The posting the question:
Is it strange that results about sets should depend on results about order?
and offers two answers: yes and no ....
Rather than go into the mathematics, the discussion of the point is fascinating from two angles: firstly, that isn't it amazing how results in one area of mathematics appear to be inextricably linked with results in seemingly unrelated areas. I guess elliptic curves and modular forms via the Taniyama-Shimura conjecture is a good example. Secondly the discussions open up quite a debate on the philosophy of mathematics and ends up discussing computer programming, data structures and the Axiom of Choice.
Then there's Einstein's letter on religion which is being auctioned on eBay in which he detailed his views and more importantly his understanding of religion; this is a quote from a 1930's essay by Einstein:
"To sense that behind anything that can be experienced there is something that our minds cannot grasp, whose beauty and sublimity reaches us only indirectly: this is religiousness. In this sense, and in this sense only, I am a devoutly religious man."Quite profound and I'll finish with an xkcd cartoon:
:-)
Sunday, 30 September 2012
Grothendieck Biography
Mathematics is full of "characters", Grigori Perelman, Pierre de Fermat, Évariste Galois, Paul Erdős, Andrew Wiles** to name just a few and each having their own, unique, wondrous story about their dedication to their mathematical work and life.
Perhaps none more so than Alexander Grothendieck exemplifies the mathematician and since 1991 lived as a recluse in Andorra. However since the body of work and his contribution to mathematics, particularly, category theory and topology has been almost legendary he retains a great mystery about him.
In order to understand Groethendieck and possibly the mind of the mathematician a series of biographies of Groethendieck are being written by Leila Schneps. The current draft and extracts can be found on her pages about this work.
I'll quote a paragraph from Chapter 1 of Volume II, that gives a flavour of Groethendieck's work and approach to mathematics:
This is truly a work at a scale a magnitude more detailed than, say Simon Singh's fascinating documentation about Wiles' work and proof of Fermat's Last Theorem. Suffice to say, I look forward to reading it. Maybe Simon Singh should make a documentary about Groethendieck?
** I credit Andrew Wiles with inspiring me to study for my PhD back in 1995
Perhaps none more so than Alexander Grothendieck exemplifies the mathematician and since 1991 lived as a recluse in Andorra. However since the body of work and his contribution to mathematics, particularly, category theory and topology has been almost legendary he retains a great mystery about him.
In order to understand Groethendieck and possibly the mind of the mathematician a series of biographies of Groethendieck are being written by Leila Schneps. The current draft and extracts can be found on her pages about this work.
I'll quote a paragraph from Chapter 1 of Volume II, that gives a flavour of Groethendieck's work and approach to mathematics:
Taken altogether, Grothendieck’s body of work is perceived as an immense tour de force, an accomplishment of gigantic scope, and also extremely difficult both as research and for the reader, due to the effort necessary to come to a familiar understanding of the highly abstract objects or points of view that he systematically adopts as generalizations of the classical ones. All agree that the thousands of pages of his writings and those of his school, and the dozens and hundreds of new results and new proofs of old results stand as a testimony to the formidable nature of the task.
This is truly a work at a scale a magnitude more detailed than, say Simon Singh's fascinating documentation about Wiles' work and proof of Fermat's Last Theorem. Suffice to say, I look forward to reading it. Maybe Simon Singh should make a documentary about Groethendieck?
** I credit Andrew Wiles with inspiring me to study for my PhD back in 1995
Sunday, 27 May 2012
Writing
A friend of mine - David Cord - has written on his blog about writing, or more importantly, how to write and keep writing. Great advice from a famous author. I'll highlight a few paragraphs of his wisdom on the matter:
I was thinking about this very thing recently and deciding that the work office is a bad place to write, open offices even more so (actually does any work get done in an open office environment?)...the best place? Somewhere with some noise with comfort and coffee/tea or chai latte (yep, Starbucks), seems like I'm just like Ernest Hemingway in this respect, if not the writing - just need to find that perfect cafe.
David doesn't talk about tools - meaning software tools - for writing (and beating writer's block). The problem with things such as Microsoft Word is that they distract you with things such as formatting the text. I find simple tools without complications such as spelling and grammar checkers such as vi (and for some emacs). Using a simple text editor you can just focus on the text itself, though with the problem that you do need to leave the editing environment to deal with graphics; though here I tend just to leave a note in the text and draw the diagram or mathematics using pen (ink pen) in my engineering notebook.
I write during the day, during working hours, and the best stuff always comes in the morning. I write almost every day, even weekends and holidays. Before I start,
I have to prepare myself. I check my email, check the stock market, read the news, and get everything out of the way that I know could draw my attention before I start.
Next, I remove all possible distractions. The internet gets turned off, and the phone gets put on silent and placed in another room. It takes a long time for me to get properly focused, but one little distraction immediately throws me off track. One text message could destroy fifteen or thirty minutes of writing time. Sometimes, if I feel weak, I will even close the curtains so I’m not tempted to look outside.
I was thinking about this very thing recently and deciding that the work office is a bad place to write, open offices even more so (actually does any work get done in an open office environment?)...the best place? Somewhere with some noise with comfort and coffee/tea or chai latte (yep, Starbucks), seems like I'm just like Ernest Hemingway in this respect, if not the writing - just need to find that perfect cafe.
David doesn't talk about tools - meaning software tools - for writing (and beating writer's block). The problem with things such as Microsoft Word is that they distract you with things such as formatting the text. I find simple tools without complications such as spelling and grammar checkers such as vi (and for some emacs). Using a simple text editor you can just focus on the text itself, though with the problem that you do need to leave the editing environment to deal with graphics; though here I tend just to leave a note in the text and draw the diagram or mathematics using pen (ink pen) in my engineering notebook.
Tuesday, 22 May 2012
Fountain pens
Good to see a move back to real writing with real instrument of writing. There really is nothing better than putting pen to paper, and a real ink pen at that.
Used a fountain pen exclusively in my engineering notebook for years - nothing better focuses the mind as the permanence of ink in such a volume.
Why are fountain pen sales rising?
22 May 2012 Last updated at 17:08 GMTBy Steven BrocklehurstYou might expect that email and the ballpoint pen had killed the fountain pen. But sales are rising, so is the fountain pen a curious example of an old-fashioned object surviving the winds of change?...But for others, a fat Montblanc or a silver-plated Parker is a treasured item. Prominently displayed, they are associated with long, sinuous lines of cursive script.Sales figures are on the up. Parker, which has manufactured fountain pens since 1888, claims a worldwide "resurgence" in the past five years, and rival Lamy says turnover increased by more than 5% in 2011....continued...
Used a fountain pen exclusively in my engineering notebook for years - nothing better focuses the mind as the permanence of ink in such a volume.
Tuesday, 16 August 2011
The Elusive Big Idea
An excellent article from the New York Times on "The Elusive Big Idea"; the thesis that in today's society we just don't care about ideas anymore...
The Elusive Big Idea
By NEAL GABLER
Published: August 13, 2011
Scary thought indeed that "thinking is no longer done"; the implications of this will extend everywhere and affect every facet of our existence.
A further quote from the article:
A direct reference to Twitter and social media overall - the nature of communication is changing such that if you can't get your idea in 140 characters no-one is going to read it - assuming that anyone reads and thinks about what is actually contained in those 140 characters.
Is this some kind of existentialist crisis for ideas? And finally as a conclusion:
As Levitt and Dubner state in the Freakonomics series of books ( see: Superfreakonomics ) humans need incentive, and if ideas don't provide any incentive in either their thinking, development or execution to our narcissism then those ideas have no value. The problem here is rather self-referential in that narcissists aren't interested in anyone else's ideas - be honest who reads anyone's Twitter feeds or Facebook comments and partakes in deep rhetoric?
The Elusive Big Idea
By NEAL GABLER
Published: August 13, 2011
...
If our ideas seem smaller nowadays, it’s not because we are dumber than our forebears but because we just don’t care as much about ideas as they did. In effect, we are living in an increasingly post-idea world — a world in which big, thought-provoking ideas that can’t instantly be monetized are of so little intrinsic value that fewer people are generating them and fewer outlets are disseminating them, the Internet notwithstanding. Bold ideas are almost passé.
...
Post-Enlightenment refers to a style of thinking that no longer deploys the techniques of rational thought. Post-idea refers to thinking that is no longer done, regardless of the style.
The post-idea world has been a long time coming, and many factors have contributed to it. There is the retreat in universities from the real world, and an encouragement of and reward for the narrowest specialization rather than for daring — for tending potted plants rather than planting forests.
...
Scary thought indeed that "thinking is no longer done"; the implications of this will extend everywhere and affect every facet of our existence.
A further quote from the article:
...we get instant 140-character tweets about eating a sandwich or watching a TV show. While social networking may enlarge one’s circle and even introduce one to strangers, this is not the same thing as enlarging one’s intellectual universe.
A direct reference to Twitter and social media overall - the nature of communication is changing such that if you can't get your idea in 140 characters no-one is going to read it - assuming that anyone reads and thinks about what is actually contained in those 140 characters.
Is this some kind of existentialist crisis for ideas? And finally as a conclusion:
We have become information narcissists, so uninterested in anything outside ourselves and our friendship circles or in any tidbit we cannot share with those friends that if a Marx or a Nietzsche were suddenly to appear, blasting his ideas, no one would pay the slightest attention, certainly not the general media, which have learned to service our narcissism.
As Levitt and Dubner state in the Freakonomics series of books ( see: Superfreakonomics ) humans need incentive, and if ideas don't provide any incentive in either their thinking, development or execution to our narcissism then those ideas have no value. The problem here is rather self-referential in that narcissists aren't interested in anyone else's ideas - be honest who reads anyone's Twitter feeds or Facebook comments and partakes in deep rhetoric?
Subscribe to:
Posts (Atom)



