Showing posts with label Linguistics. Show all posts
Showing posts with label Linguistics. Show all posts

Sunday, 10 April 2016

Information wants to be free...?

I both like and dislike the quote "information wants to be free", mainly because it opens up a very nice philosophical discussion on what 'free' means but also because - and this is part I hate - it is some damned meaningless without any grounding in any form of semantics; and we've seen this before!

For the first part, this statement treats information in an anthropomorphic manner. Is it really information itself that has the need to be free? Let's assume that it does, though in a very fairy tale like way, it seems to me.

So let's then look at the word 'free', which I assume does not mean 'free' as in 'without cost' in the sense that someone has to pay for it. Though this is a curious idea in that information is somehow prostituting itself and despite all attempts someone (the information's pimp perhaps?) insists on controlling things. I guess this is the idea that information is going through some pre-1960's sexual revolution...
Rather I think the word 'free' refers to 'freedom' albeit in a Western sense of the word. Think of the use of the concept freedom as used in the US Declaration of Independence.

"We hold these truths to be self-evident, that all information is created equal, that it is endowed by its Creator with certain unalienable Rights, that among these are Life, Liberty and the pursuit of Happiness"
...or should that be Existence, Communication and Semantics perhaps....?

Let's stick with the word 'freedom' and its naive or common-sense meaning. What does it mean to be free? We can turn further to the the UN Declaration of Human Rights and the EU Charter of Fundamental Rights for further clarification, though here I'm sure we go into more of a legal-political debate more than anything. Evidently freedom means either the right to do or be something or the right to be protected from something.

So the question is, if information want to be free:

  • What does information want the freedom to do/be?
  • What does information want the freedom from?

Under the first question, the freedom to be 'free' as in 'without cost' certainly falls. What about the freedom to be private, or the freedom not to be abused - as in excessive privacy violations? Do we further need a notion of agency - does information have an owner or provenance?

Without answering those - I don't think I can give a definitive answer anyway - here's another thought. Given that matter = energy, isn't the use of the term information quite literally another way of saying 'humans' (or 'men' as in the Declaration of Independence). In which case the question 'information wants to be free' is just an expression of man's desire to define what freedom is - ostensibly in terms of freedom to do/be and freedom from.

And here comes the practical part, which freedoms to do/be or from do we allow or deny in order to be "free"?

Monday, 17 November 2014

A Definition of PII and Personal Data

There's been an interesting discussion on Twitter about the terms PII and "personal data", classification of information and metrics.

Personally I think the terms "PII" and "personal data" are too broadly applied. Their definitions are poor at best; when did you last see a formal definition of these terms? Indeed classifying a data set as PII only comes about from the types of data inside that data set and by measuring the amount of identifiability of that set.

There now exists two problems in that a classification system underneath that of PII isn't well established in normal terminology. Secondly metrics for information content are very much defined in terms of information entropy.

Providing these underlying classifications is critical to better comprehending the data that we are dealing with. For example, consider the following diagram:

Given any set of data, each field can be mapped into one or more of the seven broad categories on the left - If we wanted we could create much more sophisticated ontologies to express this. Within each of these we can specialise more and this is somewhat represented as we move horizontally across the diagram.

Avoiding information entropy as much as possible, we can and have derived some form of metric to at least assess the risk of data being held or processed. A high 'score' means high risk and a high degree of reidentification is possible, while a low score the opposite - though not necessarily meaning that there is no risk. Each of the categories could be further weighted such as using location is twice as risky as financial data.

There could be and are some interesting relationships between the categories, for example, identifiers such as machine addresses (IPs) can be mapped into personal identifiers and locations - depending upon the use case.

I'm not going to go into a full formalisation of the function to calculate this, but a simple function which takes in a data set's fields and produces a value, say in the range 0 to 5 to state the risk of the data set might suffice. A second function to map that value to a set of requirements to handle that risk is the needed.

What about PII?  Well, to really establish this we should go into the contents of the data and the context in which that data exists. Another, rather brutal way, is to draw a boundary line across the above diagram such that things on the right-hand-side are potentially PII and those on the left not. This might then become a useful weighting metric, that if anything appears to the right of this line then the whole data set gets tagged with being potentially PII. I guess you could also become quite clever in using this division line to normalise the risk scoring across the various information classifications.

In summary, we can therefore give the term PII (or personal data) a definition in terms of what a data set contains rather than using it as a catch-all classification. This allows us then to have a proper discussion about risk and requirements.

References

Ian Oliver. Privacy Engineering: A Data Flow and Ontological Approach. ISBN 978-1497569713


Friday, 22 November 2013

Semiotic Analysis of a Privacy Policy

Pretty much every service has a privacy policy attached to it; these policies state the data collection, usage, purpose and expectations that the customer has to agree with before using that said service. But at another level they also attempt to signify (that's going to be a key word) that the consumer can trust the company providing the service at some level. Ok, so there's been huge amounts of press about certain social media and search service provides "abusing" this trust and so on, but we still use the services provided by those companies.

So this gets me thinking, when a privacy policy is written, could we analyse that text to understand better the motives and expectations from both the customer and the service provider perspective? Effectively can we make a semiotic analysis of a privacy policy.

What would we gain by this? It is imperative that any texts of this nature portray the right image to the consumer, thus this can be used in the drafting of such a text to ensure that this this right image is correctly portrayed. For example, the oft seen statement:

"Your privacy is important to us"

is a sign in the semiotic sense, and in this case probably an 'icon' in its near universal usage. Signs are a relationship between the 'object' and 'interpretant', respectively the subject matter at hand and the clarified meaning respectively.

Pierce's Semiotic Trangle
The object may be that we (the writer of the statement) are trying to state a matter of fact about how trust worthy we are, or at least we want to emphasise that we can be trusted.

The interpretant of this, if we are the customer, can of course vary from total trust to utter cynicism. I guess of late the latter interpretation tends to be true. Understanding the variation in interpretants is a clear method for understanding what is being conveyed by the policy itself and whether the right impression is being given to the consumer.

At a very granular level the whole policy itself is a sign and the very existance of that policy and its structure, is it long and full of legalese or short and simple? Then there's the content (as described above) which may or may not depend upon the size of the policy....as in the World's Worst Privacy Policy.

References:


Aside:

Found this paper while researching for this: Philippe Codognet's THE SEMIOTICS OF THE WEB, it starts with a quote:
I am not sure that the web weaved by Persephone in this Orphic tale, cited in exergue of Michel Serres’ La communication , is what we are currently used to call the World Wide Web. Our computer web on the internet is nevertheless akin Persephone’s in its aims : representing and covering the entire universe. Our learned ignorance is conceiving an infinite virtual world whose center is everywhere and circumference nowhere ...
Must admit, I find that very, very nice. Best I've got is getting quotes about existential crises and cosmological structures in a paper about the Semantic Web with Ora Lassila.


Sunday, 22 September 2013

Learning Languages

I spent part of this morning watching childrens' TV with my son, in particular we watched Unna Junna - a children's programme broadcast on the Finnish YLE network in the Sami language which was conveniently subtitled in Finnish.

Aside from the discussion about what the presenter was saying and its translation into Finnish, at least for the words I recognised or could guess, this ended up for me as an early morning exercise in comparitive linguistics.

For almost as long as I can remember linguistics and language excited me and this morning was just one of those exciting lingustic experiences so belolved of polyglots. After studying Finnish for many years I find myself getting excited about reading road signs in Estonian, children's TV in Sami, etc and mapping these to my knowledge of Finnish. I find this pretty cool.

I also came across today a video of a presentation by Anthony Lauder on "PolyNots" given at a Polyglot conference in Budapest in 2013. This video is worth watching for just for Anthony's presentation skills alone. From this video (if the YouTube Dieties are smiling on you) you'll get links to a host of other videos on multilingualism and language.

Something I've noted is that many polyglots and people who are generally interested in languages all seem to admit that they were never any good in school. Now, for me, I certainly remember been rote schooled in French and German much to my disappointment and also to much detriment of the learning process; enjoyment was quite literally a foreign concept. To this day I recall German lessons in one school as being a pedagogical nightmare: crammed into a small room with over 40 other teenagers all hell bent on not learning with a teacher whose attiude to teaching was less than exemplary.

During these times I took solice in buying dictionaries and book on languages. I still have a dogeared copy of Russian Made Simple [1] which I studied intently.

Though having no support in terms of another Russian speaker and nor as it turned out any help with certain linguistic concepts made things `difficult' to say the least. I remember one incident with a school teacher when I asked what the dative case was - sadly the answer wasn't an explanation of how indirect objects and transitive verbs work but rather a full scale dressing down of my poor performance in French and German lessons - I wonder why?

So if I were to learn another language again it obviously couldn't be on the terms of the UK secondary education system. University was quite a different matter but I spend most of my time studying computer science and mathematics. These however provided me with an interesting set of tools for natural language learning.

Aside: I wrote my bachelor's degree dissertation on machine translation - coincidence?

The first tool is that all languages follow some general patterns. At least most Indo-European and Finno-Ugric languages do. There will be numbers, there will be pronouns, there will be nouns of various kinds, there will be verbs and tenses, possibly even adjectives and adverbs too. All sentences have a mix of subject, verb and object in various orders. So at that level there isn't too much difference eh?

Actually for the most part you can make a one-to-one mapping from your mother tongue to any other langauge and get by. For example:

English:  I, He, She, We, You, Red, Blue, House, Dog,...
Welsh: Fi, Fe, Hi, Ni, Chi, Coch, Glas, Ty, Ci,...
Finnish: Minä, Hän, Hän, Me, Te, Punainen, Sininen, Talo, Koira,...
RussianЯ, он, она, мы, вы, красный, синий, дом, собака, ...
Estonian: Mina, Ta, Ta, Me, Sa, Punane, Sinine, Maja, Koer, ...

Note the similarities between Estonian and Finnish, learn one and you almost get the other for free! The only thing that makes Russian a little more difficult is the script.

Once you get by, then you can build vocabulary, attune yourself to the subtlies in expressing yourself in the new language and most importantly gain confidence.

Actually let's emphasise the latter: GAIN CONFIDENCE. The difference between a child learning a language and an adult is that children have infinite confidence and don't care about not understanding, making mistakes and playing with the language.

Learn a small core set of words: I, you, he, she, it, we, they, red, blue, green, one, two three, come, go, buy, want, please, thankyou, hello, big small, "help! I'm trying to learn!" etc etc etc

Learn words that you find interesting: if you like Formula 1 then learn the words relevant there: race, win, crash, speed, overtake etc (acutally these could be very useful in any conversation with motorsport obsessed Finns).

Don't worry about perfect or sometimes even vaguely correct grammar.

The more you use a langauge and the more you TRY to use a language the better you will become and the more accepting of your mistakes. The better you will become in terms of grammar and style.

You WILL MAKE MISTAKES ... if you analyse to two native speakers against what the books tell you then you will notice immediately that they are making huge amounts of "mistakes" with the grammar, phrasing etc. Remember what we said about how children learn languages.

Read stuff that you find interesting: many learners books and I remember one newspaper for Finnish learners are so simplified and grammatically correct they were totally uninteresting and demoralising to read. If you're trying to learn Finnish go read Mika Valtari's books from the Komisario Palmu books to his literary classics such as Sinuhe. Palmu is Finland's answer to James Bond.

Here you will learn three things:
  1. there's huge amount you don't understand
  2. the bits that you do understand greatly compensate for the bits you don't understand and you'll learn to guess and work around the bits you don't understand.
  3. you'll have fun and gain confidence
If you don't understand something GUESS! This works really well in speech as well in comprehension and reading.

In Anthony's talk he described the two step process for learning 10 languages (based on Peano's Axioms apparently!!)
  • Step 1: Learn 9
  • Step 2: Add 1
Of course, the first new language is the hardest, but once you've learnt the patterns, a core vocabulary and found out what you like reading and discussing the next one is much easier; just like applying the successor function in Peano's axioms.

Actually at the end of the day you'll be surprised what a little vocabulary and a heap of confidence will do. I'm in no way fluent or even reasonably competent in French, but a knowledge of menus, how to order beer and food gets me remarkably far in France and seems to be very appreciated by the natives. In other words a level of fluency my school language teachers could only have hoped for.

One final word, when speaking with a native in that person's language, resist as much as you can any attempts by that native to speak your language: Finns will almost invariably speak English to a foreigner because "no foreigner learns Finnish" and that "they need the practice in English anyway" (despite most Finns speak English better than most native English speakers). Don't worry about code switching, that is, mixing languages if you don't know a word, keep the flow of conversation going rather than worry about correct grammar, pronunciation etc...indeed that is the very essence of fluency.


References:

[1] Eugene Jackson, Elizabeth Gordon, Geoffrey Braithwaite, Albina Tarasova (1977) Russian Made Simple. W.H.Allen, London. 0-491-01582-B

Wednesday, 30 May 2012

Semantic Isolation (Pt.1)

Working with Ora Lassila we have been discussing and working on the definition of the term "data silo" in order to clarify our ideas of semantic isolation when applied to databases, data assets and the interoperability and integration of the information contained within.

The term "silo" when applied to databases, data and information is interesting in that it occurs in a number of statements, such as, "my app data is siloed" or "we need to break the data silos" and so on.

The meaning of the term however is mixed in its usage and thus its usage is inconsistent and misleading in many cases; quite simply it is used to cover a large number of overlapping scenarios.

Understanding the scope and meaning of this term in its various contexts is central to understanding the interoperability problem in a practical sense. In a workshop today I have heard the term used in a large number of ways and also applied to the notion of interoperability: The term "silo" has been used to mean (at least!)
  • The data is siloed because it exists in its own database infrastructure
  • The data is siloed because it is on accessible via some access control
  • The data is siloed because it is in its own representation format
  • The data is siloed because it is not understandable/translatable (semantics)
We can present these are some kind of "lock-in" or "siloing continuum", where those usages on the left are more related to physical aspects and those on the right to more semantic in the information sense:




We obviously can create a more granular continuum (indeed that's what a continuum should allow) but the point here is to at least to present some kind of ordering over the differing uses of the term. The ordering runs from physical deployment and implementation through to abstract semantics.

Now it seems that when people talk about "breaking the [data] silos" they are actually referring to enabling interoperability of the data between differing services; and often this is addressed at the physical database or access control level. Occasionally the discussion gets mixed and syntax and representation of data is addressed.

Interoperability of information starts at the semantic level and works in reverse (right to left) through the above continuum; physical, logical, access control and syntax should not prevent sharing and common understanding of data. For example, if one tackles interoperability of information by standarising on syntax or representation (eg: JSON vs XML) then the resultant will be two sets of data that can't be merged because they don't have the same meaning; similarly at the other end of the continuum centralising databases (physically or logically) doesn't result in interoperability - maybe easier system management but never interoperability of information.

Interestingly I had an extremely interesting discussion about financial systems and that interoperability between these is extremely high even at the application (local usage) level and this is simply because the underlying semantics of any financial system is unified. The notions of profit, loss, debit, credit and translations between the meanings of things such as dollars, yen, euros, pounds and the mathematics of financial values is formally defined and unambiguously understood; even if the mechanics if financial and economic systems isn't, but that's a different aspect altogether.

Also an important points here is that the link between financial concepts to real-world concepts and objects is well relatively easily definable. Indeed probably all real-world concepts and objects have their semantics defined in terms of financial transactions and concepts. Thus siloing of data probably can only occur in the financial world at the access control level.

The requirements for breaking the silos is easily understood as the ability to cross-reference two different data-sets and be sure (within certain bounds) that the meaning of the information contained there is is compatible. We want to perform things such as "1 + one equals 2" and be sure that the concept of "one" is the same as "1", the definition of "+" matches the concept of "+" applied to things such as "1","2" etc as well as things such as "one","two" etc. In this case the common semantics of "1" and "one" has been defined...fortunately.

It is vitally important to understand that if we can unify data sets through translations via common semantics then the siloing of data breaks and we get data liberation or what some call data democratisation. Unification of semantics however is faught with difficulties [1] but is the key prerequisite to integration and interoperability and ultimately a more expansive usage of that information.


References:

[1]Ian Oliver, Ora Lassila (2011). Integration "In The Large". Position paper accepted at the W3C Workshop on Data and Services Integration, October 20-21 2011, Bedford, MA, USA


Sunday, 13 May 2012

Kusunda

The BBC has a report on the last speaker of the Kusunda language of Nepal. Always sad not just to lose a language but a way of thinking. Also very interesting to see a link to the languages of the Andaman Islands and the potential of a link with the Sentinelese people and language.

Nepal's mystery language on the verge of extinction




Gyani Maiya Sen, a 75-year-old woman from western Nepal, can perhaps be forgiven for feeling that the weight of the world rests on her shoulders.

She is the only person still alive in Nepal who fluently speaks the Kusunda language. The unknown origins and mysterious sentence structures of Kusunda have long baffled linguists.

Tuesday, 25 October 2011

The Copiale Cipher

The Copiale Cipher has been decrypted - the discussion of the work can be found on a website provided by the authors of the paper describing how the process of decrypting the document was made:
The “Copiale Cipher” is a 105 pages manuscript containing all in all around 75 000 characters. Beautifully bound in green and gold brocade paper, written on high quality paper with two different watermarks, the manuscript can be dated back to 1760-1780. Apart from what is obviously an owner's mark (“Philipp 1866”) and a note in the end of the last page (“Copiales 3”), the manuscript is completely encoded. The cipher employed consists of 90 different characters, comprising all from Roman and Greek letters, to diacritics and abstract symbols. Catchwords (preview fragments) of one to three or four characters are written at the bottom of left–hand pages.
Kevin Knight, Beáta Megyesi and Christiane Schaefer (2011), The Copiale Cipher. ACL Workshop on Building and Using Comparable Corpora (BUCC). 

The New York Times has an article:
Published: October 24, 2011
A team of linguists applied statistics-based techniques to translate one of the most stubborn of codes, a German mix of letters and symbols.

While the Copiale Cipher website at Uupsala University's Department of Linguistics and Philology has everything you need to know, here's the direct link to the English language translation (as a PDF).

The contents of the document are particularly interesting, referring as they do to n 18th century secret society known as an "oculist order" - oculist coming from the Greek and referring to eyes...here's the obligatory link to Wikipedia about ophthalmology.



Wednesday, 17 November 2010

Invented Languages

Nice to see this sort of thing on El Reg:

Speak geek: The world of made-up language
By Caleb Cox
17th November 2010 11:43 GMT

The world of invented language is a difficult place to succeed and those who have the patience to create their own tend to have a hard time gathering followers.

Klingon and Elvish are notable exceptions, thanks to the huge fan bases for Star Trek and Lord of The Rings.

Society tends to regard people who learn these languages as über geeky and socially-inept but we often overlook the reasons why they’re so obsessed with the fantasies they love.

Until recently, expanding the speaker numbers was a challenge: conventions were the only place for enthusiasts to gather and sporadic publications the only other method of sharing their passion.

With the internet, mobile app markets and other techie possibilities, these languages now have easily accessible platforms to grow. While such languages thrive, constructed languages, or “conlangs”, that were created in our past generally struggled.

Wednesday, 6 October 2010

More on Koro

More on Koro

New language identified in remote corner of India; One of thousands of endangered tongues around world

Interesting fact in that article is that we're losing language every two weeks...sad

Koro

WOW!



Indian language is new to science


Koro speakers (National Geographic) The newly recognised language is spoken by between 800 and 1,200 people in north-east India
Researchers have identified a language new to science in a remote region of India.
Known as Koro, it appears to be distinct from other languages in the family to which it belongs; but it is also under threat.

Saturday, 28 August 2010

If only....

Noam Chomsky to become new X-Factor judge

 Professor of linguistics and political campaigner Noam Chomsky has been confirmed as the new judge on TV talent show The X Factor. ‘Cheryl Cole was still recovering from malaria and we needed someone who could fill the intellectual void,’ said programme creator Simon Cowell, ‘Professor Chomsky is perfect and the audience just loves him.’

 ....now go read the rest of it on Newsbiscuit.

 

Thursday, 4 February 2010

Language death...

This is incredibly sad, from the BBC; basically an way of thinking has been lost...

Last speaker of ancient language of Bo dies in India 

By Alastair Lawson
BBC News

Boa Sr
Boa Sr remained the last Bo speaker for at least 30 years
The last speaker of an ancient language in India's Andaman Islands has died at the age of about 85, a leading linguist has told the BBC.
Professor Anvita Abbi said that the death of Boa Sr was highly significant because one of the world's oldest languages - Bo - had come to an end.

 

Thursday, 5 November 2009

The Sentinelese

The Sentinelese are a very rare kind of people in that they have almost no contact with the outside World. Very little is known about these people: almost nothing about their language, culuture, rituals etc - that what do know is so utterly sparse and incomplete renders that worthless in anything other than mere guesswork.

Wikipiedia has an article about the Sentinelese and in the article The Last Island of the Savages by Adam Goodheart are more detailed exposition is made.