Showing posts with label http-range. Show all posts
Showing posts with label http-range. Show all posts

Wednesday, July 29, 2009

The Illusion of Internet Identity

You've certainly heard of Arthur C. Clarke's Third Law, "Any sufficiently advanced technology is indistinguishable from magic", which says more about magic and our perceptions of the world than it does about technology. When technology does something that is not natural to us, we of course perceive it to be supernatural. But what happens when technology approximates something so natural to us that we don't even perceive that there's anything remarkable? Then we attribute powers to the technology that just don't exist. Just as we can perceive emotions in a stuffed teddy bear, it is only with difficulty that we avoid anthropomorphizing technologies. Do you have a cute name for your car? Do you refer to your GPS as "Lola"? If you have not done so, try the web version of ELIZA, and see if you can avoid thinking of ELIZA as a real person. It's very hard for us to understand how complicated the act of carrying on a real conversation really is- we do it all the time. Even a profoundly retarded technology will be imbued with magic if its function is sufficiently mundane.

I've been reading a paper by Patrick Hayes and Harry Halpin and a presentation by Pat Hayes, both with the unfortunate title "In Defense of Ambiguity". The paper provides a wonderful review of the theory of identity. I've been living very happily, doing productive work in the world of identifiers without ever knowing that identity needed to have a theory behind it. In retrospect, I've managed to do this by staying away from the difficult bits.

After reading Hayes and Halpin, I've come to realize what a miracle human communication is. The fact that I can meet someone with whom I share no languages and that we can exchange our own names and establish names for things may seem simple, but it's something that machines cannot do. For example, I may gesture at myself and say "Eric" to establish my identity. If I then gesture at a banana and say "banana", it's very likely that my counterpart will understand that I have not established an identity for the banana, but rather I have given a name for the kind of fruit. This is possible because people have brains that are similarly wired- our brains are wired to recognize individual people but not individual bananas (though sometimes our knowledge models diverge). Our computers on the other hand, have no fruit wiring, or individual person wiring, so establishment of identity is very hard for them.

The difficulty of teaching computers to identify things has not stopped us from using them to build elaborate identity systems. Hayes and Halpin observe that internet identity can only be established by description, and description is inherently ambiguous. Attempts to make real-world-object identifiers global or to add description actually make the situation worse, by increasing ambiguity. In our daily lives, ambiguity in our communications is mostly not a problem. When I say the word "rose" a listener will almost never be confused between the flower and the verb. I can say "rows of rosebushes" and only rarely will people hear "rose of rosebushes". Our brains are so good at using context to resolve ambiguity that we don't realize how hard it is for computers to do the same thing.

That the situation becomes worse with added description was a bit hard for me to absorb, because at first it seems that the better you define something, the less ambiguous your statements about it become. But it's not true for computers. Suppose my internet identity description added a physical description of me- for example the fact that I have blonde hair. That might help to identify me under certain circumstances, but then when my hair turns gray, it makes my identification more tenuous. You could say that I had blonde hair on a particular date, but then you'd need to add a physical model for hair color to your internet identity system. In actual fact, the added description might help a human to identify me, but it hurts a computer's efferts to establish my identity.

The Hayes and Halpin paper was written in the context of the "http-range" semantic web controversy that I touched upon in my post on the semantics of redirection. They argued that the http protocol is not the right place to put establishments of identity, and that the description model is better suited to do that. As I understand it, the Hayes-Halpin view did not prevail with the W3C TAG; "ambiguity" was not a great concept for people to rally around, I guess.

The internet identity systems I've worked with revolved around identifying things in libraries- books, serials, and articles. For the most part, these sorts of objects do not usually present deep identification quandaries, and so I've not noticed my ignorance of identity theory. For example, most people imagine that computers use ISBNs to identify books (or as I imagined in my last post, that computers use ISBNs to identify items in bookstores). Most often this illusion does not get us into any trouble, just as the illusion that teddy bears have feelings is mostly harmless. Computers are wired to deal with records in data files, and they use ISBNs (with frequent success) to identify and match records in data files, that's all. The rest is just software trickery.

It's interesting to note that the ISBN was not developed by the library community. It was developed by a statistics professor named Gordon Foster for the British Publishers Association. Librarians lived without identifier systems for many years and were content with the library equivalents of an address or locator system. It's as if librarians have intuitively known something that the architects of the semantic web have only recently struggled with- that we can aspire to build description systems and access systems, but building a system that can provide identity is more difficult than it looks like to a human.

Friday, July 24, 2009

If Elvis had an OpenID and the Mome Raths Outgrabe

If the space aliens that kidnapped Elvis decided to return him tomorrow, and he decided to use Twitter and a blog to communicate that fact, how would anyone know it was really him? Would the National Enquirer even bother to report the news? Would Elvis ever be able to reclaim his public or private identities? Would he be able to remember any of his passwords?

Password proliferation has been a problem for so long that innovators have solved it over and over again. The library world came up with an single-sign-on authentication/authorization system called Shibboleth, and then implemented EZProxy so they wouldn't have to deal with it. The UK developed the single-sign-on system called Athens. The dot-com bubble came up with a bunch of single-sign-on companies; some of them, including PassLogix and Imprivata are still at it. I am still waiting for the announcement of the Single-Single-Sign-On system.

OpenID took a different approach, and is now somewhat usable for the purpose of allowing people to establish an identity with one provider that can be used on many websites. For example, I've used the OpenID identity "http://go-to-hellman.blogspot.com/" to register comments on the Semtech 2009 website and the Paul Miller's Cloud of Data Blog. On my last post, comments were left by "nicomo", and "breizhlady", whose OpenID's are http://nicomo.pip.verisignlabs.com/ and http://breizhlady.myopenid.com/ . Jodi Schneider used Blogger credentials to leave her comment. My OpenID can be used to determine with some degree of certainty that the the Eric Hellman who left a comment on Cloud of Data is the same Eric Hellman who's writing on this blog. A bit of googling will tell you who nicomo and breizhlady are, if you really want to know. If Elvis had been issued an OpenID before he left we would be able to tie his new blog to his old identity.

There are still the single-single problems with OpenID. The user experience for OpenID systems gets a bit clunky- my wife was frustrated when she tried to leave a comment on this blog. But overall there seems to be slow convergence and user acceptance of OpenID.

This brings me to the questions I wanted to raise today: What does an OpenID identify? Does http://go-to-hellman.blogspot.com/ identify me? Can these OpenIDs be used to make assertions about people to enter into the Linked Data Cloud? How should the Linked Data semantics for redirects be implemented for OpenID? Should 303 redirects be used to indicate that the "thing" being identified by OpenID is a real-world object?

To some extent, it's really the way identifiers are used that determines semantics- identification of any real-world object can never have perfect accuracy. The use of ISBN to identify a book is a good example. Although ISBN is frequently used to identify a book, ISBNs are managed in such a way that they most accurately identify items sold in a bookstore- toys and dolls often get ISBNs. Similarly, you might think that the US identifies people with Social Security Numbers (SSN), but if you think about it, the "thing" an SSN most accurately identifies is an account with the Internal Revenue Service. Similarly, I think it's pretty clear that an OpenID identifies a set of login credentials, although people might well use the OpenID to identify the person or persons behind it.

I have been guilty in the past of driving people to distraction by arguing that it can be almost impossible to decide whether something is an "information resource" (something whose essential characteristics can be conveyed in a message) or whether it is a "real-world object". It's pretty easy to blur the issue with an e-book, for example, but what about the SSN? It used to be that "an IRS account" was something on paper somewhere, but I'm pretty sure that my entire IRS account is digitized somewhere.

Section 3 of the W3C's Technical Recommendation "Cool URIs for the Semantic Web" assumes that it's easy to determine whether something is an information resource or whether it's a real-word object and that it's impossible to convey the essence of real-world objects in a stream of bits. I find this a bit unworldy. It even cites the unicorn as an example of a "real-world object". I guess that makes Elvis a real-world object, too. Conversely, even things that live completely on the internet are rarely "conveyed in a message" any more. A typical URI-addressable service today is constructed out of software, web services, content delivery networks, advertising delivery networks and clustered hardware so that the "essential characteristics" include the attributes of real world objects like me.

I've recently become aware that lots of really smart people have thought and written about the theory of identifiers and about how the Semantic Web should handle them. I've particularly enjoyed an article called "In Defense of Ambiguity" by Patrick Hayes and Harry Halpin. But to answer my questions about the semantics of OpenID, there's no sage more useful than the one who said "When I use a word, it means just what I choose it to mean - neither more nor less." Semantics do not get determined by those who mint the identifiers, but rather by those who make use of them. It helps if they are also willing to pay the IDs a bit extra.

Thursday, July 9, 2009

URL Shorteners and the Semantics of Redirection

When I worked at Bell Labs in Murray Hill, NJ, it amused me that at one end of the building, the fiber communications people were worrying that no one could ever possibly make use of all the bandwidth they could provide- we would never be able to charge for telephone calls unless they figured out how to limit the bandwidth. At the other end of the building, computer scientists were figuring out how to compress information so that they could pack more and more into tiny bit-pipes. I'm still not sure who won that battle.

When I was part of a committee working on the OpenURL standard, we had a brief discussion about the maximum length URL that would work over the internet. A few years before that, there were some systems on the internet that barfed if a URL was longer than 512 characters, but most everything worked up to 2,000 characters, and we anticipated that that limit would soon go away. So here we are in 2009, and Internet Explorer is just about the only thing that still has a length limit as low as 2083 characters. Along comes Twitter, with a 140 character limit on an entire message, and all of a sudden, the URL's we've been making have become TOO LONG! Just as fast, URL shortening services sprung up to make the problem go away.

The discussion on my last post (on CrossRef and OpenURL) got me interested in the semantics of redirection, and that got me thinking about the shortening services, which have become monster redirection engines. When we say something about a URI that is resolved by a redirector, what, exactly are we talking about?

First, some basics. A redirection occurs when you click on a link and the web server for that link tells your browser to go to another URL. Usually, the redirection occurs in the http protocol that governs how your web browser gets web pages. Sometimes, a redirect is caused by a directive in an html page, or programmed by a javascript in that page. The result may seem the same but the mechanism is rather different, and I won't get into it any further. There are actually 3 types of redirects provided for in the http protocol, known by their status codes as "301" "302" and "303" redirects. There are 5 other redirect status codes that you can safely ignore if you're not a server developer. The 301 redirect is called "Moved Permanently", the 302 is called "Found" and the 303 is called "See Other". Originally, the main reason for the different codes was to help network servers figure out whether to cache the responses to save bandwidth (the fiber guys had not deployed so much back then and the bit squeezers were top dogs). Nowadays the most important uses of the different codes are in search engines. Google will interpret a 301 as "don't index this url, index the redirect URL". A 302 will be interpreted as "index the content at the redirect URL, but use this URL for access". According to a great article on url shorteners by Danny Sullivan, Google will treat a 303 like a 302, but who knows?

Just as 301 and 302 semantics have been determined by their uses in search engines, the 303 has been coopted by the standards-setters of the semantic web, and they may well be successful in determining the semantics of the 303. As described in a W3C Technical Recommendation, the 303 is to be used
... to give an indication that the requested resource is not a regular Web document. Web architecture tells you that for a thing resource (URI) it is inappropriate to return a 200 because there is, in fact, no suitable representation for those resources.
In other words, the 303 is suppoesed to indicate that the Thing identified by the URI (URL) is something whose existence is NOT on the web. Tim Berners-Lee wrote a lengthy note about this that I found quite enjoyable, though at the end I had no idea what it was advocating. The discussion that led to the W3C Recommendation has apperently been extremely controversial, and has been given the odd designation "http-range-14". The whole thing reminds me of reading the existentialists Sartre and Camus in high school - they sounded so much more understandable in French!

As discussed in Danny Sullivan's article, most of the URL shorteners use 301 redirects, which is usually what most users want to happen. An indexing agent or a semantic web agent should just look through these redirectors and use the target resource URL in its index. The DOI "gateway" redirector at dx.doi.org discussed in my previous post uses a 302 redirect. Unless doi's are handled specially by a search engine, it means that the "link credit" (a.k.a. google juice) for a dx.doi.org link will accrue to the dx.doi.org URL rather than the target URL. This seems appropriate. Although I indicated that if you use Linked Data rules the dx.doi.org link identifies whatever is indicated by the returned web page, from the point of view of Search engines, that URI identifies an abstraction of the resource it redirects to. A redirection service similar in conception, PURL, also uses 302 redirects.

I was curious about the length limits of the popular url shorteners. Using a link to this blog, padded by characters ignored by Blogger.com, I shortened a bunch of long URLs. Here are 4 shortened 256 character links to this blog:
They all work just fine. Moving to 1,135 character links, everything still works (at least in my environment):
At 2083 characters, the limit for Internet Explorer, we start separating the redirection studs from the muffins.
When I add another character, to make 2,084 total, bit.ly and snurl.com both work, but blogger.com reports an error!
The compression ratios for these last two links is 109 to 1 for bit.ly and 95 to 1 for snurl. The bit squeezers would be happy.

Next, I wanted to see if I could make a redirection loop. Most of the shortening services decline to shorten a shortened URL, but they're quite willing to shorten a URL from the PURL service. Also, I couldn't find any way to use the shortening services to fix a link that had rotted after I shortened it. It could be useful to add the PURL service as link-rot insurance behind a shortened url if the 302 redirect is not an issue. So here's a PURL: http://purl.oclc.org/NET/backatcha that redirects to http://bit.ly/aE0od which redirects to http://purl.oclc.org/NET/backatcha etc. Don't click these expecting an endless loop- your browser should detect the loop pretty fast.

A recent article about how bit.ly is using its data stream to develop new services got me thinking again about how a shortening redirector might be useful in Linked Data. I've written several times that Linked Data lacks the strong attribution and provenance infrastruction needed for many potential applications. Could shortened URIs be used as Linked Data predicates to store and retrieve attribution and provenance information, along with the actual predicate? And will I need another http status code to do it?