I used to run my own mail server. But then came the spammers. And dictionary attacks. All sorts of other nasty things. I finally gave up and turned to Gmail to maintain my online identities. Recently, one of my web servers has been attacked by a bot from a Russian IP address which will eventually force me to deploy sophisticated bot-detection. I'll probably have to turn to Google's recaptcha service, which watches users to check that they're not robots.
Isn't this how governments and nations formed? You don't need a police force if there aren't any criminals. You don't need an army until there's a threat from somewhere else. But because of threats near and far, we turn to civil governments for protection. The same happens on the web. Web services may thrive and grow because of economies of scale, but just as often it's because only the powerful can stand up to storms. Facebook and Google become more powerful, even as civil government power seems to wane.
When a company or institution is successful by virtue of its power, it needs governance, lest that power go astray. History is filled with examples of power gone sour, so it's fun to draw parallels. Wikipedia, for example, seems to be governed like the Roman Catholic Church, with a hierarchical priesthood, canon law, and sacred texts. Twitter seems to be a failed state with a weak government populated by rival factions demonstrating against the other factions. Apple is some sort of Buddhist monastery.
This year it became apparent to me that Facebook is becoming the internet version of a totalitarian state. It's become so ... needy. Especially the app. It's constantly inventing new ways to hoard my attention. It won't let me follow links to the internet. It wants to track me at all times. It asks me to send messages to my friends. It wants to remind me what I did 5 years ago and to celebrate how long I've been "friends" with friends. My social life is dominated by Facebook to the extent that I can't delete my account.
That's no different from the years before, I suppose, but what we saw this year is that Facebook's governance is unthinking. They've built a machine that optimizes everything for engagement and it's been so successful that they they don't know how to re-optimize it for humanity. They can't figure out how to avoid being a tool of oppression and propaganda. Their response to criticism is to fill everyone's feed with messages about how they're making things better. It's terrifying, but it could be so much worse.
I get the impression that Amazon is governed by an optimization for efficiency.
How is Google governed? There has never existed a more totalitarian entity, in terms of how much it knows about every aspect of our lives. Does it have a governing philosophy? What does it optimize for?
In a lot of countries, it seems that the civil governments are becoming a threat to our online lives. Will we turn to Wikipedia, Apple, or Google for protection? Or will we turn to civil governments to protect us from Twitter, Amazon and Facebook. Will democracy ever govern the Internet?
Happy 2019!
Showing posts with label facebook. Show all posts
Showing posts with label facebook. Show all posts
Monday, December 31, 2018
Sunday, August 26, 2012
The "I Used This" Button
I frequent-flyered off to San Francisco this weekend to surprise my Ph. D. Advisor, Jim Harris, for his 70th Birthday. I was the first of his students to graduate, and he's up to 105. On thing I learned helping to start his group was the immense value of being thrown together with a group of smart people with a variety of experience. I met members of Jim's current group, which includes a student from Gunn High School, visiting scholars from around the world, and Ph. D. students bursting with ideas.
I manufactured some business-related meetings for the trip, some of which I'll relate in a to-be-written post, but I also lucked into a hackathon for Open-Access hosted by PLoS. I spent the day with a group of smart people with a wide range of experience, including software developers, product managers, film-makers and a librarian or three.
The group I ended up working with included Greg Grossmeier from Creative Commons, Cameron Neylon from PLoS, and Ana Nelson, the developer-entrepreneur behind dexy. We were interested in counting open-access things. Counting things can be harder than you think, because you have to define the things and identify them; you need to be able to tell whether a thing is the same thing as another thing, or perhaps it's three things. Counting bananas is one thing, but have you ever tried counting ideas?
Creative Commons (CC) is interested in knowing how much its licenses are used. When an Unglue.it ebook edition is released (of course under Creative Commons!), how often is it used? Does a single license apply to the entire book, or can we apply different licenses to the different resources inside the book? For example, an author may want to use a CC-BY license for the text of a book, which might contain figures that are used under CC BY-ND. And the metadata should be CC0. How should these licenses be expressed?
After some discussion, we settled down to work on some specific projects. My project turned out not to be code at all, but rather a description of a scheme for measuring Creative Commons usage, i.e. the rest of this blog post.
Creative Commons has thought about ways to measure the usage of its licenses. For example, it can track the display of its license "badges", such as the one right here. Web browsers will send a referrer header that tells the image server the web page and user IP address. But there are problems. Many web sites use their own copy of the badge. In an ebook, the badge would be embedded in the ebook file. If the page is served over a secure socket, the referrer won't be set. And do you really want to tell Creative Commons about everything you're reading?
Speaking of which, have you clicked on a Facebook "Like" button this week? Was it good for you too?
Suppose there was a button on Creative Commons licensed documents that allowed the user to express their delight at the creator's enlightened choice of license. Would you click it? I call it the "I Used This" (IUT) button, but maybe you can think of a better name.
- The IUT button would send a signal to a Creative Commons server about usage of the resource. These signals would be compiled and reported.
- IUT button would also send attribution url.
- Pressing the button would display an amusing animation to the user. Perhaps every button would have a different animation to avoid button fatigue.
- The button would be at the center of an advocacy campaign for open licenses.
- Unlike the Facebook Like button, the IUT button would respect a user's privacy. A signal would be sent only when initiated by the user, and would be optional.
- An IUT button packaged as a javascript would work in epub, html, etc.
- IUT signals would be evidence of the resource's status as a CC licensed work. A licensor attempting to revoke a CC license (you can't do that!) would have to overcome a verifiable usage trail.
- Users could create accounts at CC to provide a retrospective record of the user's Use.
- Clicking the IUT Button would put the attribution url on clipboard to ease correct citations.
- Usage information for each resource would be public- creators could easily track the usage signals for their works.
- We might need anti-ballot-stuffing measures if CC usage rankings become commercially important.
If efforts like Unglue.it are to succeed, people who appreciate the benefits of Creative Commons licensing need to stand up and be counted. We need to make it a mass movement in the minds of every lover of books, everywhere.
Sometimes you need to do more than just consume. Sometimes you need to do some SHOUTING.
Posted by
Eric
at
7:17 PM
1 comments
Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest
Labels:
Creative Commons,
facebook,
Hackathon,
Open Access
Wednesday, July 27, 2011
Liking Library Data
If you had told me ten years ago that teenagers would be spending free time "curating their social graphs", I would have looked at you kinda funny. Of course, ten years ago, they were learning about metadata from Pokemon cards, so maybe I should have seen it coming.
Social networking websites have made us all aware of the value of modeling aspects of our daily lives in graph databases, even if we don't realize that's what we're doing. Since the "semantic web" is predicated on the idea that ALL knowledge can be usefully represented as a giant, global graph, it's perhaps not so surprising that the most familiar, and most widely implemented application of semantic web technologies has been Facebook's "Like" button.
When you click a Like button, an arc is added to Facebook's representation of your social graph. The arc links a node that represents you and another node that represents the thing you liked. As you interact with your social graph via Facebook, the added Like arc may introduce new interactions.
Google must think this is really important. They want you to start clicking "+1" buttons, which presumably will help them deliver better search. (You can try following me+, but I'm not sure what I'll do with it.)
The technology that Facebook has favored for building new objects to but in the social graph is derived from RDFa, which adds structured data into ordinary web pages. It's quite similar to "microdata", a competing technology that was recently endorsed by Google, Microsoft, and Yahoo. Facebook's vocabulary for the things it's interested in is called Open Graph Protocol (OGP), which could be considered a competitor for Schema.org.
My previous post described how a library might use microdata to help users of search engines find things in the library. While I think that eventually this will be an necessity for every library offering digital services, the are a bunch of caveats that limit the short-term utility of doing so. Some of these were neatly described in a post by Ed Chamberlain:
Facebook's OGP may have more immediate benefits. Libraries are inextricably linked to their communities; what is a community if not a web of relationships? Libraries are uniquely positioned to insert books into real world social networks. A phrase I heard at ALA was "Libraries are about connections, not collections".
Libraries don't need to implement OGP to put a like button on a web page, but without OGP Facebook would understand the "Like" to be about the web page, rather than about the book or other library item.
To show what OGP might look like on a library catalog page, using the same example I used in my post on "spoonfeeding library data to search engines":
Open Graph Protocol wants the web page to be the digital surrogate for the thing to be inserted into the social graph, and so it wants to see metadata about the thing in the web page's meta tags. Most library catalog systems already put metadata in metatags, so this part shouldn't be horribly impossible.
The first thing that OGP does is to call out xml namespaces- one for xhtml, a second for Open Graph Protocol, and a third for some specific-to-Facebook properties. A brief look at OGP reveals that it's even more bare bones than schema.org; you can't even express the fact that "Paul Bryers" is the author of "Avatar".
This is less of an issue than you might imagine, because OGP uses a syntax that's a subset of RDFa, so you can add namespaces and structured data to your heart's desire, though Facebook will probably ignore it.
The next step is to add the actual like button by embedding a javascript from Facebook:
The "og:url" property tells facebook the "canonical" url for this page- the url that Facebook should scrape the metadata from.
Now here's a big problem. Once you put the like button javascript on a web page, Facebook can track all the users that visit that page. This goes against the traditional privacy expectations that users have of libraries. In some jurisdictions, it may even be against the law for a public library to allow a third party to track users in this way. I expect it shouldn't be hard to modify the implementation so that the script is executed only if the user clicks the "Like" button, but I've not been able to find a case anyone has done this.
It seems to me that injecting library resources into social networks is important. The libraries and the social networks that figure out how to do that will enrich our communities and the great global graph that is humanity.
Social networking websites have made us all aware of the value of modeling aspects of our daily lives in graph databases, even if we don't realize that's what we're doing. Since the "semantic web" is predicated on the idea that ALL knowledge can be usefully represented as a giant, global graph, it's perhaps not so surprising that the most familiar, and most widely implemented application of semantic web technologies has been Facebook's "Like" button.
When you click a Like button, an arc is added to Facebook's representation of your social graph. The arc links a node that represents you and another node that represents the thing you liked. As you interact with your social graph via Facebook, the added Like arc may introduce new interactions.
Google must think this is really important. They want you to start clicking "+1" buttons, which presumably will help them deliver better search. (You can try following me+, but I'm not sure what I'll do with it.)
The technology that Facebook has favored for building new objects to but in the social graph is derived from RDFa, which adds structured data into ordinary web pages. It's quite similar to "microdata", a competing technology that was recently endorsed by Google, Microsoft, and Yahoo. Facebook's vocabulary for the things it's interested in is called Open Graph Protocol (OGP), which could be considered a competitor for Schema.org.
My previous post described how a library might use microdata to help users of search engines find things in the library. While I think that eventually this will be an necessity for every library offering digital services, the are a bunch of caveats that limit the short-term utility of doing so. Some of these were neatly described in a post by Ed Chamberlain:
- the library website needs to implement a site-map that search engine's crawlers can use to find all the items in the Library's catalog
- the library's catalog needs to be efficient enough to not be burdened by the crawlers. Many library catalog systems are disgracefully inefficient.
- the library's catalog needs to support persistent URLs. (Most systems do this, but it was only ten years ago that I caused Harvard's catalog to crash by trying to get it to persist links. Sorry.)
Facebook's OGP may have more immediate benefits. Libraries are inextricably linked to their communities; what is a community if not a web of relationships? Libraries are uniquely positioned to insert books into real world social networks. A phrase I heard at ALA was "Libraries are about connections, not collections".
Libraries don't need to implement OGP to put a like button on a web page, but without OGP Facebook would understand the "Like" to be about the web page, rather than about the book or other library item.
To show what OGP might look like on a library catalog page, using the same example I used in my post on "spoonfeeding library data to search engines":
<html> <head> <title>Avatar (Mysteries of Septagram, #2)</title> </head> <body> <h1>Avatar (Mysteries of Septagram, #2)</h1> <span>Author: Paul Bryers (born 1945)</span> <span>Science fiction</span> <img src="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg"> </div>
Open Graph Protocol wants the web page to be the digital surrogate for the thing to be inserted into the social graph, and so it wants to see metadata about the thing in the web page's meta tags. Most library catalog systems already put metadata in metatags, so this part shouldn't be horribly impossible.
<html xmlns="http://www.w3.org/1999/xhtml"
xmlns:og="http://ogp.me/ns#"
xmlns:fb="http://www.facebook.com/2008/fbml">
<head>
<title>Avatar (Mysteries of Septagram, #2)</title>
<meta property="og:title" content="Avatar - Mysteries of Septagram #2"/>
<meta property="og:type" content="book"/>
<meta property="og:isbn" content="9780340930762"/>
<meta property="og:url"
content="http://library.example.edu/isbn/9780340930762"/>
<meta property="og:image"
content="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg"/>
<meta property="og:site_name" content="Example Library"/>
<meta property="fb:admins" content="USER_ID"/>
</head>
<body>
<h1>Avatar (Mysteries of Septagram, #2)</h1>
<span>Author: Paul Bryers (born 1945)</span>
<span>Science fiction</span>
<img src="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg">
</div>
The first thing that OGP does is to call out xml namespaces- one for xhtml, a second for Open Graph Protocol, and a third for some specific-to-Facebook properties. A brief look at OGP reveals that it's even more bare bones than schema.org; you can't even express the fact that "Paul Bryers" is the author of "Avatar".
This is less of an issue than you might imagine, because OGP uses a syntax that's a subset of RDFa, so you can add namespaces and structured data to your heart's desire, though Facebook will probably ignore it.
<html xmlns="http://www.w3.org/1999/xhtml"
xmlns:og="http://ogp.me/ns#"
xmlns:fb="http://www.facebook.com/2008/fbml"
xmlns:dc="http://purl.org/dc/elements/1.1/"
xmlns:foaf="http://xmlns.com/foaf/0.1/">
<head>
<title>Avatar (Mysteries of Septagram, #2)</title>
<meta property="og:title"
content="Avatar - Mysteries of Septagram #2"/>
<meta property="og:type"
content="book"/>
<meta property="og:isbn"
content="9780340930762"/>
<meta property="og:url"
content="http://library.example.edu/isbn/9780340930762"/>
<meta property="og:image"
content="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg"/>
<meta property="og:site_name"
content="Example Library"/>
<meta property="fb:app_id"
content="183518461711560"/>
</head>
<body>
<h1>Avatar (Mysteries of Septagram, #2)</h1>
<span rel="dc:creator">Author:
<span typeof="foaf:Person"
property="foaf:name">Paul Bryers
</span> (born 1945)
</span>
<span rel="dc:subject">Science fiction</span>
<img src="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg">
</div>
The next step is to add the actual like button by embedding a javascript from Facebook:
<div id="fb-root"></div>
<script src="http://connect.facebook.net/en_US/all.js#appId=183518461711560&xfbml=1"></script>
<fb:like href="http://library.example.edu/isbn/9780340930762/"
send="false" width="450" show_faces="false" font=""></fb:like>
The "og:url" property tells facebook the "canonical" url for this page- the url that Facebook should scrape the metadata from.
Now here's a big problem. Once you put the like button javascript on a web page, Facebook can track all the users that visit that page. This goes against the traditional privacy expectations that users have of libraries. In some jurisdictions, it may even be against the law for a public library to allow a third party to track users in this way. I expect it shouldn't be hard to modify the implementation so that the script is executed only if the user clicks the "Like" button, but I've not been able to find a case anyone has done this.
It seems to me that injecting library resources into social networks is important. The libraries and the social networks that figure out how to do that will enrich our communities and the great global graph that is humanity.
Posted by
Eric
at
3:39 PM
0
comments
Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest
Labels:
facebook,
Libraries,
library automation,
metadata,
Microdata,
RDFa,
social networks
Sunday, June 27, 2010
Global Warming of Linked Data in Libraries
Libraries are unusual social institutions in many respects; perhaps the most bizarre is their reverence for metadata and its evangelism. What other institution considers the production, protection and promulgation of metadata to be part of its public purpose?
The W3C's Linked Data activity shares this unusual mission. For the past decade, W3C has been developing a technology stack and methodology designed to support the publication and reuse of metadata; adoption of these technologies has been slow and steady, but the impact of this work has fallen short of its stated ambitions.
I've been at the American Library Association's Annual Meeting this weekend. Given the common purpose of libraries and Linked Data, you would think that Linked Data would be a hot topic of discussion. The weather here has been much hotter than Linked Data, which I would describe as "globally warming". I've attended two sessions covering Linked Data, each attended by between 50 and 100 delegates. These followed a day long, sold-out preconference. John Phipps, one of the leaders in the effort to make library metadata compatible with the semantic web, remarked to me that these meeting would not have been possible even a year ago. Still, this attendance reflects only a tiny fraction of metadata workers at the conference; Linked Data has quite a ways to come. It's only a few months ago that the W3C formed a Library Linked Data Incubator Group.
On Friday morning, there was an "un-conference" organized by Corey Harper from NYU and Karen Coyle, a well-known consultant. I participated in a subgroup looking at use cases for library Linked Data. It took a while for us to get around to use cases though, as participants described that usage was occurring, but they weren't sure what for. Reports from OCLC (VIAF) and Library of Congress (id.loc.gov) both indicated significant usage but little feedback. The VIVO project was described as one with a solid use case (giving faculty members a public web presence), but no one from VIVO was in attendance.
On Sunday morning, a meeting of the Association for Library Collections and Technical Services (ALCTS), Rebecca Guenther, Library of Congress, discussed id.loc.gov, a service that enables both humans and machines to programatically access authority data at the Library of Congress. Perhaps the most significant thing about id.loc.gov is not what it does but who is doing it. The Library of Congress provides leadership for the world of library cataloguing; what LC does is often slavishly imitated in libraries throughout the US and the rest of the world. id.loc.gov started out as a research project but is now officually supported.
Sara Russell-Gonzalez of the University of Florida then presented the VIVO which has won a big chunk of funding from the National Center for Research Resources, a branch of NIH. The goal of VIVO is to build an "interdisciplinary national network enabling collaboration and discovery between scientists across all disciplines." VIVO started at Cornell and has garnered strong institutional support there, as evidenced by an impressive web site. If VIVO is able to gain similar support nationally and internationally, it could become an important component of an international research infrastructure. This is a big "if". I asked if VIVO had figured out how to handle cases where researchers change institutional affiliations; the answer was "No". My question was intentionally difficult; Ian Davis has written cogently about the difficulties RDF has in treating time-dependent relationships. It turns out that there are political issues as well. Cornell has had to deal with a case where an academic department wanted to expunge affiliation data for a researcher who left under cloudy circumstances.
At the un-conference, I urged my breakout group to consider linked data as a way to expose library resources outside of the library world as well as a model for use inside libraries. It's striking to me that libraries seem so focused on efforts such as RDA, which aim to move library data models into Semantic Web compatible formats. What they aren't doing is to make library data easily available in models understandable outside the library.
The two most significant applications of Linked Data technologies so far are Google's Rich Snippets and Facebook's Open Graph Protocol (whose user interface, the "Like" button, is perhaps the semantic webs most elegant and intuitive). Why aren't libraries paying more attention to making their OPAC results compatable with these application by embedding RDFa annotations in their web-facing systems? It seems to me that the entire point of metadata in libraries is to make collections accessible. How better to do this than to weave this metadata into peoples lives via Facebook and Google? Doing this will require the dumbing-down of library metadata and some hard swallowing, but it's access, not metadata quality, that's core to the reason that libraries exist.
The W3C's Linked Data activity shares this unusual mission. For the past decade, W3C has been developing a technology stack and methodology designed to support the publication and reuse of metadata; adoption of these technologies has been slow and steady, but the impact of this work has fallen short of its stated ambitions.
I've been at the American Library Association's Annual Meeting this weekend. Given the common purpose of libraries and Linked Data, you would think that Linked Data would be a hot topic of discussion. The weather here has been much hotter than Linked Data, which I would describe as "globally warming". I've attended two sessions covering Linked Data, each attended by between 50 and 100 delegates. These followed a day long, sold-out preconference. John Phipps, one of the leaders in the effort to make library metadata compatible with the semantic web, remarked to me that these meeting would not have been possible even a year ago. Still, this attendance reflects only a tiny fraction of metadata workers at the conference; Linked Data has quite a ways to come. It's only a few months ago that the W3C formed a Library Linked Data Incubator Group.
On Friday morning, there was an "un-conference" organized by Corey Harper from NYU and Karen Coyle, a well-known consultant. I participated in a subgroup looking at use cases for library Linked Data. It took a while for us to get around to use cases though, as participants described that usage was occurring, but they weren't sure what for. Reports from OCLC (VIAF) and Library of Congress (id.loc.gov) both indicated significant usage but little feedback. The VIVO project was described as one with a solid use case (giving faculty members a public web presence), but no one from VIVO was in attendance.
On Sunday morning, a meeting of the Association for Library Collections and Technical Services (ALCTS), Rebecca Guenther, Library of Congress, discussed id.loc.gov, a service that enables both humans and machines to programatically access authority data at the Library of Congress. Perhaps the most significant thing about id.loc.gov is not what it does but who is doing it. The Library of Congress provides leadership for the world of library cataloguing; what LC does is often slavishly imitated in libraries throughout the US and the rest of the world. id.loc.gov started out as a research project but is now officually supported.
Sara Russell-Gonzalez of the University of Florida then presented the VIVO which has won a big chunk of funding from the National Center for Research Resources, a branch of NIH. The goal of VIVO is to build an "interdisciplinary national network enabling collaboration and discovery between scientists across all disciplines." VIVO started at Cornell and has garnered strong institutional support there, as evidenced by an impressive web site. If VIVO is able to gain similar support nationally and internationally, it could become an important component of an international research infrastructure. This is a big "if". I asked if VIVO had figured out how to handle cases where researchers change institutional affiliations; the answer was "No". My question was intentionally difficult; Ian Davis has written cogently about the difficulties RDF has in treating time-dependent relationships. It turns out that there are political issues as well. Cornell has had to deal with a case where an academic department wanted to expunge affiliation data for a researcher who left under cloudy circumstances.
At the un-conference, I urged my breakout group to consider linked data as a way to expose library resources outside of the library world as well as a model for use inside libraries. It's striking to me that libraries seem so focused on efforts such as RDA, which aim to move library data models into Semantic Web compatible formats. What they aren't doing is to make library data easily available in models understandable outside the library.
The two most significant applications of Linked Data technologies so far are Google's Rich Snippets and Facebook's Open Graph Protocol (whose user interface, the "Like" button, is perhaps the semantic webs most elegant and intuitive). Why aren't libraries paying more attention to making their OPAC results compatable with these application by embedding RDFa annotations in their web-facing systems? It seems to me that the entire point of metadata in libraries is to make collections accessible. How better to do this than to weave this metadata into peoples lives via Facebook and Google? Doing this will require the dumbing-down of library metadata and some hard swallowing, but it's access, not metadata quality, that's core to the reason that libraries exist.
Posted by
Eric
at
3:10 PM
4
comments
Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest
Labels:
ALA Annual,
facebook,
Google,
Libraries,
linked data,
metadata,
RDFa,
Semantic web
Friday, May 21, 2010
Bit.ly Preview Add-on Leaks User Activity; Referer Header Considered Harmful
This week FaceBook and MySpace had to deal with the consequences of obscure bugs that leaked personal subscriber information to advertisers. The Wall Street Journal reported that because Facebook and MySpace put user handles in URLs on their sites, these user handles, which can very often be traced back to a user identity, leaked to advertisers via the referer headers sent by browser software.
Reaction on one technology blog reminded me of Intel's missteps. Marshall Kirkpatrick, on ReadWriteWeb, called the the Journal's article "a jaw dropping move of bizarreness", going on to explain that passing referrer information was "just how the Internet works" and accusing the Journal of "anti-technology fear-mongering".
When a web browser requests a file from a website, it sends a bunch of extra information via http headers. One header gives the address of the file, which might be a web page, an image, or a script file. Other headers give the name of the software being used, the language and character sets supported by the browser. The Referer header (yes, that's how it's spelled, blame the RFC for getting the spelling wrong) reports the address of the page that requested or linked to the file. If the request is made to an advertiser's site, the Referer URL identifies the page that the user is looking at. When that page has an address that include private information, the private stuff can leak.
The controversy spurred me to take a look at some library websites to see what sort of data they might leak using referer headers. I used the very handy Firefox add-on called "Live HTTP Headers". I was astounded to see that a well known book database website seemed to be reporting the books I was browsing to Bit.ly, the URL shortening service! In another header, Bit.ly was also getting an identifying cookie. I went to another website, and found the exact same thing. This set off some alarm bells.
I soon realized that a report of EVERY web page I visit is being sent to Bit.ly. The culprit turned out to be Bit.ly's Bit.ly Preview add-on for Firefox. It turns out that for every web page I visit, this line of javascript is executed:
this.loadCss("https://s.bit.ly/preview.s3.v2.css?v=4.2");
This request for a CSS stylesheet has the side effect of causing Firefox to transmit to Bit.ly the address for each and every web page I visit in a referer header.It's ironic. My last post described how URL shortening services can be abused for evil, but my point was that these abuses were a burden for the services, not that the services were abusive themselves. In fact, Bit.ly has probably done more than any shortening service to combat abuse and the Preview add-on is part of that anti-abuse effort. With Preview installed, users can safely check what's behind any of the short URLs they encounter by hovering over the link in question.
The privacy leak in bit.ly Preview is almost certainly an unintentional product of sloppy coding and deficient testing rather than an effort to spy on the 100,000 users who have installed the add-on. Nonetheless, it's a horrific privacy leak. There are other add-ons that intentionally leak private information, but typically they disclose their activity as a natural part of the add-on's functionality. One example would be GetGlue, which I've written about, and even Bit.ly preview cannot help but leak some info when it's doing what it's supposed to do (expand and preview shortened URLs).
I'm sure that Bit.ly will fix this bug quickly; their support was amazingly fast when I reported another issue. But a larger question remains. How do we make sure that the services we use everyday aren't leaking our info all over the place? The most widely deployed services- Google, Amazon, Facebook, etc. all deserve a higher level of scrutiny because of the quantity of data at their fingertips. All the privacy policies in the world aren't worth a dime if web sites can't be held accountable for the effects of sloppy coding. It's high time for popular sites to submit to strict third-party privacy auditing, and for web users to demand it. It doesn't matter whether any advertisers actually used the personal information that Facebook sent them; what matters is whether users can trust Facebook.
It's also time for the internet technology community to recognize that referer headers are as dangerous to privacy as they are to spelling. They should be abolished. Browser software should stop sending them. The referer header was originally devised to help dispersed server admins fix and control broken links. Today, the referer header is used for "analytics", which is a polite word for "spying". The collection of referer headers helps web sites to "improve their service", but you could say the same of informants and totalitarian governments.
The pipe is rusty- that's why it leaks. We need to fix it.
Posted by
Eric
at
3:44 PM
6
comments
Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest
Labels:
bit.ly,
facebook,
Intel,
linking technology,
privacy
Thursday, April 22, 2010
Facebook vs. Twitter: To Like or To Annotate?
Facebook and Twitter each held developer conferences recently, and the conference names speak worlds about the competing worldviews. Twitter's conference was called "Chirp", while Facebook's conference was labeled "f8" (pronounced "FATE"). Interestingly, both companies used their developer conferences to announce new capability to integrate meaning into their networks.
Facebook's announcement surrounded something it's calling the "Open Graph protocol". Facebook showed its market power by rolling it out immediately with 30 large partner sites that are allowing users to "Like" them on Facebook. Facebook's vision is that web pages representing "real-world things" such as movies, sports teams, products, etc. should be integrated into Facebook's social graph. If you look at the right-hand column of this blog, you'll see an opportunity to "Like" the Blog on Facebook. That action has the effect of adding a connection between a node that represents you on Facebook with a node that represents the blog on Facebook. The Open Graph API extends that capability by allowing the inclusion of web-page nodes from outside Facebook in the Facebook "graph". A webpage just needs to add a bit of metadata into its HTML to tell Facebook what kind of thing it represents.
I've written previously about RDFa, the technology that Facebook chose to use for Open Graph. It's a well designed method for adding machine-readable metadata into HTML code. It's not the answer to all the world's problems, but it can't hurt. When Google announced it was starting to support RDFa last summer, it seemed to be hedging its bets a bit. Not Facebook.
The effect of using RDFa as an interface is to shift the arena of competition. Instead of forcing developers to choose which APIs to support in code, using RDFa asks developers to choose among metadata vocabularies to support their data model. Like Google, Facebook has created its own vocabularies rather than use someone else's. Also, like Google last summer, the documentation for the metadata schemas seems not to have been a priority. Although Facebook has put up a website for Open Graph protocol at http://opengraphprotocol.org/ and a google group at http://groups.google.com/group/open-graph-protocol, there are as yet no topics approved for discussion in the group. [Update- the group is suddenly active, though tightly moderated.]
Nonetheless, websites that support Facebook's metadata will also be making that metadata available to everyone, including Google, putting increased pressure on websites to make available machine readable metadata as the ticket price for being included in Facebook's (or anyone's) social graph. A look at Facebook's list of object types shows their business model very clearly. Here's their "Product and Entertainment" category:
Facebook clearly believes in that fate follows its intelligent design. Twitter, by contrast, believes its destiny will emerge by evolution from a primordial ooze.
At Twitter's "Chirp" conference, Twitter announced that it will add "Annotations" to the core Twitter platform. The description of Twitter annotations is characteristically fuzzy and undetermined. There will be some sort of triple structure, the annotations will be fixed at a tweet's creation, and annotations will have either 512 bytes or maybe 1K. What will it be used for? Who knows?
Last week, I had a chance to talk to Twitter's Chairman and co-Founder Jack Dorsey at another great "Publishing Point" meeting. He boasted about how Twitter users invented hashtags, retweets and "@" references, and Twitter just followed along. Now, Twitter hopes to do the same thing with annotations. Presumably, the Twitter ecosystem will find a use for Tweet annotations and Twitter can then standardize them. Or not. You could conceivably load the Tweet with Open Graph metadata and produce a Facebook "Like" tweet.
Many possibilities for Tweet annotations, underspecified as they are, spring to mind. For example, the Code4Lib list was buzzing yesterday about the possibility that OpenURL references (the kind used in libraries to link to journal articles and books) could be loaded into an annotated tweet. It seems more likely to me that a standard mechanism to point to external metadata, probably expressed as Linked Data, will emerge. A Tweet could use an annotation to point to a web page loaded with RDFa metadata, or perhaps to a repository of item descriptions such as I mentioned in my post on Linked Descriptions. Clearly, it will be possible in some way or other to put real, actionable literature references into a tweet. Whether it will happen, it's hard to say, but I wouldn't hold my breath for Facebook to start adding scientific articles into its social graph.
Although there's a lot of common capability possible between Facebook's Open Graph and Twitter's Annotations, the worldviews are completely different. Twitter clearly sees itself as a communications media and the Annotations as adjuncts to that communication. In the twitterverse, people are entities that tweet about things. Facebook sees its social graph as its core asset and thinks of the graph as being a world-wide web in and of itself. People and things are nodes on a graph.
While Facebook seems offer a lot more to developers than Twitter, I'm not so sure that I like its worldview as much. I'm much more than a node on Facebook's graph.
Facebook's announcement surrounded something it's calling the "Open Graph protocol". Facebook showed its market power by rolling it out immediately with 30 large partner sites that are allowing users to "Like" them on Facebook. Facebook's vision is that web pages representing "real-world things" such as movies, sports teams, products, etc. should be integrated into Facebook's social graph. If you look at the right-hand column of this blog, you'll see an opportunity to "Like" the Blog on Facebook. That action has the effect of adding a connection between a node that represents you on Facebook with a node that represents the blog on Facebook. The Open Graph API extends that capability by allowing the inclusion of web-page nodes from outside Facebook in the Facebook "graph". A webpage just needs to add a bit of metadata into its HTML to tell Facebook what kind of thing it represents.
I've written previously about RDFa, the technology that Facebook chose to use for Open Graph. It's a well designed method for adding machine-readable metadata into HTML code. It's not the answer to all the world's problems, but it can't hurt. When Google announced it was starting to support RDFa last summer, it seemed to be hedging its bets a bit. Not Facebook.
The effect of using RDFa as an interface is to shift the arena of competition. Instead of forcing developers to choose which APIs to support in code, using RDFa asks developers to choose among metadata vocabularies to support their data model. Like Google, Facebook has created its own vocabularies rather than use someone else's. Also, like Google last summer, the documentation for the metadata schemas seems not to have been a priority. Although Facebook has put up a website for Open Graph protocol at http://opengraphprotocol.org/ and a google group at http://groups.google.com/group/open-graph-protocol, there are as yet no topics approved for discussion in the group. [Update- the group is suddenly active, though tightly moderated.]
Nonetheless, websites that support Facebook's metadata will also be making that metadata available to everyone, including Google, putting increased pressure on websites to make available machine readable metadata as the ticket price for being included in Facebook's (or anyone's) social graph. A look at Facebook's list of object types shows their business model very clearly. Here's their "Product and Entertainment" category: albumbookdrinkfoodgamemovieproductsongtv_show
Facebook clearly believes in that fate follows its intelligent design. Twitter, by contrast, believes its destiny will emerge by evolution from a primordial ooze.
At Twitter's "Chirp" conference, Twitter announced that it will add "Annotations" to the core Twitter platform. The description of Twitter annotations is characteristically fuzzy and undetermined. There will be some sort of triple structure, the annotations will be fixed at a tweet's creation, and annotations will have either 512 bytes or maybe 1K. What will it be used for? Who knows?
Last week, I had a chance to talk to Twitter's Chairman and co-Founder Jack Dorsey at another great "Publishing Point" meeting. He boasted about how Twitter users invented hashtags, retweets and "@" references, and Twitter just followed along. Now, Twitter hopes to do the same thing with annotations. Presumably, the Twitter ecosystem will find a use for Tweet annotations and Twitter can then standardize them. Or not. You could conceivably load the Tweet with Open Graph metadata and produce a Facebook "Like" tweet.
Many possibilities for Tweet annotations, underspecified as they are, spring to mind. For example, the Code4Lib list was buzzing yesterday about the possibility that OpenURL references (the kind used in libraries to link to journal articles and books) could be loaded into an annotated tweet. It seems more likely to me that a standard mechanism to point to external metadata, probably expressed as Linked Data, will emerge. A Tweet could use an annotation to point to a web page loaded with RDFa metadata, or perhaps to a repository of item descriptions such as I mentioned in my post on Linked Descriptions. Clearly, it will be possible in some way or other to put real, actionable literature references into a tweet. Whether it will happen, it's hard to say, but I wouldn't hold my breath for Facebook to start adding scientific articles into its social graph.
Although there's a lot of common capability possible between Facebook's Open Graph and Twitter's Annotations, the worldviews are completely different. Twitter clearly sees itself as a communications media and the Annotations as adjuncts to that communication. In the twitterverse, people are entities that tweet about things. Facebook sees its social graph as its core asset and thinks of the graph as being a world-wide web in and of itself. People and things are nodes on a graph.
While Facebook seems offer a lot more to developers than Twitter, I'm not so sure that I like its worldview as much. I'm much more than a node on Facebook's graph.
Related articles
- Facebook's shot across the bow of Google's web: more on the Open Graph (digital.venturebeat.com)
- Facebook's Open Graph is Huge Potential for Who? (pamil-visions.net)
- Everything Facebook Will Announce At f8: The Definitive Guide (dccrowley.posterous.com)
- Ignore Facebook Open Graph at your peril - this is Web 3.0 (thenextweb.com)
Posted by
Eric
at
3:33 PM
5
comments
Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest
Labels:
facebook,
linked data,
OpenURL,
RDFa,
Twitter
Thursday, February 11, 2010
Blog Post Number One Hundred
In eleventh grade one of my English teachers predicted that I would become a writer. I scoffed. I was going to be a scientist, or an engineer. I attributed his prediction to projection- the same sort of thinking that led the minister of our church predict that I would become a preacher of the gospel. Of course, when I was four years old, my ambition was that I would either become a doctor or a garbage man.
I did the scientist thing and the engineer thing, but recently I've become a blogger. When I started, I resolved to write at least ten posts, but believe it or not, this is my 100th. I think that qualifies me to say a bit about the blog as a literary form, although it doesn't qualify me to say anything original. I pity the English Ph. D. student 50 years from now whose dissertation topic is "the blog as a literary form", trying to come up with something original to say. (assuming of course that the Ph. D. dissertation is still extant as a literary form come 50 years.)
The blog is perhaps the first literary form native to the web. It's not a news story, though it can be news. It's something that can't be done in print. For example, a blog post without hyperlinks is like a Superbowl without commercials. It exists as part of a web, like a conversation. My smash-hit article on "Offline Book Lending" could not have existed without a Publisher's Weekly article to bounce off of; as for my posts on Dung Beetle Armament and Dancing Parrots- well what more could you want?
The creative use of multimedia is integral to a good blog post, though much of my subtlety is rarely appreciated. I'm especially proud of the duchampian picture in The Illusion of Internet Identity and the punk rock references in The Rock-Star Librarian and Objective Selector.
The blog post is also the first literary form that is optimized for search engines, with embedded metadata that's integral to the content. I'm proud that I have a post ranking #2 in Google for "hashtags for conferences" and #4 in Google for "bird shit antenna".
Writing as many articles as I have has made me much more aware of their construction. My favorite construction pattern seems to be [odd story]-[dry exposition]-[surprising connection], but I somehow I never intend it that way to start. I also find that my best and most popular posts are the ones that I write quickly; I'm never very happy with the ones where I do a lot of research and work hard on.
If you've ever left a comment, thanks! The opportunity to have my thoughts enriched and corrected by the range of experts who have done so is a real privilege. I've also received some wonderful private comments and links from other blogs, which are deeply appreciated.
Now that the blog has had over 55,000 "unique visitors", I've started to get a few Facebook friend requests from people I don't think I know. I'm sort of old fashioned about only friending people I've met (unless they're cousins!), but it's nice to see people so interested. As a response, I've started a Facebook fan page as a place for blog readers who are active Facebookers to interact.
I'm not exactly sure where this will lead, but so far it's been both fun and worthwhile. I'll leave lucrative to the imaginary future.
That is all.
I did the scientist thing and the engineer thing, but recently I've become a blogger. When I started, I resolved to write at least ten posts, but believe it or not, this is my 100th. I think that qualifies me to say a bit about the blog as a literary form, although it doesn't qualify me to say anything original. I pity the English Ph. D. student 50 years from now whose dissertation topic is "the blog as a literary form", trying to come up with something original to say. (assuming of course that the Ph. D. dissertation is still extant as a literary form come 50 years.)
The blog is perhaps the first literary form native to the web. It's not a news story, though it can be news. It's something that can't be done in print. For example, a blog post without hyperlinks is like a Superbowl without commercials. It exists as part of a web, like a conversation. My smash-hit article on "Offline Book Lending" could not have existed without a Publisher's Weekly article to bounce off of; as for my posts on Dung Beetle Armament and Dancing Parrots- well what more could you want?
The creative use of multimedia is integral to a good blog post, though much of my subtlety is rarely appreciated. I'm especially proud of the duchampian picture in The Illusion of Internet Identity and the punk rock references in The Rock-Star Librarian and Objective Selector.
The blog post is also the first literary form that is optimized for search engines, with embedded metadata that's integral to the content. I'm proud that I have a post ranking #2 in Google for "hashtags for conferences" and #4 in Google for "bird shit antenna".
Writing as many articles as I have has made me much more aware of their construction. My favorite construction pattern seems to be [odd story]-[dry exposition]-[surprising connection], but I somehow I never intend it that way to start. I also find that my best and most popular posts are the ones that I write quickly; I'm never very happy with the ones where I do a lot of research and work hard on.
If you've ever left a comment, thanks! The opportunity to have my thoughts enriched and corrected by the range of experts who have done so is a real privilege. I've also received some wonderful private comments and links from other blogs, which are deeply appreciated.
Now that the blog has had over 55,000 "unique visitors", I've started to get a few Facebook friend requests from people I don't think I know. I'm sort of old fashioned about only friending people I've met (unless they're cousins!), but it's nice to see people so interested. As a response, I've started a Facebook fan page as a place for blog readers who are active Facebookers to interact.
I'm not exactly sure where this will lead, but so far it's been both fun and worthwhile. I'll leave lucrative to the imaginary future.
That is all.
Saturday, January 2, 2010
Ten Predictions for the Next Ten Years

I didn't do so well in 2000 when I made predictions for the coming year; a year later, I determined that only one of my seven predictions came true.
I'm ten years older and wiser, and I guarantee, triple your money back, that at least 3 of this years predictions will come true. In 2000 I didn't have Twitter to try my first draft on.
- The number of public libraries in 2020 will be less than half today's number. Addendum: the number of public library locations will be 50% more in 2020 than today.
I will write a full post about this, but I believe the driving force for this will be e-books and book digitization, and the result will be consolidation, outsourcing and shuttering of public libraries. Update: I've written a full post. - By the end of 2014, the world's largest aggregation of bibliographic metadata will not be WorldCat. By 2020, no one will care which aggregation is largest.
Currently, the growth curve for LibraryThing makes it look like it will pass WorldCat in a few years. SerialsSolutions' Summon is definitely in the running. Google can't be discounted. But by the middle of the decade, the size question will seem silly, sort of like "What's the largest computer chip in the word?" or "Who has the most powerful nuclear bomb?" In 2010, we don't care about these questions. In 2020, data quality and currency will be much more important than data completeness. Also, see my article on "When are you collecting too much data?".
Thanks, @DataG for the comments! - In 2020, general purpose quantum computers will not be useful for any purpose.
If there's one thing I learned from doing physics, it's there ain't no such thing as a free lunch. If you spend a billion dollars on quantum computing, you might be able to factor an unfactorable integer or two by 2020. - Open Linked Data will hockey-stick in 2012 on standardization of of quad (named graphs?) transport.
I've been meaning to write more about quad transport, but if you read my article on Pat Hayes' Surfaces, Leigh Dodds' article on Named Graphs, and the DERI proposal on N-quads, you'll know more than I do. - In 2020, the search engine era will be ending. Search engines will give way to less centralized "knowledge fabrics".
Search engines have a specific topology: spiders pull in data from millions of distributed sites and add it to one big pile that can be searched on. This topology works great if what you want to do is search, but have you ever noticed that Google can't count? Understanding the connections in rapidly changing data will require new topologies and new business models. In 2020, we'll know what they are. - In 2020, China will be seen as having a more modern, sensible, and practical copyright regime than the US.
In 2010, China has a poor reputation enforcement of Copyright. China will certainly mature in this respect, but to expect it to adopt the regime currently prevailing internationally is to ignore the best interests of China. I think that China will look to the original intent of the US Constitution and invent a copyright regime optimized "To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries." - In 2020, more than half of the book industry's revenue will be facilitated by a Book Rights Registry.
The Book Rights Registry that would be created by the Google Books Settlement Agreement is too good of an idea to be tied to the settlement agreement. It will happen whether the settlement is approved or not. People will complain about it... all the way to the bank. Note that my prediction uses the indefinite article. There may be more than one book rights registry! - In 2020, the New York Times will be profitable, and will not have gone bankrupt.
It's easy to predict that the newspaper industry will contract- it's already happening! But the New York Times is uniquely positioned to take advantage of the market gaps that will open when local newspapers fail. Because they do expensive original reporting, they will have little competition. Because they're family-controlled, like Ford, they won't fall victim to the stupidities of the equity markets. - In 2020, Twitter will be a distant memory; Facebook will still be with us.
Facebook has demonstrated ability to purposefully evolve and extend. Twitter seems not to understand itself. While my neighbor David Carr thinks that Twitter Will Endure, his argument applies to the idea, not the company. Twitter the company will be squeezed between multipurpose networks like Facebook on the high end and non-proprietary protocols on on the low end.
Thanks, @CodyBrown for the comments! - On January 1, 2020, when I review this list of predictions, I will use a Mac to do it.
It's been almost 25 years that I've been using a Mac. Do you really think that the mythical Apple tablet of 2020 will not be a Mac?
Friday, September 11, 2009
Public Identity and Bowerbird Privacy
My legal name is "Eric Sven Hellman". On Twitter, I'm using "gluejar". On Facebook, I took the username "eshellman", which I also use in a number of other places. For the most part, while I do try to separate my work from my personal life, I don't try to isolate my online identity from my "real world" identity. I used to use the identity "openly" in some circumstances for my work identity, but I sold that name as part of my previous company. My work identities can be easily connected to my personal identity, and as a result, the use of different identities affords me negligible privacy.
The very concept of privacy has changed a great deal over the last 20 years, in large part due to the internet-induced shrinkage of the world and the relentlessly growing power of large databases. Our traditional notions of privacy have had embedded within them an implicit equation of privacy with obscurity. Public documents with records of where we lived, what we owned, who we were married to and how much our house was worth would be available in the sense that anyone could go to a county registrar and get the information. If I walked to town, anyone who saw me and knew me would know where I was. In the not so distant future, it's not hard to imagine that an internet connected camera could see me, recognize my face, and post my whereabouts on the internet so that anyone searching for me on google could discover my whereabouts. It doesn't really matter whether that happens through Linked Data or by discovering my GPS coordinates on Twitter. As any computer security expert will tell you, security-by-obscurity is ultimately doomed to failure, and I'm pretty sure that the same is true of the privacy-by-obscurity.
Yesterday, Wired Magazine writer Evan Ratliff was found. As part of reporting an article about how hard is for someone to "disappear" in the digital age, Wired had offered $5000 to anyone who could track down Ratliff during 30 days starting August 15. Ratliff's downfall was partly that he "followed" a vegan pizza restaurant in New Orleans on Twitter. It should not be surprising that Ratliff was found, given that a Facebook group with 1,000 members formed in an effort to track him down, so the relevance to our everyday privacy is a bit tenuous. Hollywood celebrities are only too aware that privacy retention is much harder for famous people.
Partly inspired by Ratliff's article, I decided to do a bit of investigation of my own. If you've been reading blogs around the topic of e-book technology, you probably have encountered posts by someone that signs posts with the name "bowerbird". Bowerbird's posts are always on-topic, but they are written with oddly short lines, as if bowerbird was typing on a 40 character-wide terminal. Bowerbird's posts are often impolite and sometimes really insulting, and in several forums, the posts have provoked complaints of trolling or that bowerbird is "hiding behind a pseudonym". Bowerbird replies that "bowerbird" is his real identity "in many versions of reality". It's clear that bowerbird is an iconoclast. When bowerbird posted a comment on one of my recent posts, I decided to see what I could find out about him or her.
It turns out that "bowerbird" is really the first name of "bowerbird intelligentleman". He has used this as his professional name since at least the late eighties. The name is written with lowercase letters in the manner of e e cummings, and Mr. intelligentleman, as the New York Times might refer to him, is a performance poet, among other things. In fact, he claims to have started performance poetry as an art form in 1987, and was an early promoter of "poetry jams". In a charming, self-deprecating bio (PDF), he writes "bowerbird is also one of the world's worst poetry producers" and describes how his forays into computer typesetting of poetry magazines led him into the world of electronic publishing and ebooks. He was very active on the Project Gutenberg volunteer discussion list, where his talent for provocation prompted Marcello Parathoner to cathartically excerpt a collection of his postings. Much of his energy in the ebook arena was spent promoting his ideas about "Zen Markup Language" (z.m.l.) whose philosophy can be summed up as "the best mark-up is no mark-up". The short line endings in bowerbird's post appear to be his insistence on using z.m.l. for his posts. Or perhaps they're performance poetry. It's a cute idea, but personally, I find that the formatting makes the posts hard to read in their context.
When bowerbird posted his comment, he left digital footprints. He visited the blog on a link from LanguageLog. He lives in the Los Angeles area (he's posted elsewhere that he can be found in Santa Monica), uses Verizon DSL, and uses version 4.0 of Safari on the Mac as his browser. The blog uses statcounter.com to monitor usage, so a cookie has been placed in his browser so I can tell if he returns for a visit; bowerbird is able to control these cookies using privacy controls in Safari. DSL lines use a pool of IP addresses, so although I know the IP address he used, I can't use that IP address to persistently track him. However, StatCounter can follow him to other sites that use StatCounter. In principle, StatCounter could report his interest in my blog to other sites and perhaps even connect him to other identities he might have, which would bother me a lot and prompt me to stop using StatCounter.
What's interesting to me is that bowerbird has had an online public identity for over 20 years, and although his entire online life, warts and all, is open for examination (how many of us can say the same?) it appears as though he has successfully walled it off from his private life. Even if I go to the register of deeds in Santa Monica, I probably won't be able to discover whether he owns a house. I can't find out from fundrace if he has donated to a political candidate. I can find his cell phone number because he's chosen to post it, but I don't know anything he hasn't chosen to divulge. (He once owed 1-800-GET-POEM!) In the course of leading a poet's life, bowerbird has been living an experiment in public identity and privacy for 20 years!
I've previously written about the evolution and fluidity of personal names. The use of professional names for public identity is quite common in our society. Women who marry and take their husband's family name routinely retain their names professionally. Use of professional names is particularly common among authors, actors, and musicians. For them, the additional privacy afforded by the use of a professional name is particularly valuable. It strikes me that the separation and isolation of identities may become an essential privacy curtain even for people who aren't celebrities.
It's probably too late for me and most people of my generation. But "bowerbird privacy" could be a reasonable solution for the next generation. A significant number of my son's friends use Facebook under not-their-real-names, and I say more power to them. I think that privacy advocacy organizations should be working to put rules in place to prevent Facebook from enforcing its "only your real name" terms of service and prohibit companies such as Twitter and Google and Yahoo (and StatCounter) from working with ISPs to connect online identities with offline identies.
Nature's bowerbird gets its name from the bower, a structure that male bowerbirds construct to attract females. You might think of it as the bird's public identity. It's not sure why the females are attracted to the bower. Maybe it's privacy?
The very concept of privacy has changed a great deal over the last 20 years, in large part due to the internet-induced shrinkage of the world and the relentlessly growing power of large databases. Our traditional notions of privacy have had embedded within them an implicit equation of privacy with obscurity. Public documents with records of where we lived, what we owned, who we were married to and how much our house was worth would be available in the sense that anyone could go to a county registrar and get the information. If I walked to town, anyone who saw me and knew me would know where I was. In the not so distant future, it's not hard to imagine that an internet connected camera could see me, recognize my face, and post my whereabouts on the internet so that anyone searching for me on google could discover my whereabouts. It doesn't really matter whether that happens through Linked Data or by discovering my GPS coordinates on Twitter. As any computer security expert will tell you, security-by-obscurity is ultimately doomed to failure, and I'm pretty sure that the same is true of the privacy-by-obscurity.
Yesterday, Wired Magazine writer Evan Ratliff was found. As part of reporting an article about how hard is for someone to "disappear" in the digital age, Wired had offered $5000 to anyone who could track down Ratliff during 30 days starting August 15. Ratliff's downfall was partly that he "followed" a vegan pizza restaurant in New Orleans on Twitter. It should not be surprising that Ratliff was found, given that a Facebook group with 1,000 members formed in an effort to track him down, so the relevance to our everyday privacy is a bit tenuous. Hollywood celebrities are only too aware that privacy retention is much harder for famous people.
Partly inspired by Ratliff's article, I decided to do a bit of investigation of my own. If you've been reading blogs around the topic of e-book technology, you probably have encountered posts by someone that signs posts with the name "bowerbird". Bowerbird's posts are always on-topic, but they are written with oddly short lines, as if bowerbird was typing on a 40 character-wide terminal. Bowerbird's posts are often impolite and sometimes really insulting, and in several forums, the posts have provoked complaints of trolling or that bowerbird is "hiding behind a pseudonym". Bowerbird replies that "bowerbird" is his real identity "in many versions of reality". It's clear that bowerbird is an iconoclast. When bowerbird posted a comment on one of my recent posts, I decided to see what I could find out about him or her.It turns out that "bowerbird" is really the first name of "bowerbird intelligentleman". He has used this as his professional name since at least the late eighties. The name is written with lowercase letters in the manner of e e cummings, and Mr. intelligentleman, as the New York Times might refer to him, is a performance poet, among other things. In fact, he claims to have started performance poetry as an art form in 1987, and was an early promoter of "poetry jams". In a charming, self-deprecating bio (PDF), he writes "bowerbird is also one of the world's worst poetry producers" and describes how his forays into computer typesetting of poetry magazines led him into the world of electronic publishing and ebooks. He was very active on the Project Gutenberg volunteer discussion list, where his talent for provocation prompted Marcello Parathoner to cathartically excerpt a collection of his postings. Much of his energy in the ebook arena was spent promoting his ideas about "Zen Markup Language" (z.m.l.) whose philosophy can be summed up as "the best mark-up is no mark-up". The short line endings in bowerbird's post appear to be his insistence on using z.m.l. for his posts. Or perhaps they're performance poetry. It's a cute idea, but personally, I find that the formatting makes the posts hard to read in their context.
When bowerbird posted his comment, he left digital footprints. He visited the blog on a link from LanguageLog. He lives in the Los Angeles area (he's posted elsewhere that he can be found in Santa Monica), uses Verizon DSL, and uses version 4.0 of Safari on the Mac as his browser. The blog uses statcounter.com to monitor usage, so a cookie has been placed in his browser so I can tell if he returns for a visit; bowerbird is able to control these cookies using privacy controls in Safari. DSL lines use a pool of IP addresses, so although I know the IP address he used, I can't use that IP address to persistently track him. However, StatCounter can follow him to other sites that use StatCounter. In principle, StatCounter could report his interest in my blog to other sites and perhaps even connect him to other identities he might have, which would bother me a lot and prompt me to stop using StatCounter.
What's interesting to me is that bowerbird has had an online public identity for over 20 years, and although his entire online life, warts and all, is open for examination (how many of us can say the same?) it appears as though he has successfully walled it off from his private life. Even if I go to the register of deeds in Santa Monica, I probably won't be able to discover whether he owns a house. I can't find out from fundrace if he has donated to a political candidate. I can find his cell phone number because he's chosen to post it, but I don't know anything he hasn't chosen to divulge. (He once owed 1-800-GET-POEM!) In the course of leading a poet's life, bowerbird has been living an experiment in public identity and privacy for 20 years!
I've previously written about the evolution and fluidity of personal names. The use of professional names for public identity is quite common in our society. Women who marry and take their husband's family name routinely retain their names professionally. Use of professional names is particularly common among authors, actors, and musicians. For them, the additional privacy afforded by the use of a professional name is particularly valuable. It strikes me that the separation and isolation of identities may become an essential privacy curtain even for people who aren't celebrities.
It's probably too late for me and most people of my generation. But "bowerbird privacy" could be a reasonable solution for the next generation. A significant number of my son's friends use Facebook under not-their-real-names, and I say more power to them. I think that privacy advocacy organizations should be working to put rules in place to prevent Facebook from enforcing its "only your real name" terms of service and prohibit companies such as Twitter and Google and Yahoo (and StatCounter) from working with ISPs to connect online identities with offline identies.
Nature's bowerbird gets its name from the bower, a structure that male bowerbirds construct to attract females. You might think of it as the bird's public identity. It's not sure why the females are attracted to the bower. Maybe it's privacy?
Posted by
Eric
at
5:12 PM
9
comments
Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest
Labels:
Evan Ratliff,
facebook,
identifiers,
privacy,
public identity,
social practice
Monday, July 13, 2009
Dung Beetle Armament and the Real Threats to Scientific Publishing
To illustrate an article on dung beetle armament, the New York Times Science section published a graphic with a spectacular montage
of 35 animals with grotesque armaments, ranging from the Narwhal to the Giraffe Weevil. 13 of them are extinct. The reason that many dung beetles have evolved such elaborate armaments is not so much that they are effective in combat with other dung beetles, but rather that female dung beetles select mates based upon the outward display of combat fitness.
In my last post, I argued that scholarly publishers were not being threatened by imminent disruption by the same factors that have the newspaper publishing industry on the brink. I suggested that a potential vulnerability of the scholarly publishing industry would be the disintegration of the linkage between the industry's activity, publishing scholarly articles, and the industry's main revenue source- library subscriptions. I see two possible ways that this could occur. It could occur through a collapse of library funding; I hope to discuss that in a future post. This post discusses another way this could occur: I think there is a possibility that the adoption of social networking technologies will lead to a collapse of scholarly publishing as it exists today.
If this sounds a bit far-fetched, consider the parallels between scholarly publishing and dung beetle armament. The development of scholarly publishing today is driven by the selections made by authors about where and how to publish articles. The authors ultimate goal is to propagate their work and thus gain tenure, status and funding, just as "the ultimate goal" of the female dung beetle is to gain a safe tunnel to enjoy dung and raise baby dung beetles. The authors do not really know which journals do the best job of propagating their work, but they recognize prestige and the badges of prestige, and they know what sorts of publications will look best to their tenure committees. Authors do not consider the cost of journals any more than female dung beetles consider the energy cost of male armature. The size and form of today's scholarly publishing ecosystem is thus driven to a significant extent by the superficial judgments of tenure committees.
Anything that might change the way tenure committees, and thus authors, perceive journal publication has the potential to reverse the fortunes of journal publishers. To my mind, social networking technologies have that potential as do few other other things on the horizon. The reason is that tenure decisions have used journal publication records as objective measures of a candidate's social status within the scientific community. Publication in a prestigious journal has been an important way for scholars to become known, to gain speaking invitations, and to advance ideas. But publications are only part of this process. Knowing the right people, studying with the right professors, schmoozing at conferences, all of these are probably more important to the advancement of new ideas, but they have been very hard to measure in any objective way.
Social network technologies open new possibilities for the propagation of new ideas and for the assessment the impact of those ideas. Already, we see people using the number of followers they have on Twitter or the number of recommendations they have on LinkedIn as measures of social status, so it's not much of a stretch to imagine that similar measures could be used to evaluate young academics or to award grants. It's beyond dispute that Twitter is already being widely used to propagate links to interesting technical papers and posts on scientific subjects. If targeted development of social network-based evaluation methodologies were pursued by groups such as the library community who wish to re-inject usage and low-cost access into the tenure equation, the competitive environment for scholarly and scientific publishers could change radically.
Every threat is an opportunity, of course, and it's equally possible that social networking technologies could reinforce the scholarly publishing industry- after all, dung beetle armaments evolve to adapt to changing fashion choices among female dung beetles. A potential weakness- for example, the unwillingness of people to post or retweet links to subscriber-only content, could turn into strengths if publishers develop access models that grant special access for re-tweeted links or an author's Facebook friends. Publishers could also try to ward off challengers in the scholar evaluation game by developing improved and more ostentatious badges of honor- best paper prizes, awards for the most forwarded paper, etc.
In fact, I've come up with a mathematical model for how all this will evolve. First, assume a spherical dung beetle...
of 35 animals with grotesque armaments, ranging from the Narwhal to the Giraffe Weevil. 13 of them are extinct. The reason that many dung beetles have evolved such elaborate armaments is not so much that they are effective in combat with other dung beetles, but rather that female dung beetles select mates based upon the outward display of combat fitness.In my last post, I argued that scholarly publishers were not being threatened by imminent disruption by the same factors that have the newspaper publishing industry on the brink. I suggested that a potential vulnerability of the scholarly publishing industry would be the disintegration of the linkage between the industry's activity, publishing scholarly articles, and the industry's main revenue source- library subscriptions. I see two possible ways that this could occur. It could occur through a collapse of library funding; I hope to discuss that in a future post. This post discusses another way this could occur: I think there is a possibility that the adoption of social networking technologies will lead to a collapse of scholarly publishing as it exists today.
If this sounds a bit far-fetched, consider the parallels between scholarly publishing and dung beetle armament. The development of scholarly publishing today is driven by the selections made by authors about where and how to publish articles. The authors ultimate goal is to propagate their work and thus gain tenure, status and funding, just as "the ultimate goal" of the female dung beetle is to gain a safe tunnel to enjoy dung and raise baby dung beetles. The authors do not really know which journals do the best job of propagating their work, but they recognize prestige and the badges of prestige, and they know what sorts of publications will look best to their tenure committees. Authors do not consider the cost of journals any more than female dung beetles consider the energy cost of male armature. The size and form of today's scholarly publishing ecosystem is thus driven to a significant extent by the superficial judgments of tenure committees.
Anything that might change the way tenure committees, and thus authors, perceive journal publication has the potential to reverse the fortunes of journal publishers. To my mind, social networking technologies have that potential as do few other other things on the horizon. The reason is that tenure decisions have used journal publication records as objective measures of a candidate's social status within the scientific community. Publication in a prestigious journal has been an important way for scholars to become known, to gain speaking invitations, and to advance ideas. But publications are only part of this process. Knowing the right people, studying with the right professors, schmoozing at conferences, all of these are probably more important to the advancement of new ideas, but they have been very hard to measure in any objective way.
Social network technologies open new possibilities for the propagation of new ideas and for the assessment the impact of those ideas. Already, we see people using the number of followers they have on Twitter or the number of recommendations they have on LinkedIn as measures of social status, so it's not much of a stretch to imagine that similar measures could be used to evaluate young academics or to award grants. It's beyond dispute that Twitter is already being widely used to propagate links to interesting technical papers and posts on scientific subjects. If targeted development of social network-based evaluation methodologies were pursued by groups such as the library community who wish to re-inject usage and low-cost access into the tenure equation, the competitive environment for scholarly and scientific publishers could change radically.
Every threat is an opportunity, of course, and it's equally possible that social networking technologies could reinforce the scholarly publishing industry- after all, dung beetle armaments evolve to adapt to changing fashion choices among female dung beetles. A potential weakness- for example, the unwillingness of people to post or retweet links to subscriber-only content, could turn into strengths if publishers develop access models that grant special access for re-tweeted links or an author's Facebook friends. Publishers could also try to ward off challengers in the scholar evaluation game by developing improved and more ostentatious badges of honor- best paper prizes, awards for the most forwarded paper, etc.
In fact, I've come up with a mathematical model for how all this will evolve. First, assume a spherical dung beetle...
Posted by
Eric
at
11:18 AM
0
comments
Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest
Labels:
facebook,
linkedin,
scholarly publishing,
social networks,
Twitter
Wednesday, May 13, 2009
My Pathetic Life (#mpl) and Collaborative Intelligence
I'm still pretty new to Twitter, but it feels pretty familiar to me. Part of the reason for that is that some of my Facebook friends have been parallel posting their status on Facebook and on Twitter. This mostly annoyed me, because Twitterers update their statuses more often than facebookers, and a lot of the Twitter vernacular is totally inexplicable when viewed on Facebook. On Twitter, by contrast, I find that I'm annoyed when people that I follow, but don't really know, mix details from their personal lives into their otherwise interesting Twitter streams. For example, I follow dchud because I know that he will throw out some very interesting ideas. (He was the very first person to follow me on Twitter; I was the very first person to comment on his blog way back when). But I'm not really interested in his reports on the Washington Capitols. Somehow I find that Facebook is a much better place to get to know details like that- if you friend me there, you'll find that I'm a rabid fan of the Philadelphia Phillies, and I won't mind it if you update me on the triumphs of your Columbus Blue Jackets. We both knew they would lose eventually.
The thing that intrigues me about Twitter is that it does so little so well that it's really a lot easier to fix the problems that it has. The past two days I've been writing about the challenges of propagating vocabulary (and grammar for that matter!) for use in the semantic web. Yesterday, Google demonstrated one way of propagating vocabulary- be big and powerful and just tell the world what vocabulary to use. Ian Davis called Google's approach to implementing RDFa "a damp squib" which is what Americans would less colorfully call a "dud" or a wet firecracker. He lamented that Google had chosen to use their own limited vocabulary rather than adopt vocabulary already in use. R.V. Guha, who I mentioned in yesterday's post, commented on Davis' blog that we shouldn't judge too soon what Google is doing. A lot of us are hoping that "igniter fuse" will turn out to be an apter pyrotechnical analogy.
The other strategy for vocabulary propgation is based on community-based collaboration. In my post yesterday, I complained that it was hard to find vocabulary that I might use to attach an ISBN to a resource. In contrast, Twitter, together with the accessories that are built around it, seems to enable rapid propagation of vocabulary and grammar. So back to my complaint about how Twitter streams seem to annoyingly mix tweets of varying interest. One way that people deal with this is to use multiple accounts to organize their tweets in different genres, the same way you might want to have a business email and a personal email. Perhaps a better, more flexible way to address this would be to adopt a special hashtag to signal that a tweet is not a product of one's brilliant intellect, but rather just a status message about "my personal life" (#mpl). That way, your mom can easily filter out your irrelevant work stuff and and your boss can filter out your irrelevant personal stuff. Anyway, if you think this a good idea, see if you can help propagate the use of #mpl (let's call the idea #mplIdea). If you don't think it's a good idea, or if there's some better way to fix this twitter deficiency, leave a comment. Let's see if we can demonstrate the power of the collaborative approach to vocabulary propagation.
The adoption of RDFa by Google and their centralized approach to vocabulary may be a turning point in the first stage of the semantic web- that of using the web to aggregate data. I think this approach is not going to take us very far. We need to start building the second stage of the semantic web. We should be thinking about collaborative intelligence rather than about accumulating distributed sets of data. I don't expect that machines will be able to come up with ideas like #mplIdea, but I do think it's reasonable for machines to be able to help us judge wether ideas like #mplIdea are inspired (and should be propagated) or whether they're just stupid.
The thing that intrigues me about Twitter is that it does so little so well that it's really a lot easier to fix the problems that it has. The past two days I've been writing about the challenges of propagating vocabulary (and grammar for that matter!) for use in the semantic web. Yesterday, Google demonstrated one way of propagating vocabulary- be big and powerful and just tell the world what vocabulary to use. Ian Davis called Google's approach to implementing RDFa "a damp squib" which is what Americans would less colorfully call a "dud" or a wet firecracker. He lamented that Google had chosen to use their own limited vocabulary rather than adopt vocabulary already in use. R.V. Guha, who I mentioned in yesterday's post, commented on Davis' blog that we shouldn't judge too soon what Google is doing. A lot of us are hoping that "igniter fuse" will turn out to be an apter pyrotechnical analogy.
The other strategy for vocabulary propgation is based on community-based collaboration. In my post yesterday, I complained that it was hard to find vocabulary that I might use to attach an ISBN to a resource. In contrast, Twitter, together with the accessories that are built around it, seems to enable rapid propagation of vocabulary and grammar. So back to my complaint about how Twitter streams seem to annoyingly mix tweets of varying interest. One way that people deal with this is to use multiple accounts to organize their tweets in different genres, the same way you might want to have a business email and a personal email. Perhaps a better, more flexible way to address this would be to adopt a special hashtag to signal that a tweet is not a product of one's brilliant intellect, but rather just a status message about "my personal life" (#mpl). That way, your mom can easily filter out your irrelevant work stuff and and your boss can filter out your irrelevant personal stuff. Anyway, if you think this a good idea, see if you can help propagate the use of #mpl (let's call the idea #mplIdea). If you don't think it's a good idea, or if there's some better way to fix this twitter deficiency, leave a comment. Let's see if we can demonstrate the power of the collaborative approach to vocabulary propagation.
The adoption of RDFa by Google and their centralized approach to vocabulary may be a turning point in the first stage of the semantic web- that of using the web to aggregate data. I think this approach is not going to take us very far. We need to start building the second stage of the semantic web. We should be thinking about collaborative intelligence rather than about accumulating distributed sets of data. I don't expect that machines will be able to come up with ideas like #mplIdea, but I do think it's reasonable for machines to be able to help us judge wether ideas like #mplIdea are inspired (and should be propagated) or whether they're just stupid.
Friday, May 8, 2009
Death, another unsolved problem. Wait... Facebook says it's a "Known Problem"
Last September, the library world lost one of its really nice people and great resources. Johan van Halm was one of those people I would see at practically every conference I went to. I remember meeting him a bit more than 10 years ago when I had just started in the library and information business. He went out of his way to introduce me to many of his friends, and I soon realized that he knew just about everyone in the library automation world. We would make it a point to get together for a drink at least once every ALA. I really miss him.
But I'm not writing about Johan today- but I had to get that out of the way. I wanted to write about LinkedIn. You see, Johan is a "connection" of mine on LinkedIn. Every time I look at the list of my connections on LinkedIn, there I see Johan, living on forever in social network space. I could remove him as a connection, of course, but somehow that just doesn't seem right. Also, LinkedIn encourages you to treat your connections and the network they expoe to you as valuable assets and protects them accordingly. In contrast, your Facebook and Twitter "friends" are by default exposed to just about everyone. So on LinkedIn, it would not be to one's advantage to unconnect with someone just because they've left the corporeal parts of this life.
I hope you're sitting down, dear reader, because I have some serious things to discuss with you. Neither you nor I are going to be forever avoid "becoming deceased". However, to make this bleak situation a bit easier to swallow, let's assume, just for fun, that both you and are are going to live forever. In that case, I can pretty much guarantee you that everyone who is following you now on Twitter, all your Facebook Friends, all your LinkedIn Connections, all the friends of your friends, and even all the email accounts that you send jokes to, they will all either pass away into inactivity or become zombies controlled by someone else who is probably not your friend or connection.
Now it might just be that all the social networks we have become so enthralled with have just assumed that they themeselves will have gone belly-up or will have gone through their liquidity event before the death problem becomes severe. Or more likely, its just that they experence such a large volume of accounts fading away into inactivity that having users die is only a small perturbation on their services. But it is certainly the case that as these services become more important to our lives, the fact that we are not immortal increasingly must be addressed.
LinkedIn has done it this way:
Facebook is also ready with a policy:
GMail has an elaborate paper based procedure to access a deceased person's mail:
Part of my interest in the death problem stems from my interest in the Google Book Search settlement. You see, book authors and publishers have been ignoring the death problem for much longer than the Facebooks and Gmails of the world have been. The result is that many works are "orphaned", which means that the rights holders cannot be found or died without leaving instructions or documention about what to do with their intellectual property. It's worse outside the US, because the duration of copyright protection frequently depends on the death date of the author, which can be rather difficult to ascertain. Now that we have the technology to make out-of-print books in libraries generally available through the internet, the corpus of orphan works is once again important, but copyright law presents barriers to many uses, particularly those with economic value less that the cost of finding rights holders and obtaining permissions.
Do you think that perhaps Google is saving its deceased-person access requests for the day 50 years from now when they will become relevant to copyright status? I'll bet the answer is NO.
But I'm not writing about Johan today- but I had to get that out of the way. I wanted to write about LinkedIn. You see, Johan is a "connection" of mine on LinkedIn. Every time I look at the list of my connections on LinkedIn, there I see Johan, living on forever in social network space. I could remove him as a connection, of course, but somehow that just doesn't seem right. Also, LinkedIn encourages you to treat your connections and the network they expoe to you as valuable assets and protects them accordingly. In contrast, your Facebook and Twitter "friends" are by default exposed to just about everyone. So on LinkedIn, it would not be to one's advantage to unconnect with someone just because they've left the corporeal parts of this life.
I hope you're sitting down, dear reader, because I have some serious things to discuss with you. Neither you nor I are going to be forever avoid "becoming deceased". However, to make this bleak situation a bit easier to swallow, let's assume, just for fun, that both you and are are going to live forever. In that case, I can pretty much guarantee you that everyone who is following you now on Twitter, all your Facebook Friends, all your LinkedIn Connections, all the friends of your friends, and even all the email accounts that you send jokes to, they will all either pass away into inactivity or become zombies controlled by someone else who is probably not your friend or connection.
Now it might just be that all the social networks we have become so enthralled with have just assumed that they themeselves will have gone belly-up or will have gone through their liquidity event before the death problem becomes severe. Or more likely, its just that they experence such a large volume of accounts fading away into inactivity that having users die is only a small perturbation on their services. But it is certainly the case that as these services become more important to our lives, the fact that we are not immortal increasingly must be addressed.
LinkedIn has done it this way:
The profile "may need to be removed"?????
What if I see a Profile of someone who is deceased?
Unfortunately, there may be a time when you come across a Profile of a deceased colleague, classmate or connection. If this occurs you are welcome to notify Customer Service that the Profile still exists and may need to be removed. We ask that you provide any important information about the deceased member that may aid our Privacy Department in their investigations and act on the account accordingly. Items to provide in your email would be one or two of the following:
- An Obituary Link.
- A Death Notice.
- Consular Report of Death.
- Death Certificate.
Facebook is also ready with a policy:
Profile: Bugs and Known Problems"Bugs and Known Problems"???? There are some other odd results when you search Facebook help for "death", and following one of these results, you get: a link to Facebook's "Deceased" page.
I’d like to report a deceased user or an account that needs to be memorialized.
Please report this information here so that we can memorialize this person’s account. Memorializing the account removes certain more sensitive information like status updates and restricts profile access to confirmed friends only. Please note that in order to protect the privacy of the deceased user, we cannot provide login information for the account to anyone. We do honor requests from close family members to close the account completely.
GMail has an elaborate paper based procedure to access a deceased person's mail:
Accessing a deceased person's mailI'll bet that even if you have made a last will and testament, you've not included any instructions there about what your executor is to do with your email accounts, your blog passwords, your websites, your social networks. Probably your family might want access to your flickr and youtube account. Also, if you think about it, you don't really want your executor poking around in your e-mails, especially that address you use only for illicit activity.
If an individual has passed away and you need access to the content of his or her mail, please fax or mail us the following information:
Postal Mail:
- Your full name and contact information, including a verifiable email address.
- The Gmail address of the individual who passed away.
- a.The full header from an email message that you have received at your verifiable email address, from the Gmail address in question. (To obtain the header from a message in Gmail, open the message, click 'More options,' then click 'Show original.' Copy everything from 'Delivered-To:' through the 'References:' line. To obtain headers from other webmail or email providers, please refer to http://mail.google.com/support/bin/answer.py?hl=en&answer=22454#) b.The entire contents of the message.
- Proof of death.
- One of the following: a) if the decedent was 18 or older, please provide a proof of authority under local law that you are the lawful representative of the deceased or his or her estate or b) if the decedent was under the age of 18 and you are the parent of the individual, please provide a copy of the decedent’s birth certificate.
Google Inc.
Attention: Gmail User Support
1600 Amphitheatre Parkway
Mountain View, CA 94043
Fax: 650-644-0358 After we've received the above information, we'll need 30 days to process and validate the documents that you've provided. If you need access to the address sooner, in accordance with state and federal law, it is Google's policy to only provide information pursuant to a valid third party court order or other appropriate legal process. Please note that our ability to ability to comply with these requests varies according to applicable law.
Part of my interest in the death problem stems from my interest in the Google Book Search settlement. You see, book authors and publishers have been ignoring the death problem for much longer than the Facebooks and Gmails of the world have been. The result is that many works are "orphaned", which means that the rights holders cannot be found or died without leaving instructions or documention about what to do with their intellectual property. It's worse outside the US, because the duration of copyright protection frequently depends on the death date of the author, which can be rather difficult to ascertain. Now that we have the technology to make out-of-print books in libraries generally available through the internet, the corpus of orphan works is once again important, but copyright law presents barriers to many uses, particularly those with economic value less that the cost of finding rights holders and obtaining permissions.
Do you think that perhaps Google is saving its deceased-person access requests for the day 50 years from now when they will become relevant to copyright status? I'll bet the answer is NO.
Posted by
Eric
at
1:32 PM
2
comments
Email ThisBlogThis!Share to XShare to FacebookShare to Pinterest
Labels:
death,
facebook,
gmail,
Google Book Search,
linkedin,
social networks
Subscribe to:
Posts (Atom)
















