Showing posts with label Twitter. Show all posts
Showing posts with label Twitter. Show all posts

Tuesday, November 4, 2014

Reading Privacy Enables Reader Sharing

Digital privacy is a weird thing. People confuse it for digital security, but it's much more than that. Privacy isn't keeping secrets, it's controlling the information we share. What we think of as privacy depends on trusting that the people we share with won't do bad things. Privacy isn't digital at all. Maybe instead of "digital privacy" we should talk about "digital discretion".

The recent revelations of how Adobe Digital Editions was spewing the users' reading activity, unencrypted, to a logging server are an instructive example of poor digital discretion. I thought that Adobe was working on an ebook synchronization system, but it now looks like ADE was doing the logging "to support new business models" rather than for ebook sync.  It got me thinking about how ebook synchronization can and should be done.

Synchronization is a useful function. I'd like to be able to start reading a book on my iPhone while on the train in the morning, then pick up reading where I left off in the evening using my iPad. But to accomplish this function, I need to trust someone with information that discloses what I'm reading. It's easy to design a centralized sync system that requires a reader to register who they are,  what book they're reading and an activity stream of what pages are being read.

But a sync system designed for privacy doesn't need all that information. The central server doesn't even need to know the identity of the book! As Jason Griffey pointed out in his article on Adobe's spyware, the book's identifier could be hashed with a password, effectively hiding its identity from the central server.

I wrote about how Bluefire is doing sync for their apps while trying their best to respect user privacy. Rather than obscuring the identity of the book, they focus on making it hard to identify users in their system.

Adobe was justifiably criticized for sending lots of information back to its central server without encryption. Although their version 4.0.1 sends less information, mostly Adobe is just encrypting the stream and claiming the privacy problem is solved. The core privacy problem remains- when a DRM ebook is read, an encrypted activity stream is sent back to Adobe. If the information is sensitive or useful, why should Adobe get the benefit of this information at all? At the very least, providing your activity stream to Adobe should be opt-in.

There's a second privacy problem that hasn't been discussed anywhere. It may seem contradictory, but central-server synchronization systems impose TOO MUCH privacy. In many situations, a reader will want to share their reading stream. Look at GoodReads - you can share your opinions with friends. Look at Kobo Reading Life - you get awards and statistics in return for your stream. In classroom situations, students could sync their readers with the instructors'. These sorts of affordances can't be developed without access to the reading-activity stream, and won't work unless everyone participating in the stream is in the same reading ecosystem, using the same central server.

If instead of encrypting the reading-event stream, encryption were applied to the events themselves, the events could be shared over most any messaging system, and distributed according to the user's choices and desired application. In fact, you could use Twitter.

Every user of a Twitter-reading-sync system would create a Twitter feed to publish their reading activity. Other users could subscribe to the event stream. Direct messages could be used to send decryption data for private reading streams. The system could be engineered so that even Twitter would be unable to know what's being read privately. And the whole world would have access to reading that's being done publicly. In addition to page turning,  bookmarking and annotation activity could be of interest.

It's interesting to think about what might happen in a reading ecosystem where readers, not corporations, control the access to their reading activity streams. Publishers and authors might provide incentives to readers who share their reading-events with them. Social networks might match users reading the same page of the same book. Libraries could learn how to meet the needs of their communities. Teachers might be alerted to passages that students find to be difficult. Ironically, these public uses are enabled by a system design which puts a premium on privacy for the reader.

Dave Egger's novel "the Circle" gave us the expression "Privacy is Theft". The novel imagines a social norms that consider privacy to be a reflection of selfishness. But in the real world, it's the lack of discretion by companies building up vast private collections of personal information that's the true threat to social sharing. Too bad that theft is not a crime.

Thursday, March 14, 2013

Twitter Bots are Getting Stranger

I like to see what people are saying about Unglue.it, so I follow the "unglue" tweetstream. For many months, the false positives were really quite entertaining. In the words of normal human twitterers, the word "unglue" unearths all sorts of things people are stuck to and want to be unstuck. Lips, asses, couches, various electronic devices, and of course, Twitter itself. And combinations of two or more of these sticky things. Fun times!

The first bot butting in on my "unglue" stream belonged to some sort of travel agency social marketing bot named Leadify. Various destinations were being touted by this bot under by different users:
@SkyRunTelluride Can’t unglue the kids from the tv on vacations? Go camping in Telluride and leave technology behind.
@VisitGatlinburg Can’t unglue the kids from the tv on vacations? Go camping in Gatlinburg and leave technology behind.
@CrestedButteMt Can’t unglue the kids from the tv on vacations? Go camping in Crested Butte and leave technology behind.
@TimberlineVaca Can’t unglue the kids from the tv on vacations? Go camping in Breckenridge and leave technology behind.
@lodgingdeals Can’t unglue the kids from the tv on vacations? Go camping in Snowmass and leave technology behind.
@vailmountain Can’t unglue the kids from the tv on vacations? Go camping in Vail and leave technology behind.
You get the idea. At least it's clear what "service" this bot is providing.

But recently, a new bot has started getting in on the "unglue" action. What bugs me is I can't figure out why it does what it does:

@IsabellaMariah3 Unglue latent clubhouse derby: .cpP
@MervinSacco1 Unglue official pass250 607-109 audition check over guides: .BCO http://bit.ly/ZMz6fj
@MacDonaldBoswor Dancery unglue rounders online: .Fhi
@ClarenceWither1 Charley 95010 online until unglue straight a rich conjunction unripe forethink: .daw http://bit.ly/WckiGL
@GoldmanLarry1 Unglue fund online casinos: .pyT 525471
@IsaacShade1 Unglue swop 185-113 prelim niagara mopes: .AQb http://bit.ly/10OwHmW
@AllenHoffmann1 Unglue pos software - baksheesh in preference to forthright pos software, guides with acquit pos software: .hFg http://bit.ly/Xv9k0W
@BootmanRussel1 Betting parlor online unglue green stuff extant professional athlete bribe: .YnL 050623
@PassBobby Unglue liquid assets repudiation cash upon be unfaithful online poolroom: .mNG 263810
@GladysSavannah Current unglue contribute nonobservance stationing show biz: .Obd
@EricksonCarter Tavern unglue do tool motion hiatus: .wrm 894279
@JohnsonAllison2 Amusement park unglue participate: .Sby 584823
@BarnesIsabelle Flat reputable volume dvd unglue: .QfH 362486
@JustinCharlie1 261 sporting house unglue tropez aggrandize: .cln
@FrederickWilli4 The hard-and-fast fender in relation with unglue online auction: .RqZ http://bit.ly/Xr2Tw3
@AlanLindley2 How head an neutral advisor grant-in-aid myself pick the uppermost glamour issue unglue racket: .mcY http://bit.ly/W8lrz9
@PaigeCarol1 Thereupon the album's unglue, the belt stirred drummer chouse health resort but salaried nicholas dingley, go ... 121898
@PaigeSandra1 Theater bootlegging unglue surface structure: .nxz 968560
@MiloVelasquez1 332 gambling hall unglue online volutation: .otR
@PeregrinBoles Unglue downloadlot-054exam chamber music guides: .qxf http://bit.ly/10KfJGp
@LeapmanRebecca Entertainment industry unglue contract bridge toad: .iyG 836459
@MakaylaCooper15 Cafe chantant coupon unglue gamut: .eqG 184945
What could possibly be the purpose of these tweets?

My first thought was this is some SEO scam. About 1/3 of the tweets have a bit.ly link to the http://promotion-web.tk/ or related websites; these pages contain more nonsense text under the title "First-class portal". (The "dot .tk" top level domain is a free domain registry based in Amsterdam) But that doesn't make much sense. Why would nonsense tweets point at nonsense websites? And why would most of the tweets come without links?

And if it's an SEO scam, why add things like ".nxz 968560"? Who's going to click on a tweet like that? Even search engines aren't that dumb.

My next thought was that these accounts are those "followers" that social marketing bozos buy for their twitter feeds. But no, many of these these accounts aren't following more than two or three other accounts, though they may have 250 or so of their fellow robots following them, along with a surprising number of apparently human social media consultants.

It's puzzling, and I don't take kindly to unsolved puzzles. This army of zombie twitter accounts must be assembling for some sort of mischief.

So here's my best guess. I think these twitter bots are hiding information in plain view. Suppose you were a terrorist organization or a criminal network, and you needed to publish communications to large numbers of people world wide. What better way to do this than to publish encrypted information on twitter. Or even better, put the encrypted information on a network of websites, and use a distributed network of twitter accounts to distribute the decryption keys? Or maybe this is where Wikileaks is storing its secret files.

The data publishing rate appears to be about 100 tweets per minute, or about 230 bytes/sec. That's 20 MBytes per day. Maybe the three letter codes are the intended recipients, and the 6 digit numbers are constantly changing keys (like RSA's SecurID) for files posted on 2-factor secured websites.

Or maybe its just garbage ungluing latent clubhouse derby.

Notes:
1. Just to be clear, it's not just the word "unglue" that zombie bot is attacking.
2. I can't wait to get head-desked by a simple explanation in the comments.
3. So you don't have to try one of those bitly links yourself, heres a sample of text from one of the garbage websites:
Oneabe is a free online bidding site offering best auctions and known as beat penny auction site , here we conduct Free Online Auction and oneabe is one of the best Online Auction Sites. We also offer free international auction. Presently we are bidding on thunder-Quadband Dual SIM Wifi Touchscreen World and on superb LCD Home theatre media projector and so on. We do our auctions category wise. As here you would see a plethora of options and catalogs within which you can choose whatever is of your choice and need and participate in bidding as well as can buy them. We offer categories like Antiques and art, automobile & bikes, survival kits, businesses for sale, clothing and accessories, coins and collectibles and much more. We are known as a penny auction site worldwide.Under the category of antique and art we offer 20th century antiques ranging from 1920’s till modern , architectural antiques like garden antiques and others, under the wing of Art we offer contemporary art, drawings, paintings, general, photographic images, prints as well as sculptures. We also sell books and manuscripts those are rare and precious. We offer a plethora of ceramic goods and also clocks, decorative items to decorate your home and your office. Our folk art is very unique and our foreign arts are all master pieces. We also do bidding on furniture, map or atlas as well as on metal ware such as brass, copper, bronze, gold, silver, and silver plated goods also we sell music instruments. We also offer here to our customers a very good quality of textiles and linens that includes fabric, embroidery, linens and quilts and much more. Under the gaming option we offer...
4. (update) More theories being discussed on Hacker News https://news.ycombinator.com/item?id=5373161 
Enhanced by Zemanta

Wednesday, November 10, 2010

Infochimps and the scaling of dataset value

Image representing Infochimps as depicted in C...Image via CrunchBaseSure, a picture is worth a thousand words, but what is a thousand words worth? How about a million? If I had a dataset of the most recent trillion words spoken by humanity, (anonymized and randomized of course!) would that be worth any more than the set of words in this blog post?

These are real questions. A Texas company called Infochimps has datasets quite similar to these, ready for you to use. Some of the datasets are free, others you have to pay for. More interesting is that if you have a dataset you think other people might be interested in, or even pay for, InfoChimps will host it for you and help you find customers. (Infochimps just announced they had raised $1.2 million in its first round of institutional funding.)

One of the datasets you can get from Infochimps for free is the set of smileys used on twitter in tweets sent between March 2006 and November 2009. It's free. It tells you that the smiley ":)" was used 13,458,831 times, while ";-}" was only used 1,822 times.

If you're willing to fork over $300, you can get a 160MB file conatining a month-by-month summary of all the hashtags, URLs and smiley's used on twitter during the same period. That dataset wil tell you that during September of 2009, the hashtag #kanyeisagayfish was used 11 times while #takekanyeinstead was used 141 times.

If you're a scrabble player, you can spend $4 for a list of the 113,809 official words, with definitions. Or you can get them free, without the definitions.

courtesy of Infochimps, Inc. CC-BY-A
I had a great talk with Infochimps President and Co-Founder Flip Kromer a few weeks ago before his presentation to the New York Data Visualization Meetup. I fell in love with one of the visualizations he showed in his presentation, and he's given me permission to reproduce it here. (Creative Commons Attribution License) It's derived from the same Twitter data set you can get from Infochimps, and shows networks of characters that are found in the same tweet. So if ♠ and ♣ appear in the same tweet over and over again, the two characters will have a strong connection in the network of characters.

The character connection data was fed to a program called Cytoscape, which is an open source visualization program used in bioinformatics; Mike Bergmann has a nice article about its use for large RDF graphs. The networks are laid out using a force-directed algorithm (which is pretty much the simplest thing you can do). Coloring is applied arbitrarily.

As you might expect, the main character networks that show up are associated with languages, but there are some anomalies. For example, the katakana character ツ (tu) sticks out. Katakana is a set of phonetic characters used in Japanese for non-Japanese words. The reason "tu" is set apart from all the other katakana is that people use it on Twitter as a smiley.

The other anomalous character subnet is labeled "???" in the graph. A closer look reveals this to be the set of characters that look like upside down roman text.

Kromer has noticed that the price (or perhaps cost) of a partial data set follows a non-monotonic curve (see graphic). Small amounts of data are essentially free, but a peak value is reached when portions of the data set are extracted from the full data set. If we were discussing book metadata, for example, peak value might accrue for a set of the 100,000 top selling books.

There's much less value, according to Kromer, in having a large incomplete chunk of a data set. Data for 10,000,000 books, for example, would have less value than the 100,000 book data set, because it's not complete. Complete data sets become extremely expensive because of the logistics involved, and because of the value of having the complete set.

This pattern seems plausible to me, but I'd like to see some clearer examples. I've previously written about having too much data, but that article looked at the effect of error rates on data collection; Kromer's curve is about utility.

For me, the most interesting thing about Infochimps is the idea that the best way to make data flow in large volumes and create new types of knowledge is to provide the right incentives for data producers through the establishment of a market. This makes a lot of sense to me; however I'm not sure that the Infochimps market has also established incentives needed for data set maintenance; the world's most valuable and expensive data sets are one that change rapidly.

Kromer contrasted the Infochimps approach to that of Wolfram, whose Alpha service is produced by "putting 100 PhDs and data in a lab". He also feels that much of the work being put into the semantic web is a "crock" because its technology stack solves problems that we don't have. Humans are pretty good at extracting meaning from data, given a good visualization.

We can even recognize upside-down text.
Enhanced by Zemanta

Thursday, April 22, 2010

Facebook vs. Twitter: To Like or To Annotate?

Facebook and Twitter each held developer conferences recently, and the conference names speak worlds about the competing worldviews. Twitter's conference was called "Chirp", while Facebook's conference was labeled "f8" (pronounced "FATE"). Interestingly, both companies used their developer conferences to announce new capability to integrate meaning into their networks.

Facebook's announcement surrounded something it's calling the "Open Graph protocol". Facebook showed its market power by rolling it out immediately with 30 large partner sites that are allowing users to "Like" them on Facebook. Facebook's vision is that web pages representing "real-world things" such as movies, sports teams, products, etc. should be integrated into Facebook's social graph. If you look at the right-hand column of this blog, you'll see an opportunity to "Like" the Blog on Facebook. That action has the effect of adding a connection between a node that represents you on Facebook with a node that represents the blog on Facebook. The Open Graph API extends that capability by allowing the inclusion of web-page nodes from outside Facebook in the Facebook "graph". A webpage just needs to add a bit of metadata into its HTML to tell Facebook what kind of thing it represents.

I've written previously about RDFa, the technology that Facebook chose to use for Open Graph. It's a well designed method for adding machine-readable metadata into HTML code. It's not the answer to all the world's problems, but it can't hurt. When Google announced it was starting to support RDFa last summer, it seemed to be hedging its bets a bit. Not Facebook.

The effect of using RDFa as an interface is to shift the arena of competition. Instead of forcing developers to choose which APIs to support in code, using RDFa asks developers to choose among metadata vocabularies to support their data model. Like Google, Facebook has created its own vocabularies rather than use someone else's. Also, like Google last summer, the documentation for the metadata schemas seems not to have been a priority. Although Facebook has put up a website for Open Graph protocol at http://opengraphprotocol.org/ and a google group at http://groups.google.com/group/open-graph-protocol, there are as yet no topics approved for discussion in the group. [Update- the group is suddenly active, though tightly moderated.]

Nonetheless, websites that support Facebook's metadata will also be making that metadata available to everyone, including Google, putting increased pressure on websites to make available machine readable metadata  as the ticket price for being included in Facebook's (or anyone's) social graph. A look at Facebook's list of object types shows their business model very clearly. Here's their "Product and Entertainment" category:
  • album
  • book
  • drink
  • food
  • game
  • movie
  • product
  • song
  • tv_show
Whether you "Like" it or not, Facebook is creating a new playing field for advertising by accepting product pages into their social graph.

Facebook clearly believes in that fate follows its intelligent design. Twitter, by contrast, believes its destiny will emerge by evolution from a primordial ooze.

At Twitter's "Chirp" conference, Twitter announced that it will add "Annotations" to the core Twitter platform. The description of Twitter annotations is characteristically fuzzy and undetermined. There will be some sort of triple structure, the annotations will be fixed at a tweet's creation, and annotations will have either 512 bytes or maybe 1K. What will it be used for? Who knows?

Last week, I had a chance to talk to Twitter's Chairman and co-Founder Jack Dorsey at another great "Publishing Point" meeting. He boasted about how Twitter users invented hashtags, retweets and "@" references, and Twitter just followed along. Now, Twitter hopes to do the same thing with annotations. Presumably, the Twitter ecosystem will find a use for Tweet annotations and Twitter can then standardize them. Or not. You could conceivably load the Tweet with Open Graph metadata and produce a Facebook "Like" tweet.

Many possibilities for Tweet annotations, underspecified as they are, spring to mind. For example, the Code4Lib list was buzzing yesterday about the possibility that OpenURL references (the kind used in libraries to link to journal articles and books) could be loaded into an annotated tweet. It seems more likely to me that a standard mechanism to point to external metadata, probably expressed as Linked Data, will emerge. A Tweet could use an annotation to point to a web page loaded with RDFa metadata, or perhaps to a repository of item descriptions such as I mentioned in my post on Linked Descriptions. Clearly, it will be possible in some way or other to put real, actionable literature references into a tweet. Whether it will happen, it's hard to say, but I wouldn't hold my breath for Facebook to start adding scientific articles into its social graph.

Although there's a lot of common capability possible between Facebook's Open Graph and Twitter's Annotations, the worldviews are completely different. Twitter clearly sees itself as a communications media and the Annotations as adjuncts to that communication. In the twitterverse, people are entities that tweet about things. Facebook sees its social graph as its core asset and thinks of the graph as being a world-wide web in and of itself. People and things are nodes on a graph.

While Facebook seems offer a lot more to developers than Twitter, I'm not so sure that I like its worldview as much. I'm much more than a node on Facebook's graph.
Reblog this post [with Zemanta]

Thursday, April 1, 2010

License Agreement for Go To Hellman Blog

PLEASE READ THIS LICENSE AGREEMENT ("LICENSE") CAREFULLY BEFORE USING THE GO TO HELLMAN BLOG ("THE BLOG").  BY USING THE BLOG, YOU ARE AGREEING TO BE BOUND BY THE TERMS OF THIS LICENSE. IF YOU DO NOT AGREE TO THE TERMS OF THIS LICENSE, CEASE READING AT ONCE. IF YOU DO NOT AGREE TO THE TERMS OF THE LICENSE, YOU MAY APPLY FOR A REFUND. IF THE BLOG WAS ACCESSED ELECTRONICALLY, CLICK "DISAGREE/DECLINE" BELOW. IF NOT, YOU MUST RETURN THE ENTIRE BLOG PACKAGE IN ORDER TO OBTAIN A REFUND.

IMPORTANT NOTE: This Blog may be used to induce synapse connection patterns in neural networks. It is licensed to you only for induction of non-copyrighted connection patterns, connection patterns in which you own the copyright, or connection patterns you are authorized or legally permitted to induce or have induced. This Blog may also be used for linking to music and video files for listening or viewing via your reading/listening/viewing device. Linked access of copyrighted music or video is only provided for lawful personal use or as otherwise legally permitted. If you are uncertain about your right to access to any linked material you should contact your legal advisor at once. 

1. General. The Blog, commentary and any graphics accompanying this License whether transmitted electronically, resident on disk or in memory, on any other media or in any other form (collectively the "Go To Hellman Blog") are licensed, not sold, to you by Gluejar Inc. ("Gluejar") for use only under the terms of this License, and Gluejar reserves all rights not expressly granted to you. The rights granted herein are limited to Gluejar's and its licensors' intellectual property rights in the Go To Hellman Blog and do not include any other patents or intellectual property rights. You may own the chair you are sitting on, subject to other ownership claims, and nothing else remotely connected to this Blog. The terms of this License will govern any correction or additional posts provided by Gluejar that replace and/or supplement the original Gluejar product, unless such changes are accompanied by a separate license in which case the terms of that license will govern.  

2. Permitted License. Uses and Restrictions. This License allows you to read and mentally process the Go To Hellman Blog. The Blog may be used to inspire new thoughts so long as such use is limited to thoughts which are lawful in your country of domicile. You may not make the Blog available over a network except as permitted under this License. You may make one copy of the Blog in machine-readable form for backup purposes only; provided that the backup copy must include all copyright or other proprietary notices contained on the original. You may also extract quotes, limited to forty words or less, and include them in your lawfully constructed derivative works, always providing that this License accompanies such derivative works. Except as and only to the extent expressly permitted in this License or by applicable law, you may not copy, decompile, reverse engineer, disassemble, modify, or create derivative works of the Blog or any part thereof. THE GO TO HELLMAN BLOG IS NOT INTENDED FOR USE IN THE OPERATION OF NUCLEAR FACILITIES, AIRCRAFT NAVIGATION OR COMMUNICATION SYSTEMS, AIR TRAFFIC CONTROL SYSTEMS, LIFE SUPPORT MACHINES OR OTHER EQUIPMENT IN WHICH THE SILLY THINGS THE BLOG SAYS COULD LEAD TO DEATH, PERSONAL INJURY, OR SEVERE PHYSICAL OR ENVIRONMENTAL DAMAGE.  

3. Hyperlinking. You may not rent, lease, lend, redistribute or sublicense the Go To Hellman Blog. You may, however, make a one-time permanent hyperlink of all of your license rights to the Go To Hellman Blog to another party, provided that: (a) the hyperlink must include all of the Blog, including all its component parts, original media, printed materials and this License; (b) you do not retain any copies of the Blog, full or partial, including copies stored on a computer or other storage device;  and (c) the party activating the hyperlinked Blog reads and agrees to accept the terms and conditions of this License.  

4. Consent to Use of Data. You agree that Gluejar and its millions of subsidiaries, present and future, real and imagined, may collect and use technical and related information, including but not limited to technical information about your computer, system and application software, and peripherals, that is gathered periodically to facilitate the embarrassment of bashful users. Gluejar may use this information, as long as it is in a form that does not personally identify you or your grotesque bodily deformities, to improve our stories or to provide services or technologies to you.  

5. Other Services. You understand that by using this Blog, you may encounter Content that may be deemed offensive, indecent, or objectionable, which content may or may not be identified as having explicit language. Nevertheless, you agree to use the Blog at your sole risk and that Gluejar shall have no liability to you for content that may be found to be offensive, indecent, objectionable, obscene, pornographic or just plain stupid. Content types (including tags, links, videos, comments and the like) and descriptions are provided for convenience, edification and amusement, and you acknowledge and agree that Gluejar does not guarantee their accuracy or grammar.

Certain Content may include materials from third parties or links to certain third party web sites. You acknowledge and agree that Gluejar is not responsible for examining or evaluating the content or accuracy of any such third-party material or web sites, notwithstanding that it's pretty much the whole point. Gluejar does not warrant or endorse and does not assume and will not have any liability or responsibility for any third-party materials or web sites, or for any other materials, products, or services of third parties. Links to other web sites are provided solely as a convenience to you and your chattel. You agree that you will not use any third-party, fourth party, or even nth-party materials in a manner that would infringe or violate the rights or put a damper on any other party and that Gluejar is not in any way responsible for any such use by you, especially insofar as we have not been invited to such parties.  

6. Termination. This License is effective until terminated. Your rights under this License will terminate automatically without notice from Gluejar if you fail to comply with any term(s) of this License. Upon the termination of this License, you shall cease all use of the Blog and delete all copies of the Blog in your possession as well as all derivative works you have made using quotes from the Blog.  

7. Disclaimer of Warranties. YOU EXPRESSLY ACKNOWLEDGE AND AGREE THAT USE OF THE BLOG (AS DEFINED ABOVE) IS AT YOUR SOLE RISK AND THAT THE ENTIRE RISK AS TO SATISFACTORY QUALITY, PERFORMANCE, ACCURACY AND EFFORT IS WITH YOU. TO THE MAXIMUM EXTENT PERMITTED BY APPLICABLE LAW, THE GO TO HELLMAN BLOG IS PROVIDED "AS IS", WITH ALL FAULTS AND WITHOUT WARRANTY OF ANY KIND, AND GLUEJAR AND GLUEJAR'S LICENSORS (COLLECTIVELY REFERRED TO AS "GLUEJAR" FOR THE PURPOSES OF SECTIONS 7 AND 8) HEREBY DISCLAIM ALL WARRANTIES AND CONDITIONS WITH RESPECT TO THE GO TO HELLMAN BLOG, EITHER EXPRESS, IMPLIED OR STATUTORY, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES AND/OR CONDITIONS OF READABILITY OR INTELLIGIBILITY, OF SATISFACTORY QUALITY, OF FITNESS FOR A PARTICULAR PURPOSE, OF ACCURACY, OF QUIET ENJOYMENT, AND NON-INFRINGEMENT OF THIRD PARTY RIGHTS. GLUEJAR DOES NOT WARRANT AGAINST INTERFERENCE WITH YOUR ENJOYMENT OF THE GO TO HELLMAN BLOG, THAT THE FUNCTIONS CONTAINED IN THE GO TO HELLMAN BLOG WILL MEET YOUR REQUIREMENTS, THAT THE OPERATION OF THE GO TO HELLMAN BLOG WILL BE UNINTERRUPTED OR ERROR-FREE, OR THAT DEFECTS IN THE GO TO HELLMAN BLOG, NUMEROUS AND EGREGIOUS AS THEy MAY BE, WILL BE CORRECTED. NO ORAL OR WRITTEN INFORMATION OR ADVICE GIVEN BY GLUEJAR OR AN GLUEJAR AUTHORIZED REPRESENTATIVE SHALL CREATE A WARRANTY. SHOULD THE GO TO HELLMAN BLOG PROVE DEFECTIVE, YOU ASSUME THE ENTIRE COST OF ALL NECESSARY SERVICING, REPAIR OR CORRECTION. SOME JURISDICTIONS DO NOT ALLOW THE EXCLUSION OF IMPLIED WARRANTIES OR LIMITATIONS ON APPLICABLE STATUTORY RIGHTS OF A CONSUMER, SO THE ABOVE EXCLUSION AND LIMITATIONS MAY NOT APPLY TO YOU BUT WE REALLY DONT GIVE A DAMN.  

8. Limitation of Liability. TO THE EXTENT NOT PROHIBITED BY LAW, IN NO EVENT SHALL GLUEJAR BE LIABLE FOR PERSONAL INJURY, OR ANY INCIDENTAL, SPECIAL, INDIRECT OR CONSEQUENTIAL DAMAGES WHATSOEVER, INCLUDING, WITHOUT LIMITATION, DAMAGES FOR LOSS OF PROFITS, LOSS OF DATA, DISTRACTION FROM MORE IMPORTANT THINGS, BUSINESS INTERRUPTION OR ANY OTHER COMMERCIAL DAMAGES OR LOSSES, ARISING OUT OF OR RELATED TO YOUR USE OR INABILITY TO USE THE GO TO HELLMAN BLOG, HOWEVER CAUSED, REGARDLESS OF THE THEORY OF LIABILITY (CONTRACT, TORT OR OTHERWISE) AND EVEN IF GLUEJAR HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES. SOME JURISDICTIONS DO NOT ALLOW THE LIMITATION OF LIABILITY FOR PERSONAL INJURY, OR OF INCIDENTAL OR CONSEQUENTIAL DAMAGES, SO THIS LIMITATION MAY NOT APPLY TO YOU. In no event shall Gluejar's total liability to you for all damages (other than as may be required by applicable law in cases involving personal injury) exceed the amount of fifty cents ($0.50). The foregoing limitations will apply even if the above stated remedy fails of its essential purpose.  

9. Trademarks. “Gluejar”, “Go To Hellman”, “Hellman” and "ebook" are trademarks/service marks of Gluejar, Inc.  Gluejar and/or its suppliers own all rights, title and interest, including, without limitation, all intellectual property rights, in and to the Blog.  Except as expressly provided for in these Terms, You acquire no rights in or to Gluejar Trademarks. If your name happens to be "Hellman", doesn't that really suck? Lillian, you're dead, so stop twittering. Warren, wouldn't "We Shall Be Friedman" be a much better name for your blog? And did Frances ever tell you I inherited her desk at Ginzton Lab? Speaking of Twitter, don't you hate mayonnaise misspellers? The Best has two Ns, idiots! Marty, you can have my second key any time, keep up the good work. Jakob, do a new album already. Lauren, may you rot in Hel for not doing Googling your blog name first.  

10. Export Control. You may not use or otherwise export or reexport the Go To Hellman Blog except as authorized by United States law and the laws of the jurisdiction in which the Go To Hellman Blog was obtained. In particular, but without limitation, the Go To Hellman Blog may not be exported or re-exported (a) into any U.S. embargoed countries or (b) to anyone on the U.S. Treasury Department's list of Specially Designated Nationals or the U.S. Department of Commerce Denied Person’s List or Entity List. By using the Go To Hellman Blog, you represent and warrant that you are not located in any such country or on any such list. You also agree that you will not use this Blog for any purposes prohibited by United States law, including, without limitation, the development, design, manufacture or production of missiles, or nuclear, chemical or biological weapons.  

11. Government End Users. The Go To Hellman Blog and related commentary are "Commercial Items", as that term is defined at 48 C.F.R. §2.101, as such terms are used in 48 C.F.R. §12.212 or 48 C.F.R. §227.7202, as applicable.  Consistent with 48 C.F.R. §12.212 or 48 C.F.R. §227.7202-1 through 227.7202-4, as applicable, the Blog is being licensed to U.S. Government end users (a) only as Commercial Items (b) with only those rights as are granted to all other end users pursuant to the terms and conditions herein and (c) can you believe this inpenetrable language is contained in the iTunes License agreement that you probably accepted this week.  

12. Controlling Law and Severability. This License will be governed by and construed in accordance with the laws of the State of New Jersey. If you gotta problem with that, I know a guy. This License is Born in the USA and shall not be governed by the United Nations Convention on Contracts for the International Sale of Goods, the application of which is expressly excluded by Bruce Himself. If for any reason a court of competent jurisdiction finds any provision, or portion thereof, to be unenforceable, the remainder of this License shall continue in full force and effect, so go stuff it.  

13. Complete Agreement; Governing Language. This License constitutes the entire agreement between the parties with respect to the use of the Go To Hellman Blog licensed hereunder and supersedes all prior or contemporaneous understandings regarding such subject matter. No amendment to or modification of this License will be binding unless in writing and signed by Gluejar. Any translation of this License is done for local requirements and in the event of a dispute between the English and any non-English versions, the English version of this License shall govern, because we're Americans and what other language do you expect us to understand. In particular, any declarations made in robots.txt or meta tag assertions files are null and void. Googlebot, this means you.

    Last updated, April 1, 2010.







    Notwithstanding the foregoing, CC BY-SA works just fine.

    Saturday, January 2, 2010

    Ten Predictions for the Next Ten Years


    I didn't do so well in 2000 when I made predictions for the coming year; a year later, I determined that only one of my seven predictions came true.

    I'm ten years older and wiser, and I guarantee, triple your money back, that at least 3 of this years predictions will come true. In 2000 I didn't have Twitter to try my first draft on.
    1. The number of public libraries in 2020 will be less than half today's number. Addendum: the number of public library locations will be 50% more in 2020 than today.

      I will write a full post about this, but I believe the driving force for this will be e-books and book digitization, and the result will be consolidation, outsourcing and shuttering of public libraries. Update: I've written a full post.

    2. By the end of 2014, the world's largest aggregation of bibliographic metadata will not be WorldCat. By 2020, no one will care which aggregation is largest.

      Currently, the growth curve for LibraryThing makes it look like it will pass WorldCat in a few years. SerialsSolutions' Summon is definitely in the running. Google can't be discounted. But by the middle of the decade, the size question will seem silly, sort of like "What's the largest computer chip in the word?" or "Who has the most powerful nuclear bomb?" In 2010, we don't care about these questions. In 2020, data quality and currency will be much more important than data completeness. Also, see my article on "When are you collecting too much data?".

      Thanks, @DataG for the comments!

    3. In 2020, general purpose quantum computers will not be useful for any purpose.

      If there's one thing I learned from doing physics, it's there ain't no such thing as a free lunch. If you spend a billion dollars on quantum computing, you might be able to factor an unfactorable integer or two by 2020.

    4. Open Linked Data will hockey-stick in 2012 on standardization of of quad (named graphs?) transport.

      I've been meaning to write more about quad transport, but if you read my article on Pat Hayes' Surfaces, Leigh Dodds' article on Named Graphs, and the DERI proposal on N-quads, you'll know more than I do.

    5. In 2020, the search engine era will be ending. Search engines will give way to less centralized "knowledge fabrics".

      Search engines have a specific topology: spiders pull in data from millions of distributed sites and add it to one big pile that can be searched on. This topology works great if what you want to do is search, but have you ever noticed that Google can't count? Understanding the connections in rapidly changing data will require new topologies and new business models. In 2020, we'll know what they are.

    6. In 2020, China will be seen as having a more modern, sensible, and practical copyright regime than the US.

      In 2010, China has a poor reputation enforcement of Copyright. China will certainly mature in this respect, but to expect it to adopt the regime currently prevailing internationally is to ignore the best interests of China. I think that China will look to the original intent of the US Constitution and invent a copyright regime optimized "To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."

    7. In 2020, more than half of the book industry's revenue will be facilitated by a Book Rights Registry.

      The Book Rights Registry that would be created by the Google Books Settlement Agreement is too good of an idea to be tied to the settlement agreement. It will happen whether the settlement is approved or not. People will complain about it... all the way to the bank. Note that my prediction uses the indefinite article. There may be more than one book rights registry!

    8. In 2020, the New York Times will be profitable, and will not have gone bankrupt.

      It's easy to predict that the newspaper industry will contract- it's already happening! But the New York Times is uniquely positioned to take advantage of the market gaps that will open when local newspapers fail. Because they do expensive original reporting, they will have little competition. Because they're family-controlled, like Ford, they won't fall victim to the stupidities of the equity markets.

    9. In 2020, Twitter will be a distant memory; Facebook will still be with us.

      Facebook has demonstrated ability to purposefully evolve and extend. Twitter seems not to understand itself. While my neighbor David Carr thinks that Twitter Will Endure, his argument applies to the idea, not the company. Twitter the company will be squeezed between multipurpose networks like Facebook on the high end and non-proprietary protocols on on the low end.

      Thanks, @CodyBrown for the comments!

    10. On January 1, 2020, when I review this list of predictions, I will use a Mac to do it.

      It's been almost 25 years that I've been using a Mac. Do you really think that the mythical Apple tablet of 2020 will not be a Mac?

    Reblog this post [with Zemanta]

    Friday, November 13, 2009

    The New York Times Gets It Right; Does Linked Data Need a CrossRef or an InfoChimps?

    I've been saying this long enough that I don't remember whether I was quoting someone else: whenever the internet disintermediates a middleman, two new intermediaries pop up somewhere else. It's disintermediation whack-a-mole, if you will. The reasons for this are:
    1. The old middlemen became fat on mark-ups an order of magnitude larger than needed by internet-enabled middlemen.
    2. Internet-enabled middlemen add value in ways that the old ones didn't.
    My last business functioned as an intermediary that aggregated linking data. We'd get data from publishers, clean it up and add it to our collection, then provide feeds of that data to our customers (libraries and library systems vendors). Our customers got good data and support if was a problem. The companies who provided the data didn't have to deal with hundreds of libraries or system vendors, and they came to understand that we would help their customers link to their content.

    Some companies, especially the large ones, were initially uncomfortable with the knowledge that we were selling feeds of data that they were giving out for free. They felt that somehow there was money left on the table. Other companies were fearful of losing control of the information, even though they didn't really have control of it in the first place. Once we explained to them how their data contained mangled character encodings, fictitious identifiers, stray column separators and Catalan month names, they began to see the value we provided.

    While my company focused on the data needs of libraries (and did pretty well), a group of the largest academic publishers put up some money and formed a consortium to pool a different type of linking data in a way that let the publishers have more control of the data distribution. This consortium, known as Crossref, just celebrated its 10th anniversary. Crossref has not only paid back the money that its founders invested in it; it has arguably done more to push academic publishing into the 21st century than any other organization on the planet.

    As academic publishing companies began to understand the benefits of distributing linking data through Crossref, my company, and others like it, they became more comfortable opening up their content and reaping the financial benefits. Despite the global recession, and despite predictions of its impending collapse, STM publishing has been financially healthy with companies such as Elsevier reporting increased profits. This is rather unlike the newspaper industry, for example.

    Before I get to the newspaper industry, I should note yesterday's news that InfoChimps are publishing a collection of token data harvested from Twitter.
    Today we are publishing a few items collected from our large scrape of Twitter’s API. The data was collected, cleaned, and packaged over twelve months and contains almost the entire history of Twitter: 35 million users, one billion relationships, and half a billion Tweets, reaching back to March 2006.
    InfoChimps is positioning itself as a marketplace to buy, sell, and share data sets of any size, topic or format. Yet another intermediary has popped up!

    Two weeks ago, I wrote a somewhat alarmist article about problems in an exciting set of Linked Data being released by the New York Times. I am pleased to be able to be report that the New York Times is now getting it right! The most important thing that they're doing right is that they're listening to the people who want to consume their data. They've started a Google Group based community for the specific purpose of understanding how best to deliver their data. They've also corrected the problems pointed out by myself and others. It's not perfect, but it's not reasonable to expect perfect. The New York Times has set a very hopeful example for other companies that want to start publishing semantic linking information on the open web.

    If, as many of us hope, many publishers decide to follow the lead of the Times and make more data collections available, will more intermediaries such as InfoChimps arise to facilitate data distribution, as happened with linking data in scholarly publishing? Will ad hoc groups such as "the Pedantic Web" become key participants in a less centralized data distribution environment? Or maybe large companies will turn off the spigots as "the suits" grow increasingly worried about their ability to control data once it is let out into the web of data.

    Perhaps the time is ripe for a set of forward-looking publishers to emulate the nervous-but-smart journal publishers who started Crossref 10 years ago and start a similar consortium for the distribution of Linked Data.
    Reblog this post [with Zemanta]

    Sunday, October 11, 2009

    Wave is Better on the iPhone

    The more I play with Google Wave, the less I like the user interface. In contrast, the more I learn about the Wave technology stack, the less I care about the user interface, because it's clear to me that the important innovations are underneath the skin. That became very clear to me when I got on the train on Friday morning and decided to see whether Wave worked on the iPhone. I've decided that wave is better on the iPhone than it is in a full screen browser.

    On a full screen browser, Wave takes up 3 columns. When you get to try Wave, the first thing you should do is get rid of the leftmost column. This gives the rest of Wave a bit of room to breathe. On the iPhone, by contrast, you get only one column of content per screen. The result is much, much easier to digest. You don't have waves yipping at you while you read another wave.

    The iPhone version on Wave is implemented as a web app running inside Safari. Inexplicably, when you log in, there's a message that says that the iPhone browser is not fully supported. This is inexplicable for 2 reasons.
    1. The iPhone implementation is quite a bit more mature than the Firefox 3.5 implementation I run on my laptop.
    2. The main bug in the iPhone implementation is that the screen telling you the iphone browser is not fully supported breaks links into Wave.
    I would expect to see a iPhone native Wave app very soon.

    I think most people running Wave will eventually choose to run Wave in native applications that talk the Wave protocol, just as most people (i.e. me) use native email clients to do their email. I just hope that Google's choice to launch Wave inside browsers is not just a plan to establish Chrome as a new web operating system. Twitter owes its success in no small part to the ecosystem of client applications that have sprung up around it. Although I use the Nambu client myself; I'm sure that many satisfied Tweetdeck users would be unhappy if the only way they could use Twitter was through Nambu.

    I've read the assertion that Google Wave is the Segway of email, but I think that's wrong. The better analogy is the videophone. Although it's often thought that AT&T's videophone failed because there was no one to call, that's only half of the story. The reason that there was no one to call was that the videophone didn't fit into any social practice- no one really want to have people popping up on little screens in their living rooms.

    I'm betting that Wave will turn out to be more popular than the Picturephone, but the real turning point will be third-party clients.

    This article is also posted inside Wave.

    Tuesday, October 6, 2009

    Telephone Dread and Wave Inbox Puppies

    It took me 30 years to get over telephone dread. Before my cure, I could only make a telephone call in a quiet environment without anyone to disturb me. Having decided to call someone, I would screw up my courage, pre-decide what to say to every possible answerer, and then pick up the phone. The dial tone would taunt me until I started dialing. With every digit of the phone number, I would experience doubt and anxiety. The touch-tone made it better; in the days of pulse-tone dialing, I hated numbers with lots of zeroes. Maybe the number was wrong. Maybe the person I was calling would be out. Maybe I would have to talk with someone's mother. Maybe the person wouldn't know who I was, and I'd have give a humiliating explanation of who I was. Half the time I would click the hook part way through the number after thinking of some unconsidered terror. By the time I got through, I would be an emotional wreck.

    It was the mobile phone that cured me. I could just key in the number, double check it, stare at it for a minute, and then I could press the send button and I was launched into the call. Ten buttons of anxiety was reduced to a single button, and that, combined with a small bit of emotional maturity, has enabled me to use the phone like a normal human being.

    I've always been comfortable with email. There's a single send button, and once an email has been sent, I can forget about it completely until I get a reply. And there's no need to reply immediately- the ease with which email can get lost is perhaps its best attribute!

    The tele-presence afforded by Instant Messaging is quite comfortable. You can see who's "around" and strike up casual conversation. The intermittent immediacy can be very stimulating.

    Twitter was very un-threatening to start out with. No one was ever going to read that first status message, or so you thought. Gradually, Twitter will reel you in because once in a while, without rhyme or reason, people will react to what you say. Retweets are like the random rewards that kept pigeons pecking in the famous random reinforcement schedule experiment.

    Google Wave is different, and I still can't tell if it will terrorize me as the telephone dial did or whether it will comfort me the way IM status messages do. Although it's conceived as email re-invented, it moves beyond familiar modes of communication. My first impressions post got a commenter who pointed to a blog post from June that included a note about the unexpectedness of one of Wave's features.
    Wave is changing paradigms. People can no longer take back what is released. Even if someone deletes part of the document, the deleted part can be seen in playback. While this "permanent memory" was there almost since the beginning of the Internet, it was never before real-time. How could we take back an information from a Wave? Imagine you have misplaced your password to the wave instead of password input box. It will always be visible. OK, I could change my password, but what about unfortunate copy&paste event with a credit card number?
    I had never really considered the importance of forgettability in communication.

    Wave's waves are described as "living things" because they can be continually added to. They also live somewhere that's not on your computer. Google's conception is that there will eventually be many Wave Service Providers running the Wave open-source service platform, but for now, the waves all reside on Google's servers. I can't think of any form of human communication that "lives" in remotely the same way. Maybe people playing music or games together is the closest example.

    Email was easy to start out with because it wasn't so different from the kind of mail that requires stamps and envelopes. You wrote a message and sent it off. Instant messaging was not so different from having a real-word or telephone chat. Wave, by contrast, maps most closely to other things you do on the internet, like IM and Wiki editing, each of which is a degree removed from things you do in real life. Wave is two degrees removed from normal human intercourse.

    Since it's a living thing you can't just send a wave and then forget it. Your "inbox" is like a litter of puppies that are yelping for attention and growing in front of your eyes. Or maybe they just sit there dead, waiting to be "archived". It will take a while before I'll know whether the Wave window will be filling me with dread or delight. At least I've learned how to give my puppies a bit more breathing room.

    After four days of Wave, I'm most excited to see how the Wave ecosystem is developing, even though many features are still hidden because important things aren't working yet. A sorely needed feature has been a way to do ranking of public waves. Today, that feature is being born. This is a great example of the benefits to having a platform open to developers from the very beginning. Needed functionality is being quickly added using the robot mechanism. I'm starting to think about how to use Wave robots to do useful things in libraries and scholarly communication.

    This post is posted in Wave.

    Reblog this post [with Zemanta]

    Friday, October 2, 2009

    Google's Wave is a Tsunami of Conversation

    Human communication is a magical thing, and over the last decade, we've acquired some super-powers. We've had the telephone for a hundred and thirty three years and it is still developing. Nowadays you can hardly set your twelve-year-old loose in the mall without packing a cell phone, and people are dying because they just have to send texts while driving. Just think about all the "conversation" tools we've added since the birth of the internet: e-mail, instant messaging, chat rooms, mailing lists, bulletin boards, blogs, various and sundry social networks, Facebook, Skype, Twitter. Our use of these tools will continue to evolve, but at some point there's going to be a consolidation. Do we really need both Twitter and Facebook?

    Yesterday, I was lucky enough to get an invitation to the Google Wave Preview. I have no idea whether Wave will be the next big thing, but I can tell you it won't fail for lack of ambition. Unlike Twitter, which started with a concept so simple that it sounded really stupid, Wave starts by imagining what email would look like if its designers were to start from scratch, knowing what we know today. The result is a daunting attempt to roll all of our communication superpowers into a single user interface. At first, I couldn't really figure out how to get Wave to do what I wanted it to do. Part of the problem was that some of the features mentioned in the help videos had not been activated in the Preview version, though they are still working in the "Sandbox" version which has been available to developers for a few months now. Other things just don't work yet- the contacts module seems very buggy, which is a big problem because you need that to connect with other "Wavers". Luckily, my multitool communications mischmash is still working, and I was able to find some friends (via Twitter and Facebook) who were in the same situation as myself, still stumbling around a dark room of functionality, looking for someone to connect with.

    With some help here and there, I gradually figured things out. Wave introduces several new (to me) user-interface widgets; it struck me that web applications rarely introduce new interface widgets, unlike applications like Excel or Photoshop that put powerful capabilities in inteface objects. I still haven't figured out the sliders. The threaded discussions have little colored boxes that indicate who is typing something. The overall effect is that gMail has been tricked out and turbocharged. After a day of playing with it, I'm not sure I like the user interface, but I can't really think af any way to make it better given the scope of what Wave is trying to do.

    The focus of Wave is editable, threaded conversations, i.e. waves. (There's already a convention to use lower case w for the conversations and upper case W for the platform as a whole.) These are quite well done. They support hierarchical threading, styled text, auto-linking, distributed editing, history, history playback, insertion of images, videos, and software objects known as extensions (there is a poll widget pre-installed). The one thing that's missing is undo. I can think of all sorts of uses for these capabilities; I intend to try a few over the next week or so.

    The intro videos plug a tool called Bloggy that allows you to publish a wave to your blog. Bloggy is a robot that you add as a participant in your conversation. I spent about an hour trying to figure out how to get Bloggy to work, until I discovered (via Twitter) that Bloggy had not been activated on the Preview version yet.

    Having failed with Bloggy, I started thinking "Is that all there is?" and envying Google's white-hot overhype machine. But then @jillmwo pointed me to the way to search on public waves (put "with:public" before a search). Immediately the little ripply waves I was seeing turned into a tsunami, and I began to see the enormous possibilities of Wave. To make a wave public, you add a participant called "Public" (public@a.gwave.com) to the wave; "Bloggy" does the same thing.

    The ability to search public waves (and to make your own public wave) is a feature qualitatively different from anything I've seen before. A public wave is sort of the bastard offspring of a Twitter hashtag mated with a Wikipedia topic page. When it matures, this creature will surely be a monstrous beast; what no one can tell yet is whether the beast will be a tame workhorse or whether it will be a velociraptor requiring a good strong cage. New technology is always easy to create and control compared to new social practice.

    If it was me trying to invent email from scratch, I would spend 80% of my effort on spam prevention. It looks to me as though many of Wave's feaures have been hidden in the Preview because the functions needed for spam prevention are not ready yet. This is not to say that the Wave team is not spending 80% of its effort on spam prevention- these are incredibly hard things to get right. As an example, it's currently not possible to remove a participant from a wave, and any public wave participant can see the addresses of all the other participants. The problems with the contacts module are also probably related- waves from people you might not know just appear in your inbox, even though it looks like the intent is for them to first appear in the "Requests" folder in the "Navigation" module. Groups are also not yet implemented. It's difficult to know how this will play out until we see everything working.

    Since the Preview roll-out, there have been quite a number of negative reviews by people pointing to all the problems in Google Wave. I think these are missing the point to some extent. Sure, Wave could end up being a complete failure, but even if that happens, Wave is giving us today a glimpse of what the future could look like. If not Wave, then surely there will be a SuperTwitter or a SpaceBook or maybe even a MagicForce that will contend for the consolidation of our one hundred flowers of conversation.

    Note: this post will also be published as a public wave.

    Reblog this post [with Zemanta]

    Thursday, September 24, 2009

    Nambu Gets Better and Shortener User Tracking is Undermined


    When I last wrote about the tribulations of tr.im and the business of bit.ly, our heroes had just stepped away from their nose-to-nose struggle, with Nambu founder Eric Woodward having announced the shut-down of tr.im, his URL shortener, only to vow its revival a few days later. Bit.ly, with a cozy relationship with Twitter, seemed to have taken a dominant position in the URL shortening business, whatever that turned out to be. I speculated that Bit.ly would use its position in the ocean of usage data to build psychographic profiles of users to help target advertising.

    Since then, Woodward decided not to sell the tr.im business and has instead released the tr.im software as free open source, making it that much easier for websites to do their own URL shortening. He's also focused his company's attention on its Nambu Twitter clients (Mac and iPhone), the development of which suffered a major setback when his Chinese developers left for richer opportunities as soon as their contract was up. Nambu for Mac OS X has been my preferred Twitter client; it has a much more Mac-like user interface than others I've tried. When I updated my system to Snow Leopard last week, I was disappointed to find that Nambu had not survived the system update.

    After unhappily revisiting Tweetdeck, I decided to try the beta version of Nambu, even though it's described as being not quite done. So far, it looks pretty solid. One change in particular pleased me, and that's the way the new Nambu works with URL shorteners. It seems that by surrendering URL shortening to bit.ly, Nambu is now freer to innovate in the user experience. Nambu now pre-expands all the shortened links so that the user can see the hostname that the links are pointing to. This has a number of consequences:
    1. The user can tell where a link will go. This will help avoid wasted clicks, and will help the user avoid spam and malware sites.
    2. Because all of the links are dereferenced before use, the URL shortening sites will no longer be able to track the user's reading preference. The business model I previously suggested for bit.ly will be defeated, and the user's reading privacy will be protected.
    3. The URL shortener will have to deal with an increased load. Nambu's going to make bit.ly work harder for the privilege of domination the URL shortening space.
    Now I understand why bit.ly has been registering a bunch of instantaneous hits whenever I tweeted a link- it was robot agents, not people, that were clicking the links.

    I was curious to see if Nambu was querying the URL shorteners directly or whether Nambu was trying to aggregate and cache the expanded links. I installed a nifty program called "Little Snitch" to see the outbound connections being made by programs on my laptop. It turns out that Nambu is doing a direct check for redirection on ALL of the links that it shows me, not just the shortened ones. Although this could break links that are routed as part of a redirect chain, I imagine that sort of link occurrs rarely in a Twitter stream.

    The new behavior of Nambu and its effects on usage tracking points up a general problem faced by any system designed to measure and track internet usage. In my post on "bowerbird privacy", I mentioned that I use StatCounter to measure usage on this blog. StatCounter works quite well for now, but I imagine that its methods (based on javascript) might well stop working so well as web client technology evolves. That's one reason I expect that efforts to standardize measurements of usage in the publishing community, such as Projects "COUNTER" and "USAGE FACTOR" are doomed to rapid obsolescence.

    Will bit.ly ever get a business model? Will Nambu find peace with the chilly kitty? Find out in next months installment of... As th URL Trns