Saturday, January 8, 2011

Inside the Dataculture Industry

wild blueberries
I don't really know how all the food gets to my table. Sure, I've gathered berries, baled hay, picked peas, baked bread and smoked fish, but I've never slaughtered a pig, (successfully) milked a cow or roasted coffee beans. In my grandparents generation, I would have seemed rather ignorant and useless. Agriculture has become an industry as specialized as any other modern industry; increasingly inaccessible to the layperson or small business.

I do know a bit about how data gets to my browser. It gets harvested by data farmers and data miners, it gets spun into databases, and then gets woven into your everyday information diet. Although you've probably heard of the "web of data", you're probably not even aware of being surrounded by data cloth.

The dataculture industry is very diverse, reflecting the diversity of human curiosity and knowledge. Common to all corners of the industry is the structural alchemy that transmutes formless bits into precious nuggets of information.

In many cases, this structuring of information is layered on top of conventional publishing. My favorite example of this is that the publishers of "Entertainment Week" extract facts out of their stories and structure them with an extensive ontology. Their ontologists (yes, EW has ontologists!) have defined an attribute "wasInRehabWith" so that they can generate a starlet's biography and report to you that she attended a drug rehabilitation clinic at the same time as the co-star of her current movie. Inquiring minds want to know!

If you look at location based services such as Facebook's "places", Foursquare, Yelp, Google Maps, etc, they will often present you with information pulled from other services. Often, a description comes from Wikipedia and reviews come from Yelp or Tripadvisor and photos come from Panoramio or Flickr. These services connect users to data using a common metadata backbone of Geotags. Data sets are pulled from source sites in various ways.

Some datasets are produced in data factories. I had a chance to see one of these "factories" on my trip to India last month. Rooms full of data technicians (women do the morning shift, men the evening) sit at internet connected computers and supervise the structuring of data from the internet. Most of the work is semi-automated, software does most of the data extraction. The technicians act as supervisors who step in when the software is too stupid to know when it's mangling things and when human input is really needed.

There's been a lot of discussion lately about how spammers are using data scraped from other websites and ruining the usefulness of Google's search results. There are plenty of companies that offer data scraping services to fuel this trend. Data scraping is the use of software that mimics human web browsing to visit thousands of web pages and capture the data that's on them. This works because large websites are generated dynamically out of databases; when machines assemble web pages, machines can disassemble them.

A look at the variety of data scraping companies reveals a broad spectrum. Scraping is an essential technology for dataculture; as with any technology, it can be used to many ends. One company boasts of their "massive network of stealth scrapers capable of downloading massive amounts of data without ever getting blocked. Some companies, such as Mozenda, offer software to license. Others, such as Xtractly and Addtoit are strictly service offerings.

I spoke to Addtoit's President, Bill Brown, about his industry. Addtoit got its start doing projects for Reuters and other firms in the financial industry; their client base has since become more "balanced". Companies such as Bloomberg, Reuters and D&B get paid premiums to provide environments rich in structured data by customers wanting a leg up on competitors. Brown's view is that the industry will move away from labor intensive operations to being completely automated, and Addtoit has developed accordingly.

A small number of companies, notably Best Buy, have realized that making their data easily available can benefit them by promoting commerce and competition. They have begun to use technologies such as RDFa to make it easy for machines to read data on their web sites; scraping becomes superfluous. RDFa is a method of embedding RDF metadata in HTML web pages; RDF is the general data model standardized by the W3C for use on the semantic web, which has been discussed much on this blog.

This doesn't work for many types of data. Brown sees very slow adoption of RDFa and similar technologies but thinks website data will gradually become easier to get at. Most websites are very simple, and their owners see little need or benefit in investing in newer website technologies. If people who really want the data can hire firms like Addtoit to obtain the data, most of the potential benefits to website owners of making their data available accrue without needing technology shifts.

The library industry is slowly freeing itself from the strictures of "library data" and is broadening its data horizons. For example, many libraries have found that genealogical databases are very popular with patrons. But there is a huge world of data out there waiting to be structured and made useful. One of the most interesting dataculture companies to emerge over the last year is ShipIndex. As you'd expect from the name, ShipIndex is a vast directory of information relating to ships. Just as place information is tied together with geoposition data, ShipIndex ties together the world of information by identifying ships and their occurrence in the world's literature. The URIs in ShipIndex are very suitable for linking from other resources.

The Götheborg
ShipIndex is proof that a "family farm" can still deliver value in the dataculture industry. The process used to build ShipIndex. Nonetheless, in coming years you should expect that technologies developed for the financial industry will see broader application and will lead to the creation of data products that you can scarcely imagine.

The business model for ShipIndex includes free access plus a fee-for-premium-access model. One question I have is how effectively libraries will be able leverage the premium data provided with this model. Imagine for example the value you might get from a connection between ShipIndex and a geneological database bound by passenger manifests. I would be able to discover the famous people who rode the same ship that my parents took to and from the US and Sweden (my mom rode the Stockholm on the crossing before it collided with the Andrea Doria). For now though, libraries struggle to leverage the data they have; better data licensing models are way down on the list of priorities for most libraries.

Peter McCracken
ShipIndex was started by Peter and Mike McCracken, who I've known since 2000. Their previous company (SerialsSolutions) and my previous company (Openly Informatics) both had exhibit tables in the "Small Press" section of the American Library Association exhibit hall, where you'll often find the next generation of innovative companies serving the library industry. They'll be back in the Small Press Section at this weekend's ALA Midwinter meeting. Peter has promised to sing a "shanty" (or was that a scupper?) for anyone who signs up for a free trial. You could probably get Mike to do a break dance if you prefer.

I'll be floating around the meeting too. If you find me and say hello, I promise not to sing anything.
Enhanced by Zemanta

Thursday, January 6, 2011

Fundamental Constant Numerology

My father was obsessed with units of measurement and fundamental constants. He got his engineering degree at the Royal Institute of Technology (Kungliga Tekniska Högskolan) in Stockholm, Sweden. His favorite professor there was Erik Hallén, who was famous for his work on antenna theory and for laying groundwork for the world's most widely used system of measurement, the SI system.

My dad nearly failed Hallén's class, which could be one reason for his lifelong obsession with units. The other reason was that Dad was convinced that he could explain some of physics' deepest questions about the nature of matter by applying the electromagnetic theory he learned in Hallén's class. Dad   explained the structure of the electron by modeling it as a circulating charge wave in a resonant cavity formed by general relativistic warping of space. In his notes, he wrote
To me it looks like all the puzzle is defined and ripe to be put together and the extension to other particles will not be difficult- only time consuming.
Dad used these insights to come up with a relationship between the gravitational constant and the mass and charge of the electron. Here's his equation:
                     G = 6/π 10-44 Z0 c (e/m)2
where G = the gravitational constant (which defines the force that holds the universe together), Z0 is the impedance of free space, c is the speed of light, and e and m are the charge and mass of the electron.

Here's a prettier version of that equation, made using Roger's Equation Editor from the following TeX code:
                      G = \frac{6}{\pi} 10^{-44} Z_0 c (\frac{e}{m})^2
(TeX is the most commonly used formatting language for mathematics.)


A derivation of this equation would easily have earned my dad a Nobel Prize, but without a derivation and explanation of the underlying physics, it was just numerology. If you plug in the numbers, Dad's equation is within 0.025% off of the consensus value for the gravitational constant, whose experimental uncertainty is about 0.01%. Dad understood that his equation was worthless without an explanation, so he spent endless hours studying Bessel equations and all sort of obscure mathematics. He was sure that somehow, somewhere, there existed a solution to Maxwell's equations combined with general relativity to explain the 6/π and confirm his resonant cavity. (He said the 10-44 was just using the right units- I never understood that!)

The Nature of the Physical WorldMy dad was not alone in physics numerology. The fine structure constant, which is very close to 1/137, has attracted all sorts of numerological explanations. (Dick Lipton calls it a "miracle number") Arthur Eddington, one of the most famous physicists of his time, had an explanation for why the fine structure constant should be exactly 1/137 involving the number of protons in the universe. A more modern numerological result is that of James Gilson, whose suggested value for the fine structure constant is only 30 parts per trillion off.

While fundamental constant numerology has deservedly been on the fringes of science, new internet search technologies may soon change that. Last year, scientific publisher Springer introduced a beta service called LaTeX Search that allows researchers to search for LaTeX formatted equations in all of Springer's journals. (LaTex is a dialect of TeX most widely used for scientific publishing.) That's something you can't do with Google, or any other search engine. The ability to connect obscure mathematical discoveries from disparate fields of science could soon be facilitating new avenues of research, perhaps even new methodologies.

For example, I can search for a fragment of my dad's equation and get at least one result that seems relevant. I don't know of any meaningful discoveries that have been made so far with LaTeX search, but if my dad had been able to search all of the mathematical literature to connect his numerological result with a mathematical solution, perhaps he would have explained the gravitational constant and structure of the electron and would have won his Nobel Prize.

He would have been 83 today. Happy Birthday, Dad! We miss you.

Friday, December 31, 2010

2010 Summary: Libraries are Still Screwed

In mathematics, catastrophe theory is the study of nonlinear dynamical systems which exhibit points or curves of singularity. The behavior of systems near such points is characterized by sudden and dramatic changes resulting from even very small perturbations. The simplest sort of catastrophe is the fold catastrophe.

When a fold catastrophe occurs, a system that was formerly characterized by a single stable point evolves to a system with no stability. The point where stability disappears is known as the tipping point.

One of my goals for this past year was to raise awareness of the tipping point for libraries that will accompany the obsolescence of the print book. In January, I noted that Hal Varian's equation describing the economic value of libraries also predicts that libraries of the current sort won't exist for ebooks.

In March, I put the question directly to John Sargent, Macmillan's CEO. His response, that ebooks in libraries were a "thorny problem" got quite a bit of notice. Unfortunately, the big trade publishers have yet to actually do much to address the thorns.

In May, I was pleased that the editors of Library Journal were putting together an "eBook Summit" virtual meeting to address some of these issues, and even more pleased to be invited to write a series of articles to help frame issues for the Summit. The event ended up being titled ebooks: Libraries at the Tipping Point. For me, the highlight of the summit was Eli Neiburger's talk on How eBooks Impact Libraries. This talk is destined to be known forever as the "Libraries are Screwed" talk, and if you've not viewed it I urge you to do so forthwith.



Several other contributions raising awareness of the library-ebook catastrophe are worth noting. Emily Williams' commentary on Eli's talk is worth reading and her attention to the issue has been consistent. Library Journal's Heather McCormack is another persistent voice- I particularly loved the story she told in a column written for O'Reilly Radar. Tim Spalding's post on "Why are you for killing libraries is another favorite.

So at the end of the year, what have we accomplished? One disappointment for me was that although Library Journal's eBook Summit was quite popular with librarians, it appears that very few publishers took notice. On the rare occasions when publishers took notice of the role libraries could play in the ebook future, they tended to be depressingly reactionary, such as when the UK's Publishers Association set out their plan to marginalize libraries while apparently thinking they were boosting them.

Similarly, Amazon announced yesterday the addition of a lending feature for the Kindle. This feature seems designed to compete with a similar feature in the Nook, but nowhere in the announcement is there any mention of libraries as being anything other than the books on a user's Kindle.

Meanwhile, adoption of ebooks and ebook readers has accelerated. Amazon announced that the third-generation Kindle is the bestselling product in Amazon's history. Barnes&Noble fired back, reporting that the NOOKcolor is the best selling product in its history. In comparison, this month's announcement by Overdrive that it (finally!) has released apps for reading its library ebooks on Android devices and iPhone seems a bit too-little-too-late. (sorry, no iPad version!).

Perhaps the time is over for raising awareness about the catastrophic future of libraries. In 2011, let's build things that change the system dynamics.

Tuesday, December 28, 2010

Why does Sarah Palin Charge for eBooks?

A neighbor of mine, a professor at Rutgers, has several published books to his name. A few years ago, the publication costs for one of his books were partially underwritten by an environmental organization, because they thought it was important work which ought to see the light of day, or at least the light of a few libraries.

Digital publishing makes possible new ways for works like my neighbor's to be seen by as many readers as want to read them. Since the cost of making digital copies is almost zero, it makes sense for them to be given out for free once the first copy costs have been paid for, especially if that cost is being borne by an advocacy group.

Which brings me to Sarah Palin, and her big-time new book, America by Heart : Reflections on Family, Faith, and Flag, published by HarperCollins (a division of Rupert Murdoch's NewsCorp). It costs $12.99 for the Kindle version and the same on iBooks, Nook, and Kobo, thanks to agency pricing. If Palin is running for president, as is widely assumed, you'd think that she'd want as many people as possible read the book. So far, sales have been disappointing, but for argument's sake, let's assume that it's a brilliantly persuasive thrill of a read. Why bother charging for the ebook, since it costs almost nothing to make copies?

Well, there's the small matter of the big advance (she got $1.25 million for Going Rogue). It all boils down to money. Most of the people buying the book are probably already political supporters, so writing (or ghost-writing, as the case may be) a book like America at Heart is as much a mechanism to monetize political support as a genuine contribution to our nation's political discourse. For the publisher, a Sarah Palin book is a win-win-win. Not only can the publisher make money if the book turns out to be a blockbuster like Going Rogue, which sold over 2 million copies, but the prestige of a popular figure rubs off on the publisher's brand. And when access to the corridors of power comes along with the advance, how can the publisher lose?

Bill Clinton, speaking to the AAP in March 2009
I don't mean to be partisan; the argument applies equally to Dreams of My Father and The Audacity of Hope; the Obama administration has so far been very friendly to the book publishing industry on its "special issues". And don't think that Bill Clinton would have been speaking at the Association of American Publishers Annual Meeting, as he did last year, if he wasn't also an author with a big deal.

It's surprising that these arrangements don't get more scrutiny. It would be illegal for a politician to accept money into a personal bank account from a big company with major political issues before him, but it's OK for a publishing company to write 7-figure checks as advances to those same politicians if there's a book being written to serve as a fig leaf.

Publishing books has always been a double-edged sword for politicians and advocacy groups; a good book bolsters a cause both financially and intellectually. The publicity tour surrounding a book is as much about publicizing the cause as it is about selling the book. But looking to the future, what sort of digital business models will be most effective for authors whose primary goal is to advance a cause?

While the current business model offers many advantages such as access to a publisher's marketing and distribution channels, the author receives a rather small percentage of the money paid by book purchasers. In 2008, the year he was elected, Barack Obama's books sold about a million copies; he reported income from these of $2.5 million. Most authors get a significantly smaller percentage of sales.

If you've been reading this blog, you probably won't be surprised at the alternative I suggest:  bounty markets for open-access ebooks will be the ideal way to accomplish the twin aims of advocacy and fundraising. By first raising money from core supporters, and then releasing an ebook for free, a cause-oriented author can use the open access ebook to gain new converts.

Perhaps ebook bounty markets will even need to put in safeguards to avoid being misused as a naked money conduit to launder campaign contributions.

A guy can dream, can't he?

Friday, December 24, 2010

Christmas and Cooking

In Lalbagh Garden, in Bangalore, I found this plaque establishing the link between Christmas (tree) and the Cook (pine):
"Christian community treat this tree as sacred one & worship during Christmas."
Note the Muslim women having a photo snapped, with "Christmas Trees" in the background.
In my house, many holiday rituals surround cooking. One of my most precious possessions is this copy of Stora Kokboken, which I haul out every December to help with my liver pâté, my award-winning Swedish coffeebread and whatever else sparks happy memories. I say hello to my mom's notations and feel her presence.
(Stora Kokboken, Edith Ekegårdh and Britta Hallman-Haggren, Wezäta Förlag, Göteborg, 1957)
Opposite the title page is this:
My parents were married in December of 1957; the cookbook was a Christmas present from my mom's father, stepmother and stepsister. The note says "En god jul önskar vi er!" which translates as "We wish you a Merry Christmas".

What they said.