Showing posts with label business models. Show all posts
Showing posts with label business models. Show all posts

Friday, August 25, 2023

Let's pretend they're ebooks

In days of yore, back when people were blogging, I described the way that libraries were offering ebooks as being a "Pretend It's Print" model. At the time, I felt that this model was designed to sustain and perpetuate the model that libraries and publishers had been using since prehistoric times, and that it ignored most of the possibilities inherent in the ebook. Ebooks could liberate the book from the shackles of their physical existences!
 
I was right, and I was wrong. The book publishing world seized on digital technology to put even heavier shackles on their books. In turn, technology companies such as Amazon locked down innovation in the ebook world so that libraries could no longer be equal contributors to the enterprise of distributing books, all the while pretending to their patrons that the ebooks they licensed were just like the print books sitting on their shelves.
 
Somehow libraries and publishers have survived. Maybe they've even thrived with the "pretend it's print" model for ebooks. There are plenty of economic problems, but whenever I talk to people about ebooks, the conversation is always some variation of "I love reading ebooks through my library". Most library users are perfectly happy pretending that their digital ebooks are just like the printed books.
 
robot writing on an ipad
A decade later, we need to change our perspective. It's time we seriously started pretending that printed books are just like ebooks, not just the other way around. The library world has been doing something called "Controlled Digital Lending" (CDL) , which flips the "pretend it's print" model and pretends that print is just like digital. The basic idea behind controlled digital lending is that owning a print book should allow you to read it any way you want, even if that involves creating a digital substitute for it. A library that owns a print book ought to be able to lend it, as long as it's lent to only one person at time. It's as if books were printed and sold in order to spread ideas and information!
 
Of course radical ideas such as spreading information have to be stopped. And so we have the Hachette v. Internet Archive lawsuit and its assorted fallout. I'm not a lawyer, so I won't say much about the legal validity of the arguments on either side. I'm an ebook technologist, so I will explain to you that whole lawsuit was about whether the other side was sufficiently serious about pretending that print books are just like ebooks and that ebooks are just like print books. Also that the other side doesn't understand how print books are completely different things than ebooks. Those lawyers really take to heart the White Queen's recommendation to believe 6 impossible things before breakfast.
 
The magic of technology is that it can make our pretendings into something real. So let's think a bit about how we can make the pretense of print-ebook equivalency more real, and if the resulting bargain makes any sense.
 
Here are some ways that we could make these ebooks, derived from printed books, more like print books:
  1. Speed. It takes me an hour or so to get a print book from a library. Should I be able to get the digital substitute in a minute? Should I be able to read a chapter and the "return" it so that someone else can use it the next seconf? CDL already puts some limits on this, but maybe there could be a standard that makes the digital surrogate more like the real thing?

  2. Geography. Printed books need to be transported to where the reader is. Once digitized they could go anywhere!. Maybe something like a shipping fee could be attached to a loan or other transfer. Maybe part of the fee could accrue to creators? Academic libraries have long done interlibrary loan of journal articles by copying and mailing the article, so why not do something equivalent for books?

These two attributes matter a lot in defining commercial markets for books and ebooks, and will become increasingly important as distribution technologies scale up and improve. Although publishers today make most of their money on the most popular books, book sales and usage of books in libraries have very long tails. There are millions of books for which global demand could be met by aggressive CDL of just a few copies. The CDL system instituted by Internet Archive also has a countervailing effect - the world-wide availability combined with so-so EPUB quality and usability probably result in stimulation of demand for print copies. This effect is likely to diminish as technologists like me smooth out the DRM speedbumps in CDL and begin to apply machine learning to EPUB generation.
 
It's worth noting that the "long tail" in book publishing also applies to authors and publishers. It's likely that the Internet Archive's CDL service has a larger market effect (whether positive or negative) on these market participants.
 
Here are some ways that we shouldn't make ebooks more like  print books:
  1. Search. Ebooks make search much easier than in print books. Maybe search should be disabled in CDL ebooks? Or maybe, we could enable search in print books. Google Books already sort of does this, if you have the right edition, but the process of making an ebook from a print book should give you an easy way to enable search in the print!

  2. Accessibility. Many reading-disabled users rely on ebooks for access to literature, science and culture. Older adults such as myself often find that flowable text with adjustable font size is easier on our eyes. In addition to international treaties that treat accessible text as an exception to copyright, most authors and publishers don't want to be monsters.

  3. Smell. Let's not go there.

  4. Privacy. The intellectual property world seems to think that copyright gives them the right to monitor and data-mine the behavior of readers on digital platforms. In some cases, copyright extremists have required root access to our devices so they can sniff out infringing files or behavior. (While they're at it, they might as well mine some bitcoin!) It is an outrage to think anyone who makes ebooks from print books would wire them with surveillance tools; the strong privacy policies of Internet Archive should be codified for CDL.

  5. Preservation. Publishers do a terrible job of preserving the lion's share of the printed books they publish, and society has always relied on libraries for this essential service. In this digital age, any grand bargain on copyrights has to provide libraries with the rights and incentives needed to do digital preservation of both printed and digital books.

The bottom line is that if we're going to continue to pretend that intellection property is a real thing, we need to start pretending that printed books are like ebooks, and vice versa. A grand bargain that benefits us all can eventually make these illusions real.

Notes: 

  1. Copyability. CDL books, like publisher-created ebooks, rely on device-enforced restrictions on duplication (DRM). Printed books rely on the expense of copying machines and paper to limit reproduction. In both cases, social norms and legal strictures discourage unauthorized reproduction. Building those social norms is what creating a grand bargain is all about.
  2.  Simultaneous use. Allowing simultaneous use of library ebooks during the pandemic is what really got the publishers mad at Internet Archive. A lot of people went mad during the lockdown, to be honest, and we're still recovering. 
  3.  Comments. I encourage comment on the Fediverse or on Bluesky. I've turned off commenting here.

Saturday, March 22, 2014

eBook ILL is silly. The reason why will bore you.

When we try to think about digital things as if they are still the real things they used to be, we can lose touch with the parts of reality that are important. It's silly.

If you're not of the library world, let me explain what ebook ILL is and why it's not silly per se. ILL stands for Inter-Library Loan. In the print world, libraries have finite collections and they depend on other libraries to make sure that even if they don't have a book that a user needs, another library will step in and fill the gap. For the user, it means that their small library can provide them with books from a huge virtual collection. A book might take a few days to arrive.

There are significant costs involved in ILL. Most libraries charge the borrowing library a fee to cover the expenses of packaging the book and sending it to the recipient library. The fee might be 10 or 20 dollars, and it might be waived for closely cooperating libraries. At the same time, libraries pay the same fees to other libraries, so in the end, it all evens out. But many libraries run a significant surplus, rewarding them for smart acquisition policies of the past.

Library lending cooperatives have figured out that the combination of Amazon and modern warehouse logistics have partly upset the economics of ILL; a library can often purchase a used copy of a needed book on Amazon for less than ILL transaction costs (Especially with Amazon Prime!). But ILL is still an important part of the library ecosystem.

For digital content, the buy vs. borrow equation shifts back a bit. In principle, there's no shipping cost and modern databases can retrieve a digital item in milliseconds. But if a library can do digital ILL, what is to prevent libraries from sharing a resource so widely that only one library in the world needs to buy the item?

The solution that e-journal publishers typically use is the "print-and-ship" solution. In other words, a library is allowed to send articles from a subscribed journal only if they print it out first. The transaction is thus identical to what it was back in the dark ages of ink and paper and xerox machines. For publishers, the friction of print-and-ship discourages libraries from canceling subscriptions; besides, the big-deal model of bundling many subscriptions into one has been much more advantageous for publishers than the document-delivery model that ILL competes with. (Also, when they first went digital, journal publishers were poorly equipped to do article-by-article e-commerce.)

Printing article PDFs and mailing them is a stretch, but mapping this model into ebooks is a farther stretch. The book ecosystem has never included libraries creating copies of in-print printed books. And why should library A ever acquire a book if the copy owned by library B works just as well? Since most ebooks never really go "out-of-print", the inter-library loan system will be competing directly with publisher sales.

To see why it still makes sense for publishers to allow ebook ILL, consider what it is competing against: "patron-driven acquisition" (PDA). The core idea behind PDA is that a library doesn't buy an ebook until a patron shows up that wants to use it. For many books that libraries buy, this means that they don't buy the book at all, and for the rest, there might be no purchase until many years after publication.

It's often better for the publisher to encourage "just-in-case" acquisition, because the resulting revenue can be put to work immediately to publish more books. For books with low demand, inter-library loan encourages just-in-case acquisition by increasing the likelihood that somewhere, sometime, someone will need the library's copy of even the most obscure book.

eBook licensing with ILL has very similar economic characteristics as licensing to a library with many users. The larger the library, the more demand can be aggregated, and thus books can remain economically viable even at very low levels of user demand. In the limit of large user bases, ILL looks very much like the Open Access collective funding such as has been demonstrated by Knowledge Unlatched.

"Just-in-case" acquisition has benefits for libraries, too. Coupled with an effective archiving strategy, the library can make sure a resource doesn't disappear if a publisher has to withdraw an ebook title. Or perhaps the publisher goes out of business, or decides to change their business model.

But ebook ILL is still silly. Admittedly, the one-user-at-a-time licensing model has proven to be a useful conceit for selling ebooks. People are used to paying for a copy of a book, so it seems natural to buy a copy of an ebook. But stretching that model to inter-library lending turns the conceit into an outright lie. Just because one library has bought an ebook copy doesn't mean that they should be able to lend it instantly to anyone in the world.

Clinging to the pretend-its-print conceit when developing licensing models for "just-in-case" acquisition results in harmful misunderstandings for both publishers and for libraries. Publishers focus on sales substitution, and libraries misunderstand what they're paying for. Pricing and terms for such licenses will better benefit both libraries and publishers if the license is seen for what it really is rather than something it's pretending to be. There's no sense in locking in the negative attributes of the old when developing the new.

So maybe we need a new acronym. How about "Interacting Libraries License"?

ILL is dead, long live ILL.

Saturday, February 1, 2014

Crowd-Frauding: Why the Internet is Fake

Power in human societies derives from the ability to get people to act together. Armies, religions, governments, and businesses have dominated societies using weapons, beliefs, laws and money to exert collective effort. In modern societies, mass media have emerged as a similar organizing power.

A new kind of collective organization, mediated by the internet, code and connections, is emerging as another avenue of power. It's no longer ridiculous to think that social networks, crowd-sourcing and crowd-funding could achieve the social consensus, action and compulsion that were the province of governments, armies and religions. As the founder of a crowd funding site for ebooks, I'm naturally optimistic that the crowd, connected by the social internet, will be an immensely powerful force for good.

I'm also continually reminded that the bad guys will use the crowd, too. And it won't be pretty.

In December, our site, Unglue.it, began to get a surge of new users. But there was something fishy. The registrations were all from hotmail, outlook, and various dodgy-sounding email hosts. The names being registered were things like Linette, Ophelia, Rhys, Deanne, Agueda, Harvey, Darcy, Eleanore and Margene. Nothing against the Harveys of the world, but those didn't look like user handles. Of course it was registration bots coming from many different IP addresses.

But they were stupid registration bots- they never complete the registration, so the fake accounts can't leave comments or anything. It wasn't causing us any harm except it was inflating our user numbers. It was mystifying. So I started studying why bots around the world might start making inactive accounts on our site.

And that's how I learned about the dark side of the crowd-force. The best example of this is a program called Jingling, also known as FlowSpirit. It's been around for 5 years or so.

Jingling is an example of a "cooperative traffic generation" tool. It's software-organized crime. Crowd-frauding, if you will.

It works like magic. You download the Jingling software and install it on your computer. You then enter the URLs for four webpages that you want to promote. (or more, if you have a good internet connection.)  Although the user interface is in Chinese, you can get annotated instructions in English on YouTube or websites like theBot.net. Once you've activated Jingling, the webpages you want to promote start getting hundreds of visitors from around the world. The visitors look real, they click around your page, they click on the advertisements, they register accounts on websites, they click "like" buttons and follow you on Twitter.

Meanwhile, your computer starts running a website-visiting, ad-clicking daemon. It visits websites specified by other Jingling users. It clicks ads, registers on sites, watches videos, makes spam comments and plays games. In short, your computer has become part of a botnet. You get paid for your participation with web traffic. What you thought was something innocuous to increase your Alexa- ranking has turned you into a foot-soldier in a software-organized crime syndicate. If you forgot to run it in a sandbox, you might be running other programs as well. And who knows what else.

The thing that makes cooperative traffic generation so difficult to detect is that the advertising is really being advertised. The only problem for advertisers is that they're paying to be advertised to robots, and robots do everything except buy stuff. The internet ad networks work hard to battle this sort of click fraud, but they have incentives to do a middling job of it. Ad networks get a cut of those ad dollars, after all.

The crowd wants to make money and organizes via the internet to shake down the merchants who think they're sponsoring content. Turns out, content isn't king, content is cattle.

Jingling is by no means alone; there are all sorts of bots you can acquire for "free". Diversity of bots is enforced because the click fraud countermeasures only attack the most popular bots; new bots are being constantly developed and improved.

What does this mean for advertising, ad-supported websites, and the internet in general?

It means that the internet rich will get richer and power will concentrate. Let me explain.

I used to run my own mail server. It was a small process on one of my old machines. I was a small independent internet entity. The NSA couldn't scan my emails and I could control my mail service. But as spammers cranked up their assault, it became more and more complicated to run a mail server. At first, I could block some bad ip addresses. But when dictionary attacks could be run by script kiddies, running my own email server got to be a real drag. And because I was nobody, other people running mail servers started blocking the email I tried to send. So I gave up and now I let big brother Google run my email. And Google gets to decide whether email reaches me or gets blocked by spam. That doesn't make me happy, and I still get a fair amount of spam. (Somehow the SEO and traffic generation scammers still get through!)

It's probably the same way that kingdoms and countries arose. Farmers farmed and hunters hunted until some bad guys started making trouble. People accepted these kings and armies because it was too much trouble for farmers and hunters to deal with the bad guys on their own. But sooner or later the bad guys figured out how to be the kings. Power concentrated, the rich got richer.

So with the crowd-frauders attacking advertising, the small advertiser will shy away from most publishers except for the least evil ones- Google or maybe Facebook. Ad networks will become less and less efficient because of the expense of dealing with click-fraud. The rest of the the internet will become fake as collateral damage. Do you think you know how many users you have? Think again, because half of them are already robots, soon it will be 90%. Do you think you know how much visitors you have? Sorry, 60% of it is already robots.

Sooner or later the VCs will catch on to the fact that companies they've funded are chasing after bot-clicks and bot views. They'll start demanding real business models; those of us older than forty may remember those from college. And maybe reality will have a renaissance. But more likely, the absolute power of Google, Amazon, Apple and the rest will corrupt them absolutely and we'll suffer through internet centuries of dark ages (5 solar years at least) before the arrival of an internet enlightenment.

Until then, let's not give in to the dark side of the force, OK?

Notes:
  1. Last year, I wrote about some strange Twitter bots. I now think it's likely that the encoded messages I saw are part of a cooperative traffic generation scheme. If you're trying to orchestrate a vast network of click-bots, what better way to communicate with them than twitter?
  2. There are now disposable email hosts that will autoclick confirmation links. These email hosts are the registration spammer's best friends. A list is here.

Saturday, March 9, 2013

African Drummers Invented an Internet

Maybe in 50 years we'll reminisce how there used to be one internet that covered the globe. But even before the telephone was invented, there were internets of a sort that covered regions in Africa. Probably there were others all around the globe, maybe even now the dolphins have their own version of an internet, disconnected from ours.

An internet, for the purposes of this article, is "a digital communications network that connects intelligent nodes distributed thoughout a region".

Even civilizations without written languages needed to communicate with each other. If you lived in a rain forest where villages were separated by miles of bush, the best way of communication with a neighboring village was to use drums. The "talking drums" of Africa used digital codes that could be understood from distances of as much as 5-10 kilometers. The codes weren't at all like Morse code, but were based on the tones of spoken languages.

A "talking drum" has two tones, so the signal is essentially binary. To overcome the lack of consonants, drum languages would add habitual phrases to words to disambiguate one word for another, resulting in an error-correcting code. Ruth Finnegan's chapter on "Drum Language and Literature" in Oral Literature in Africa gives some wonderful examples.
In the Kele language the words meaning, for example, ‘manioc’, ‘plantain’, ‘above’, and ‘forest’ all have identical tonal and rhythmic patterns. By the addition of other words, however, a stereotyped drum phrase is made up through which complete tonal and rhythmic differentiation is achieved and the meaning transmitted without ambiguity. Thus ‘manioc’ is always represented on the drums with the tonal pattern of ‘the manioc which remains in the fallow ground’, ‘plantain’ with ‘plantain to be propped up’, and so on. Among the Kele there are a great number of these ‘proverb-like phrases’ to refer to nouns. ‘Money’, for instance, is conventionally drummed as ‘the pieces of metal which arrange palavers’, ‘rain’ as ‘the bad spirit son of spitting cobra and sunshine’, ‘moon’ or ‘month’ as ‘the moon looks down at the earth’, ‘a white man’ as ‘red as copper, spirit from the forest’ or ‘he enslaves the people, he enslaves the people who remain in the land’, while ‘war’ always appears as ‘war watches for opportunities’. Verbs are similarly represented in long stereotyped phrases. 
So that's how information was transmitted digitally, but it's not an internet yet. For that, you need a network. It turns out that drummed messages of note would be retransmitted to the next village. In modern terminology, packet switching. I imagine the drummers used a protocol similar to that used by Ethernet to ensure a clear channel for retransmission. Thus announcements, warnings, poetry (maybe even advertisements!) were packet-switched between nodes based on topic and relevancy.

A lot of expressive power of the drum language was used for names. From Oral Literature in Africa:
Personal drum names are usually long and elaborate. In the Benue-Cross River area of Nigeria, for instance, they are compounded of references to a man’s father’s lineage, events in his personal life, and his own personal name . Similarly among the Tumba of the Congo, all-important men in the village (and sometimes others as well) have drum names: these are usually made up of a motto emphasizing some individual characteristic, then the ordinary spoken name; thus a Belgian government official can be alluded to on the drums as ‘A stinging caterpillar is not good disturbed’. Carrington describes the Kele drum names in some detail. Each man has a drum name given him by his father, made up of three parts: first the individual’s own name; then a portion of his father’s name; and finally the name of his mother’s village. Thus the full name of one man runs ‘The spitting cobra whose virulence never abates, son of the bad spirit with the spear, Yangonde’. Other drum names (i.e. the individual’s portion) include such comments as ‘The proud man will never listen to advice’, ‘Owner of the town with the sheathed knife’, ‘The moon looks down at the earth / son of the younger member of the family’, and, from the nearby Mba people, ‘You remain in the village, you are ignorant of affairs’. (citations omitted)
So the drum languages seem to put importance on uniquely identifying individuals, something that our Internet is just starting to figure out. (See ORCID.) Reputation of individuals was important; I wonder if creators of particularly compelling drum poems were identified by custom, as we're starting to learn how to do with Attribution licenses.

It goes without saying that the literary forms transmitted by drumming were not copyrighted and the there was no notion of paying a creator for "copies" of a drummed message. But certainly the practitioners of this early digital literature were valued by their societies.
drumming tends to be a specialized and often hereditary activity, and expert drummers with a mastery of the accepted vocabulary of drum language and literature were often attached to a king’s court. 
Masters of unwritten literatures found many ways of making a living. The "court poet" is a familiar role to us; modern writers often find wealthy patrons. In addition, Finnegan relays another way that creators of oral literature earned their livings:

The singer arrives at a village and finds out the names of the important and wealthy individuals in the area. Then he takes up his stand in public and calls out the name of the individual he has decided to apostrophize. He proceeds to his praise songs, punctuated by frequent and increasingly direct demands for gifts. If they are forthcoming in sufficient quantity he announces the amount and sings his thanks in further praise. If not, his innuendo becomes gradually sharper, his delivery harsher and more staccato. This is practically always effective—all the more so as the experienced singer knows the utility of choosing a time when all the local people are likely to be within hearing, in the evening, the early morning before they have left for the farm, or on the occasion of a market which leaves no escape for the unfortunate object singled out for these ‘praises’. The result of this public scorn is normally the victim’s surrender. He attempts to silence the singer with gifts of money or, if he has no ready cash, with clothes or a saleable object like a new hoe.
So even "astroturfing" is not an exclusively modern phenomenon.

I learned all this reading Ruth Finnegan's Oral Literature in Africa which was the first book made free to the world by Unglue.it, working with Open Book Publishers. Download and enjoy. Open Book Publishers have just launched an ungluing campaign for a second book, called Feeding the City, a translation of a seminal work from the original Italian, about the dabawallahs of Mumbai, a subject Internet entrepreneurs could learn a lot from. Support the campaign to make it free to the world!
Enhanced by Zemanta

Wednesday, February 29, 2012

The Slow Books Movement

An interesting idea is a powerful thing, but you can't eat it. You really have to give it away to make it worth anything, and even then you can't eat it. So what if you're good at having interesting ideas? You can go arounds giving talk at conferences. If the idea is incomprehensible enough, maybe you can even get tenure and a salary out of it. But what if it's an easy idea, one that resonates with a lot of people? Then your best option for making a living off the idea for a while is to write a best-selling book.


 Clay Johnson has an interesting idea, and I encountered his talk at this month's Tools of Change for Publishing Conference. I can even explain it for you, but of course Johnson does it much more entertainingly, and the TOC talk is on the web. In short, Johnson's thesis is that the things that are right and wrong with out information diets are the same things that are right and wrong with our food diets.

Humans are evolved to prefer fat and sugar over the things that are good for us because these foods were scarce in the environment we evolved in. Fat and sugar aren't scarce anymore, and the result is that we eat too much of them and are not as healthy as we could be. Worse, the economic structure of our society has created incentives for corporations to create huge complex systems that efficiently feed us sweet fatty food at a low cost. It's what we want. Fast Food.

Thus the "slow food movement". Slow food is the opposite of fast food. Inefficiently and locally produced with few of the benefits of technology or high fructose corn syrup from Iowa. Cooked from raw ingredients, crafted by a grandmother in a small kitchen. High in fiber, vitamins, protein and minerals. Low in antibiotics and fossil fuel byproducts. I exaggerate, but I'm smelling a frozen pie in the oven, and my zeal for slowness is tempered by my desire... to finish writing before it cools!

We suffer similarly from our "information diets". People want information that confirms what they already believe. Who wants to be intellectually challenged, really? We want sweet stories that entertain us, make our lives seem justified, and feed the emotions we most want to feel. And so we get information products designed to give us what we want, not what's good for us. Our news is designed and selected for search engine optimization rather than human wisdom maximization. It's Fox News and MSNBC, not Walter Cronkite and the New York Times.

Johnson's idea and his presentation of it are so tasty, so delectable, so comforting, that once you absorb it you become suspicious that it's self-referential. Are we liking the idea because it's confirming a bias we already have? Is it a "Fast Idea"? Aaaaack!

Luckily there's a book, The Information Diet, published by O'ReillyThe book's cover looks like a cereal box, the kind that's supposed to be good for you. The book is thin and lean, and has little of the fat or sugar of Johnson's presentation. It has footnotes and ScienceDirect URLs in all their unshortened bleakness. There are no pictures. It's the ideal souvenir from a Clay Johnson talk. It grounds Johnsons talk in tight argument and factual background. It's a "Slow Book" for a "Fast Idea".

And this leads Johnson to a concept that he calls "Infoveganism" in the book, but more appealingly, "Slow Information" in his talk.
A healthy Information diet always starts locally-and your political information should be no different. The goings-on of your state representatives and city and county governments, along with your school boards, and other local government offices are the best, healthiest forms of content for political news, and should be consumed over the national or global news.
Really. The main problem I have with Johnson's book is its "veganism". It's a sermon that, as much as I might agree with it, is not going to change my behavior.

It's my view that economic and social incentives are what ultimately determine the behavior of organisms, not information, whether slow or fast. If we want a better information diet, we need to alter the economics of information availability. That's why a society with libraries can be better that a society with only bookstores, Harry Potter notwithstanding. It's also why I'm developing Unglue.it. If authors and publishers are rewarded for the value that people have received from books rather than for the catchiness of the cover art or the placement in the bookstore, we'll have better books to read. The books will last longer, and they'll connect people in richer ways. They'll be "slow books" because the economic model will have a long view.

I think I'll skip the whipped cream on my apple pie!

Wednesday, September 21, 2011

Can JSTOR Solve the Course-Assigned eBook Problem?


Just imagine if libraries were like airlines. Your flight to Information could cost $20 or $2,000 depending on how far in advance you book, whether you're connecting through Seattle or going direct, how full the flight is, the day of the week, the phase of the moon, and who knows what else. Airlines employ operations research Ph.D.s who create and maintain elaborate systems to "manage the load". These systems make sure that people who have little choice of when and where to fly pay a lot more than others who have flexibility to seek low fares.

You don't have to imagine very hard to realize that academic libraries are part of a system that's a lot like the airline industry. Consider college students. Once they've signed up for a class, they have little flexibility with their assigned reading. Even a college library with a 10 million volume collection won't be of much use to them, because someone else will have checked out the assigned reading – if it's not on reserve. The student is sent to the college bookstore, where it's a pretty good bet the book won't be on sale. If the student has planned in advance, they might bought a copy used on Amazon, or borrowed it from an upperclassman.

Academic book "load management" was not devised by financial engineers, it arose as a by-product of a paper-based distribution system. Students don't perceive librarians as evil just because they haven't bought a hundred copies to supply the whole class. But the environment is very different when the books are digital. Scarcity of digital books is a manufactured fiction, so it's easy for a student to perceive the library as being complicit in a money-extraction machine, never mind the benefits that accrue globally when publishers are rewarded for producing books worthy of being assigned in a course.

Libraries hate saying “No” to their users, so ebook models that allow unlimited simultaneous access are very attractive to them. Even limited simultaneous access would help meet the needs of students. The same models are very scary for publishers, because they count on profits from selling many copies of course-assigned texts to cover their losses on slow-selling titles. Academic libraries don't have the money to replace that revenue, because the embedded practice is for students to pay for their own textbooks.

A few universities, notably Indiana University, have started to reexamine the role of academic libraries in the digital provision of course-assigned textbooks. Libraries have a better bargaining position with publishers than students, and are in a unique position to integrate digital texts with databases and other types of information available from the library. It makes little sense to tell students to come to the library for quality information- unless they need it for a class!

Publishers also have incentives to find library-based digital solutions for course-assigned reading. If the library can be charged based on the number of students enrolled in a course, no revenue is lost to used-book sellers, peer-to-peer networks or students who don't bother to read the texts. Solutions that work have nonetheless been elusive, partly because it's very difficult to have print and digital models coexisting, but also because budgets are tight all around.

At Ithaka's Sustainable Scholarship meeting this week, JSTOR's Bruce Heterick told the gathered publishers and librarians that JSTOR's ebook program, to be launched in 3rd quarter of 2012, would include a pilot program to address the course-assigned title problem. This news was warmly received by the entire audience. Books at JSTOR will make over 15,000 books available on multiple platforms (including ebrary and NetLibrary) and on multiple devices (iOS and Android). At launch, the ebooks will be offered on a sales model, and all the ebooks will be archived in Ithaka's Portico service so that libraries can be assured that their access will really be perpetual. The icing on the cake is that the ebooks will be integrated with JSTOR's discovery and crosslinking platform, which is very popular with the students and faculty at subscribing institutions.

Although the details of the course-assigned title pilot were not yet decided, JSTOR Managing Director Laura Brown responded to questions by saying that as many as 3 different models would be tried in the pilot. JSTOR's subscriber outreach has suggested that different approaches are needed in different sorts of institutions. I suppose you could think of these models as “charter flights” that can be added to package tours.

Regular readers of the Go To Hellman blog will know that I'm working on a new approach to selling ebooks. I view every market problem as a possible opportunity for the new model, which we call "unglued ebooks". Course-assigned titles are no exception.

My analysis indicates that the market dynamics for some books may favor the ungluing model. It works by aggregating donations by many people and institutions with a stake in a particular book and then paying the book's rights holders who “name their own price” to issue a Creative Commons licensed digital edition. The ebook can then be used without limit by everyone, everywhere. (OK, it’s backwards from Priceline, but we totally have to get Leonard Nimoy as our spokesperson!)

In situations where two titles compete with each other to be put on a course list, the ungluing model introduces a severe form of price competition. A title that is successfully unglued, even at a price equal to the present value of its entire future revenue stream, would have a huge advantage over a competing title that remained on the per-copy selling model. It remains to be seen, of course, whether students and libraries will be able to organize an ungluing campaign to meet the high ungluing prices commanded by books with steady recurring sales. Still, there is a huge variety of courses and books, and thus a reasonable chance that the model will work in at least a few cases.

There’s a great need for experimentation and risk-taking by libraries and publishers in the transition to the digital environment. It's good to see JSTOR step up to that plate. But don’t take my airline industry analogy too seriously. When developing new models for ebooks and libraries, let’s omit the pat-down security searches and checked-baggage fees, OK?
Enhanced by Zemanta

Sunday, May 29, 2011

Unbound wants to be the Kickstarter for Books

... and what they really are is the editor-curated, agent-filtered Gluejar for books that haven't been written.

Ignoring the fact that tomorrow is Memorial Day (and why aren't you Brits out barbecuing anyway?), I feel compelled to write promptly about today's unveiling of Unbound, because many of the words being used to describe Unbound are similar to things I've written about Gluejar.

One of my early attempts to describe our model of "Ungluing eBooks" was that Gluejar would be "like Kickstarter for ten million books". I found that approximately 50% of the subjects tested were mystified by that description, and the other 50% got totally the wrong idea.

So the introduction of Unbound allows me the chance to compare and contrast Unbound, Kickstarter, and GlueJar.

Business model

Unbound is a conventional publisher that asks readers to pre-fund some or all of the fixed cost of producing a book that hasn't been written yet. Unbound tells you how many supporters a book needs, but not how much cash. Unbound doesn't tell you how much money they get or how much the authors get, and once a project is subscribed, Unbound publishes the book, and splits net profits 50/50 with the author.

Kickstarter is not a publisher at all. They just let creators ask for a specific amount of money to support their projects, including projects that might result in the production of a book. Kickstarter takes a 5% fee from funds raised; 100% of subsequent profits from a book go to the creator.

Gluejar won't be a publisher as the term is currently understood.  Gluejar will allow book lovers to pledge support for making books (that already exist) free to the world in a creative-commons licensed ebook edition. Authors retain commercial rights for print and other subsidiary rights. The price is set by the rights holder to match or exceed the income they would expect for future sales of the ebook; Gluejar takes a fee similar to Kickstarter's from funds raised.

Selection Process

Despite their slogan "Books are now in your hands", Unbound is using a selection process that's pretty much identical to how it already works. Book proposals will be carefully curated, and Unbound is only going to deal with submissions coming from literary agents. So if you're Monty Python's Terry Jones, great. I feel so empowered.

Kickstarter also reviews projects rather carefully to ensure quality. But anyone can propose a project, and it's clear from the projects on the site that it's relatively open to newcomers and nobodies with good ideas. It's really the crowd that decides what flies.

Gluejar will allow patrons to pick books for themselves. Although there are a huge number of books out there, you already know which ones you love. We're not sure how to extend the concept to new books or new authors.

Risks

Because Unbound acts as the publisher, supporters have a reasonable assurance that a completed project will actually deliver a book. On the downside, a book that has already been funded might turn out to be less than promised. The incentives encourage the author to split a narrative into multiple volumes, and if a book turns out to be bad, or perhaps just dull, the supporters don't get their money back. I don't know what Unbound means when it says that "All unused credits expire after 30 days."

Kickstarter doesn't do anything to assure that projects get completed. Supporters have to judge for themselves whether the creator is honest and worth supporting.

Gluejar will act as a trusted third party to make sure that good quality, Creative Commons editions are delivered to patrons of a successful pledge campaign. What Gluejar can't guarantee is that a rights holder with all the needed rights for relicensing a particular book will exist. That's why we'll let supporters spread their pledges onto lists of books.

Bottom Line

Unbound has launched with a very nicely done website. They've done a nice job of setting up reward levels and website features matched to book publishing. But Unbound is profoundly timid about putting publishing into the hands of the reader. It's more of a brilliant marketing gimmick than a publishing revolution; they've mapped out a healthy way to pre-sell an ebook for £10.


Thanks to @pablod, @julietalionetti and @muttinmall for a great discussion bringing out some of these issues.

Saturday, March 5, 2011

eBook Carrots for Libraries

"Provide a great service and charge a lot for it" was the advice of an old friend who became a successful businesswoman. I frequently think of this advice; I have sometimes failed to follow the second part and have mostly regretted it. If you provide what your customers value, you should have no qualms about asking them to pay a premium. If you don't give the customers what they value, they won't be happy even if you give them a big discount.

The results of the dual survey I posted on Monday are confirming my guesses about HarperCollins' new strategy for limiting checkouts of ebooks they license to libraries though Overdrive, which sparked the so-tagged #HCOD furor. (The limitations are in addition to a one user at a time limitation imposed on these ebooks.) The results indicate that HarperCollins' new service terms don't give the customers what they value. They'll be unhappy, even if they're offered big discounts.
At what price discount would your library opt for a 26-check-out ebook?
At what premium would your publishing company offer an unlimited-check-out ebook?
The survey for publishers has only attracted 28 responses so far, not enough to make anything other than very broad statements. The survey for libraries has attracted 155 responses, and thus has much better statistics. The poll is in no way scientific; there is sure to be significant sampling bias. In other words, the survey only measures the opinions of librarians and publishers who are motivated to answer.

Significantly, 37% (±5%) of librarians indicated they would not purchase limited-check-out ebooks at any price. I would characterize this response as arising from non-quantitative considerations, which might be practical, ideological or philosophical. A similar percentage of publishers, 28% (±12%) indicated that no amount of money would convince them to offer an unlimited-check-out ebook (which is the most common type today). So it seems that publishers also have considerations that transcend math, which I find a bit surprising.

If we compare the rest of the responses, omitting the non-quants, we see that the librarians perceive a much lower value for limited-check-out ebooks than do publishers. 52 of these 97 librarians would purchase limited-check-out ebooks only if the they were priced at a quarter or a tenth of the ebooks offered without checkout limitations. In contrast, only 1 of 18 quantitative publishers thought the relative value of limited-check-out ebooks was so small.

What's clear is that even omitting the non-quant responses, librarians are perceiving the new HarperCollins licenses as being worth a small fraction of the previous licenses, offered at the same price. It's not surprising that they think it's an awful deal. It's a stick, not a carrot.

Publishers SHOULD be valuing the two licenses based on revenue lift, and they don't seem to expect a huge revenue lift by limiting check-outs. 10 of 18 quantitative publisher respondents seem to expect a revenue difference of 50% or less. My guess is that they're roughly right; I will do some modeling based on library check-out statistics and report on that next week or so.

Looking at the survey results from the other side, librarians are reporting that they put a huge value on the "durability" of the ebooks they license. They don't want books of any kind that wear out! Publishers that want to deliver the highest perceived value (and thus justify the highest prices) should consider finding ways to add to this quality.

One way to increase an ebook's durability is to use standard formats, such as ePub or PDF. This increases a library's confidence that the ebooks will survive into the future; ePub and PDF are the formats used by Overdrive. Unfortunately the DRM ("Digital Right Management") systems that wrap these files are proprietary, and there is a risk that a library's "purchases" will disappear if their ebook platform vendor (Overdrive) or DRM provider (Adobe) disappear in the future. Libraries are used to thinking with long time horizons, and it's a rare library that doesn't have books over 50 years old, much older than either Overdrive or Adobe.

The simplest way to add to the long-term durability for ebooks is to provide libraries with DRM-free, not-for-circulation files in addition to the  DRM wrapped files for circulation. Libraries are used to dealing with license restrictions and have a good record of compliance in this sort of matter; it's likely they would opt to delegate the safekeeping of such files to third-parties. They'd also want to be able to use the files to replace the statutory copying of print books allowed to libraries under US copyright law and to aid discovery in their catalog systems.

Another way to increase the value of an ebook license to libraries without reducing publisher revenue is to selectively allow those uses that are most likely to create publicity and lead to sales. Imagine what would happen if most library ebooks allowed simultaneous use in the first month after a book's publication. This would help libraries attract patrons with "hot" items, and would likely increase total sales by building buzz. Many library readers would want to purchase the book once their loan period expired. More patrons for libraries translates into stronger funding, (or at least less cuts!) which in turn allows for better acquisition budgets.

Andy Woodworth has some more ideas on making ebook rights packages that would be attractive to libraries, and I'm sure there are be many more ways for publishers to offer ebook carrots to libraries. Or at least a parsnip.

Updates: The polls remain open. Gluejar is still hiring, but it's looking like the team will be awesome!
Enhanced by Zemanta

Saturday, January 8, 2011

Inside the Dataculture Industry

wild blueberries
I don't really know how all the food gets to my table. Sure, I've gathered berries, baled hay, picked peas, baked bread and smoked fish, but I've never slaughtered a pig, (successfully) milked a cow or roasted coffee beans. In my grandparents generation, I would have seemed rather ignorant and useless. Agriculture has become an industry as specialized as any other modern industry; increasingly inaccessible to the layperson or small business.

I do know a bit about how data gets to my browser. It gets harvested by data farmers and data miners, it gets spun into databases, and then gets woven into your everyday information diet. Although you've probably heard of the "web of data", you're probably not even aware of being surrounded by data cloth.

The dataculture industry is very diverse, reflecting the diversity of human curiosity and knowledge. Common to all corners of the industry is the structural alchemy that transmutes formless bits into precious nuggets of information.

In many cases, this structuring of information is layered on top of conventional publishing. My favorite example of this is that the publishers of "Entertainment Week" extract facts out of their stories and structure them with an extensive ontology. Their ontologists (yes, EW has ontologists!) have defined an attribute "wasInRehabWith" so that they can generate a starlet's biography and report to you that she attended a drug rehabilitation clinic at the same time as the co-star of her current movie. Inquiring minds want to know!

If you look at location based services such as Facebook's "places", Foursquare, Yelp, Google Maps, etc, they will often present you with information pulled from other services. Often, a description comes from Wikipedia and reviews come from Yelp or Tripadvisor and photos come from Panoramio or Flickr. These services connect users to data using a common metadata backbone of Geotags. Data sets are pulled from source sites in various ways.

Some datasets are produced in data factories. I had a chance to see one of these "factories" on my trip to India last month. Rooms full of data technicians (women do the morning shift, men the evening) sit at internet connected computers and supervise the structuring of data from the internet. Most of the work is semi-automated, software does most of the data extraction. The technicians act as supervisors who step in when the software is too stupid to know when it's mangling things and when human input is really needed.

There's been a lot of discussion lately about how spammers are using data scraped from other websites and ruining the usefulness of Google's search results. There are plenty of companies that offer data scraping services to fuel this trend. Data scraping is the use of software that mimics human web browsing to visit thousands of web pages and capture the data that's on them. This works because large websites are generated dynamically out of databases; when machines assemble web pages, machines can disassemble them.

A look at the variety of data scraping companies reveals a broad spectrum. Scraping is an essential technology for dataculture; as with any technology, it can be used to many ends. One company boasts of their "massive network of stealth scrapers capable of downloading massive amounts of data without ever getting blocked. Some companies, such as Mozenda, offer software to license. Others, such as Xtractly and Addtoit are strictly service offerings.

I spoke to Addtoit's President, Bill Brown, about his industry. Addtoit got its start doing projects for Reuters and other firms in the financial industry; their client base has since become more "balanced". Companies such as Bloomberg, Reuters and D&B get paid premiums to provide environments rich in structured data by customers wanting a leg up on competitors. Brown's view is that the industry will move away from labor intensive operations to being completely automated, and Addtoit has developed accordingly.

A small number of companies, notably Best Buy, have realized that making their data easily available can benefit them by promoting commerce and competition. They have begun to use technologies such as RDFa to make it easy for machines to read data on their web sites; scraping becomes superfluous. RDFa is a method of embedding RDF metadata in HTML web pages; RDF is the general data model standardized by the W3C for use on the semantic web, which has been discussed much on this blog.

This doesn't work for many types of data. Brown sees very slow adoption of RDFa and similar technologies but thinks website data will gradually become easier to get at. Most websites are very simple, and their owners see little need or benefit in investing in newer website technologies. If people who really want the data can hire firms like Addtoit to obtain the data, most of the potential benefits to website owners of making their data available accrue without needing technology shifts.

The library industry is slowly freeing itself from the strictures of "library data" and is broadening its data horizons. For example, many libraries have found that genealogical databases are very popular with patrons. But there is a huge world of data out there waiting to be structured and made useful. One of the most interesting dataculture companies to emerge over the last year is ShipIndex. As you'd expect from the name, ShipIndex is a vast directory of information relating to ships. Just as place information is tied together with geoposition data, ShipIndex ties together the world of information by identifying ships and their occurrence in the world's literature. The URIs in ShipIndex are very suitable for linking from other resources.

The Götheborg
ShipIndex is proof that a "family farm" can still deliver value in the dataculture industry. The process used to build ShipIndex. Nonetheless, in coming years you should expect that technologies developed for the financial industry will see broader application and will lead to the creation of data products that you can scarcely imagine.

The business model for ShipIndex includes free access plus a fee-for-premium-access model. One question I have is how effectively libraries will be able leverage the premium data provided with this model. Imagine for example the value you might get from a connection between ShipIndex and a geneological database bound by passenger manifests. I would be able to discover the famous people who rode the same ship that my parents took to and from the US and Sweden (my mom rode the Stockholm on the crossing before it collided with the Andrea Doria). For now though, libraries struggle to leverage the data they have; better data licensing models are way down on the list of priorities for most libraries.

Peter McCracken
ShipIndex was started by Peter and Mike McCracken, who I've known since 2000. Their previous company (SerialsSolutions) and my previous company (Openly Informatics) both had exhibit tables in the "Small Press" section of the American Library Association exhibit hall, where you'll often find the next generation of innovative companies serving the library industry. They'll be back in the Small Press Section at this weekend's ALA Midwinter meeting. Peter has promised to sing a "shanty" (or was that a scupper?) for anyone who signs up for a free trial. You could probably get Mike to do a break dance if you prefer.

I'll be floating around the meeting too. If you find me and say hello, I promise not to sing anything.
Enhanced by Zemanta

Saturday, December 11, 2010

The Most Important e-Reader Company You've Never Heard Of

Imagine you're the head of a big print media company. Sales of your print products have been eroding steadily and you find yourself competing with internet businesses that have very different business models and cost structures. Your own website revenues have been growing steadily, but it will be many years before they match your print subscription revenue. You know in your heart that digital subscriptions will somehow be the answer, but your otherwise loyal reader base is resisting web subscription rates that would allow you to sustain your business.

What do you do?

One answer is something I've speculated about, based on cost reduction trends for tablets and ebook readers. Bundle a reader device into your subscriptions! The consumer gets a device for "free", and you've retained a premium subscriber. You could even bundle a shopping application and collect a commission on purchases of content made through your estore.

At first glance, the obstacles to a successful implementation of the bundled e-reader strategy might seem profound. Here's an incomplete list of what you'd have to do:
  1. You'd need an inexpensive device, of course. But it can't be something cheesy, because it will be carrying your brand.
  2. You'd need an operating system for your device. Lucky for you, Google seems to have a good solution with Android. But you still need people who can customize and adapt Android to make it shine.
  3. You'd need to figure out the logistics of getting your content onto the device. Maybe you've already started this process with Apps for iPad or Android.
  4. If you're serious about an estore, someone has to build that, too. It's not only an engineering project it's also a big business development task to gather businesses willing to sell their content and goods through your estore.
  5. You'll need to build a customer support operation.
  6. You'll need some big marketing power and a sales channel.
Most big media companies have only the last item in their portfolio of competencies, so you might expect this isn't going to happen anytime soon.

But you'd be wrong.

On Thursday, my rounds in Bangalore took me to Ninestars Information Technologies Ltd. Ninestars specializes in back-end infrastructure for publishing companies. They started out doing newspaper backfile digitization, and today they work with a Who's Who of the newspaper and publishing world. It's Ninestars, an Indian company,  that's actually doing much of the heavy lifting in the preservation of the entire world's history. Early on, Gopal Krishnan, the company's Chairman, made a crucial decision to steer away from the brute-force, labor intensive approach to digitization in favor of developing technology and automation so that jobs could be done faster and with fewer people. That decision has paid off as publishing has become more and more dependent on advanced technology, whether it's website and e-commerce development, content delivery, or mobile and tablet apps.

Ninestars operates under a "white-label" business model. If you go to their website, you'll be surprised at its small size and inertness. But don't judge a book by its cover- Ninestars and Gopal are major players. Playing the white-label role to the hilt, Ninestars wants the spotlight to shine on their customers, as if Ninestars didn't even exist.

Ninestars' e-reader strategy is similarly white-label. Tablets and e-readers are being developed by Ninestars' associated company DisplayTronics Reader Devices Ltd.. DisplayTronics CEO Dr. Somanath V S gave me a peek inside the DisplayTronics development laboratory. I saw a range of e-reader prototypes- both EPD and color LCD, all with touchscreens. These will be sold under brands that are already familiar to consumers.

DisplayTronics is also building a white-label "estore" that its customers will use to sell their content, whether on their branded e-readers, or through apps on other device platforms.

While I've written a fair amount here about business models for ebooks, I haven't really covered business models for e-readers. Everyone knows about Apple's hardware-oriented business model and Amazon's commerce oriented business model, but there's not been much in the way of business models that put content at the center -yet. With e-reader pricing dropping at a relentless pace, big media companies are potentially in an excellent position to remake the e-reader market in a ways that sustain their businesses. Companies that try to make money by selling devices might find it difficult to compete against free or subsidized devices bundled into media subscriptions.

We'll find out how well the bundled-e-reader business models work when they launch - in 2011.
Enhanced by Zemanta

Tuesday, December 7, 2010

Lots of Markets, Lots of Business Models

from Wikipedia
The book industry is a lot like the Soviet Union. The Soviet Union consisted of fifteen ethnically divergent states (soviets) stitched together by a highly centralized government model. When that government model weakened, it turned out that there was little holding the soviets together. The Soviet Union no longer exists.

Discussions of the ongoing transition from print to digital books frequently make reference to corresponding transitions that have occurred in the music industry and in the film industry. As I've learned more and more about how the book industry has worked in the age of print, it's been made clear to me that unlike music and film, the book industry has consisted of a large number of disparate industries thinly stitched together by a common delivery mechanism and shared supply chain.

Predicting what will happen to the book industry in its shift to digital by looking at music is like making predictions about the Soviet Union by looking at Germany.

Consider the different ways that books can be of value. The utility of a cookbook has very little in common with the utility of a romance novel. A dictionary and a travel guide are very different, even though you might take both on a trip. A manual on object oriented programming and a superman comic book might both be useful for squashing bugs, but the former should be open and the latter should be rolled up when doing so. I won't even mention coffee table books.

Compare the music industry. No matter the genre, most everyone uses music in the same way. Whether rap or raga, Beethoven or Roll Over Beethoven, the variations in consumption patterns are relatively small.

When utility variations of goods are small, the business models behind their distribution tend to converge. The film and television industries had very different pre-digital distribution models, but their business models are converging today because they are valuable in much the same way. Scholarly journals are another good example. 10 years ago, when journals began the transition from print to digital, publishers started out with a variety of business models for different fields. Today, however, the vast majority of journal publishers use roughly the same business model.

That's why I think the book industry of the past is fragmenting into many smaller industries based on business models that deliver the most value. We've already seing this begin to happen. The old business model of the encyclopedia industry has already died;  Wikipedia has almost wiped out the incumbent encyclopedia publishers with a public charity business model.

Whenever O'Reilly Media and their willingness to do free-on-the-web versions of their books come up in a conversation among publishers, one of them will say something like "well that's a very different market segment from ours". And it's true. The utility of a book such as Programming in Perl is of a sort that free-on-the-web + ebook + print works well.

While I was in London I had a chance to chat with Frances Pinter, who I've mentioned on the blog. Her imprint, Bloomsbury Academic, is experimenting with another business model, one that resembles the "freemium" models common on the web. Basic ebooks will be available under a creative commons license, but print and enhanced digital versions will be available for purchase. Other publishers, such as Springer, are going with subscription based models for institutions. The subscription model, without DRM, works just fine for ejournals, and Springer is confident that it will work just fine for eBooks in libraries. I've previously written about other business models for ebooks, such as PDA (Patron Driven Acquisition) and of course, "collective acquisition" and "bounty markets". Other models that are likely to work in at least some of the book industry fragments include advertising and what I will henceforth call the PIP (Pretend It's Print) business model.
London Online Information

As I learned at October's ICv2 Digital Comics Conference, digital publishers of graphic novels are finding that a serial mini-book business model works well for their industry. By pricing smallish chunks of content at under a dollar, a series can build readership momentum and do well financially without DRM.

Geography and language are also likely to define book industry fragments. At the London Online Information show, I met a group of entrepreneurs from Estonia. Their company, LibroCS OÜ , has developed an ebook delivery platform that includes a storefront and nifty color-video capable ebook reader devices manufactured in Shenzhen. They hope to provide infrastructure for retailers and publishers, especially those in smaller countries like Estonia.

I was interested to hear how the market dynamics in Estonia might be very different from those in a larger country. There are only about 1.1 million speakers of Estonian in the world, but it's a very well connected group. (The one Estonian I went to school with is brother to the President, so it must be true!) As a result, any Estonian is only one Facebook friend away from any other Estonian, or so it's claimed. If you ask an Estonian publisher to allow lending of an ebook, they immediately think they'll only sell 2 copies, 3 if the author's mother is living. To address these concerns, LibroCS emphasizes their use of Adobe Content Server DRM to enable PIP business models. Just because the print model is old doesn't mean it's dead.

Physics tells us that interacting systems eventually reach their lowest energy state, although there may be some heat required for them to reach it. The corresponding rule in economics is that a superior business model will sooner or later drive out its inferiors. In the context of the Former Book Industry, "superior" means that it offers the best value relative to its cost. Business model experimentation is what provides the heat. An English chemist friend of mine pointed out to me this week that when the experiment releases energy, you can get an explosion.

For the record, many Estonians refuse to acknowledge that Estonia was part of the Soviet Union. I don't know what that means for book publishing.

Wednesday, November 10, 2010

Infochimps and the scaling of dataset value

Image representing Infochimps as depicted in C...Image via CrunchBaseSure, a picture is worth a thousand words, but what is a thousand words worth? How about a million? If I had a dataset of the most recent trillion words spoken by humanity, (anonymized and randomized of course!) would that be worth any more than the set of words in this blog post?

These are real questions. A Texas company called Infochimps has datasets quite similar to these, ready for you to use. Some of the datasets are free, others you have to pay for. More interesting is that if you have a dataset you think other people might be interested in, or even pay for, InfoChimps will host it for you and help you find customers. (Infochimps just announced they had raised $1.2 million in its first round of institutional funding.)

One of the datasets you can get from Infochimps for free is the set of smileys used on twitter in tweets sent between March 2006 and November 2009. It's free. It tells you that the smiley ":)" was used 13,458,831 times, while ";-}" was only used 1,822 times.

If you're willing to fork over $300, you can get a 160MB file conatining a month-by-month summary of all the hashtags, URLs and smiley's used on twitter during the same period. That dataset wil tell you that during September of 2009, the hashtag #kanyeisagayfish was used 11 times while #takekanyeinstead was used 141 times.

If you're a scrabble player, you can spend $4 for a list of the 113,809 official words, with definitions. Or you can get them free, without the definitions.

courtesy of Infochimps, Inc. CC-BY-A
I had a great talk with Infochimps President and Co-Founder Flip Kromer a few weeks ago before his presentation to the New York Data Visualization Meetup. I fell in love with one of the visualizations he showed in his presentation, and he's given me permission to reproduce it here. (Creative Commons Attribution License) It's derived from the same Twitter data set you can get from Infochimps, and shows networks of characters that are found in the same tweet. So if ♠ and ♣ appear in the same tweet over and over again, the two characters will have a strong connection in the network of characters.

The character connection data was fed to a program called Cytoscape, which is an open source visualization program used in bioinformatics; Mike Bergmann has a nice article about its use for large RDF graphs. The networks are laid out using a force-directed algorithm (which is pretty much the simplest thing you can do). Coloring is applied arbitrarily.

As you might expect, the main character networks that show up are associated with languages, but there are some anomalies. For example, the katakana character ツ (tu) sticks out. Katakana is a set of phonetic characters used in Japanese for non-Japanese words. The reason "tu" is set apart from all the other katakana is that people use it on Twitter as a smiley.

The other anomalous character subnet is labeled "???" in the graph. A closer look reveals this to be the set of characters that look like upside down roman text.

Kromer has noticed that the price (or perhaps cost) of a partial data set follows a non-monotonic curve (see graphic). Small amounts of data are essentially free, but a peak value is reached when portions of the data set are extracted from the full data set. If we were discussing book metadata, for example, peak value might accrue for a set of the 100,000 top selling books.

There's much less value, according to Kromer, in having a large incomplete chunk of a data set. Data for 10,000,000 books, for example, would have less value than the 100,000 book data set, because it's not complete. Complete data sets become extremely expensive because of the logistics involved, and because of the value of having the complete set.

This pattern seems plausible to me, but I'd like to see some clearer examples. I've previously written about having too much data, but that article looked at the effect of error rates on data collection; Kromer's curve is about utility.

For me, the most interesting thing about Infochimps is the idea that the best way to make data flow in large volumes and create new types of knowledge is to provide the right incentives for data producers through the establishment of a market. This makes a lot of sense to me; however I'm not sure that the Infochimps market has also established incentives needed for data set maintenance; the world's most valuable and expensive data sets are one that change rapidly.

Kromer contrasted the Infochimps approach to that of Wolfram, whose Alpha service is produced by "putting 100 PhDs and data in a lab". He also feels that much of the work being put into the semantic web is a "crock" because its technology stack solves problems that we don't have. Humans are pretty good at extracting meaning from data, given a good visualization.

We can even recognize upside-down text.
Enhanced by Zemanta