Monday, March 15, 2010

The Starbucks Library, Version 1.0

On Saturday afternoon, a gust of wind blew down a big branch from a tree in front of my house. A spectacular arc of electricity made me think we had been struck by lightning, but it was just the branch knocking out electric service on my block. So here I am on Monday afternoon, sitting in a local Starbucks powering up the batteries of my laptop and phone, listening to jazz music along with other power refugees. The WiFi tells me it's Fireheads, by Emiliana Torrini, from the Me And Armini album. I'm offered a download button, and a click later, I'm in iTunes, with a chance to buy.

When I said in my last post that libraries and publishers need to work out a new way to work together in the ebook future, I was thinking of something like a cross between Starbucks and the Storefront Library. Here and now, on a rain day in the wake of a nasty Nor'easter,  I can see it quite clearly. There are already bookshelves (filled with coffee) and plenty of outlets. Add more comfy chairs and a room for the kids, and this Starbucks could work quite nicely as an ebook library. If it can work for music, why not do the same thing for ebooks? Give patrons access to a huge library of books that can be read for free on the library premises; if they want to take it home on the reader device, they have to buy it. It seems like a reasonable quid pro quo for both libraries and publishers.

But then I started thinking. What if it's NOT libraries that do this? What if Starbucks decides someday to get into the ebook distribution business? What if publishers decide that they're more comfortable doing a guaranteed-revenue-stream deal with a single for-profit entity with over 16,000 locations worldwide than trying to roll out service in thousands of public libraries, each of them different in their own special way. Would public libraries become marginalized? Would people without money to spend on an iPad and expensive coffee be able to find ways to access information? What if a Starbucks-iBookstore connection is ALREADY in the works, set for debut on April 3???

I caught my breath and decided to take a few pictures. When I returned to my seat I saw that in the window just behind where I had been sitting was Starbucks Library version 1.0. Next to the window, a sign said simply: "Read 'em, share 'em, return 'em". And can you believe it, the children's section featured Go, Dog. Go!.

I'm looking forward to version 2. Starbucks will need to work on their metadata- the song I was listening to wasn't EmilĂ­ana Torrini at all. Still, when the wind blows after a hard rain, some trees will come down.
Reblog this post [with Zemanta]

Wednesday, March 10, 2010

eBooks in Libraries a Thorny Problem, Says Macmillan CEO

John Sargent, CEO of Macmillan, one of the US's "Big Six" publishers, is not afraid of new business models. Over the past year, Macmillan has been trying to figure out how to push ebook pricing above the $9.99 level that Amazon had set as a standard on the Kindle. They had explored "enhanced" ebooks- ebooks that come with extra content- and were about to implement "windowing" (holding back ebook release to protect hardcover pricing, something that Sargent felt was "completely stupid").

Instead, Sargent decided to take advantage of Apple's announced entrance into the ebook distribution game to force a change of Macmillan's business relationship with Amazon. Instead of using the same discount model for both ebooks and print books, Macmillan wanted Amazon to change to a "agency model" where pricing would be controlled by Macmillan and Amazon would take a percentage. Amazon (which is Macmillan's 2nd largest customer) balked, and stopped selling Macmillan books entirely. But two days later, Amazon gave in. As a result, Sargent has been called publishing's "new hero".

Sargent spoke with the "Publishing Point" Meetup Group today in New York City, and I got to participate in the questioning. Michael Healy, Executive Director Designate of the Book Rights Registry, did a great job of leading the conversation. I was very impressed with Sargent, who dressed in jeans and had a casual, down-to-earth manner that matched. Sargent clearly understands all the challenges his industry faces- disintermediation, shifting distribution, the need to develop technology expertise, but at the same time he's very optimistic about publishing's prospects. He understands the assets at his disposal, in his words, "a lot of extremely good people who know how to obtain manuscripts and who know what people want to read", and who know how to gather enthusiasm around a piece of writing, a process that's "magic".

The most amusing comments by Sargent came in response to Healy's questions about whether the large, generalist, publishing houses would continue to be viable. Sargent seemed to think that in the near term (5-10 years) the big 6 would likely remain intact. (HarperStudio's Robert Miller has predicted the Big 6 could shrink to 3) His reason was not what I expected. The Big 6 are in no danger of implosion- they survived a very hard economic stretch quite well, but no private equity firm or bank would go near them because of "disastrous" balance sheets. They "suck cash, and have terrible profits." "We're disastrous but stable" quipped Sargent.

When my turn came to ask a question, I asked Sargent if he had thought about the role of libraries, and particularly public libraries, in ebook distribution.  His answer indicated that just as he was not afraid of changing the relationship with Amazon, Sargent is not afraid of changing the publisher's relationship with libraries. In fact, change may well be required.

"That is a very thorny problem", said Sargent. In the past, getting a book from libraries has had a tremendous amount of friction. You have to go to the library, maybe the book has been checked out and you have to come back another time. If it's a popular book, maybe it gets lent ten times, there's a lot of wear and tear, and the library will then put in a reorder. With ebooks, you sit on your couch in your living room and go to the library website, see if the library has it, maybe you check libraries in three other states. You get the book, read it, return it and get another, all without paying a thing. "It's like Netflix, but you don't pay for it. How is that a good model for us?"

"If there's a model where the publisher gets a piece of the action every time the book is borrowed, that's an interesting model."

Sargent has clearly thought about libraries, but perhaps he's not talked much to them. His points are valid- the existing business relationship between publishers and libraries won't work for ebooks the way it has worked for print books and the "frictions" that exist for print materials could disappear for ebooks. But he has gaps in his knowledge of libraries. The patron-on-the-couch scenario wouldn't work for libraries either- why would a town support its library's ebook purchasing if everyone could get the ebook from a library 3 states away? The fee-per-circulation model would be a disaster for most libraries, which have fixed annual budgets, and can't just close in September if they've spent their circ budget.

On the other side, the models preferred by libraries are not necessarily going to work for publishers. While the subscription model will probably work for academic institutions, it would turn public libraries into unnecessary intermediaries. The "perpetual access" model would be suicide for publishers if applied to their most profitable top-line books.

Now is the time for publishers and libraries to sit down together and develop new models for working together in the ebook economy. Executives like John Sargent are not afraid of change, but they need to better understand the ways that they can benefit from working with libraries on ebook business models. Libraries need to recognize the need for change and work with publishers to build mutually beneficial business models that don't pretend that ebooks are the same as print.

Sunday, March 7, 2010

After 25 Years, My Mac Plus Still Works

On this day 25 years ago, I got my first Mac.
It had 128KB of RAM, a single-sided 3.5 inch internal floppy drive and a Motorola 68000 microprocessor running at 8 MHz. The black and white 9-inch CRT screen had a resolution of 512×342 pixels. My purchase was a bundle that included a 1200 baud modem, an Imagewriter printer, an external floppy drive, and a copy of MacPascal. As a Stanford student, I was eligible for a discount, so the whole package cost me $2,051.92, including sales tax. About a year later I got it upgraded to a Mac Plus.

I'm currently typing on the 8th Mac that I've used as my main computer. It's a MacBook Pro.  It has 4 GB of RAM, a 320 GB hard drive, An Intel Core 2 Duo Microprocessor running at 2.53 GHz, and a 15 inch color LCD screen with a resolution of 1440x900 pixels.

I never got rid of my original Mac. To celebrate its 25th birthday, I went up to the attic to bring it out for some air. My kids were excited to get a look at the antique. It still works.

What was interesting to me is that apart from being alarmed at the disk drive noises, and asking "is this what they called a floppy disk?", my teenagers sat down and immediately knew how to use MacPaint, MacDraw and Word 3.0. They understood how to interact with Ultima II. The graphical user interface notions introduced with the Mac are still alive and well.

This got me thinking about the longevity of user interfaces. For example, the rotary dial telephone that I grew up with was an interface introduced in the US in 1919. It lasted about 60 years. The Model T Ford that I wrote about last July had the same basic driver interface as my car does today and is still going strong, but the television I grew up with has almost nothing in common with the one I own today.

My all-time favorite YouTube video is taken from a Norwegian comedy show. It imagines what it might have been like for users when the new-fangled "book" came along:


The book's "user interface" (more precisely, the Codex) has had a pretty good run; it's in its third millenium. Kids 25 years from now will know how to use the codex interface, though I'm guessing they'll consider books to be hopelessly out of date, like the vinyl LPs that I had to move around to get at the Mac in my attic.

It won't be the Nook that replaces the book, though. I got to play with one the other night, and while it has some pretty interesting features, the user-interface, which uses a small touch screen and a larger e-ink display, is not long for this world.

It's possible that my long run of Macs will eventually end with a touch oriented device, such as the iPad. Its hard to imagine the devices that, 25 years from now, will make my very nice MacBook Pro seem as much an antique as my Mac Plus.

The kids lost interest in the Mac Plus after about 20 minutes. It had no internet.

More pictures of my Mac are on the Facebook fan page.
Reblog this post [with Zemanta]

Friday, March 5, 2010

Business Idea Number 3: Gluejar Book Search

A few years ago, I was invited to give a talk about the future of libraries at a library staff retreat. After the talk, the speakers were given a special tour of the library, which had recently undergone renovation. I was struck by the loneliness of the stacks. So many books, so much knowlege, so little usage.

As OCLC's Lorcan Dempsey has recently observed, the lawsuit over Google Book Search and its proposed settlement has highlighted the limitations on libraries' ownership of their book collections. There are many things that libraries would like to do with their books that they are prevented from doing by copyright law. The possibility that the Google Books service will enable libraries to reanimate their lonely book collections is the reason that libraries have, for the most part, been sympathetic to Google's digitization program.

One session at last week's Code4Lib conference sharpened my awareness of how libraries are struggling to acheive this reanimation on their own. There were 3 different presentations, from Stanford, NC State (3.65 MB ppt), and University of Wisconsin, Oshkosh, on "virtual bookshelves". The virtual bookshelf tries to enliven the presentation of an electronic library catalog by trying to reproduce part of the experience of browsing a physical library- sometimes the book you really need is sitting there next to the book you're looking for. It's an idea based on a sound user-interface design principle: try to present information in ways that that look familiar to the user.

The virtual bookshelf is not a new idea. Google has even been awarded a patent on virtual bookshelves- see the commentary here and here. Given that Naomi Dushay (who presented the Stanford work) wrote about Virtual Bookshelves in 2004, it appears to unlikely that the Google patent (filed in 2006) will apply broadly at all.

While the virtual bookshelf is a sensible and practical incremental improvement on the library catalog interface, it's also backward looking. People looking for information today want to search inside the books, not just "browse the stacks". But libraries don't have the ability (today) to search inside the books that they think they own.

Google Books could enable libraries to do just that. Google is spending huge sums of money to digitize books in libraries and make them searchable. When they got sued for doing this, the library community looked forward to having questions surrounding the fair use of digitized books settled in court. For example, while it's pretty clear that using digitization to create an full-text index of a book would be allowed as fair use, the display of "snippets" (as done by Google) may or may not be held to be a fair use of the page scans. When a settlement of the lawsuit was announced, much of the library community was disappointed that these fair-use questions would not be settled.

Google Books already allows users to set up book collections of their own and search them. The results come with snippets (see pictures), but if the settlement is approved, Google's ability to show snippets with vastly reduced infringement liability would leave it with a dominant position in libraries because of its ability to search inside huge numbers of books. If the settlement is not approved, Google's dominance would be similar, except that a copyright decision could shut down Google Books at some time in the distant and irrelevant future.

Some aspects of the settlement create holes in Google's index. As part of the settlement, rights holders can exclude their works from Google's index. Google's publisher partner program allows publishers to create these holes today. For example, even if you add Tolkein's "the Two Towers" in your Google library, Google won't let you search inside it. Only limited research uses can be made of the digitized works; as the Open Book Alliance's Peter Brantley has argued, it's very hard to tell what sort of innovations might arise from the availability of large numbers of digitized texts as data; the same goes for indices of these works.

Many other works have been excluded from the settlement. Works published only outside the US, Canada, UK and Australia, as well as works published in the US, but not registered with the copyright office, are not covered by the settlement. Works other than books, such as newspapers, magazines, and other periodicals are also excluded.

For these reasons and others, I've begun talking to people about "Gluejar Book Search". Gluejar Book Search would be a business focused on collecting, aggregating and redistributing full-text indices of copyrighted material. To comply with copyright law, it would focus on indices that can be distributed without infringinging copyright, and would help provide libraries and publishers with tools  to produce copyright-safe index documents.

I've frequently encountered the assertion that digitizing all the books in libraries is prohibitively expensive, and that only Google (or possibly the government) could possibly have the financial resources to do it. For example, Ivy Anderson reports an estimate by the California Digital Library that digitization of the 15 million books in the libraries University of California would take a half a billion dollars and one and a half centuries. There are two coutervailing arguments. First, the cost of book digitization software and equipment has rapidly fallen, and will continue to fall. Last year, I wrote about the Dan Reetz' DIY book scanner, but even commercial devices capable of both image aquisition and OCR are currently available for as little as $1,400. I described how it could cost as little as $10,000,000 to put scanners in 10,000 libraries to enable scanning of 5,000,000 books per year.

The other factor that could drastically lower the cost of producing digital full-text indices of all types of copyrighted materials is the drastically lower technical demands of an indexing system compared to that of an archival imager. Archival imagers produce huge scanned image files because of the need for high resolution in an archival image. The resulting demands on storage hardware are significant and expensive. In contrast, an index file can be quite small; the laptop I'm typing on could store indices for 3,000,000 books; I estimate that full-text indices of all the worlds books would today require at most ten commercially available hard drives.

Gluejar Book Search would be fueled by two main revenue streams. The first stream would come from customized search services to enable library patrons to search inside the library's books. The second stream would be to provide aggregated feeds of index files to mass-market and specialized search providers- Google's competitors, and book retailers such as Amazon and its competitors. Google may even want to acquire index files for works it has been asked to remove from its own index, such as the Tolkein book mentioned above.

A possible third revenue stream would come from partnerships with rightsholders willing to permit page or snippet display in exchange for link traffic. If a Book Rights Registry comes into existence, it's possible that many business models could be arranged without prohibitive transaction costs.

Part of the revenue from Gluejar Book Search could be returned to libraries, publishers and other institutions that have contributed index files to the aggregation. Libraries could choose to use these funds to fund further digitization; alternatively, they may prefer to contribute to an Open-Access index.

The success of Gluejar Book Search would depend to a significant extent on its ability to reach critical mass. If it could reach index 80% of a library's book collection, it would deliver significant value to the library. (That statement is based purely on conjecture- email me or leave a comment if you agree or disagree!) Critical mass might be rapidly attained by working closely with publishers and by partnering with low-cost digitization providers and existing content aggregators so obtain indices for the most widely held books. Once critical mass is obtained, the "long tail" could be addressed by encouraging the particpation of large numbers of libraries around the world.

A Gluejar Book Search business would require a significant but not huge raise of capital, if for no other reason than to address litigation risk. Although I believe the legal position of building copyright-safe book indices is secure, there are bound to be litigious rightsholders with a poor grasp of fair use under copyright. The other big risks involve Google. Google might well develop services that greatly undercut Gluejar Book Search's revenue streams. Finally, the "copyright-safe" approach might be completelyundermined if courts in many countries were to rule decisively for an expansive view of fair-use.

If you want to know more about Gluejar, read this post. I have been exploring many possibilities about "what to do next", and I've written about other ideas, as well. As always, I'm interested in feedback of all kinds. Over the next few months, I hope to develop this and other ideas in more depth, so stay tuned.
Reblog this post [with Zemanta]

Monday, March 1, 2010

eBook Pricing Calculus and A/B Testing

You've probably read about how book publisher Macmillan has won a big battle with Amazon over the pricing of ebooks. By shifting to an "agency" model, publishers will gain the ability to control the price that consumers pay for ebooks. A much discussed question has been whether this is really a win or a pyrrhic victory for publishers.

My question is a bit different. How will book publishers determine the correct pricing?

In Econ 101, we learned that markets set pricing by matching supply and demand curves. The publisher's task in the ebook economy is to find a price that will maximize their profits. Too high a price will result is low unit sales, while too low a price will leave money on the table.

One of the frustrations you encounter trying to apply Econ 101 lessons to the real world is that you quickly find that most supply and demand curves are completely hypothetical. When W. W. Norton & Company set a retail price of $13.95 for The Blind Side (Movie Tie-in Edition) they didn't solve a set of equations that told them their profit would be maximum at this value. Norton doesn't know how many copies they would sell at $99.95, and they don't know how many they would sell at 99¢. It's likely they know how many total copies they're selling, but they probably don't have solid numbers telling them how many of those are selling at $9.81, the current price on Amazon.

But Amazon does.

Booksellers like Amazon can map out a large part of a demand curve using A/B testing. In A/B testing, website visitors are divided into two groups. The A group sees one version of a website and the B group gets another. The behavior of the two groups is then measured and compared. For example, the two groups could be shown different pricing for The Blind Side, and the rate that they purchase the book would be measured. Using repeated measurements of purchase rate vs. price, a dominant retailer such as Amazon is able to measure the consumer demand curve for a book or group of books. Pricing and profit can be optimized accordingly.

Amazon's pricing calculus will be somewhat different from the publisher's calculus, however. If the price they pay publishers is fixed (as it is for books), then the optimum price for Amazon will be higher that the optimum pricing for the publisher. You can do the math.

Amazon is well known for doing A/B Testing- see Bryan Eisenberg's description of the evolution of the Amazon shopping cart for a great example. Google is also notorious for depending on the technique. It even tested 41 shades of blue when it couldn't decide on a color for a design element.

Book publishers, on the other hand, have little experience with running e-commerce websites. A successful web merchant will optimize their site for search engine ranking, and will make it simple for users to find and get what they want.

Try a Google search for "The Blind Side". Since the book has become an Oscar-nominated major motion picture starring Sandra Bullock, it's not surprising that the top hits relate to the movie, not the book, but the complete absence of publisher results is striking. Here are the links my Google search pulls up:
  1. Movie times
  2. IMDB (Amazon property, links to Amazon)
  3. the movie web site (Warner Bros., with move commerce links)
  4. Wikipedia (film)
  5. Google News Results (no book links)
  6. Google Image Search results (First one a book cover at AOL shopping)
  7. YouTube (official trailer) (no book links)
  8. Amazon page for the book At last, a place to buy the book!
  9. Rotten Tomatoes (movie reviews, no book links)
  10. Yahoo Movies (no book links)
  11. Apple iTunes Movie Trailers (no book links)
  12. Fandango (no book links)
  13. Moviephone (no book links)
  14. Google video search results. The second result is a link to a YouTube interview with The Blind Side Author Micheal Lewis, labeled "WW Norton: The Blind Side". It seems the publisher ponied up for some promotional video! But are there any links from the video to a book related page? Of course not!
On the second page of google results, the book gets a Wikipedia link and another Amazon link. On page 3, there's a book link to Powell's. On page 5, there's an  excerpt from the book on the NPR website.

Perhaps the publisher web presence for The Blind Side has been swamped by the movie pages.  If we add "book" to the search term we might expect to see a publisher presence for the book. On the third page of that search, there it is: a result from WW Norton. It's their home page, and no mention of The Blind Side at all. A message there tells me that WW Norton has
"recently relaunched our website, and many things have moved around. If you're looking for a book, try the search field above, or browse all books by subject." 
Oh, and when I search Norton for "the blind side", I find this page, which says the book is out of stock! If a competant merchant were running the site, it would tell me that the version without the movie-tie-in cover was in stock, but no such luck. However, there's a tiny link way on the other side of the page that says the book is available on the iPhone/iPod Touch iTunes App Store! Although no one has submitted a review on iTunes, I'm told I can buy it there from Kiwitech for $13.99.

There are so many things wrong with Norton's attempt at e-commerce that pricing is almost the last thing you would want to test with an A/B study.

So the funny thing about the shift to an "agency" model for the selling of ebooks is that the power to call the plays (set prices) now belongs to the one player (Norton, Macmillan, Random House, etc.) that has the poorest view of the ebook playing field; in fact, I'm not sure they all know the rules. The big huge left guard (Amazon) has just been benched even though he blocks like a superstar, because he's urged Norton to run the ball. Norton wants to pass the ball, to his stylish wide receiver, Apple, but the other team's blitzing, and a speedy right defensive end named Google is bearing down on Norton from his blind side.

I'm not sure I want to look.