Showing posts with label EPUB. Show all posts
Showing posts with label EPUB. Show all posts

Friday, October 14, 2016

Maybe IDPF and W3C should *compete* in eBook Standards

A controversy has been brewing in the world of eBook standards. The International Digital Publishing Forum (IDPF) and the World Wide Web Consortium (W3C) have proposed to combine. At first glance, this seems a sensible thing to do; IDPF's EPUB work leans heavily on W3C's HTML5 standard, and IDPF has been over-achieving with limited infrastructure and resources.

Not everyone I've talked to thinks the combination is a good idea. In the publishing world, there is fear that the giants of the internet who dominate the W3C will not be responsive to the idiosyncratic needs of more traditional publishing businesses. On the other side, there is fear that the work of IDPF and Readium on "Lightweight Content Protection" (a.k.a. Digital Rights Management) will be a another step towards "locking down the web". (see the controversy about "Encrypted Media Extensions")

What's more, a peek into the HTML5 development process reveals a complicated history. The HTML5 that we have today derives from a a group of developers (the WHATWG) who got sick of the W3C's processes and dependencies and broke away from W3C. Politics above my pay grade occurred and the breakaway effort was folded back into W3C as a "Community Group". So now we have two, slightly different versions of HTML, the "standard" HTML5 and WHATWG's HTML "Living Standard". That's also why HTML5 omitted much of W3C's Semantic Web development work such as RDFa.

Amazon (not a member of either IDPF or W3C) is the elephant in the room. They take advantage of IDPF's work in a backhanded way. Instead of supporting the EPUB standard in their Kindle devices, they use proprietary formats under their exclusive control. But they accept EPUB files in their content ingest process and thus extract huge benefit from EPUB standardization. This puts the advancement of EPUB in a difficult position. New features added to EPUB have no effect on the majority of ebook user because Amazon just converts everything to a proprietary format.

Last month, the W3C published its vision for eBook standards, in the form on an innocuously titled "Portable Web Publications Use Cases and Requirements".  For whatever reason, this got rather limited notice or comment, considering that it could be the basis for the entire digital book industry. Incredibly, the word "ebook" appears not once in the entire document. "EPUB" appears just once, in the phrase "This document is also available in this non-normative format: ePub". But read the document, and it's clear that "Portable Web Publication" is intended to be the new standard for ebooks. For example, the PWP (can we just pronounce that "puup"?) "must provide the possibility to switch to a paginated view" . The PWP (say it, "puup") needs a "default reading order", i.e. a table of contents. And of course the PWP has to support digital rights management: "A PWP should allow for access control and write protections of the resource." Under the oblique requirement that "The distribution of PWPs should conform to the standard processes and expectations of commercial publishing channels." we discover that this means "Alice acquires a PWP through a subscription service and downloads it. When, later on, she decides to unsubscribe from the service, this PWP becomes unavailable to her." So make no mistake, PWP is meant to be EPUB 4 (or maybe ePub4, to use the non-normative capitalization).

There's a lot of unalloyed good stuff there, too. The issues of making web publications work well offline (an essential ingredient for archiving them) are technical, difficult and subtle, and W3C's document does a good job of flushing them out. There's a good start (albeit limited) on archiving issues for web publications. But nowhere in the statement of "use cases and requirements" is there a use case for low cost PWP production or for efficient conversion from other formats, despite the statement that PWPs "should be able to make use of all facilities offered by the [Open Web Platform]".

The proposed merger of IDPF and W3C raises the question: who gets to decide what "the ebook" will become? It's an important question, and the answer eventually has to be open rather than proprietary. If a combined IDPF and W3C can get the support of Amazon in open standards development, then everyone will benefit. But if not, a divergence is inevitable. The publishing industry needs to sustain their business; for that, they need an open standard for content optimized to feed supply chains like Amazon's. I'm not sure that's quite what W3C has in mind.

I think ebooks are more important than just the commercial book publishing industry. The world needs ways to deliver portable content that don't run through the Amazon tollgates. For that we need innovation that's as unconstrained and disruptive as the rest of the internet. The proposed combination of IDPF and W3C needs to be examined for its effects on innovation and competition.

Philip K. Dick's Mr. Robot is
one of the stories in Imagination
Stories of Science and Fantasy
,
January 1953. It is available as
an ebook from Project Gutenberg
and from GITenberg
My guess is that Amazon is not going to participate in open ebook standards development. That means that two different standards development efforts are needed. Publishers need a content markup format that plays well with whatever Amazon comes up with. But there also needs to be a way for the industry to innovate and compete with Amazon on ebook UI and features. That's a very different development project, and it needs a group more like WHATWG to nurture it. Maybe the W3C can fold that sort of innovation into its unruly stable of standards efforts.

I worry that by combining with IDPF, the W3C work on portable content will be chained to the supply-chain needs of today's publishing industry, and no one will take up the banner of open innovation for ebooks. But it's also possible that the combined resources of IDPF and W3C will catalyze the development of open alternatives for the ebook of tomorrow.

Is that too much to hope?

Wednesday, September 7, 2016

Start Saying Goodbye to eBook Pagination

Book pages may be the most unfortunate things ever invented by the reading-industrial complex. No one knows the who or how of their invention. The Egyptians and the Chinese didn't need pages because they sensibly wrote in vertical lines. It must have been the Greeks who invented and refined the page.
Egyptian scroll held in the UNE Antiquities Museum
CC BY 
unephotos
In my imagination, some scribes invented the page in a dark and damp scriptorium after arguing about landscape versus portrait on their medieval iScrolls. They didn't worry about user experience. The debate must have been ergonomics versus cognitive load. Opening the scroll side-to-side allowed the monk to rest comfortably when the scribing got boring, and the addition brainwork of figuring out how to start a new column was probably a great relief from monotony. That and drop-caps. The codex probably came about when the scribes ran out of white-out.
Scroll of the Book of EstherSevilleSpain
Technical debt from this bad decision lingers. Consider the the horrors engendered by pagination:
  • We have to break pages in the middle of sentences?!?!?!? Those exasperating friends of yours who stop their sentences in mid-air due to lack of interest or memory probably work as professional paginators. 
  • When our pagination leaves just one line of a paragraph at the top of a page, you have what's known as a widow. Pagination is sexist as well as exasperating.
  • Don't you hate it when a wide table is printed sideways? Any engineer can see this is a kludgy result of choosing the wrong paper size.
To be fair, having pages is sometimes advantageous.
  • You can put numbers on the pages. This allows an entire class of students to turn to the same page. It also allows textbook companies to force students to buy the most recent edition of their exorbitantly priced textbooks. The ease of shifting page numbers spares the textbook company of the huge expense of making actual revisions to the text.
  • Pages have corners, convenient for folding.
  • You can tell often-read pages in a book by looking for finger-grease accumulation on the page-edges. I really hope that stuff is just finger-grease.
  • You can rip out important pages. Because you can't keep a library book forever.
  • Without pages in books, how would you press flowers?
  • With some careful bending, you can make a cute heart shape.
Definition of love by Billy Rowlinson, on Flickr; CC-BY 

While putting toes in the water of our ebook future, we still cling to pages like Linus and his blankie. At first, this was useful. Users who had never seen an ebook could guess how they worked. Early e-ink based e-reading devices had great contrast and readability but slow refresh rates. "Turning" a page was giant hack that turned a technical liability of slow refresh into a whizzy dissolve feature. Apple's iBooks app for the iPad appeared at the zenith of skeuomorphic UI design fashion and its too-cute page-turn animation is probably why the DOJ took Apple to court. (Anti-trust?? give me a break!)

But seriously, automated pagination is hard. I remember my first adventures with TEΧ, in the late '80s, half my time was spent wrestling with unfortunate, inexplicable pagination and equation bounding boxes. (Other half spent being mesmerized at seeing my own words in typeset glory.)

The thing that started me on this rant is the recent publication of the draft EPUB 3.1 specification, which has nothing wrong with it but makes me sad anyway. It's sad because the vast majority of ebook lovers will never be able to take advantage of all the good things in it. And it's not just because of Amazon and its Kindle propriety. It's the tug of war between the page-oriented past of books and the web-oriented future of ebooks. EPUB's role is to leverage web standards while preserving the publishing industry's investment in print-compatible text. Mission accomplished, as much as you can expect.

What has not been part of EPUB's mission is to leverage the web's amazingly rapid innovation in user interfaces. EPUB is essentially a website packaged into a compressed archive. But over the last eight years, innovations in "responsive" web reading UI, driven by the need of websites to work on both desktop and mobile devices, have been magical. Tap, scroll and swipe are now universally understood and websites that don't work that way seem buggy or weird. Websites adjust to your screen size and orientation. They're pretty easy to implement, because of javascript/css frameworks such as Bootstrap and Foundation. They're perfect for ebooks, except... the affordances provided by these responsive design frameworks conflict with the built-in affordances of ebook readers (such as pagination). The result has been that, from the UI point of view,  EPUBs are zipped up turn-of-the-century websites with added pagination.

Which is what makes me sad. Responsive, touch-based web designs, not container-paginated EPUBs, are the future of ebooks. The first step (which Apple took two years ago) is to stop resisting the scroll, and start saying goodbye to pagination.

Saturday, January 2, 2016

The Best eBook of 2015: "What is Code?"

When the Compact Disc was being standardized, its capacity was set to accommodate the length of Beethoven's Ninth Symphony, reportedly at the insistence of Sony executive Norio Ohga. In retrospect it seems obvious that a media technology should adapt to media art wherever possible, not vice versa. This is less possible when new media technologies enable new forms of creation, but that's what makes them so exciting.

I've been working primarily on ebooks for the past 5 years, mostly because I'm excited at the new possibilities they enable. I'm impressed - and excited -  when ebooks do things that can't be done for print books, partly because ebooks often can't capture the most innovative uses of ink on paper.

Looking back on 2015, there was one ebook more than any other that demonstrated the possibilities of the ebook as an art form, while at the same time being fun, captivating, and awe-inspiring, Paul Ford's What Is Code?

Unfortunately, today's ebook technology standards can't fully accommodate this work. The compact disc of ebooks can store only four and a half movements of Beethoven's Ninth. That makes me sad.

You might ask, how does What Is Code? qualify as an ebook if it doesn't quite work on a kindle or your favorite ebook app? What Is Code? was conceived and developed as an HTML5 web application for Business Week magazine, not with the idea of making an ebook. Nonetheless, What Is Code? uses the forms and structures of traditional books. It has a title page. It has chapters, sections, footnotes and a table of contents which support a linear narrative. It has marginal notes, figures and asides.

Despite its bookishness, it's hard to imagine What Is Code? in print. Although the share buttons and video embeds are mostly adornments for the text, activity modules are core to the book's exposition. The book is about code, and by bringing code to life, the reader becomes immersed in the book's subject matter. There's a greeter robot that waves and seems to know the reader, showing the ebook's "intelligence". The "how do you type an "A" activity in section 2.1 is a script worth a thousand words  and the "mousemove" activity in section 6.2 is a revelation even to an experienced programmer. If all that weren't enough, there's a random, active background that manages to soothe more than it distracts.

Even with its digital doodads, What Is Code? can be completely self contained and portable. To demonstrate this, I've packaged it up and archived it at Internet Archive; you can download with this link (21MB).  Once you've downloaded it, unzip it and load the "index.html" file into a modern browser. Everything will work fine, even if you turn off your internet connection. What Is Code? will continue to work after Business Week disappears from the internet (or behind the most censorious firewall). [1]

I was curious how much of What is Code? could be captured in a standard EPUB ebook file. I first tried making a EPUB version 2 file with Calibre. The result was not a lame as I thought it would be, but stripped of interactivity, it seemed like a photocopy of a sticker book - the story's there, but the fun, not so much. Same with the Kindle version .

I hoped that more of the scripts would work with an EPUB 3 file. This is more or the same as the zipped html file I made but I was unable to get it to display properly in iBooks despite 2 days of trying. Perhaps someone more experienced with javascript in EPUB3 could manage it. The display in Calibre was a bit better. Readium, the flagship software for EPUB3, just sat there spinning a cursor. It seems that the scripts handling the vertical swipe convention of the web conflict with the more bookish pagination attempted by iBooks.

The stand-alone HTML zip archive that I made addresses most of the use cases behind EPUB. The text is reflowable and user-adjustable. Elements adjust nicely to the size of the screen from laptop to smartphone. The javascript table of contents works the same as in an ebook reader. Accessibility could be improved, but that's mostly a matter of following accessibility standards that aren't specific to ebooks.

My experimentation with the code behind What Is Code? is another exciting aspect of making books into ebooks. Code and digital text can use a open licenses [2] that permit others to use, re-use and learn from What Is Code?. The entire project archive is hosted on GitHub and to date has been enhanced 671 times by 29 different contributors. There have been 187 forks (like mine) of the project. I think this collaborative creation process will be second nature to the ebook of the future.

There have been a number of proposals for portable HTML5 web archive formats for ebook technology moving forward. Among these are "5DOC"  and W3C's "Portable Web Platform".   As far as I can tell, these proposals aren't getting much traction or developer support. To succeed, the format has to be very lightweight and useful, or be supported by at least 2 of Amazon, Apple, and Google. I hope someone succeeds at this.

Whatever happens I hope there's room for Easter Eggs in the future of the ebook. There's a "secret" keyboard combination that triggers a laugh-out-loud Easter Egg on What is Code? And if you know how to look at What Is Code?'s javascript console, you'll see a message that's an appropriate ending for this post:


Best of 2015, don't you agree?

[1] To get everything in What Is Code? to work without an internet connection, I needed to add a small number of remotely loaded resources and fix a few small javascript bugs specific to loading from a file. (If you must know, pushing to the document.history of a file isn't allowed.) The YouTube embed is blank, of course, and a horrible, gratuitous Flash widget needed to be excised. You can see the details on GitHub.

[2] In this case, the Apache License and the Creative Commons By-NC-ND License.

Friday, January 17, 2014

EPUB has a Steep Road Ahead (Notes from Open Book Hack 2014)

People I talk to about ebook technology belong to one of two camps.
  1. EPUB3 is the future of ebooks. 
  2. EPUB3 has lost the ebook format war because no one is supporting it. 

Here I'm helping Fortitude fix a bug.
(photo by Ray Schwartz)
So it was very enlightening to join a group of developers at New York Public Library last weekend at Open Book Hack 2014. Sponsored by NYPL Labs and the Readium Foundation, the event was convened to examine the challenges of the "Open Book" on the Web. Since I had written that another publishing hackathon "pretty much ignored ebooks", I felt compelled to attend, and write a report.

Open Book Hack was the first event I've been to where developers were actually working on EPUB, the format that is emerging as the underpinning of the digital book industry. This was both encouraging and discouraging. Encouraging because you could start to glimpse possibilities that EPUB will enable, and discouraging because the road ahead looks so steep.

Open Book Hack attracted a very stimulating mix of developers from around the world (who were in New York for Digital Book World) and local developers looking for interesting problems. Notably, there was a very impressive contingent of students from New York's Flatiron School. The projects are listed on a github wiki.

I was very eager to meet the folks from Berkeley's iSchool that have been working on epub.js, a javascript package that lets you read EPUBs from the web in your browser. One of them, Jake Hartnell, is also on the team building Hypothes.is, a tool designed to let everyone annotate the web (and one of the sponsors of the event). (The others were AJ and Fred.) Sooner or later you'll see these tools added to Unglue.it. But there's still a lot of work left to do.

For example, the project at Open Book Hack with the highest ratio of usefulness to impressive-sounding-ness was the project to add scrolling to epub.js (check it out, it works!). That's right, they added an option (mostly) to make chapter 2 come below chapter 1. You will soon be able to use ebooks on the web using an open-source 9th century reading interface.

There was one prize awarded. The winning project, "See and Read" allowed two people with two screens to interact via an ebook. Very nicely executed and well-deserved, but if I told you that I had invented something that allows two people to interact with a book, you would say "oh?" because you don't really need 100 billion transistors and two glowing screens to interact with a book.

Another project "Breadcrumbs and Beanstalks", extended Harvard's StackLife interface to enable ebook browsing similar to that found in a physical collection of books, with 2 dimensional browsing (the 2 axes being publication date and the general-to-specific axis of subject headings. It looked a bit clunky, but thar's gold in there somewhere.

Hugh MacGuire, Max Fenton, Jean Kaplansky, and Fendi M were doing something brave with linking in  PressBooks, but having already done too much linking in my day, I decided not to understand it. A team from Sophia University in Japan was doing something clever with textbooks and Readium. There was work on converting PDF to EPUB and a project to use Phonegap to make apps from EPUBs.

The project I joined up with focused on making a book club application around a shared EPUB reader. (It works, too!) We based it on epub.js and Hypothes.is. I wasn't very useful to the effort, because the three Flatiron-trained Ruby-on-Rails developers in our group were too awesome for my plodding-but-powerful python to compete with. So I helped with exploration and documentation of the Hypothes.is API, and finding bugs and deployment gotchas in epub.js. I now know how to configure CORS for buckets in S3. Yeah that was my weekend. I also fixed a bug by staring, in an intimidating way, over someone's shoulder. Ah, good times.

What became apparent to me in working with these tools was that the freshly trained developers got everything to work by un-EPUB-ing everything. The web platform just works, with the one exception being that centering text blocks in CSS just doesn't, unless you look away from the screen. The EPUB platform always throws something in your way, for reasons that even StackOverflow doesn't explain. Ruby Zips won't unzip. Cross Sites won't request. Java Scripts won't bind.

EPUB's competition isn't Amazon and KF8 fixed layout, it's the web and HTML5 and its huge gravitational pull. For 90% of ebooks, the benefits of EPUB over HTML are scant (because EPUB is based on HTML!) and the development barriers are significant. It's been years, and still EPUB authoring tools aren't mature or mass market. Deployment tools are barebones.

Don't get me wrong. I'm still betting big on EPUB, but dammit, Publishing Industry, for an $80 billion pillar of modern society, you're investing a nanoscopic amount on your basic infrastructure (i.e. EPUB), despite the herculean efforts of the people I met last weekend.

Notes:
  1. I was really impressed with the Flatiron School students I worked with. If the rest of the students are anything like Edina, Tiff, and Dan (Ivan helped a bit, too), they are going to have a huge impact on the New York area economy. Maybe I should learn RoR.
  2. Bill McCoy has done an amazing job bringing people together under the IDPF and Readium umbrellas. Imagine what he could do with financial support commensurate to his task. Perhaps he should take up bootlegging.
  3. Jake wrote up his impressions, too.
  4. As have Virginie Clayssen and Camille Pène, in French.
  5. Would have posted sooner, but MILESTONE IN UNLUE.IT.

Enhanced by Zemanta

Friday, November 15, 2013

Blogifying a Book

On Flatland the Blog, I'm turning a book into a blog. It's Flatland.

The ostensible reason is to promote our test campaign of Unglue.it's Buy-to-Unglue campaigns. It'll be a while before it's ready to launch for real, but I've found that there's no substitute for having real users try things out. One ungluer managed to find 2 different bugs within three minutes, partly by virtue of the 'ě' [LATIN SMALL LETTER E WITH CARON] in his username. Another user was the first to ever try changing their e-mail address to something invalid while having a username containing '@'. For some reason our unit tests didn't foresee these possibilities.

But really, I've been fascinated by the possibilities of the read-write book. We now have lots of ways to save and share annotations, but in most cases, this is done as a networked overlay on top of user-immutable texts. The annotation layers in Readmill, or in Kindle, live in their respective network.

Another effort to spread annotations over the web is Hypothes.is, which is trying to use standards to break the annotation layer out of closed networks.

I don't think there's anything wrong with networked annotation layers, but there's another technical direction that's been largely unexplored. What if a user's annotations are stored in the digital file that packages the ebook? This has the effect of restoring the individuality to copies of a book. The annotations could then be shared by sharing the file, the same way that pencilled annotations in a printed book might be shared privately. An anti-facebook, if you will, for an era when everything in the network layer is sure to be scanned by the NSA. And it also changes the dynamic of sharing a file in a library.

So what I want to do is collect comments and put them into the Flatland ebook that we're producing. I spent a fair amount of time producing a clean, attractive EPUB file from public domain scans by Google and Project Gutenberg, but I'd like to do more. The idea of turning the book into a blog occurred to me because Flatland's chapters were the right length, and, well they're curious, and need comment in the modern context.

So read along with me on Flatland the Blog, leave comments and suggestions, and at the end maybe something interesting will come out of the experiment!
Enhanced by Zemanta

Monday, October 7, 2013

NYLSLR: The eBook Copyright Page is Broken

Somehow it slipped my mind that my article "The eBook Copyright Page is Broken" was published in the New York Law School Law Review in April. And I am still not a lawyer! Here's the meat of the article:
The traditional copyright statement is thoroughly and fundamentally broken. Consider the simplest possible case of a single copyright holder:
                         © Eric S. Hellman, 2013. All Rights Reserved. 
This is broken in the following ways:
  1. Since there currently are not any copyright formalities, the copyright symbol means nothing. The work is subject to copyright with or without the copyright symbol.
  2. The work may also not be subject to copyright, for example, if Eric S. Hellman is a government employee, a robot, or a non-creative compiler of factual information. In these cases there is no copyright even if there is a copyright symbol present. There is no legal duty for a publisher to put a copyright symbol only on a copyrightable work. How is the ebook user supposed to know the true copyright status of a digital work? 
  3. “Eric S. Hellman” is an uncommon name. But suppose the author is named ”John Smith.” What use, then, is the copyright statement? It does not specify which Eric S. Hellman or which John Smith is the author.
  4. The asserted name of the copyright holder can’t be relied on because text in a digital file can be altered without a trace. It’s simple to take a digital copy of Merchants of Culture and change its asserted copyright holder to “John Smith,” then redistribute it. This is a negligible problem in the print world.
  5. The asserted date of publication may be unrelated to the date of the underlying copyright. For purposes of copyright (for example, when a work is produced as a work-for-hire), re-publication of a book does not change the copyright expiration date of the underlying text.
  6. There is no specification of the work being copyrighted. In print there’s not much ambiguity, but digital books are composite objects (text and graphics are always separate entities in a digital book file) and are frequently distributed in pieces. Some ebooks even have front matter distributed as a pdf file completely separate from the chapters. In other cases, an ebook may be displayed on a website that has a separate set of copyright statements.
  7. If the digital book is legally on your ebook reader, then, somehow, the rights holder has granted you some rights, perhaps under the terms of an explicit license or with the license implicit in its availability on a website. Either way, “all rights” have not been reserved. Licenses are not needed for printed books, but they may be needed for ebooks.
In February, I wrote about ebook front matter and back matter and there's more work to be done in this vein.

The last footnote deserves some glossing. In it, I assert that the ccREL submission for marking Creative Commons status of web pages is currently in conflict with the EPUB 3 standard for ebooks. While that's technically true, it's a bit misleading. A better way to say it is that developments in HTML5 and EPUB3 have made ccREL's approach archaic. The metadata machinery in EPUB3 and HTML5 is fully up to the task of expressing and applying Creative Commons licenses. What's lacking is consensus around which of the available mechanisms to use. Since the RDFa vs. Microdata in HTML5 controversy has not yet fully shaken out, you can't really follow ccREL as written, so we'll need to have some patience.
Enhanced by Zemanta

Sunday, May 19, 2013

Publishing Hackathon Pretty Much Ignores eBooks

The "First Annual" Publishing Hackathon was this weekend. As advertised, I participated and worked on an EPUB backmatter project. My awesome team consisted of me, Javascript/Ruby developer Max Jacobson (who's going to be even more highly sought-after when he finishes Rails school this summer), and TLC librarian Dianne Coan.

Here's our demo video:

 

Here's how we described the project:

Book Discovery INSIDE the eBook

When is a reader most receptive to reading suggestions? Right when they’ve finished a book of course! That’s why printed books have information about other books by the same author, the first chapter of the next book in the series and similar material at the end as part of the back matter.

Back matter has existed pretty much as long as books have. This includes the appendix, glossary, index, and bibliography. Back matter for digital books needs to be optimized to serve the needs of the digital reader. An informal survey by @suw indicates the most popular endmatter desires were other books by the same author and some information about the author.

Digital back matter for ebooks is not constrained by having to proceed the publication; unlike print, digital back matter can be kept up to date with the release of new content. For instance, if an author publishes a sequel, that title could be included in previously published ebooks.

It’s easy to insert a page listing an author’s other books at the end of an ebook, but how do you keep that list up-to-date? What if you’ve developed a great recommendation system to do “if you liked Pride and Prejudice, you’ll like X”? (or maybe “if you hated...”!)

The answer is to make use of the javascript capability of emerging ebook environments. Our project explores means of connecting to APIs from within an EPUB for the purpose of suggesting the user’s next read.

An existence proof is the “widget” capability of the iBooks iAuthor platform. It allows the insertion of html snippets into extended EPUB. Unfortunately, the javascript capability of ebook reading platforms, like the future, is unevenly distributed.

For this demo, we tested three reading EPUB environments, Readium, Readmill, and iBooks. We modified the Project Gutenberg EPUB version of Pride and Prejudice to include hooks and data to other books by Jane Austen.

Readium, which has been built as an EPUB3 reference environment, is the most capable for our purposes. It supports both javascript and connections to external web resources. In Readium, our EPUB displays the set of books by Jane Austen returned by the ReadMill API.

Apple iBooks has full javascript capability, but doesn’t allow connections to external resources (except perhaps via iBooks Author hooks- this deserves further investigation.) In iBooks, our EPUB displays a result page that we generated and embedded based on Jane Austen works published in 1813, when Pride and Prejudice released. We imagine that such embedded resources could be inserted at download time in a future production bookstore or library environment.

The Readmill environment does not support javascript at all at this time, so ironically, we’re not able to display the Readmill API results, or the iframe embedded resource.

Offline reading in Readium displays the resource embedded in the EPUB, similar to the iBooks version.
There were 30 projects in total presented at the end. Here's the list, along with my one sentence summary.
Banned Books in America
Website that maps book banning incidents and links them to Openlibrary
Book Discoverability: A Graphical Solution
Concept for browsing books as nodes on a graph.
Book Discovery INSIDE the eBook
This was us! Our demo crashed and burned. The popup screens from the wifi messed up the ebook reader display of embedded dynamic content.
BookCity Finalist!
Website that recommends books by connecting them to cities.
BookieGoer
Website that helps you lend the books you've borrowed from the library.
Booklvrs: Read. Discover. Meet.
App that advertises the ebook you're reading to the people around you.
bookmatchup
Website that multi-factor-matches you to books.
BookMob
Website that aggregates book recommendations from your twitter followers.
bookshelf.me
Website that displays books as if they were on a bookshelf. I'm pretty sure there was more to it.
Publy.io
Website that recommends books to users based on books they've liked.
Captiv Finalist!
App and Website that uses machine learning algorithms and your tweet about last night's party to combat the short attention span of Today's Readers. I may not have understood this one.
Coverlist Finalist!
Website that believes in judging books by their cover.
Evoke Finalist and clear judging favorite!
Pinteresty website that recommends books based on emotions categorization.
Happy Chapter
App that recommends books based on tags you click.
I read your Brain
Brain-sensing rabbit ears that wiggle depending on your response to a book from a website.
IGNITE
Website that lets users rate romance novels for steaminess.
KooBrowser Finalist!
Browser plugin that analyses what you read to better sell you books.
Library Atlas Finalist!
Mobile app that sends you geographically appropriate quotes depending on where you are. My favorite.
Literary Trinket with Book Wish
3D printed QR-ish code baubles. Cooler than it sounds.
Meadows
Website that turns reading into a game where you earn points.
Meme a book
Website that turns books into lolcats. (I may not have described this accurately.)
MovieReader
Website that recommends books connected to the movie you just saw.
NYPL Reinvent
Analysis of NYPL metadata advocating a divorce of the library from its classification system.
OkLetsRead!
Website offering crowd-funded serial fiction (ebooks).
Quiply
Website that recommends books based on a user's video viewing.
Reading Tollbooth: A Gateway to Book Discovery
Website to match kids to books.
Something2Read
Website that recommends books based on tags you click.
Valerie's Baby App
App that promotes literacy to a girl named Valerie by making sliding block puzzles and defining words at her.
Visibrary
Website that uses library data to make graphical book circles.
Vookstore
Website that turns ex-bookstore owners into book curation engines.
Interestingly, only 3 of the 30 projects addressed ebooks at all, which seems a bit odd to me, considering the industry's ongoing transition from print to digital. The emphasis on apps (7) and websites (21) is partly due to Hackathon's theme of book discovery, but it also says something about the tech industry. Apps and websites are what the NY tech industry is doing in 2013, not ebooks. Clearly, the publishing community developing ebooks and ebook standards needs to do more outreach to developers; the hackathon was a good first step.

It's also worth noting the growing importance of geo-tagging and other non-traditional metadata. In the new world of publishing discovery, readers want books that fit their mode right where they want to be. Neither MARC nor ONIX know enough to help.

My library friends should rest assured that the hackers did not at all ignore libraries. Although $1000 prize from NYPL was a factor, the ease of connecting to NYPL and OpenLibrary helped a lot. The RDA prize, it should be noted, went unclaimed.

Update: Sorry, Coverlist, I omitted your finalist status. Corrected!
Enhanced by Zemanta

Tuesday, June 19, 2012

"Open Access eBooks" eBook is on GitHub


When I try to explain to book industry people why ebooks can and should be free, I often get a look that says "What planet are you from?" In contrast, many of my software developer friends take it as dogma that ebooks should not only be free, but also "Free". And so I seem to spend a lot of time explaining one point of view to the other.

What we're trying to do with unglue.it is to skip over the theory, and just show everyone that it works.

First, an explanation for the 99% of real people who haven't encountered the Free vs. free distinction. An ebook that's Free means more than just not having to pay for the ebook, it means that the ebook is not locked up in any way. You can do things with it without needing permission. Copy it, distribute it, convert it, print it, slice it and dice it. Extract it, analyze it, translate it compute it, archive it. In the software world, that's the essence of Free Open Source Software (FOSS). But the 99% just wants to read the book. So why should it bother with Free?

The book that is on the brink of having a successful ungluing campaign at Unglue.it, Oral Literature in Africa, has a Free license proposed for it, CC BY (Creative Commons Attribution). We can't be certain how much the Free license has been responsible for the success of the campaign (you HAVE pledged, haven't you?), but it certainly adds to the appeal. A successful conclusion to the campaign will do more than just let people read the book. It will allow scholars of African culture to add to the book, to use chapters as course material, to use large excerpts in their own work, to make corrections and translations. And the media handling capabilities of new ebook formats will allow the addition of audio to a work about material that deserves to be audible.

What frustrates me, though, is how difficult it is to actually do all the things that you would want to do with a not-locked-up ebook, even the things that don't require it to be Free. Something as simple as correcting a typo is hard for 99.9% of the public. It shouldn't be that way. There should be tools that make this easy. If I want to add my voice into Oral Literature in Africa, there should be an application that allows me to click and speak.

The software world has developed a wealth of tools that allow distributed teams of developers to work together on free software. Source control systems help to track and manage changes in software. We need the same sort of tools that work for books. Wikis do part of the job, but we need more.

So as a first step, I'm putting the short book I've written using this blog, Open Access eBooks, on GitHub, the service we use to track and manage the software behind Unglue.it. All the book's source code is there, mistakes and all. Its CC BY license allows you to take it, branch it, fix it, translate or modify it, redesign and recode it, whatever. You can send me a pull request if you want to merge your changes with my branch. Maybe you want to update the references or add a chapter. Maybe you want to embed metadata or improve accessibility. Maybe you want to fuse it with Moby Dick for some bizarre art project. Whatever. The future of books is all of ours to create.

Notes


  1. Other factors contributing to the imminent success of the Oral Literature in Africa Campaign have been its academic nature, its modest ungluing fee, and its inherent coolness. What, you haven't contributed yet?
  2. Among the Creative Commons Licenses usable at Unglue.it, CC BY and CC BY-SA are considered by Free Culture advocates to be "Free". The Public Domain Dedication (CC0) is not a license, but is another way to make a work "Free". The SA (Share Alike) restriction is a form of "copyleft" which requires derivative works to be similarly made available.
  3. Other CC licenses may add conditions including NC (Non-Commercial) and ND (No Derivatives).
  4. Wikipedia is a good example of a site that won't allow posting of NC or ND licensed content.
  5. The license used for an unglue.it campaign is specified by the rightsholder who may be constrained by  publishing contracts and byzantine international licensing regimes.
  6. I was disappointed by the lack of good tools to create ebooks. I did everything by hand and was surprised at the mess of shifting standards, conflicting ereader implementations and insular documentation.
  7. Whenever I encounter a roadblock in python or django, Google sends me to StackOverflow for the answer. With ebook production, I always end up at MobileRead, ThreePress or Liz Castro's blog. These are wonderful resources, but they're not StackOverflow.
  8. Despite my struggles with EPUB, Amazon's MOBI tools were painless. I felt so naughty!
  9. Because Git is line oriented, I put every sentence in the content file on its own line. Hope that makes sense!
  10. For an example of another ebook with source on GitHub, check out Structure and Interpretation of Computer Programs, Second Edition (SICP). It's not Free, though.
  11. If you want a really nice "free" dinner next Saturday in Anaheim California, make a $100 unglue.it pledge and ask me for an invite. Space is limited!


Enhanced by Zemanta

Wednesday, June 22, 2011

EPUB 3 Beefs Up Metadata, but Omits Semantic Enrichment

Ironic amusement fills me when I hear book industry people say things like "metadata has become cool", or "context is everything". Welcome to the 20th century and all that. Meanwhile, in the library industry, metadata has been cool long enough to coat everything with a thick rind of freezer burn.

There's good news and notsogood news for ebook metadata. The revision to the EPUB standard, published just a month ago, includes metadata tools that could eventually lead to a new era of metadata cooperation between publishers and the entire book supply chain, including libraries. At the same time, the revision fails to take advantage of ready-made vehicles for semantic enrichment of content, a move that could still provide new types of revenue for publishers while giving libraries new opportunities to remain relevant as books become digital.

Since I'm incurably optimistic, I'll start with the half-full glass: Publication-level metadata. EPUB 3 includes a whole bunch of ways to include publication-level metadata in an EPUB container. As an example, imagine an EPUB3 for "Emma" with this mark-up in its package document (essentially the navigation directory for the book):
<metadata>
...
<meta property="dcterms:identifier"
id="pub-id">urn:uuid:A1B0D67E-2E81-4DF5-9E67-A64CBE366809</meta>
<link rel="marc21xml-record" href="http://www.archive.org/download/cihm_29722/cihm_29722_marc.xml" />
<link rel="marc21xml-record"
href="/cihm_29722_marc.xml" />
<link rel="foaf:homepage" href="http://openlibrary.org/books/OL24234129M/Emma" />
...
</metadata>

In this example, the first link element points to a MARC 21 xml record (MARC 21 is a blattarian standard for library metadata (look it up)) at the Internet Archive. The second link element points to the same record included in the EPUB container itself. There is also built-in vocabulary that allows the link element to point to ONIX, MODS, and XMP metadata records.

The example also shows that other vocabularies (such as FOAF) can be added for use in metadata elements. So, if you're a believer in RDA, you can put that in an EPUB file as well.

The meta element can also be used in the EPUB package document's metadata block. It's defined quite differently from HTML5's empty meta element, with an about attribute and allowed text content. In principle, it can be used to encode arbitrary RDF triples, thanks to a prefix extension mechanism borrowed from RDFa which allows EPUB authors to add vocabularies to their documents.

These capabilities, on their own, could support major changes in the way that books are produced, delivered and accessed. In a publisher workflow, the EPUB file could serve as the carrier for all the components and versions of a book, even bits that today might be left out or lost in the caverns of so-called "content management systems". A distributor would no longer need to match up content files with records in a separate metadata feed. EPUB books for libraries could be preloaded with cataloging and enrichment data, greatly simplifying the process of making the ebooks accessible in libraries.

Given the great advances for "package-level" metadata, it's a bit disappointing that semantic mark-up of content documents missed the EPUB 3 boat. The story is a bit complicated, and it's far from over. Imagine that you want to add mark-up to a book's citations- perhaps you want to embed identifiers to support library linking systems. Or perhaps you're a medical publisher and you want to embed machine readable statements about drugs and diseases in a pharmaceutical textbook. Or perhaps you want to publish a travel guide and you want search engines to pick out the places you're describing. These applications are not really supported by the current version of EPUB 3.

EPUB content documents have a feature that you might think would do the trick, but doesn't really. The epub:type attribute supports "semantic inflection" of elements. This attribute can be used to mark a paragraph as a bibliographic citation, for example, and supports many of the requirements imposed by conversion of content from legacy or specialized formats into the HTML5 dialect used by EPUB. It's an important feature, but not enough to support semantic enrichment.

Part of the problem is EPUB 3's dependence on HTML5, which is not yet a stable spec and is enmeshed in some surprisingly raw W3C politics. W3C has been the home of HTML standards development since the very early stages of the web, and has also been the home of semantic web standards development. HTML5 started outside of W3C in the WHATWG, an initiative to develop HTML in a way that would be backwards compatible with good-old fashioned non-XML HTML. W3C was convinced to fold WHATWG into its development efforts because of WHATWG's corporate backing. Even so, the WHATWG version of the HTML5 spec drips with sarcasm towards W3C HTML Working Group decisions.

During part of the development of EPUB 3, the HTML5 draft included "Microdata", a method of embedding semantic mark-up in HTML. RDFa, a standard that competes with Microdata, was developed by W3C channels, and within W3C, it was decided in February of 2010 to move Microdata out of the HTML spec so as to give it equal footing with RDFa. Some participants in the EPUB working group wanted to include RDFa in the standard; others thought this would impose too much of a complexity burden on publisher-implementers. The EPUB draft ended up being released without either RDFa or Microdata.

The recent endorsement of Microdata by the Google-Yahoo-Bing cooperation has changed the competitive landscape for embedded semantics. It's now apparent that Microdata will get priority implementation in HTML development tools, leaving RDFa as a niche technology. For most use cases of EPUB semantic markup, the differences between RDFa and Microdata are small compared to the advantages of piggybacking on the technology investment supporting website creation.

According to members of the EPUB working group, it is expected that a dot release will follow relatively quickly behind EPUB 3.0. It seems to me that picking a semantic markup technology for content documents should now not be so hard. If you work for a publishing company that has ever mentioned semantic markup in a product plan, you should probably be making sure that the EPUB working group is aware of your needs. If you are a librarian who can imagine the possibilities of a semantically enriched EPUB collection, you should similarly be making your concerns known.

Although the EPUB working group includes representatives from tools vendors that might conceivably benefit from the adoption of EPUB-only constructs, the group's track record for adopting wider web standards has been very encouraging. By adopting HTML5 as a stack component, the group has ensured that cheap or free tools to produce and author EPUB 3 content will be readily available.

Once semantic enrichment of ebooks becomes routine, libraries will play a vital role in their use. Libraries provide a copyright-friendly DRM-free community commons in which users can access and build on the information contained in licensed content. (Of course, I see "unglued" books as playing an equally important role in the library commons.)

The EPUB metadata glass is half full, and there's more wine in the bottle!

Note: This is one thing I'll be talking about on Saturday at the American Library Association meeting in New Orleans. (The program is somewhat inaccurate; the program will end at 10:30 AM at the latest. Ross Singer from Talis will lead off with an overview of semantic web technologies in libraries; I'll follow with discussions of RDFa, the Facebook "Like" button and of course, EPUB.
Enhanced by Zemanta

Thursday, June 2, 2011

EPUB Really IS a Container

"It's OK for libraries to put things in their EPUB books." That's what Bill Kasdorf, a member of the EPUB Working Group, told me last week at the IDPF Digital Book 2011 Meeting. He checked with EPUB Revision Co-Editor Markus Gylling to make sure. I had been curious if libraries could put all their cataloging information inside an EPUB file instead of siloing it in their catalog system.

It may seem an odd question if you don't know a few things about EPUB. EPUB is a standard format for ebooks. It's used by Apple, Barnes and Noble, Kobo, Overdrive and many others not named Amazon. EPUB is near the end of a revision process that will result in EPUB 3.0.

The EPUB specs define a lot more than just a file format. Both EPUB 2 and EPUB 3 define a container format (in EPUB 3 it's called the EPUB Open Container Format (OCF) 3.0, and then go on to define a number of file formats for files that go inside this container. These files are the resources- texts, graphics, etc. that make up the ebook.

OCF uses the ubiquitous ZIP format to wrap up all a book's resource files into a neat, transportable package. That's pretty much standard these days. Java ".jar" and ".war" files use the same mechanism, as do MacOS' ".app" files.  As a consequence, you can use any unzip utility to look inside an EPUB file and manipulate its contents.

There's even a reserved name for a file to contain book level metadata in OCF: META-INF/metadata.xml, as well as another file for rights information, META-INF/rights.xml. Another file, META-INF/signatures.xml can be used to prove who made parts of the file and determine whether anyone has mucked with them. When Gluejar issues Creative Commons editions of newly relicensed works, we'll use the rights.xml file to make sure the CC declaration is explicit.

The new EPUB revision is coming fast. Last Monday, Bill McCoy, Executive Director of the International Digital Publishing Forum (IDPF) announced the release of the full EPUB 3 proposed specification. My guess is that when we look back on this event 10 years hence, we'll recognize this as the moment EPUB began to revolutionize the world of information, and with it, the book industry.

Although Amazon still uses the aging MOBI format on its kindle devices, it seems only a matter of time before the infrastructure accumulating behind EPUB pushes them into the embrace of the IDPF. Already, most of the content flowing into the Amazon system is being produced in EPUB and converted to MOBI. Don't expect this shift to happen soon though; in his IDPF presentation, Joshua Tallent of eBook Architects described rumors that this would happen soon as "bunk"- but it will happen sometime.

EPUB 3 comes with lots of goodies. The revision adds several modules of sorely needed capability. It includes MathML, SVG and JavaScript over a substrate of HTML5 and CSS2.1. While MathML and SVG are essential for education and technical markets, JavaScript has been somewhat controversial because of the difficulty of making sure things work securely and without connections. Most of the reading systems inherit javascript capability from the WebKit rendering engine they're based on, so a lot of javascript functionality will work in ebook readers regardless.

(left) Autography Founder and Author  T. J. Waters
All this capability will remain latent unless people find compelling uses for it. I'm not worried. As the BookExpo itself got started, I met two different companies who were manipulating ebook files to solve the same problem: how can an author sign a book when the book is digital? Both companies, Autography and InScribed Media, create personalized experiences that leave artifacts of an author-consumer interaction inside ebook container files. Both of these companies have compelling solutions; they differ in their business models. Autography is structured as an author focused bookstore; InScribed is developing partnerships with existing bookstores.

InScribed Media Founder and Author Alivia Tagliaferri
To some extent, InScribed and Autography are forced to be a bit convoluted in the way they deliver their product because they need to live inside DRM green zones; users don't have access to the files inside books without cracking the DRM (which is rather easy, by the way!). It's unfortunate, because personalization of ebooks could be a good way to encourage responsible use. I certainly don't want that picture of me torrenting around the world!

Libraries face a similar dilemma. The insides of an EPUB file could be greatly enriched by  libraries, which have every motivation to enhance discovery both of the book and the information inside of it. But DRM gives the publisher and its delivery agents the exclusive ability to build context inside ebook containers. Libraries and readers are locked out. I think that for DRM systems to survive they will need to accommodate a more diverse set of user manipulations; author signatures are just the tip of the iceberg.

Coming soon, I'll report on EPUB 3 metadata.
Enhanced by Zemanta

Wednesday, May 18, 2011

The Object-Oriented Book

To most people, objects are things you can touch, see, maybe even smell. They have existence on their own. Software developers talk about objects as well. Although they're more abstract, software objects can also be touched- programs can interact with them, and they exist on their own as packages of code and data.

In some recent conversations about books and content containers, I've been hit in the face with the fact that most people in publishing haven't been steeped in Object-Oriented Programming (OOP) the way I once was, and as a result, some of the things I've written about the evolution of the book into digital form have sounded a bit strange to many people. So I've decided to write a bit here about how books are becoming software objects, and why it matters.

Object orientation is a style of programming that models problems as spaces of objects from various classes. The programmer solves problems by manipulating objects; the objects communicate among themselves by passing messages. The messages that objects pass are governed by interfaces; every class of objects is defined by the interfaces it supports. If that doesn't make sense to you, don't worry, I'll have some examples.

Let's think about how we might model the book as a software object. With a physical book, you know how to get the title and name of the author. You open up the book to the title page, and there you find the title, probably the words in the largest type size, and the author's name, probably printed below the title, perhaps with a designator word such as "by".

In the prehistory of programming before OOP, a book program might define data structures containing tables of book titles and author names. The program would look in these tables for the book data. An object-oriented program would instead send the book-object messages saying "what is your name?" and "What person was your author?" An object-oriented approach binds the code and data together, so that objects of the book class know what their title is, how many chapters they have, and what the 20th word of the 32nd paragraph of their 3rd chapter is. The set of messages that an object can respond to defines its class. A programmer knows that any object in the Book class will be able to tell you its title.

Another key concept in object orientation is inheritance. A cookbook is a book and inherits from the Book class the ability to tell you its title. But you expect a Cookbook to have recipes, and you should be able to ask it how many recipes it contains.

The reason I think this is important for non-coders to understand is that very soon, the book industry will become focused on producing lots and lots of these software objects. And I'm not talking about some far-fetched digital utopia.

The third revision of the EPUB standard is very soon to become a reality, and I believe its use will quickly become pervasive in the book industry. It would be a mistake to think of EPUB3 as yet another document format. With the adoption of EPUB3, the book industry will, for the first time ever, have standardized a software object model for the book. This comes along with EPUB3's use of HTML5 as a foundational layer.

An object model became associated with HTML documents very early in its evolution. Called the DOM, or Document Object Model, it was developed by programmers working with HTML documents, and it quickly became the basis for most software that works with HTML documents. With the development of Javascript, HTML documents delivered over the web could bind to code that accesses and manipulates their data via the DOM. It's only with HTML5, however, that the DOM is officially becoming part of the HTML standard.

With HTML5 as its basis, EPUB3 becomes a very capable "container" of content. The whole discussion of how containers limit the ways in which content can interact with consumers becomes completely moot, and a bit silly. EPUB3 binds a complete "API" (application programming interface) onto the content, and provide many mechanisms for the extension of that interface. The "API" and the "container" are one and the same.

If we look at the immense infrastructure that arose around the book as a physical object, from book bags and compact shelving, to printing plants, warehouses, libraries and used bookstores, we can get an inkling of the infrastructure that will grow up around the book as a software object. In the coming weeks, I'll try to write about some of the implications of EPUB3 for the industry as a whole.
Enhanced by Zemanta

Tuesday, February 15, 2011

How Apple May Inadvertently Boost eBook Linking

"The net interprets censorship as damage and routes around it." John Gilmore, 1993
The official word from Apple finally came out today, in their press release announcing in-app subscriptions.
In addition, publishers may no longer provide links in their apps (to a web site, for example) which allow the customer to purchase content or subscriptions outside of the app.
Assuming that this limitation will be applied to the Kindle App, it means that the "Shop in Kindle Store" button will disappear, and similar features in other ebook reader software, such as Nook, Sony, and Kobo will disappear as well.

You can go elsewhere if you want to read apocalyptic whining about Apple's imperious ways. What I want to focus on is how the net will route around this damage. The net will route around this damage by making more links.

At O'Reilly's Tools of Change for Publishing Conference, I was able to spend some time with Keith Fahlgren, a partner at ThreePress Consulting. He's part of a group that has worked on the improvement of linking capability in EPUB 3. A Public Draft of the specification was released today by IDPF.

Since EPUB3 is based on HTML5, all the outbound linking that you would expect from a web page is already built into EPUB3 (as well as earlier versions of EPUB). Ebook reader apps available on iOS and Android use the "Webkit" webpage renderer for ebooks in EPUB. (Kindle devices use Webkit to render web pages and WebKit is used by Amazon to render Kindle ebooks (in mobi format) on  hardware other than their own.) So it's clear to me, at least, that even if ebook reader apps can't have "Kindle Store" buttons, the apps will be able to present "Kindle Store" links inside the ebook content. I'll bet you anything that Amazon is loading up ebook content with Kindle Store links: "If you like this book, perhaps you'd like this one". They'll even have specialized shop-books containing Kindle store links available for free. Ditto the others.

Publishers aren't going to like Apple's power-play. But neither will they like having their content getting hijacked to promote individual ebook stores. There will therefore be a great deal of pressure for the creation of vendor-neutral, customer friendly ways to link to ebooks from within ebooks, one that Apple can't ban because doing so would break Safari.

Here's where it gets tricky. If a customer has already purchased the linked-to book, it's pointless to send them out to a ebook store, they should connect their copy of the ebook. But figuring out whether a consumer already has the book is messy, given the state of ebook identification. There are many other use cases for linking to a specific chapter or paragraph inside an ebook.

Unfortunately, doing this sort of linking is not a solved problem. EPUB3 adds one tool that will help. A new required metadata property, dcterms:modified,  will help identify the epub in the case where it has been modified- in the past it was poorly specified what should happen to the epub identifier if the file was modified. With EPUB3, it's now clear that EPUB documents are identified internally at a level above the ISBN (different DRM wrappings of the same EPUB file often require different ISBNs) but below the "work".

There's still a lot of apparatus that will need to be built, both inside and outside of EPUB, for linking to work the way it should. Being able to decide which ebook to target will require external mechanisms. Perhaps some linking organization along the lines of Crossref be formed; perhaps a more wikipedia-ish database collaboration will suffice. In any case something like xISBN supercharged for ebooks will be needed. Fahlgren told me that without a strong use case to drive the solution, the EPUB group has had a hard time going very far in their linking development.

A true ebook linking solution would need to include Amazon, of course, and since they've not been using EPUB, it seems to me that ebook linking won't get done by the EPUB group itself. Amazon hasn't had much use for EPUB in the past, but now Apple may have handed the ebook technology community a giant use case for interoperable ebook linking.

Happy Day-After-Valentines-Day, EPUB!