Showing posts with label isbn. Show all posts
Showing posts with label isbn. Show all posts

Tuesday, December 22, 2015

xISBN: RIP


When I joined OCLC in 2006 (via acquisition), one thing I was excited about was the opportunity to make innovative uses of OCLC's vast bibliographic database. And there was an existence proof that this could be done, it was a neat little API that had been prototyped in OCLC's Office of Research: xISBN.

xISBN was an example of a microservice- it offered a small piece of functionality and it did it very fast. Throw it an ISBN, and it would give you back a set of related ISBNs. Ten years ago, microservices and mashups were all the rage. So I was delighted when my team was given the job of "productizing" the xISBN service- moving it out of research and into the marketplace.

Last week,  I was sorry to hear about the imminent shutdown of xISBN. But it got me thinking about the limitations of services like xISBN and why no tears need be shed on its passing.

The main function of xISBN was to say "Here's a group of books that are sort of the same as the book you're asking about." That summary instantly tells you why xISBN had to die, because any time a computer tells you something "sort of", it's a latent bug. Because where you draw the line between something that's the same and something that's different is a matter of opinion and depends on the use you want to make of the distinction. For example, if you ask for A Study in Scarlet, you might be interested in a version in Chinese, or you might be interested to get a paperback version, or you might want to get Sherlock Holmes compilations that included A Study in Scarlet. For each  question you want a slightly different answer. If you are a developer needing answers to these questions, you would combine xISBN with other information services to get what you need.

Today we have better ways to approach this sort of problem. Serious developers don't want a microservice, they want richly "Linked Data". In 2015, most of us can all afford our own data crunching big-data-stores-in-the-cloud and we don't need to trust algorithms we can't control. OCLC has been publishing rather nice Linked Data for this purpose. So, if you want all the editions for Cory Doctorow's Homeland, you can "follow your nose" and get all the data you need.

  1. First you look up the isbn at http://www.worldcat.org/isbn/9780765333698
  2. which leads you to http://www.worldcat.org/oclc/795174333.jsonld (containing a few more isbns
  3. you can follow the associated "work" record: http://experiment.worldcat.org/entity/work/data/1172568223
  4. which yields a bunch more ISBNs.

It's a lot messier than xISBN, but that's mostly because the real world is messy. Every application requires a different sort of cleaning up, and it's not all that hard.

If cleaning up the mess seems too intimidating, and you just want light-weight ISBN hints from a convenient microservice, there's always "thingISBN". ThingISBN is a data exhaust stream from the LibraryThing catalog. To be sustainable, microservices like xISBN need to be exhaust streams. The big cost to any data service is maintaining the data, so unless maintaining that data is in the engine block of your website, the added cost won't be worth it. But if you're doing it anyway, dressing the data up as a useful service costs you almost nothing and benefits the environment for everyone. Lets hope that OCLC's Linked Data services are of this sort.

In thinking about how I could make the data exhaust from Unglue.it more ecological, I realized that a microservice connecting ISBNs to free ebook files might be useful. So with a day of work, I added the "Free eBooks by ISBN" endpoint to the Unglue.it api.

xISBN, you lived a good micro-life. Thanks.

Monday, April 1, 2013

Introducing the Invalid ISBN System


We're pleased to help introduce a revolutionary new identifier system for ebooks, based on a new numeric identifier, the Invalid ISBN (or InvIS BN, for short). InvIS BN works together with the legacy ISBN system (ISBN Classic) to extend identity to ebooks without all the hassle and expense of the real thing.

InvIS BN takes advantage of two of the ISBN system's fatal flaws
  1. ISBN Classic wastes nine out of ten perfectly good numbers.
  2. ISBN Classic ignores the error-producing power of real users.
Humans have been shown to have a 1% error rate in transcribing digits. As a result, they have a 9.6% error in transcribing 10-digit ISBNs; the shift to 13 digit ISBN has increased this rate to 12.3%. As a result, many ISBNs in circulation are incorrect. While the so-called "check digit" helps to expose these errors, there is no way to correct an error once made.

The InvIS BN system, by contrast, collects these human errors in a registry, allowing them to be fixed. But most potential ISBNs are still unused, despite a growing need to identify the proliferating digital versions of books. For example, the recent acquisition of Goodreads by Amazon has caused the death of millions of ISBNs, all of which will have to be replaced somehow.

Toxic leftovers from ISBN mining.
Similarly, the Economist has noted that ebooks will cause the demise of ISBN, resulting in renewed demand for viable identifiers. Invis BN will help meet this demand.

InvIS BNs are salvaged from the 77.7% of numbers which are unused by valid or mistaken Classic ISBNs. They will be available, for free, from the invisbn.org website, now under construction.

The InvIS BN system is expected to have huge environmental benefits. The current ISBN system leaves huge piles of toxic numerical "tailings" in the regions where ISBNs are mined. These tailings are now being recycled into useful identifiers.

Arual Noswad, Mayor of Invalid, Texas, the town at the center of the American ISBN mining region and namesake of the new system, welcomes the new developments. "If you don't use this, I will break you" she threatened an innocent reporter.

Invalid, Texas was incorporated in 1948 due to a data entry error. By a quirk of our modern technology it cannot be found using modern GPS systems, which has resulted in a boom for data-security related business.


Tuesday, February 15, 2011

How Apple May Inadvertently Boost eBook Linking

"The net interprets censorship as damage and routes around it." John Gilmore, 1993
The official word from Apple finally came out today, in their press release announcing in-app subscriptions.
In addition, publishers may no longer provide links in their apps (to a web site, for example) which allow the customer to purchase content or subscriptions outside of the app.
Assuming that this limitation will be applied to the Kindle App, it means that the "Shop in Kindle Store" button will disappear, and similar features in other ebook reader software, such as Nook, Sony, and Kobo will disappear as well.

You can go elsewhere if you want to read apocalyptic whining about Apple's imperious ways. What I want to focus on is how the net will route around this damage. The net will route around this damage by making more links.

At O'Reilly's Tools of Change for Publishing Conference, I was able to spend some time with Keith Fahlgren, a partner at ThreePress Consulting. He's part of a group that has worked on the improvement of linking capability in EPUB 3. A Public Draft of the specification was released today by IDPF.

Since EPUB3 is based on HTML5, all the outbound linking that you would expect from a web page is already built into EPUB3 (as well as earlier versions of EPUB). Ebook reader apps available on iOS and Android use the "Webkit" webpage renderer for ebooks in EPUB. (Kindle devices use Webkit to render web pages and WebKit is used by Amazon to render Kindle ebooks (in mobi format) on  hardware other than their own.) So it's clear to me, at least, that even if ebook reader apps can't have "Kindle Store" buttons, the apps will be able to present "Kindle Store" links inside the ebook content. I'll bet you anything that Amazon is loading up ebook content with Kindle Store links: "If you like this book, perhaps you'd like this one". They'll even have specialized shop-books containing Kindle store links available for free. Ditto the others.

Publishers aren't going to like Apple's power-play. But neither will they like having their content getting hijacked to promote individual ebook stores. There will therefore be a great deal of pressure for the creation of vendor-neutral, customer friendly ways to link to ebooks from within ebooks, one that Apple can't ban because doing so would break Safari.

Here's where it gets tricky. If a customer has already purchased the linked-to book, it's pointless to send them out to a ebook store, they should connect their copy of the ebook. But figuring out whether a consumer already has the book is messy, given the state of ebook identification. There are many other use cases for linking to a specific chapter or paragraph inside an ebook.

Unfortunately, doing this sort of linking is not a solved problem. EPUB3 adds one tool that will help. A new required metadata property, dcterms:modified,  will help identify the epub in the case where it has been modified- in the past it was poorly specified what should happen to the epub identifier if the file was modified. With EPUB3, it's now clear that EPUB documents are identified internally at a level above the ISBN (different DRM wrappings of the same EPUB file often require different ISBNs) but below the "work".

There's still a lot of apparatus that will need to be built, both inside and outside of EPUB, for linking to work the way it should. Being able to decide which ebook to target will require external mechanisms. Perhaps some linking organization along the lines of Crossref be formed; perhaps a more wikipedia-ish database collaboration will suffice. In any case something like xISBN supercharged for ebooks will be needed. Fahlgren told me that without a strong use case to drive the solution, the EPUB group has had a hard time going very far in their linking development.

A true ebook linking solution would need to include Amazon, of course, and since they've not been using EPUB, it seems to me that ebook linking won't get done by the EPUB group itself. Amazon hasn't had much use for EPUB in the past, but now Apple may have handed the ebook technology community a giant use case for interoperable ebook linking.

Happy Day-After-Valentines-Day, EPUB!

Wednesday, January 19, 2011

eBook Identifier Confusion Shakes Book Industry

Taipei 101
I've only felt a strong earthquake once. I was on the second floor of an engineering building at Stanford, and as soon as the initial jolt shook the building I thought "cool, it's an earthquake!". Then the rolling started. It was only after the shaking was over that I started shaking myself. The feeling of solid ground beneath my feet had been wrenched out of my psyche, leaving me standing on a big bowl of jelly that could start jiggling again any moment.

Big earthquakes can cause building damage and collapse. Sometimes, it's because a builder hasn't followed code, and the violations are exposed by the stress of a quake. Other times, it's because the building code didn't properly anticipate the stresses of the earthquake. Either way, after a severe earthquake, buildings need to be inspected to assess damages and to determine if changes need to be made in the building code.

Modern technology allows buildings to soar through traditional limitations. For example, the engineers of Taipei 101, which was the worlds' tallest building from 2004 to 2010, put a huge tuned mass damper system at the top of the tower. They made a virtue out of necessity, and the damper is now on display as a dramatic part of the Taipei 101 tourism experience, well worth the visit if you go to Taipei. (I was there in 2006.)

the tuned mass damper in Taipei 101
The Book Industry has been experiencing tectonic shifts as it moves from the solid foundation of print-based production and distribution to digital forms. The so-called "supply chain" is a long-standing edifice of the book industry being shaken by the resulting quakes. One of the strings holding the supply chain together is the ISBN, and it has proven to be reasonably robust. Still, there's been enough "damage" to the ISBN and the supply chain it holds together that many participants in the book industry have been concerned for its integrity. (I wrote about the situation in July.)

Last Thursday, I was fortunate to be at a presentation of the Book Industry Study Group (BISG) about identification of eBooks. BISG hired Michael Cairns, the principal of Information Media Partners, to do a study of the use, issues and practice surrounding assignment of ISBNs in the US book industry. Think of him as a structural engineer hired to inspect the damage to the supply chain's supporting infrastructure after an earthquake. Cairns conducted 55 separate interviews with a total of 75 industry experts from all facets of the industry. (I was interviewed for my expertise in the use of ISBN in library linking systems).
Cairns (@personanondata on Twitter) is an industry veteran- he's held senior executive positions at Bowker and other companies. His presentation was clear and direct, and he quickly went to the heart of the matter. He found very little support for the policy set forth by the 2005 revision of the ISBN standard regarding when to assign a new ISBN to an ebook. Not surprisingly, he found that implementation of that policy is all over the map, with little coherence between one company and another in ISBN assignment practice. What's more, he found that the industry is almost unable to communicate with itself due the wide variations in the practical definitions of terms such as "format", "product", "version" and "work".

Despite the difficulties created by the uneven application of the standard, there's no collective desire in the industry to "fix" the problem. Everybody has patched their systems to make them work in spite of a damaged infrastructure. The result is that poor practice has been structurally incorporated into the ebook supply chain, such that it doesn't help any more to do things correctly. If everyone started following the rules tomorrow, the supply chain might stop working.

It's as if an addition to a building needed to be built during an earthquake, even as things continued to shake. The framework is crooked, but that's needed to keep the building from falling over. You shouldn't expect such an addition to be perfect; it's something of a miracle that it can be built at all.

One example of how supply chain tremors putting stress on the supply chain edifice was raised in the discussion after Cairn's talk. At BN.com, they are enhancing some ebooks for the Nook. The enhanced ebooks are then offered at a different price than unenhanced ebooks. Normally, this would not affect ISBN assignment, because the modified ebooks are sold only by BN in the Nook store, and no one else would be affected. But last year, the supply chain was shaken when 5 of the big 6 publishers moved their ebooks to the "agency model". All of a sudden, the ebooks sold in the Nook store were being set by the publisher. The publisher was now pulling price strings for each version of the ebook, and the string being used was, you guessed it, the ISBN. So the result of the shift to an agency model was that a whole bunch of ebooks suddenly needed their own ISBNs.

While everybody seems to be scraping by for now, there may be severe problems lying ahead. Cairns pointed to libraries as a supply chain participant that was already experiencing ebook ISBN dystopia, and he suggested that the experiences of libraries today may presage the sort of problems which may spread to consumer markets as the ebook industry matures.

Libraries have historically had a different relationship to metadata than  publishers and other supply chain participants. They KEEP their books. Publishers pay a lot of attention to metadata when a book is created because it helps them sell books. Then, they're pretty much done with the metadata. If the data rots (goes out of date), it's not really a publisher problem. So libraries have maintained their own metadata to allow them to manage their collections.

eBook metadata is forever. Because ebooks are licensed, not sold, the licensor retains a relationship with the purchaser extending beyond the sale, and must maintain metadata surrounding the license for much longer than in the case of printed books. There are new sets of intermediaries and many more possibilities for business models. This is already playing out in library distribution channels, where ebooks are being licensed, lent, rented, printed, viewed, bundled into packages and purchased. If multiple sets of licensing terms are used for an ebook, resulting in multiple products with different prices attached, are new ISBNs needed? In the past, the answer would be a clear "no"; things like the agency model have changed that to a clear "I don't know".

Another issue laid out by Cairns was the low profile and negative perception of the US ISBN Agency (and by extension, ISBN International) in the ebook industry. Many of his interviewees had the impression that the assignment policies were being driven by the agency's business model (basically, the selling of ISBNs and related databases). If only it were so simple!

Brian Green, Executive Director of ISBN International spoke briefly about a similar study his group had commissioned. Although this study (PDF, 509KB) focused less on the US situation, many of its findings were similar to those of the Cairns report. At least one recommendation in that report has been acted on- ISBN International has released an updated FAQ (PDF, 363 KB) on assignment of ISBNs to e-books. You can help with another recommendation by helping to disseminate it widely!

The BISG's role in all of this is to serve as a place where the book industry can sit together and figure out how to function more effectively. The work of the BISG committee that sponsored the Cairns study will be to develop new consensus around practices and resources that will help to solve problems. Clearly, the committee has a lot of work do, building on the structural assessment laid out by the Cairns report. Development of a common vocabulary and set of definitions may be a very productive starting point for the group.

Perhaps the book industry will need the standards equivalent of a tuned mass damper. I can't wait to visit that skyscraper.

Wednesday, July 14, 2010

What IS an eBook, Anyway?

One of my secret pleasures at American Library Association meetings is going to Standards sessions. Now before you think I have a completely hopeless case of nerdiness, let me explain myself.

There's never just one Standards session at ALA, there are at least two and often three or more. I'm not sure why, but I think it's because librarians feel that standards are Important, and because there are so many Standards in the library world that people forget which ones were the subject of a Standards session at the last meeting. Its not that librarians are interested in Standards, it's just that they have lots of data problems that might magically go away, if only there were a Standard. Or not.

Because there are so many session on Standards, each one tends to be sparsely attended. That's why I like them. You can go and sit in a room with some really smart and influential people (the panelists), ask them bizarre Standards questions, have some other really smart audience member join in the discussion, and feel like you're a member of some hidden clique of powerful numerologists.

One of the things that has the Standards people concerned this year is the way ISBNs are being applied to ebooks. At the session I went to, Brian Green, the Executive Director of the International ISBN Agency, was giving his standard ISBN Standards update. Brian has been doing this long enough that he expects and parries my pestering questions with aplomb.

So here's this year's burning question: How many ISBN's should be issued when ebooks are published in different formats? Should the ebook have the same ISBN as the print book? If a different ISBN, should different file formats get separate ISBNs?

And here's the burning answer from ISBN International: each ebook file format for a book should get its own ISBN:
Do different formats of an electronic or digital publication (e.g., .pdf, .html) need separate ISBNs?

Different formats of an electronic or digital publication are regarded as different editions and therefore need different ISBNs in each instance when they are made separately available
And here's the language of the Standard itself, (ISO 2108:2005) adopted through the international standards process in 2005:
Each different format of an electronic publication (e.g. ".lit", ".pdf", ".html", ".pdb") that is published and made separately available shall be given a separate ISBN.
So forgive me for having been confused in March, when I read that the “E-book ISBN Mess Needs Sorting Out,” Say UK Publishers. Why are the publishers still talking about this, more than ten years after the question was raised and thoroughly discussed? Why are we having panels at ALA to learn about this? Has the numeracy of the world's book industry been entirely depleted during ISBN's switch to 13 digits???

ISBN stands alone in the world of identifiers because of its widespread pre-internet adoption and success. Even the Internet Engineering Task Force set aside some URI space for it back in the days before "HTTP" became a religious invocation. But most people outside the book industry have had no idea of what it really identified- they usually think it identifies a book or perhaps a book version.

If the book industry had a Facebook profile, it would list its relationship with ISBN as it's complicated.  Consider ISBN 978-1593967574. It is a "Year 5 Harry Potter Bust" manufactured by Diamond Comics. It has no author, pages or even words; it is not a book in any sense. Yet it is well-behaved in the ISBN world, because it is (or was) an item distributed by the world's book supply chain to bookstores and ultimately consumers.

In the print world, it is more or less understood that a paperback has a different ISBN from the hardcover, which has a different ISBN from the library-bound version, and may have a different set of ISBNs when issued in a different country. At the deepest level, the ISBN is just a solution to a problem: "How does an item get tracked through the book supply chain?"

If you see a book on the shelves of a bookstore, you can be pretty sure that it got there through the "supply chain". Book publishers don't sell books to book stores, they mostly sell to distributors such as Ingram and Baker & Taylor. Bookstores use ISBNs to order books, and the distributors use the ISBN to report sales back to the publishers. When books don't sell, they get shipped back to warehouses, which track them using...ISBN.

When there's a question about whether a different ISBN should or should not be issued, the overriding principle is "a product needs a separate identifier if the supply chain needs to separately identify it." This clarity about the function of an ISBN is what has resulted in its overwhelming success. When people try to use the ISBN for other things, it's less successful. Supplemental services such as xISBN (which I helped put into production at OCLC), thingISBN, and emerging identifiers such as ISTC are useful for filling in the gaps between what ISBN really is and what people would like it to be.

Let's look at ebooks with the prism of the supply chain. If an ebook is issued in print, PDF and EPUB formats, it's important to the publisher to know how many of each are sold, thus the separate ISBN's. Similarly, if different DRM wrapping is used by two different channels, in many cases the publisher will need to track sales or manage the product separately. Although in many cases the DRM could be tracked by retailer, and thus wouldn't need a separate ISBN, the ISBN Standard says to give it a different ISBN. As Green has written previously,
Where publishers are selling e-books exclusively from their own websites or through another single channel and do not wish to have them listed in books in print databases then [...] publishers may not wish to bother with ISBNs. However, publishers should beware of taking a short-term view that makes them reliant on a single channel.
Unfortunately some publishers have obstinately refused to give separate ISBNs to ebooks in different formats. The US division of Random House is perhaps the most prominent example. There are excellent arguments for the "single ISBN" approach, but the worst possible situation for the emerging supply chain is for each publisher to use their own inconsistent rules for applying ISBN to ebooks. However strong the argument is for "single ISBN", its inconsistent application negates the advantages and threatens the ISBN system as a whole.

The ultimate problem with ISBN and ebooks is that ebooks are sufficiently adaptable that they expose  ambiguities and limitations of the ISBN identification architecture. For example, suppose you're in the business of selling customized digital coursebooks. You allow professors to choose 10 chapters from 100 available. That means there are exactly 17,310,309,456,440 different ebooks that you could sell. That's about 9,000 times more books than can be identified by all the ISBNs in the galaxy. But you don't need to give them ISBNs, because you sell direct and the ebooks never touch the supply chain. The chapters themselves may need to be tracked so you can pay author royalties, but you need only 100 ISBNs to do that.

How about if a retailer changes (or eliminates) the DRM wrapping an ebook? Do the ISBN's of the ebooks on a consumer's ebook reader magically change? (Transubstantiation is one of my favorite words!) The answer is no, and that's because the the supply chain is not involved.

Are there enough ISBNs for the ebooks that could be sold? The EPUB format is actually an archive file format that uses a dialect of XHTML for its insides, so you might imagine that any website or portion thereof can be packaged as an ebook. In fact, BookGlutton has a tool that (sort of) does this. As of May 2009, over 100 million websites operated, so you can easily imagine that ebooks could use up all available ISBN's almost overnight.

The "supply chain" for ebooks is rapidly mutating. The adoption of an "agency model" is an example of a change that has put new demands on ISBN; "agency" requires a retailer to identify an item's publisher before the moment of sale so that the correct sales tax can be applied. The agency model shift won't be the last or biggest change to the ebook supply chain, either. As one example, I've previously written about ebook pay-per-view and demand-driven acquisition. Another huge change would occur if  a substantial advertising revenue stream for ebooks, such as Apple's iAd system, emerges. Advertising would put new demands on reporting systems and thus on the ISBNs that enable them.

What is an ebook anyway? Ten years ago, a committee of the American Association of Publishers came up with this not-so-useful definition:
An ebook is a literary work in the form of a digital object consisting of one or more standard unique identifiers, metadata, and a monographic body of content, intended to be published and accessed electronically.
I'll bet you never realized that blog posts were really ebooks!

The truth is that we really have no idea what an ebook is or what it will become. There are certainly e-things that correspond to print books, and these are easy to recognize as ebooks. But don't be surprised if there comes a flood of things to read on our connected devices that are too long to be called "articles" or "posts". For these, "eBook" may be the best label we can come up with.

Unless of course they get shackled by a supply chain.

Wednesday, April 7, 2010

The Library IS the Machine

When librarians catalog a book, they do their best to describe a thing they have in their hands. The profession has been cataloging for a long time, and it tends to think that it's reduced the process to a science. When library catalogs became digital in the 1970's, the descriptions moved off of paper cards and into structured database records using a data format called MARC. That stands for MAchine Readable Cataloging, and as one Google engineer recently complained, "the MAchine Readable part of the name is a lie". The problem that Google's machines are having with these records is that the descriptions have always been meant for humans to read, not for computers to parse and understand.

Cataloging librarians are not stupid, and they've been working since the very beginning of digital cataloging to make their descriptions more useful to computers. They've introduced "name authority files" to bring uniformity to things like subject headings and author and publisher names. Unicode has brought uniformity to the encoding of non-roman characters and diacritics. XML has replaced some of the ancient delimiters and message length encoding. And perhaps most importantly, for a long time they've been embedding identifiers in the catalog records. Despite all this, library catalog records are still not as computer-friendly as they should be.

The move towards identifiers is worth special note. The use of identifiers in libraries dates to the first industrialization of libraries that took place in the 19th century. The classification systems of Melvil Dewey, Charles Ammi Cutter and the Library of Congress were all efforts to make library catalogs more friendly to machines.  Except the machines weren't digital computers, the machines were the libraries themselves. From the shelves to the circulation slips, libraries were giant, human-powered information storage and retrieval machines. The classification codes are sophisticated identifier systems upon which the entire access system was based. So maybe MARC isn't a lie after all!

The rest of the world took a while to catch up on the use of identifiers. The US began issuing social security numbers in 1936, but it wasn't until the 60's with the adoption of ISBN in the 1966 and ISSN in 1971 that the entire publishing industry began to use identifiers to more efficiently manage their sales, delivery and tracking of products.

The same properties that made identifiers useful in physical libraries make them essential for digital databases. Identifiers serve as keys that allow records in on table to be precisely sorted and matched against records in other tables. Well designed identifier systems provide assurances of uniqueness: there may be many people with the same name as me, but I'm the only one with my social security number.

Nowadays, it sometimes seems that almost any problem in the information industries is being solved by the introduction of a new identifier. Building on the success of ISBN and ISSN, there are efforts to identify works (ISTC),  authors (ORCID, ISNI), musical notations (ISMN), organizations (SAN), recordings (ISRC), audio-visual works (ISAN), trade items (UPC) and many other entities of interest. We live in an age of identifiers.

The apotheosis of indentifiers has been achieved in the Linked Data movement. The first rule of Linked Data is to give everything- subject, objects, and properties, their own URI (Uniform Resource Identifier). By putting EVERYTHING in one global space of identifiers, it is expected that myriad types of knowledge and information can be made available in uniform and efficient ways over the internet, to be reused, recombined, and reimagined.

What's often glossed over during the adoption of identifiers is their fundamental pragmatism. The association between any identifier and the real-world object it purports to identify is a thinly veneered but extremely useful social fiction which doesn't approach mathematical perfection. Even very good identifier systems can fail as much as 1% of the time, and automated systems that fail to recognize and accommodate the possibility of identifier failure exhibit brittleness and become subject to failure themselves. Still 99% of perfect works perfectly fine for a lot of things.

A decade ago, the world of libraries and the publishers that supply them embarked on an effort to link together the citations in journal articles and the bibliographic databases essential to libraries with the cited articles in e-journals and full text databases. Two complementary paths were pursued. One effort, OpenURL, sent bibliographic descriptions inside hyperlinks, and relied on intelligent agents in libraries to provide users with institutional specific and relevant links. The other, CrossRef, built identifiers for journal articles into a link redirection system. Together, OpenURL and CrossRef built on the strengths of the description and identification approaches and do a reasonably good job serving a wide range of users, including those in libraries.

Now, however, the slow but sure development of semantic web technologies and deployment of Linked Data has spurred both CrossRef's Geoff Bilder and the OCLC's Jeff Young (OCLC runs the OpenURL Maintenance Agency) to examine whether CrossRef and OpenURL need to make changes to take advantage of wider efforts. In another post, I'll look at this question more closely, but for now, I'd like to comment on what we've learned in the process of building article linking systems for libraries.

1. Successful linking requires both identification and description. The use of CrossRef by itself did not have the flexibility that libraries needed; CrossRef addressed this by making its bibliographic descriptions available to OpenURL systems. Similarly, the OpenURL's ability to embed CrossRef identifiers (DOIs) inside hyperlinks has made OpenURL linking much more accurate and effective.

2. Successful linking is as much about knowing which links to hide as about link discovery. Link discovery and link computation turn out not to be so hard. Keeping track of what is and isn't available to a user is much harder.

3. Bad data is everywhere. If a publisher asks authors for citations, 10% of the submitted citations will be wrong. If a librarian is given a book to catalog, 10% of the records produced will start out with some sort of transcription error. If a publisher or library is asked to submit metadata to a repository, 10% of the submitted data will have errors. It's only by imposing the discipline of checking, validating and correcting data at every stage that the system manages to perform acceptably.

Linking real world objects together doesn't happen by magic. It's a lot of work, and no amount of RDF, SPARQL, or URI fairy dust can change that. The magic of people and institutions working together, especially when facilitated by appropriate semantic technologies, can make things easier.

Reblog this post [with Zemanta]

Monday, January 18, 2010

Google Exposes Book Metadata Privates at ALA Forum

At the hospital, nudity is no big deal. Doctors and nurses see bodies all the time, including ones that look like yours, and ones that look a lot worse. You get a gown, but its coverage is more psychological than physical!

Today, Google made an unprecedented display of its book metadata private parts, but the audience was a group of metadata doctors and nurses, and believe me, they've seen MUCH worse. Kurt Groetsch, a Collections Specialist in the Google Books Project presented details of how Google processes book metadata from libraries, publishers, and others to the Association for Library Collections and Technical Services Forum during the American Library Association's Midwinter Meeting.

The Forum, entitled "Mix and Match: Mashups of Bibliographic Data", began with a presentation from OCLC's Renée Register, who described how book metadata gets created and flows though the supply chain. Her blob diagram conveyed the complexity of data flow, and she bemoaned the fact that library data was largely walled off from publisher data by incompatible formats and cataloging practice. OCLC is working to connect these data silos.

Next came friend-of-the-blog Karen Coyle, who's been a consultant (or "bibliographic informant") to the Open Library project. She described the violent collision of library metadata with internet database programmers. Coyle's role in the project is not to provide direction, but to help the programmers decode arcane library-only syntax such as "ill. (some col)". The one instance where she tried to provide direction turned out to be something of a mistake. She insisted that, to allow proper sorting, the incoming data stream should try to keep track of the end of leading articles in title strings. So for example, "The Hobbit" should be stored as "(The )Hobbit". This proved to be very cumbersome. Eventually the team tried to figure out when alphabetical sorting was really required, and the answer turned out to be "never".

Open Library does not use data records at all, instead, every piece of data is typed with a URI. This architecture aligns with W3C web standards for the semantic web, and allows much more flexible searching and data mining than would be possible with a MARC record.

Finally, Groetsch reported on Google's metadata processing. They have over 100 bibliographic data sources, including libraries, publishers, retailers and aggregators of review and jacket covers. The library data includes MARC records, anonymized circulation data and authority files. The publisher and retailer data is mostly ONIX formatted XML data. They have amassed over 800 million bibliographic records containing over a trillion fields of data.

Incoming records are parsed into simple data structures which looked similar to Open Library's, but without the URI-ness. These structures are than transformed in various ways for Googles use. The raw metadata structures are stored in an SQL-like database for easy querying.

Groetsch then talked about the nitty-gritty details of data. For example, the listing of an author on a MARC record can only be used as an "indication" of the authors name, because MARC gives weak indications of the contributor role. ONIX is much better in this respect. Similarly, "identifiers" such as ISBN, OCLC number, LCCN, and library barcode number are used as key strings but are only identity indicators with varying strengths. One ISBN with a chinese publisher prefix was found on records for over 24,000 different books; ISBN reuse is not at all uncommon. One librarian had mentioned to Groetsch that in her country, ISBNs are pasted onto a book to give it a greater appearance of legitimacy.

Echoing comments from Coyle, Groetsch spoke with pride of the progress the Google Books metadata team has made in capturing series and group data. Such information is typically recorded in mushy text fields with inconsistent syntax, even in records from the same library.

The most difficult problem faced by the Google Books team is garbage data. Last year, Google came under harsh criticism for the quality of its metadata, most notably from Geoffrey Nunberg. (I wrote an article about the controversy.) The most hilarious errors came from garbage records. For example, certain Onix records describing Gulliver's Travels carried an author description of the wrong Jonathan Swift. Most of these errors come from garbage records, and when one of these is found, almost always, the same problems can be found in other metadata sources. Google would like to find a way to get corrected records back into the library data ecosystem so that they don't have to fix them again, but that there have been issues with data licensing agreements that still need to be worked out. Article like Nunberg's have been quite helpful to the Google team. Every indication is that Google is in the metadata slog for the long term.

One questioner asked the panel what the library community should be doing to prevent "metadata trainwrecks" from happening in the future. Groetsch said without hesitation "Move away from MARC". There was nodding and murmuring in the audience (the librarian equivalent of an uproar). He elaborated that the worst parts of MARC records were the free text data, and normalization of data would be beneficial whereever possible.

One of the Google engineers working on record parsing, Leonid Taycher, added that the first thing he had had to learn about MARC records was that the "Machine Readable" part of the MARC acronym was a lie. (MARC stands for MAchine Readable Cataloging) The audience was amused.

The last question from the audience was about the future role of libraries in production of metadata. Given the resources being brought to bear on the book metadata by OCLC, Google and others, should libraries be doing cataloguing at all? Karen Coyle's answer was that libraries should concentrate their attention on the rare and unique material in their collections- without their work, these materials would continue to be almost completely invisible.
Reblog this post [with Zemanta]

Wednesday, July 29, 2009

The Illusion of Internet Identity

You've certainly heard of Arthur C. Clarke's Third Law, "Any sufficiently advanced technology is indistinguishable from magic", which says more about magic and our perceptions of the world than it does about technology. When technology does something that is not natural to us, we of course perceive it to be supernatural. But what happens when technology approximates something so natural to us that we don't even perceive that there's anything remarkable? Then we attribute powers to the technology that just don't exist. Just as we can perceive emotions in a stuffed teddy bear, it is only with difficulty that we avoid anthropomorphizing technologies. Do you have a cute name for your car? Do you refer to your GPS as "Lola"? If you have not done so, try the web version of ELIZA, and see if you can avoid thinking of ELIZA as a real person. It's very hard for us to understand how complicated the act of carrying on a real conversation really is- we do it all the time. Even a profoundly retarded technology will be imbued with magic if its function is sufficiently mundane.

I've been reading a paper by Patrick Hayes and Harry Halpin and a presentation by Pat Hayes, both with the unfortunate title "In Defense of Ambiguity". The paper provides a wonderful review of the theory of identity. I've been living very happily, doing productive work in the world of identifiers without ever knowing that identity needed to have a theory behind it. In retrospect, I've managed to do this by staying away from the difficult bits.

After reading Hayes and Halpin, I've come to realize what a miracle human communication is. The fact that I can meet someone with whom I share no languages and that we can exchange our own names and establish names for things may seem simple, but it's something that machines cannot do. For example, I may gesture at myself and say "Eric" to establish my identity. If I then gesture at a banana and say "banana", it's very likely that my counterpart will understand that I have not established an identity for the banana, but rather I have given a name for the kind of fruit. This is possible because people have brains that are similarly wired- our brains are wired to recognize individual people but not individual bananas (though sometimes our knowledge models diverge). Our computers on the other hand, have no fruit wiring, or individual person wiring, so establishment of identity is very hard for them.

The difficulty of teaching computers to identify things has not stopped us from using them to build elaborate identity systems. Hayes and Halpin observe that internet identity can only be established by description, and description is inherently ambiguous. Attempts to make real-world-object identifiers global or to add description actually make the situation worse, by increasing ambiguity. In our daily lives, ambiguity in our communications is mostly not a problem. When I say the word "rose" a listener will almost never be confused between the flower and the verb. I can say "rows of rosebushes" and only rarely will people hear "rose of rosebushes". Our brains are so good at using context to resolve ambiguity that we don't realize how hard it is for computers to do the same thing.

That the situation becomes worse with added description was a bit hard for me to absorb, because at first it seems that the better you define something, the less ambiguous your statements about it become. But it's not true for computers. Suppose my internet identity description added a physical description of me- for example the fact that I have blonde hair. That might help to identify me under certain circumstances, but then when my hair turns gray, it makes my identification more tenuous. You could say that I had blonde hair on a particular date, but then you'd need to add a physical model for hair color to your internet identity system. In actual fact, the added description might help a human to identify me, but it hurts a computer's efferts to establish my identity.

The Hayes and Halpin paper was written in the context of the "http-range" semantic web controversy that I touched upon in my post on the semantics of redirection. They argued that the http protocol is not the right place to put establishments of identity, and that the description model is better suited to do that. As I understand it, the Hayes-Halpin view did not prevail with the W3C TAG; "ambiguity" was not a great concept for people to rally around, I guess.

The internet identity systems I've worked with revolved around identifying things in libraries- books, serials, and articles. For the most part, these sorts of objects do not usually present deep identification quandaries, and so I've not noticed my ignorance of identity theory. For example, most people imagine that computers use ISBNs to identify books (or as I imagined in my last post, that computers use ISBNs to identify items in bookstores). Most often this illusion does not get us into any trouble, just as the illusion that teddy bears have feelings is mostly harmless. Computers are wired to deal with records in data files, and they use ISBNs (with frequent success) to identify and match records in data files, that's all. The rest is just software trickery.

It's interesting to note that the ISBN was not developed by the library community. It was developed by a statistics professor named Gordon Foster for the British Publishers Association. Librarians lived without identifier systems for many years and were content with the library equivalents of an address or locator system. It's as if librarians have intuitively known something that the architects of the semantic web have only recently struggled with- that we can aspire to build description systems and access systems, but building a system that can provide identity is more difficult than it looks like to a human.

Friday, July 24, 2009

If Elvis had an OpenID and the Mome Raths Outgrabe

If the space aliens that kidnapped Elvis decided to return him tomorrow, and he decided to use Twitter and a blog to communicate that fact, how would anyone know it was really him? Would the National Enquirer even bother to report the news? Would Elvis ever be able to reclaim his public or private identities? Would he be able to remember any of his passwords?

Password proliferation has been a problem for so long that innovators have solved it over and over again. The library world came up with an single-sign-on authentication/authorization system called Shibboleth, and then implemented EZProxy so they wouldn't have to deal with it. The UK developed the single-sign-on system called Athens. The dot-com bubble came up with a bunch of single-sign-on companies; some of them, including PassLogix and Imprivata are still at it. I am still waiting for the announcement of the Single-Single-Sign-On system.

OpenID took a different approach, and is now somewhat usable for the purpose of allowing people to establish an identity with one provider that can be used on many websites. For example, I've used the OpenID identity "http://go-to-hellman.blogspot.com/" to register comments on the Semtech 2009 website and the Paul Miller's Cloud of Data Blog. On my last post, comments were left by "nicomo", and "breizhlady", whose OpenID's are http://nicomo.pip.verisignlabs.com/ and http://breizhlady.myopenid.com/ . Jodi Schneider used Blogger credentials to leave her comment. My OpenID can be used to determine with some degree of certainty that the the Eric Hellman who left a comment on Cloud of Data is the same Eric Hellman who's writing on this blog. A bit of googling will tell you who nicomo and breizhlady are, if you really want to know. If Elvis had been issued an OpenID before he left we would be able to tie his new blog to his old identity.

There are still the single-single problems with OpenID. The user experience for OpenID systems gets a bit clunky- my wife was frustrated when she tried to leave a comment on this blog. But overall there seems to be slow convergence and user acceptance of OpenID.

This brings me to the questions I wanted to raise today: What does an OpenID identify? Does http://go-to-hellman.blogspot.com/ identify me? Can these OpenIDs be used to make assertions about people to enter into the Linked Data Cloud? How should the Linked Data semantics for redirects be implemented for OpenID? Should 303 redirects be used to indicate that the "thing" being identified by OpenID is a real-world object?

To some extent, it's really the way identifiers are used that determines semantics- identification of any real-world object can never have perfect accuracy. The use of ISBN to identify a book is a good example. Although ISBN is frequently used to identify a book, ISBNs are managed in such a way that they most accurately identify items sold in a bookstore- toys and dolls often get ISBNs. Similarly, you might think that the US identifies people with Social Security Numbers (SSN), but if you think about it, the "thing" an SSN most accurately identifies is an account with the Internal Revenue Service. Similarly, I think it's pretty clear that an OpenID identifies a set of login credentials, although people might well use the OpenID to identify the person or persons behind it.

I have been guilty in the past of driving people to distraction by arguing that it can be almost impossible to decide whether something is an "information resource" (something whose essential characteristics can be conveyed in a message) or whether it is a "real-world object". It's pretty easy to blur the issue with an e-book, for example, but what about the SSN? It used to be that "an IRS account" was something on paper somewhere, but I'm pretty sure that my entire IRS account is digitized somewhere.

Section 3 of the W3C's Technical Recommendation "Cool URIs for the Semantic Web" assumes that it's easy to determine whether something is an information resource or whether it's a real-word object and that it's impossible to convey the essence of real-world objects in a stream of bits. I find this a bit unworldy. It even cites the unicorn as an example of a "real-world object". I guess that makes Elvis a real-world object, too. Conversely, even things that live completely on the internet are rarely "conveyed in a message" any more. A typical URI-addressable service today is constructed out of software, web services, content delivery networks, advertising delivery networks and clustered hardware so that the "essential characteristics" include the attributes of real world objects like me.

I've recently become aware that lots of really smart people have thought and written about the theory of identifiers and about how the Semantic Web should handle them. I've particularly enjoyed an article called "In Defense of Ambiguity" by Patrick Hayes and Harry Halpin. But to answer my questions about the semantics of OpenID, there's no sage more useful than the one who said "When I use a word, it means just what I choose it to mean - neither more nor less." Semantics do not get determined by those who mint the identifiers, but rather by those who make use of them. It helps if they are also willing to pay the IDs a bit extra.

Tuesday, May 12, 2009

Google, RDFa, and Reusing Vocabularies

Yesterday, I wrote about one difficulty of having machines talk to other machines- propagation and re-use of vocabularies is not something that machines being used today know how to do on their own. I thought it would be instructive to work out a real example of how I might find and reuse vocabulary to express that a work has a certain ISBN (international standard book number). What I found (not to my great surprise) was that it wasn't that easy for me, a moderately intelligent human with some experience at RDF development, to find RDF terminology to use. I tried Knoodl, Google, and SchemaWeb to help me.

Before I complete that thought, I should mention that today Google announced that they've begun supporting RDFa and microformats in what they call "rich snippets". RDFa is a mechanism for embedding RDF in static HTML web pages, while microformats are a simpler and less formalized way to embed metadata in web pages. Using either mechanism, this means that web page authors can hide information in structures mean to be read by machines in the same web pages that humans can read.

Concentrating on just the RDFa mechanism, it's interesting to see how Google expects that vocabulary will be propagated to agents that want to contribute to the semantic web: Google will announce the vocabulary that it understands, and everyone else will use that vocabulary. Resistance is futile. Not only does Google have the market power to set a de facto standard, but it has the intellectual power to do a good job of it- one of the engineers on the Google team working on "rich snippets" is Ramanathan V. Guha, who happens to be one of the inventors of RDF.

You would think that It would be easy to find an RDF property that has been declared to use in assertions like "the ISBN of 'digital Copyright' is 1-57392-889-5". No such luck. Dublin Core, a schema developed in part by the library community, has an "identifier" element which can be modified to indicate the element contains an isbn, but no isbn property. Maybe I just couldn't find it. Similarly, MODS, which is closely related to library standards, has an identifierType element type that can contain an ISBN, but you have to add type=isbn to the element to make it an ISBN. Documentation for RDFa wants you to use the ISBN to make a urn and to make this the subject of your assertion, not an attribute (ignoring the fact the ISBN identifies things that you sell in a bookstore (for example, the paperback version of a book) rather than what most humans think of as books. I also found entries for isbn in schemes like The Agricultural Metadata Element Set v.1.1 and a mention in the IMS Learning Resource Meta-Data XML Binding. Finally I should note that while OpenURL (a standard that I worked on) provides an XML format which includes an ISBN element, it's defined in such a way that it can't be used in other schemas.

The case of ISBN illustrates some of the barriers to vocabulary reuse, and although there are those who are criticizing Google for not reusing vocabulary, you can see why Google thinks it could work better if they just define vocabulary by fiat.