Friday, January 29, 2010

Who Owns Koha?


In New Zealand, Maori customs are taken seriously. For example, in 2002, the route of a new highway through a swamp had to be altered because it was believed that three taniwha - Karutahi, Waiwai, and Te Iaroa - lived there, and were being disturbed by the road, causing an unusual number of accidents. Taniwha are mythical beings that act as guardian spirits. Many taniwha arrived in New Zealand as guardians of specific ancestral canoes and then took on a protective role over the descendants of the canoe's crew.

Another Maori custom, one that has crossed over into general New Zealand culture, is that of "koha".  "Koha" is often translated as "gift", but according to Chris Cormack, one of the original developers of the Free Open-Source Software (FOSS) Library System with the same name, a more accurate translation would be a "gift with expectations".   Cormack got his B.A. degree in both mathematics and Maori Studies, so he should know. A koha is gift that is offered with an expectation that it will be reciprocated.

In the U.S., (and to a lesser extent, in Europe) it's our lawyers that we seem to take seriously. And so when we let people use software that we've developed, its not enough to offer it as a koha, we have to use a legal license to spell out the terms of the release. The license that has become popular because of the expectations of reciprocity built in to it is the Gnu Public License (GPL).

Software released under the GPL cannot be thought of as an unconditional gift by its developer. GPL software is not in the public domain, it is copyrighted. A copyright owner can exert control over the use of the copyrighted material; the GPL uses that power to require licensees to publish any modifications they make if they want to redistribute the modified work.  When the Horowhenua Library Trust was choosing a license to use for Koha (the library system software), they chose GPL (version 2) because they thought it would prevent Koha from being further developed as non-open software.

Included in the recently announced (but not yet completed) acquisition of LibLime by PTFS were Koha-related assets, including source code copyrights, trademarks and the koha.org website. It may be difficult for the casual observer to understand what value these assets have, especially in light of the GPL license attached to Koha. Could these assets be used to privatize Koha in some way? The short answer is "No", but it gets complicated.

Some FOSS companies use "dual licensing" as a business model. They release their code under GPL, but if another company wants to use the software in a way that would not be allowed under GPL, it has the option to pay for a commercial license. Dual-licensing can only be done by the original copyright owner. Index Data has used this model for years to the great benefit of libraries everywhere. MySQL AB is an example of a company which was extremely successful with this model; it was acquired by Sun Microsystems for about a billion dollars. (See this article for an overview of how GPL licensing fared in Oracle's subsequent acquisition of Sun.)

The GPL makes it very difficult, however, for anyone to change licensing terms after software has been released into the world. That's because of the way copyright law determines the copyright holder of derivative works. In general, if I take a piece of software that you have written, and modify it in such a way that involves creative effort on my part, then I own the copyright to the changes that I've made and the resulting work is a derivative work that both you and I have a copyright interest in. Even though you are the original copyright holder, you would need my permission to release the derivative work under any license other than the GPL.

Unlike Index Data's software, Koha has included significant contributions from many developers, including many who have never worked for LibLime or Katipo Communications (which sold its copyrights to Koha source code to LibLime in 2007). So although LibLime probably owned clear copyright to a majority of Koha at some point, Koha is still a collective work locked into GPL, version 2, and it is unlikely that LibLime or PTFS would be able to distribute Koha under terms other than GPL without doing a thorough rewrite of the software.

Trademarks are a different story. LibLime owns the US trademark for Koha; a European trademark is held by BibLibre. Trademarks are frequently used by open source projects to prevent splintering. The excellent primer on legal issues by the Software Freedom Law Center puts it this way:
FOSS applications develop reputations over time as users come to associate an application’s name with a particular standard of quality or set of features. Trademark law can help protect this relationship of trust and reliance that a project develops with its users; it allows the project to maintain a certain amount of control over the use of its brand.
Since GPL and other FOSS licenses allow anyone to modify and distribute software as long as the license conditions are met, they frequently spawn variants. The owner of a trademark can prevent these variants from using the trademarked name, and thus enforce unity in a project.

In the case of Koha, there are currently two parallel tracks of development being pursued, one inside LibLime, and the other by the community of developers outside LibLime. I will have to postpone a discussion of the issues surrounding  these development tracks to yet another article, but for now, let's just assume there will be two main versions of Koha, LibLime Koha and Community Koha. In the US, LibLime could theoretically prevent anyone from  using the name "Koha" without its authorization, and could strip Community Koha of the right to use "Koha" in its name. In fact, LibLime and BibLibre threatened to use this power a year ago to regulate PTFS's use of Koha trademarks in the marketing of its Koha support services. Liblime could even apply the Koha name to non-Open Source software. Similarly, BibLibre could regulate the use of the Koha name in Europe, preventing LibLime from marketing LibLime Koha there.

Based on discussions I've had with the leaders of almost every open-source library system company, I think it is unlikely that there will be any such "trademark war". Even if the development of Koha continues on separate but related tracks, the success of every Koha-based company is tied to the success Koha as a whole, and vice versa. It would be advantageous for every stakeholder if the two trademark owners develop some sort of "big tent" system of Koha trademark governance. Assuming PTFS's acquisition of LibLime is completed, such governance will need to be acceptable to both PTFS and BibLibre, and will need to accommodate differing styles of software development.

Until a general agreement on the use of Koha trademarks is reached, Koha stakeholders would be well advised to recognize that collective copyrights tie them into the same canoe and that they should avoid disturbing the taniwha that guards and protects them.

This article is the second part of a series. Part 1 is here. Part 3 is here
Reblog this post [with Zemanta]

Tuesday, January 26, 2010

Deconstructing the Attributor Book Piracy Study

It was a lot of fun to have a post slashdotted and read by 100 times more people than any of my previous posts, however, the fact that it made fun of the way a study on piracy was "spun" overshadowed some serious points.
  1. The study in question, from Attributor, in fact has some substance beneath the spin.
  2. The number of books people get from libraries is comparable to the number they get from bookstores.
First, the Attributor study. Looking past the silly projections of book industry lost sales, there is some useful information there. In the study, illicit copies of 913 books were found on four one-click download sites which display download counts. The reported download counts for these 913 books totaled 3.2 million over 90 days. On average, each book was downloaded 3,500 times total, or 39 times per day.

These numbers are in rough agreement with a much smaller sample I collected for myself. I looked at only 10 books in a single category on a single site, and observed that download totals ranged for 0 to 5000, with an average of 578.

I followed up with some questions for Rich Pearson, the General Manager at Attributor. Most interesting to me, and probably of likely concern to publishers, was this: of 913 titles, chosen from Amazon lists to get a broad distribution rather than to focus on popularity, Attributor was able to find illicit copies of 90% of the books they looked for. I had expected that the most popular titles would be easy to find, but this breadth surprised me.

The extrapolations made to translate these download numbers into economic impact, while understandable from an marketing point of view, are multiply lacking in rigor.

The first assumption made by the study is that the "downloads" reported by  download sites represent potential readers of books. Attributor does not have any special relationship with the download sites to gain access to download statistics, they simply rely on the download counters published by the sites.  Assuming that the counts are not simply manufactured (and I've seen download sites that do just this) it's unlikely that the download sites bother to distinguish between robot activity and human activity. Robot activity is important because many of the download sites do not offer search themselves. This is a tactic that allows them to be unaware of, and thus not liable for,  the content that they host. The sites depend on other sites linking to them to drive traffic; they monetize the traffic by throttling downloads and offering "premium" subscriptions to people want to remove the throttles.

The second assumption made in the Attributor study is the way that they extrapolate numbers reported by 4 sites to the 25 sites that were monitored. What they did was to weight the sites based on the distribution of the 52,000+ takedown notices that Attributor has sent out since launching their monitoring service in July of '09. One problem with this is that while the scope of takedowns was limited to Attributor customers, the 913 monitored books were limited to non-customers. If Attributor's customers were focused in textbooks, for example, this would result in a bias towards sites used to host illicit textbooks. The other problem is "lamp post" bias- the extrapolation is based on what Attributor can find; sites hidden from Attributor would result in undercounting.

Freakonomics: A Rogue Economist Explores the Hidden Side of Everything (P.S.)Third, the sampling extrapolation to cover the entire industry has significant uncertainty. I worry most about price bias. The 1,132 downloads of the book Freakonomics were reported, which seems insignificant against sales of 2.5 million copies. Freakonomics, ranked #77 at Amazon,  can be bought new for $9.35. By contrast, Architect’s Drawings, which was downloaded over 10,000 times and ranks #547,491 at Amazon,  is a $60 hardcover. (The high download count is likely due to use as a textbook). Thus the "lost sales" for Architect’s Drawings will be hugely overweighted in the extrapolation.

Architect's Drawings: A selection of sketches by world famous architects through historyFinally, while the Attributor study does note that the actual effect of downloads on sales is speculative, there is a significant question as to whether the people downloading the book copies could have purchased the books even if they wanted to. It is evident from the chatter around the book downloads that many of the downloaders speak Spanish, French, Portuguese, Arabic, and Indonesian. According to Pearson, Attributor scans sites in many languages linking to illicit books, including Chinese, Japanese, Korean, French, Spanish, German, Italian, Czech, Polish, Russian, Portuguese and Austrian. The impact of piracy is likely to be very different in export markets. Not that this is shocking news to anyone.

Attributor is currently targetting its service (which starts at around $10,000) at large publishers, though it is working with a reseller called Author Guard to service smaller publishers. It points to its superior ability to find and take down illicit content as its comparative advantage. (I've previously surveyed other companies in this space; somehow I managed to overlook Attributor.)

Although many publishers will look at the numbers and conclude that services like Attributor are not worth the expense, I think the real danger from piracy is collective rather than individual, and thus the response should be collective rather than individual. The book publishing industry will only suffer if piracy becomes so widespread that piracy gains cultural acceptance. To prevent this from happening, book publishers need to lead the culture with both carrot and stick. Legal, allmost-free access to books must be provided such as occurs today in libraries, while illicit content should be taken down in as efficient manor as possible. This means that services such as Attributor's should be deployed on behalf of the entire industry, perhaps through a consortium, and not just by individual publishers. Finally, the book industry needs to figure out how to use ebooks to effectively address the needs of the developing world, or else huge markets will be forever out of reach.

The book publishing industry is entering some scary times and needs to decide who its friends are. A don't know whether technology companies with scary marketing will prove to be reliable friends or not. Amazon and Apple might end up being saviours, but publishers shouldn't expect them to be buddies. I'm pretty sure of one thing, though. Libraries are definitely not the enemy.
Enhanced by Zemanta

Monday, January 25, 2010

8 One-Way Business Models for Linked Data

Real Wheels - Travel Adventures (There Goes a Train/Plane/Bus)One of the videos that I was forced to watch many times when my boys were younger was There Goes a Train. It's a pretty good video. I learned many things, including the fact that locomotives don't have steering wheels. Yeah. Pretty obvious if you think about it for even a moment. It's so unfair. Without the switches, the rail network would be pretty useless, but no one will ever make a video entitled There Sits a Switch.

In electronics though, the switch is the star. With a switch, you can modify and route information; with a wire you can only send it from one point to another. You need good switches to make a computer or a network; even though photons are faster and easier to move from one place to another, computers are still based on electrons because electronic switches are so much better than optical switches.

Linked Data is a label for a set of technologies that are trying to make information move around the internet more easily and with more meaning. The Linked Data vision is one where many entities acting cooperatively and globally create a web of data much more powerful and meaningful than any single entity could bring about.

For the Linked Data vision to become a reality, each entity must have a strong motivation to cooperate; each entity must have a viable business model. If the business models were easy, the Linked Data vision would already be a vibrant reality.

Scott Brinker recently launched a round of discussion about seven business models that can make Linked Data viable. Leigh Dodds contributed some important insights in his followup, prompting Brinker to add an eighth model.

Here are Brinker's eight business models for Linked Data (somewhat relabeled based on who's writing checks):
  1. Subsidy. Entities such as governments with a mandate to make information available will pay to have it linked into a global web of information.
  2. Subscription. People will pay for valuable data, and will pay more for data that has been linked to a global web of information.
  3. Advertising. Advertisers will pay to information in raw data feeds.
  4. Authority. People will pay for the validation and certification of data.
  5. Affiliate marketing. Merchants will pay sales commissions on sales resulting from affiliate links in embedded in the global web of data
  6. Service Enhancement. People will pay for services which have been enhanced by data from a global web.
  7. Search Engine Optimization. Search engines will send you more traffic if you give them more meaningful data.
  8. Brand Enhancement. Your reputation will be burnished if you emit lots of good information.
(I should note that Brinker describes each model a bit differently so that he can add a dimension that characterizes whether data is delivered raw or as an application.  I find that this dimension is not at all orthogonal. A data driven subscription service is a service that makes use of data, but the core business model is not to sell a data subscription.)

There are difficulties with all of these business models, but it strikes me that each of them will only work in one direction, like a train track without switches. Either they work for emitting data, or they work for consuming data, but none of the models work in both directions at the same time. If you're providing a service that's either based on Linked Data or enhanced by it, you can pay for the data, but if you send that data back out, your competitors get the data for free. Conversely, if you're emitting data, it's hard for you to pay for it.

Imagine you're in the book metadata business. You can use several of these models to support creation of book metadata, or you can consume book-related metadata to provide book-related services. But what if you want to support an activity of aggregating book data or fixing errors in book metadata? None of these models will work for you because you'll either be competing with the entities you get data from, or you'll be competing with entities you send data to.

What's missing from this list is a business model for the Linked Data switch. Entities that take in Linked Data, improve it or otherwise add value and reemit it as Linked Data have no solid business model to run on. Everyone active so far in the Linked Data business is either a data sink or a data source. To realize the full potential of Linked Data, there need to be viable switches, both collecting and emitting Linked Data.
Reblog this post [with Zemanta]

Thursday, January 21, 2010

PTFS to Acquire LibLime and Move to Library Systems Premier League

Update Feb.12 - the acquisition is not happening
Update Mar. 16- the acquisition closed after all.

In 2009, the New York Yankees had a payroll almost ten times that of the Florida Marlins. The reason that baseball lives with that disparity is that the financial interests of many owners do not align with their fans- they take in roughly the same amount of money no matter what the team's performance.

It's different in the English Football Leagues. Teams which fall to the bottom of the standings in the Premier League are relegated to the second division, the equivalent of baseball's minor leagues. At the same time, the best teams in the second division are promoted to the Premier League, giving them a chance to make much more money. There is a clear alignment between the interests of the fans and the owners.

The library industry has likewise been troubled by misalignment of interests between the owners of the companies and their customers. That's why it's important for libraries to pay close attention to the frequent mergers and acquisitions of the companies that serve them. These transactions are often announced just before an ALA meeting, and this past weekend's ALA Midwinter Meeting was no exception.

The big story of the weekend was the pending acquisition of Koha support vendor LibLime by PTFS (Progressive Technology Federal Systems, Inc.). (The acquisition is still in the due diligence phase and is expected to close in early February; terms were not disclosed.) The surprising part of the announcement was the sudden emergence of PTFS, which has had a very low profile in the library industry, into the top tier of integrated library system vendors.

Here I must digress to discuss a bit about business models in the library industry. Libraries have traditionally viewed their catalog system vendors as long term partners; the migration of data from one system to another is a major project, not lightly taken, and preferably not attempted more than once a decade. The choice of a new system touches almost all the library's processes, and thus involves many consultations and lengthy RFPs.

From the vendor's point of view, the sales process is very expensive. Promises to customize the system to address customer peculiarities are common, and these add to the cost of system maintenance. Once the system has been sold, a proprietary system vendor has a guarantee of continuing profits from support contracts. Only the vendor has the system knowledge (and sometimes even the system access) to make even the most trivial changes. It's in the support phase that the vendor and customer interests can become misaligned. The vendor has every incentive to do the least work at the highest price possible. The customer is locked into whatever system they have chosen.

Companies with strong cash flow have been attractive acquisition targets for private equity firms. Once acquired the company's new management focuses on eliminating expenses by cutting support staff and cleaning up the balance sheet by offloading liabilities such as unfinished development, thus making the company very profitable. The company can them be resold at a good mark-up. Customers often become very unhappy during the process. The company they "hired" during their system selection process transforms into something different.

The recent popularity of open source library management systems is in large part a search for business models that better align the interests of vendor and customer during the support phase. If the support vendor doesn't perform to the library's expectations, the library can hire a new support vendor without ditching their automation system. If a library wants to add a new feature to their system, or integrate it with a system from another vendor, they can hire a developer based on qualifications rather than access to source. The important thing to the library is not so much the access to source or the cost of the license, it's the absence of vendor lock-in.

The reason that PTFS is not widely known is that it specializes in an obscure segment of the market- it supports libraries predominantly in the government and the military. Founded in 1995, PTFS has been installing ILS systems, doing conversions and supporting systems in the unique security environment of government systems. John Yokley, a co-founder and the CEO of PTFS, spent 13 years as a Sirsi system administrator and programmer at the U.S. Courts, NASA, and University of Virginia Health Sciences Library has spent 20 years, 5 more than the age of PTFS, working in the library industry. Yokley himself worked in a government library for a short period in the early 90’s designing and building virtual library technology.. The company has experienced steady 20% per year growth and today has 120 employees. PTFS is particularly proud of their development of the US Government Printing Office's Federal Digital System (FDsys) which supports over a thousand libraries, but the company also has a library staffing component and a digitization facility.

Although PTFS has had a strategic partnership with SirsiDynix to market ArchivalWare, a digital content management system that grew out of technology developed for FDsys the Naval Research Laboratories (NRL) TORPEDO project, it found itself hamstrung in supporting its customers because of the lack of access to source code of the proprietary systems it was supporting. About 18 Months ago, PTFS decided that Koha was the Integrated Library System that it could most easily integrate with ArchivalWare, and it began to offer support for Koha. Koha is generally considered to be the first open source integrated library system; it was initially developed in New Zealand by Katipo Communications Ltd. and first deployed in January of 2000 for Horowhenua Library Trust.

LibLime (which is actually a trade name of Columbus, Ohio based Metavore, Inc.) was started in 2005 by Joshua Ferraro, Tina Berger and two others. LibLime has been the hardest-charging and fastest-growing proponent of the Koha Library System in the world. Over the intervening years, LibLime has acquired key Koha-related assets, including the US trademark, copyrights to Koha source code, and the Koha website. The combination of PTFS and LibLime will be supporting 640 installations of Koha under 123 contracts. The combined business will have Koha-related development contracts totaling $1.7 million. Despite the state of the economy, LibLime has actually had an increase in business over the past few months.

Recently, Ferraro and his co-principals at Metavore became very interested and excited by an opportunity outside of the library space. As the LibLime business grew, they recognized that they couldn't pursue both the new opportunity and LibLime, and they began to look for an acquisition partner. PTFS was the first company they went to. Given the reasons for the sale, only Ferraro among the Metavore principals will remain with LibLime; and he will stay only for 18 months to oversee the completion of planned development.

PTFS will keep the LibLime name and fold its own Koha support business into LibLime, which will be run by Patrick Jones. At the press conference held at ALA Midwinter in Boston, PTFS CEO John Yokley indicated that PTFS was committed to the concept of user-driven development and the open source concept, but also emphasized that he was still learning about open source and he was reviewing the LibLime business model; there is much left to be decided about how the LibLime business will move forward.

I spoke with Yokley afterwards. In his conversations with LibLime customers, he has found that their top priority for adopting Koha was to avoid vendor lock-in: their systems should be expandable by LibLime, the library or by another vendor. He sees Koha as a component of a fully capable integrated library system, and vowed that in two years, Koha will be fully capable of running a major academic library. The integration of Koha and ArchivalWare will be only the first phase. Although his team has discussed making ArchivalWare into an open source project, there are issues with third party components used which may prevent that from happening.

Yokley's clarity on avoiding vendor lock-in will be reassuring to customers, particularly with respect to LibLime Enterprise Koha (LLEK), a service announced by LibLime in September of 2009. LLEK is perhaps the most exciting asset being acquired by PTFS, and also the most controversial. The controversy deserves another article entirely, as it represents a break between LibLime and other developers supporting Koha. I plan to write that article in the coming week; please e-mail me if you wish to comment.

LLEK represents the evolution of LibLime's entry into cloud computing (also known as "software-as-a-service". Unlike vendors whose idea of cloud computing is simply to offer fully hosted services, LibLime's implementation of the cloud is more in line with that of modern "lean startups" who don't even own their own servers. By using Amazon EC2, LibLime has access to instantly expandable, low cost computing resources. LibLime is able to provision, configure, and implement a new Koha server in less than an hour. To accomplish this, LibLime has developed sophisticated deployment software (which it does not intend to release).

PTFS is already doing a sort of software-as-a-service, building private clouds for its military customers who don't have the option of going out on the open internet. As Yokley explained to me, "Economies of scale are an interesting thing. We've had a few large customers, but now with LibLime, we can provide services to large numbers of small libraries."

Welcome to the big leagues, PTFS!

This article is the first part of a series. Part 2 is here. Part 3 is here.

Reblog this post [with Zemanta]

Monday, January 18, 2010

Google Exposes Book Metadata Privates at ALA Forum

At the hospital, nudity is no big deal. Doctors and nurses see bodies all the time, including ones that look like yours, and ones that look a lot worse. You get a gown, but its coverage is more psychological than physical!

Today, Google made an unprecedented display of its book metadata private parts, but the audience was a group of metadata doctors and nurses, and believe me, they've seen MUCH worse. Kurt Groetsch, a Collections Specialist in the Google Books Project presented details of how Google processes book metadata from libraries, publishers, and others to the Association for Library Collections and Technical Services Forum during the American Library Association's Midwinter Meeting.

The Forum, entitled "Mix and Match: Mashups of Bibliographic Data", began with a presentation from OCLC's Renée Register, who described how book metadata gets created and flows though the supply chain. Her blob diagram conveyed the complexity of data flow, and she bemoaned the fact that library data was largely walled off from publisher data by incompatible formats and cataloging practice. OCLC is working to connect these data silos.

Next came friend-of-the-blog Karen Coyle, who's been a consultant (or "bibliographic informant") to the Open Library project. She described the violent collision of library metadata with internet database programmers. Coyle's role in the project is not to provide direction, but to help the programmers decode arcane library-only syntax such as "ill. (some col)". The one instance where she tried to provide direction turned out to be something of a mistake. She insisted that, to allow proper sorting, the incoming data stream should try to keep track of the end of leading articles in title strings. So for example, "The Hobbit" should be stored as "(The )Hobbit". This proved to be very cumbersome. Eventually the team tried to figure out when alphabetical sorting was really required, and the answer turned out to be "never".

Open Library does not use data records at all, instead, every piece of data is typed with a URI. This architecture aligns with W3C web standards for the semantic web, and allows much more flexible searching and data mining than would be possible with a MARC record.

Finally, Groetsch reported on Google's metadata processing. They have over 100 bibliographic data sources, including libraries, publishers, retailers and aggregators of review and jacket covers. The library data includes MARC records, anonymized circulation data and authority files. The publisher and retailer data is mostly ONIX formatted XML data. They have amassed over 800 million bibliographic records containing over a trillion fields of data.

Incoming records are parsed into simple data structures which looked similar to Open Library's, but without the URI-ness. These structures are than transformed in various ways for Googles use. The raw metadata structures are stored in an SQL-like database for easy querying.

Groetsch then talked about the nitty-gritty details of data. For example, the listing of an author on a MARC record can only be used as an "indication" of the authors name, because MARC gives weak indications of the contributor role. ONIX is much better in this respect. Similarly, "identifiers" such as ISBN, OCLC number, LCCN, and library barcode number are used as key strings but are only identity indicators with varying strengths. One ISBN with a chinese publisher prefix was found on records for over 24,000 different books; ISBN reuse is not at all uncommon. One librarian had mentioned to Groetsch that in her country, ISBNs are pasted onto a book to give it a greater appearance of legitimacy.

Echoing comments from Coyle, Groetsch spoke with pride of the progress the Google Books metadata team has made in capturing series and group data. Such information is typically recorded in mushy text fields with inconsistent syntax, even in records from the same library.

The most difficult problem faced by the Google Books team is garbage data. Last year, Google came under harsh criticism for the quality of its metadata, most notably from Geoffrey Nunberg. (I wrote an article about the controversy.) The most hilarious errors came from garbage records. For example, certain Onix records describing Gulliver's Travels carried an author description of the wrong Jonathan Swift. Most of these errors come from garbage records, and when one of these is found, almost always, the same problems can be found in other metadata sources. Google would like to find a way to get corrected records back into the library data ecosystem so that they don't have to fix them again, but that there have been issues with data licensing agreements that still need to be worked out. Article like Nunberg's have been quite helpful to the Google team. Every indication is that Google is in the metadata slog for the long term.

One questioner asked the panel what the library community should be doing to prevent "metadata trainwrecks" from happening in the future. Groetsch said without hesitation "Move away from MARC". There was nodding and murmuring in the audience (the librarian equivalent of an uproar). He elaborated that the worst parts of MARC records were the free text data, and normalization of data would be beneficial whereever possible.

One of the Google engineers working on record parsing, Leonid Taycher, added that the first thing he had had to learn about MARC records was that the "Machine Readable" part of the MARC acronym was a lie. (MARC stands for MAchine Readable Cataloging) The audience was amused.

The last question from the audience was about the future role of libraries in production of metadata. Given the resources being brought to bear on the book metadata by OCLC, Google and others, should libraries be doing cataloguing at all? Karen Coyle's answer was that libraries should concentrate their attention on the rare and unique material in their collections- without their work, these materials would continue to be almost completely invisible.
Reblog this post [with Zemanta]