Showing posts with label Fair use. Show all posts
Showing posts with label Fair use. Show all posts

Monday, November 18, 2013

Google Books and Black-Box Copyright Jurisprudence

Last week, eight years after the first lawsuit was filed to stop the Google Books Project, Judge Denny Chin finally ruled on the core merits of the case. The decision is being widely hailed on one side as a "tremendous victory for fair use" and on the other side as a "fundamental challenge to copyright". But these are short-term perspectives. I think that the long term impact of the decision may turn on the acceptance of Chin's approach to technology's transformation of copyright, which I would characterize as Black-Box Jurisprudence.

In my view, the core holdings about fair use were never in much doubt. The argument saying that indexing or lexical analysis or data-mining of books always requires the permission of a rights holder was never very defensible, or even seriously argued. A holding that display of snippets was not fair use would have made scholarly writing in the digital age impossible; a decision the other way on snippets would have been swimming up a judicial stream. But fair use is always a weighing of factors, and the untold story in the Google Books case is about the factors that didn't get weighed.

The reason that Google got sued in the first place was less about "what Google did" than about "how Google did it".  Google made huge numbers of copies of books without permission of the rights holders. Judge Chin's ruling said, effectively, that all those copies were incidental to the fair use.
[I]f there is no liability for copyright infringement on the libraries' part, there can be no liability on Google's part.
In the end, it didn't matter how Google did what it did. In Judge Chin's analysis, copyright is concerned only with the ends, not the means. Copyright seems not to be concerned with what happens inside the black box.

Chin is not alone in this approach. His opinion follow's Judge Baer's ruling in the Hathitrust case, which featured a ringing endorsement of the Library's fair use
I cannot imagine a definition of fair use that would not encompass the transformative uses made by Defendants' [Mass Digitization Project] and would require that I terminate this invaluable contribution to the progress of science and cultivation of the arts that at the same time effectuates the ideals espoused by the [Americans with Disabilities Act].
But for, me, the surprise in Baer's opinion was his transformation of the Arriba Soft case into a broad license for infringement. In that case, display of thumbnail images by a search engine was held to be fair use, and the copying of the images in the course of producing thumbnails was held to be necessary for the protected use. Judge Baer wrote that the fact that the images were on websites available for anyone anywhere to download was not relevant to the analysis, which he then applied to Google's scanning and OCR of physical books.
Although Plaintiffs assert that the decisions in Perfect 10 and Arriba Soft are distinguishable because in those cases the works were already available on the internet, Aug. 6, 2012 Tr. 19:2–4, I fail to see why that is a difference that makes a difference. As with Plaintiffs’ attempt to bar the availability of fair use as a defense at all, this argument relies heavily on the incorrect assumption that the scale of Defendants’ copying automatically renders it unlawful.
Baer thus reduces and equates Google's million dollar scanning operation with Arriba Soft's one line of code because they're in a fair use black box.

The Black Box approach to copyright can cut both ways. In Chin's dissenting opinion in the Aereo case, he wrote that it didn't matter that Aereo had engineered a way to use completely legal technical means to stream television signals over the internet.
In my view, by transmitting (or retransmitting) copyrighted programming to
the public without authorization, Aereo is engaging in copyright infringement in clear violation of the Copyright Act. [...] The system employs thousands of individual dime-sized antennas, but there is no technologically sound reason to use a multitude of tiny individual antennas rather than one central antenna; indeed, the system is a Rube Goldberg-like contrivance, over-engineered in an attempt to avoid the reach of the Copyright Act and to take advantage of a perceived loophole in the law.
In the Aereo case, Chin argued that since the end result of Aereo's engineering was a system with copyright infringing intent, the under-the-hood details of Aereo's system were not compelling. (Read James Grimmelmann for more on this case and copyright arbitrage in general.)

So when presented with cases where copyright law and technology collide, Chin has more or less adopted a consistent approach that isn't inherently pro-copyright or pro-fair-use.

If Chin's ruling had focused on the infringing means (i.e. massive copying) rather than on the fair-use ends in the Google Books case, Google could have gone back to the drawing board to devise a non-infringing means to accomplish the same ends. It would have been more expensive (à la Aereo), but the plain fact is that ten engineers can run technical circles around a thousand lawyers. In the end, Google would have lost the battle but would be far ahead in the war.

As the case now stands, while Google has a free hand to go back and improve and expand its scanning operations, it is still constrained in what it can deliver. For example, since Chin's decision cites the lack of advertising on snippet result pages in his fair use analysis, Google can't put advertising there without risking another $100 million lawsuit. Another innovator in the space can't go and do things differently without worrying about another judge's fair-use analysis.

The advantages of a black box legal approach is its practicality. Judges don't have to understand the intricacies of technology in order to decide legal questions. Technical processes are opaque for business reasons, too. But perhaps more importantly, a black-box approach to copyright law means that engineers can't use clever hacks to get around copyright.

The danger of the black box is that it pretends that technology doesn't matter, that code isn't law. Copyright law is rooted in technology, that of the printing press, and turning it into an abstraction that can also govern digital media wile ignoring what goes on behind the curtain is a dubious project. A complex enterprise like Google Books is a long journey from inception to delivery. Imagine if highway safety was addressed by regulating total travel times. Does it make sense to regulate a new technology like airplane travel in the same way?

Perhaps there ought to be a fifth factor in fair use analyses of systems more complex than a printing press. In addition to the usual four factors, Judges could also be weighing whether the steps involved in accomplishing a fair use would stand under their own 4 factor analysis. In the Google Books case, the analysis could have incorporated a weighing of the scanning operation by itself. Similarly, Aereo's meticulous adherence to legal means could weigh in favor of a fair-use determination.

My worry is that in other situations, perhaps with technologies we haven't imagined yet, the black box legal approach will end up with very wrong technical results. And then we'll be stuck, waiting for Congress to fix things. Look at what's happening as digital surveillance collides with crypto-security. There, the courts have uniformly refused to look inside the black box of the NSA, and the results may end up being disastrous.

(Gary Price has a thorough opinion round-up at Infodocket.)

Enhanced by Zemanta

Thursday, February 4, 2010

Copyright-Safe Full-Text Indexing of Books

As the February 18 hearing on the revised Google Books Settlement Agreement draws near, I think its timely to explore some issues surrounding full-text indexing of books. It's important to realize that when Google began its program of scanning books in libraries, it chose to do so in a way that entered the gray zone of fair use. Google continues to maintain that its scanning activities are perfectly legal, and fair use advocates welcomed the Publishers' and Authors' lawsuit because it had the potential to clarify ambiguities around fair use. No matter where the court decided to draw the line, the both fair use and rightsholder control would be able to extend into the zone of current uncertainty.

Overlooked in the controversy is the fact that Google could have chosen a safer course in its effort to make full-text indices of books. In this article, I'll argue that it's possible to make full-text indices of books in a way that steers well clear of copyright infringement. But first, I should note that playing it safe would not have been a good plan for Google. By pushing fair use to its limits, Google assured itself a favorable competitive position. In a lawsuit, Google could have lost on 90% of the fair use they were claiming and would still have ended up 10% ahead of where a safe course would have taken them. Google is large enough that even a 10% victory in court would have paid off in the long run. As it is, Google chose to settle the lawsuit under terms that put them in a better position than they would have occupied by playing it safe, and potential competitors don't gain the benefits of a fair-use precedent.

I make two assumptions about copyright in devising an copyright-safe indexing method:
  1. You can't infringe the copyright to a work if you don't copy the work.
  2. If you can't reconstruct a work from its index, then distributing copies of the index doesn't infringe on the work's copyright.
Just in case these assumptions are weak, my fall-back position is that indexing is clearly a fair use under US copyright law.

First, the fall-back assumption: full-text indexing is allowed as fair use under US copyright law. Indices are allowed as "transformative uses". Judge Robert Patterson's decision (pdf, 195K) in the "Harry Potter Lexicon" case gives an excellent background of this jurisprudence and concludes:
The purpose of the Lexicon’s use of the Harry Potter series is transformative. Presumably, Rowling created the Harry Potter series for the expressive purpose of telling an entertaining and thought provoking story centered on the character Harry Potter and set in a magical world. The Lexicon, on the other hand, uses material from the series for the practical purpose of making information about the intricate world of Harry Potter readily accessible to readers in a reference guide. To fulfill this function, the Lexicon identifies more than 2,400 elements from the Harry Potter world, extracts and synthesizes fictional facts related to each element from all seven novels, and presents that information in a format that allows readers to access it quickly as they make their way through the series. Because it serves these reference purposes, rather than the entertainment or aesthetic purposes of the original works, the Lexicon’s use is transformative and does not supplant the objects of the Harry Potter works.
The author of the Lexicon lost his case not because his indexing was not allowed, but rather because he copied too much of J. K. Rowling's creative expression in doing so.

Second, you have to copy to infringe copyright. A more accurate statement is this: You have to either make a copy or a derivative work to infringe copyright. The second piece of this can be a bit more confusing, because "derivative work" has a specific meaning in copyright law. A translation into another language is an example of a derivative work. Indices are not derivative works. The law considers indices to be more akin to metadata. I might need access to a book to count the number of figures it contains, but a report of the number of figures in a book and what page they're on is in no way a derivative work. The copyright act defines a derivative work as
a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted.
If you make copies by scanning, however, as Google is doing, you must also establish that your use is allowed as fair use. If you don't, then you don't even need to reach the fair use provision.

The last assumption gets more technical. The simplest form of a word index is a sorted list of words with pointers to the occurrence of the word within the text. So an index of that last sentence might look like this:
a    5,9
form    3
index    7
is    8
list    11
occurrence    18
of    4,12,19
pointers    15
simplest    2
sorted    10
text    24
the    1,17,20,23
to    16
with    14
within    22
word    6,21
words    13
It doesn't take a computer science degree to see that it's easy to reconstruct the sentence from this index. For that reason this form of index is equivalent to a copy. If you remove the position pointers, however, the index loses enough information that the sentence cannot be reconstructed. So if we take the words on a page of text and sort the words in each sentence, then sort the word-sorted sentences, we get an index of a page that can't be used to reconstruct text, but can be used to build a useful full-text index of a book.

The trickiest step of completely copyright-safe indexing is producing the page index from a book without producing intermediate copies of the pages. In a conventional scanning process, a digital image of a page is stored to disk and the copy is passed to OCR software. Indexing software then works on the OCR text. A scanning process that was fastidious about copyright, however, could scan lines of text word by word and never acquire an image large enough to be subject to copyright.

US courts have considered the loading of a copyrightable work into a computer's RAM storage to constitute copying, but scanning sufficient to produce an index can in principle be done without requiring that to occur. (For an excellent law review article on the RAM-copying situation, read Jonathan Band and Jeny Marcinko's article in Stanford Technology Law Review.) Also, even sentences of more than a few words can be considered copyrightable works, as I discussed in an article from November.

Another possible way to avoid copying is to build a black-box indexer. A closer look at the RAM-copying precedent, MAI SYSTEMS v. PEAK COMPUTER suggests that a non-copying scanning indexer can be built even if page images exist somewhere in RAM. In that case, the court reasoned that the software copy could be viewed via terminal readouts, system logs, and that sort of thing. If a closed-box indexing system were built so that page images resident in RAM could never be "perceived, reproduced, or otherwise communicated", then there is a fair chance that a court would find that copying was not occurring.

I'm a technologist, not a lawyer. I would welcome comment and criticism from experts of all stripes on this analysis. For example, I've not considered international aspects at all. There are many technical aspects of copyright-safe indexing that would need to be sorted out, but doing so could open the way to countless transformative uses of all the books in the world.
Enhanced by Zemanta