Showing posts with label J. K. Rowling. Show all posts
Showing posts with label J. K. Rowling. Show all posts

Wednesday, April 3, 2019

Fudge, and open access ebook download statistics

If you found out that the top 50 authors born in Gloucestershire, England average over 10 million copies sold, you might think that those authors are doing pretty well. But it's silly to compute averages like that. When you compute an average over a population, you're making an assumption that the quantity you're averaging over is statistically distributed somehow over the population. Unless of course you don't care if the average means anything, and you just want numbers to help justify an agenda.

Most folks would look at the list of Gloucestershire authors and say that one of the authors is an outlier, not representative of Gloucestershire authors in general. And so J.K. Rowling, with more than 500 million copies sold, would get removed from the data set, revealing the presumably unimpressive book selling record of the "more representative" authors. Scientists refer to this process as "fudging the data". It's done all the time, but it's not honest.

There's a better way. If a scientific study presents averages across a population, it should also report statistical measures such as variance and standard deviation, so the audience can judge how meaningful the reported averages are (or aren't!).

Other times, the existence of "outliers" is evidence that the numbers are better measured and compared on a different scale. Often, that's a logarithmic scale. For example, noise is measured on a logarithmic scale, in units of decibels. An ambulance siren has a million times the noise power of normal conversation, but it's easier to make sense of that number if we compare the 60 dB sound volume of conversation to the 90 dB of a hair dryer, the 120 dB of the siren and the 140 dB of a jet engine. Similarly, we can understand that while J.K. Rowling's sales run into 8 figures, most top Gloucestershire-born authors are probably 3, 4 and or maybe 5 figure sellers.

Over the weekend, I released a "preprint" on Humanities Commons, describing my analysis of open-access ebook usage data. I worked with a wonderful team including two open-access publishers, University of Michigan Press and Open Book Publishers, on this project, which was funded by the Mellon Foundation. To boil down my analysis to two pithy points, the preprint argues:

  1. Free ebook downloads are best measured on a logarithmic scale, like earthquakes and trade publishing sales.
  2. We shouldn't average download counts.

If you take the logarithm of book downloads, the histogram looks like a bell curve!
For example, if someone tries to tell you that "Engineering, mathematics and computer science OA books perform much better than the average number of downloads for OA books across all subject areas" without telling you about variances of the distributions and refusing to release their data, you should pay them no mind.

Next week, I'll have a post about why logarithmic scales makes sense for measuring open-access usage, and maybe another about how log-normal statistics could save civilization.

Tuesday, June 25, 2013

Magic Rights Management for eBooks

The Fraunhaufer Institute in Germany is apparently marketing some "new" technology they're calling SiDiM which embeds digital data into texts by changing words in the text. They're telling publishers that it will fight piracy by making it easy to track files uploaded to torrents and file lockers back to their reprehensible sources.

It sounds kinda dumb, doesn't it?

The internet has over-reacted of course. Perhaps everyone is hypersensitive because of the revelations about the NSA and its data collection practices. Nick Harkaway jumped the shark a bit and called it "surveillance". The idea of changing words in books is easy to ridicule and deserves to die, but let's please take a deep breath.

There are lots of ways to put information in an ebook file, and license information is no different. For example, I've been advocating that Creative Commons licensed books should embed a digitally signed license so that the license can be relied upon. When you buy an ebook, an embedded license could protect you from accusations of infringement. Digital signatures can also tell you that a books content hasn't been tampered with.

When you buy a Harry Potter ebook direct from Pottermore, your identifying information gets digitally stamped into the ebook. According to the The Digital Reader, Pottermore uses watermarking technology from Booxtream, and I've been evaluating this technology myself for an Unglue.it project. So far, I'm impressed.

The rationale behind Pottermore's watermarking is that it prevents people from sharing the book beyond what their license allows. If the book gets on a public filesharing site, it can be traced back to the purchaser, and consequences could ensue.

Booxtream claims to be using 9 different watermarking techniques to make the embedded data hard to remove. For example, Booxtream adds digital codes to the names of the content files inside the EPUB, and adds data into image files. Although it's straightforward to strip some of the embedded info, Booxtream needs only to make it uncertain that stripping has been complete to retain some deterrence value.

For the user, the bottom line is that nothing the purchaser does or would want to do is impeded by the Booxtream watermarking. Nothing visible to the user is altered except for an ex libris page that tells the user that the copy has been personally licensed to him or her- it's customizable by the vendor.

A close analogy to ebook watermarking is the bullet serialization that's been proposed as an alternative to gun control. If every bullet was traceable to a purchaser, investigation of weapons related crime would be reduced to finding the bullet and looking it up in a database. Law abiding gun owners shouldn't notice the difference. Or maybe it would be de facto ban on ammunition. YMMV.

The argument against digital watermarking is that there will always be ways to remove the embedded data, no matter how clever you are at hiding it. Someone will make a one-click watermark stripper, and the value in watermerking will be diluted. But almost two years after Pottermore launched their digitally watermarked ebooks, it's quite hard to find watermark stripping tools. Why would anyone bother? There's nothing that the vast majority of ebook purchasers want to do that's impeded by the watermarking. Contrast that with the ease of finding tools to strip PDF watermarks, which are annoying.

You might wonder why SiDiM would be selling their technology with such a scary-dumb sounding marketing pitch. It's because publishers are the customers. I've been talking to a lot of publishers, and they're very clear that they want DRM. Or at least they THINK they want DRM. What they really want is magic. The want their ebooks to come with a magic bullet that stops piracy and over-sharing dead in its tracks. They don't understand the technology behind DRM, but many of them swallow the story that comes with it- that nobody would pay for digital files if they can get them for free from piracy sites.

The truth is that if there's magic in the kind of DRM that comes with Adobe, Apple and Kindle, it's of the variety that Voldemort would use. If there's magic in the watermarking techniques used by Pottermore, it's of the Dumbledore variety. If there's magic in SiDiM, it's like Neville Longbottom's Switching Spell that put ears on a cactus.

I'm here to tell you that magic is real. There's real magic in the stories that authors tell. There's real magic in communities and in relationships between people, between authors and readers. There's real magic in libraries. It's that real magic that will stop piracy and help authors earn a good living in the digital future.

Dumbledore's fictional magic can help make the real magic manifest, and that's what we should work towards
Enhanced by Zemanta

Sunday, December 16, 2012

Einstein's Never-Ending Copyright


Photo by Miroslav Duchacek CC-BY-SA-3.0 
In my post on Quantum Copyright, I promised, in my following post, to cover the impact of Special Relativity on Copyright. I was joking. I had no intention of putting words in Einstein's mouth about our copyright laws. How silly would that be?

Have you ever tried NOT THINKING ABOUT GIRAFFES? It's just hopeless. So here you go:

In special relativity, the passage of time depends on your frame of reference. Time is relative, and simultaneity of events can't be defined except relative to their respective reference frames.

So suppose I take a book with me on a spaceship that moves at 99.99% the speed of light relative to your reference frame. Then every day that elapses for me is about 71 days for you. In two years or so, the book goes out of copyright, and the next planet I visit, I can make copies for every sentient being I can find.

Seems a lot of trouble when I can just put it on BitTorrent.

Ah, but imagine that I'm a world-famous trillionaire author, and I'm worried about the day when my best-selling novel goes out of copyright, and everyone can just rip me off? All I have to do is buy myself a spaceship and go for a vacation. Since my copyright won't expire till 70 years after my death, my hypervelocity excursion will dilate my copyright term for a long, long time. When I get back a year from now (in my reference frame), 71 years will have elapsed on earth, and with the royalties I'll have earned (plus interest) I can probably acquire every other book on the planet. And both houses of Congress. I won't have aged much, so I'll just go on another interstellar jaunt. Rinse and repeat.

Start saving up, Jo Rowling.

For the rest of us, the bright side of this is that we can be pretty sure that copyright law will get be updated at least before interstellar drives are perfected.

Wednesday, November 11, 2009

The Uniqueness of Sentences and J. K. Rowling's (Non)Infringement of Tanya Tucker

Have you ever heard someone say something unusual and wonder to yourself if anyone in the history of humanity had ever said that before, ever? It happens a lot more than you might think.

In the discussion of my article on copyright salami, I suggested that copyright based on content as short as a sentence would not be very robust. I had reasoned that if the sentences were short enough, the would be a high probability that the same sentence had already appeared in a copyrighted work, or even in a work that was in the public domain. I imagined building huge databases of sentences that had already been used so as to clear them for reuse.

I decided to do some testing first. I chose a page at random (p. 447) from my (print) copy of J. K. Rowling's Harry Potter and the Deathly Hallows. I extracted the sentences, and put each sentence into Google and into Google Book Search. The results surprised me.

My first test sentence was
"Get - off - her!" Ron shouted.
With only 5 words, none of them uncommon, I expected to get a a few close matches. The book search produced zero hits, and no results at all close. The general Google search was more interesting. Of the 7 hits, all of them exact matches, the top two of seven hits appear to be properly attributed fair use quotations from the book. Two other hits were to complete, unauthorized copies of the book. One of these, on SlideShare, offers this disclaimer:
"hey here i got this book in pdf format .. am i violating anything after .. uploading this stuff over here ... just let me know .. if any issue come in existence, will remove it
Although the item has had 34,000 views, it pdf itself appears to have been removed from SlideShare. The pdf posted by a Filipino web designer on his web site, though, is still available (and has been since August) and is of quite good quality.

The oddest hits are to a site which masquerades as a "game ranking" portal site.
RPGRank is a real-time online game ranking system which provide a best MMORPG ranking portal for both players and games of all genre with the exclusive news, press release, review, preview, interview, trailer and vedio. RPGRank strive to provide all gamers things that they never experienced before by newest game beta keys, live-event, and online tournamentsa with attractive giveaways from games.
It appears that this site generates pages of random text for the benefit of search engines by extracting sentences from books and feeding the sentences to Google in a random order. This site has convinced Google to index "about 318,000" pages of its meaningless "content", and offers to sell "background" advertising space on the site at $1200 per month.

The last hit appears to be to a site which is presenting a Vietnamese translation of the book alongside the complete English text. Although I can't read Vietnamese, I doubt very much that it is authorized use. Vietnam joined the Berne convention only 5 years ago, so this is certainly an illegal infringement.

Of the 26 sentences on page 447, I could find only three that had been used in places that Google knows about. The first, "Leave him alone, leave him alone!" is a line from a Tanya Tucker song. The second, "Harry's stomach turned over.", has been used in James Edward Amesbury's "bloody but weakly conceived thriller", A Sporting Chance and in D. Edwards Bradley's Harry's War.

The third,"Harry did not answer immediately." is firmly in the public domain, having done duty as a complete sentence in Smith Hempstone's A Tract of Time, as a fragment in Frances Elizabeth G. Carey-Brock's 1867 My father's Hand: and Other Stories, and in Adam Williams' 2007 gripping adventure of modern China, The Dragon's Tail.


Three sentences comprising bits of dialog: "Been Stung", "And your first name?", and "Vernon Dudley", turned up numerous matches to fragments of sentences in Google. It was also amusing to see matches for the sentence "What happened to you, ugly?" This phrase matched two people-search sites which specialize in feeding Google pages with text like "What happened to Joe Smith?" Apparently there is someone who uses the screen name "you_ugly", and the people search engines just leapt to the wrong conclusions!

Most of the sentences on page 447 appear to be purely original to J. K. Rowling. Was she lucky, or were the odds stacked in her favor? Word frequencies for English have been measured, so we can easily generate a simplistic estimate of sentence occurrence rate. Ignoring the proper name "Ron", the words "Get", "off", "her" and "shout" have occurrence frequencies of 0.22%, 0.046%, 0.22%, and 0.0055%, respectively. Multiplying these occurrence rates gives us a weighted occurrence probability of this combination of 1 per 8 trillion. If you had the entire population of earth speaking random four-word English sentences they might come up with this combination in a day or two. Add "Ron" into the mix, and they might take the greater part of a year to generate the sentence J. K. Rowling wrote.

For context, it's interesting to guess at the total number of sentences that humanity has written or spoken. It's estimated that 100 billion humans have lived so far. If those humans spent 16 hours a day for an average of 65 years generating 3 sentences per minute, we'd be up to about 20 million trillion sentences. The real number is probably a factor of 100 to a thousand less (half of us are men, after all!). This estimate roughly agrees with estimates of others that all the words ever spoken could be archived using 10 exabytes of storage.

Ten exabytes is not as much storage as it used to be. The Internet Archive currently has 0.003 exabytes; although Google is quite secretive about its hardware deployment, it seems likely that their current storage capacity is in excess of 10 exabytes. Yesterday, Google announced a pricing plan where they'll rent you 0.000016 exabytes for $4096 per year. I'll do the math for you. If you want to store everything anyone has ever said, Google will rent you the space for only $2.5 billion dollars per year!

Given that Google will soon have digitized a large fraction of the world's books, there are a few things we can learn from this exercise.
  • It will soon be very easy for Google to detect unauthorized copies of books in its index, and presumably to remove them. The benefit to publishers of doing this would hugely outweigh any damages they're suffering from the Google Books digitization program. Why have publishers overlooked getting this to happen as part of the agreement settling their lawsuit?
  • It will not be difficult for Google to accurately de-duplicate the Google Books index.
  • J.K. Rowling's hesitancy to release her books in ebook format is really, really stupid.
Before you get distracted with something useful, do this: pick about 5 random words, make a sentence from them, and become the first human ever to say that sentence. Depending on what you do next, you may also be the last!