Saturday, March 1, 2014

The DMCA Takedown of a Feynman Lectures eBook Converter


The Feynman Lectures on Physics was one of my favorite textbooks in college. It wasn't the assigned textbook, it was recommended reading. I think the reason it doesn't work as a textbook is that every chapter is so deep that students would get sucked so far into every topic that they would never finish the course. It's the sort of book that transforms your life and way of thinking about the physical world. When I started Unglue.it, The Feynman Lectures was one of the first books I investigated for ungluing.

My friends at Caltech informed me that the rights situation with the Feynman Lectures was exceedingly complicated, and it would be a cold day in hell before the Feynman Lectures would be free to the world in digital form. It seems that Caltech and the book publishing world had made an awful hash of the rights, with print rights being owned by Pearson, and the audiovisual rights being owned by competing publisher Perseus. Heroic efforts by Caltech lawyer Adam Cochrane and some dedicated physicists and educators resulted in the untangling of rights, leading to a revised edition available through Perseus imprint Basic Books.

And last year, a miracle happened. An authorized free digital version of the lectures appeared on the web! There is sanity in the world! The Feynman Lectures had been unglued!

Vikram Verma, a software developer in Singapore, wanted to be able to read the lectures on his kindle. Although PDF versions can be purchased at $40 per volume, no versions are yet available in Kindle or EPUB formats. Since the digital format used by kindle is just a simplified version of html, the transformation of web pages to an ebook file is purely mechanical. So Verma proceeded to write a script to do the mechanical transformation – he accomplished the transformation in only 136 lines of ruby code, and published the script as a repository on Github.

Despite the fact that nothing remotely belonging to Perseus or Caltech had been published in Verma's repository, it seems that Perseus and/or Caltech was not happy that people could use Verma's code to easily make ebook files from the website. So they hauled out the favorite weapon of copyright trolls everywhere: a DMCA takedown.

I am not a lawyer, but I think that this use of a DMCA takedown was improper and possibly illegal. I'm pretty certain that use of Verma's script for personal use would be protected fair use in the United States, under Betamax. There are no terms of use at the Feynman Lectures website for Verma's script to violate; there wasn't even a robots exclusion. So even a legal theory that Verma's code was inducing others to violate website terms falls flat on its face.  But alas, there's no penalty for abusive DMCA takedowns, so Perseus' main downside is having to read annoying blog posts like this one. And Perseus does need to look out for their authors' rights – they probably aren't in a position to asses what some ruby code does.

Luckily, Github has a policy of publishing every DMCA takedown notice it receives, which is how I found out about Perseus' action, and Verma's counternotice. Perseus had 10 days to respond to the counter-notice and since they failed to do so, Github has re-opened the repository.

In the meantime, the Feynman Lectures website has taken some steps to break Verma's script. For example, instead of a link to http://www.feynmanlectures.caltech.edu/II_28.html (my favorite chapter), the table of contents now has a link to javascript:Goto(2,18). This will take about 10 minutes for Verma to work around. In addition, the website now has a robot exclusion (except for Googlebot).

Michael Gottlieb, the editor of The Feynman Lectures on Physics New Millennium Edition added this issue to the repo:
The online edition of The Feynman Lectures Website posted at www.feynmanlectures.caltech.edu and www.feynmanlectures.info is free-to-read online. However, it is under copyright. The copyright notice can be found on every page: it is in the footer that your script strips out! The online edition of FLP can not be downloaded, copied or transferred for any purpose (other than reading online) without the written consent of the copyright holders (The California Institute of Technology, Michael A. Gottlieb, and Rudolf Pfeiffer), or their licensees (Basic Books). Every one of you is violating my copyright by running the flp.mobi script. Furthermore Github is committing contributory infringement by hosting your activities on their website. A lot of hard work and money and time went into making the online edition of FLP. It is a gift to the world - one that I personally put a great deal of effort into, and I feel you are abusing it. We posted it to benefit the many bright young people around the world who previously had no access to FLP for economic or other reasons. It isn't there to provide a source of personal copies for a bunch of programmers who can easily afford to buy the books and ebooks!! Let me tell you something: Rudi Pfeiffer and I, who have worked on FLP as unpaid volunteers for about a decade, make no money from the sale of the printed books. We earn something only on the electronic editions (though, of course, not the HTML edition you are raping, to which we give anyone access for free!), and we are planning to make MOBI editions of FLP - we are working on one right now. By publishing the flp.mobi script you are essentially taking bread out of my mouth and Rudi's, a retired guy, and a schoolteacher. Proud of yourselves? That's all I have to say personally. Github has received DMCA takedown notices and if this script doesn't come down pretty soon they (and very possibly you) might be hearing from some lawyers. As of Monday, this matter is in the hands of Perseus's Domestic Rights Department and Caltech's Office of The General Counsel. 
Michael A. Gottlieb
Editor, The Feynman Lectures on Physics New Millennium Edition
www.feynmanlectures.info
www.feynmanlectures.caltech.edu

(Note: Gottlieb's description of the website copyright notice is inaccurate- it says nothing about "downloaded, copied or transferred for any purpose")

This is kind of sad. Here Caltech did the right and noble thing and made the Feynman Lectures free as a website. That they can make money from the work via sales of print and other versions is great. But having done that, trying to control what people do with the free digital version (other than sell it) is a hopeless endeavor, and they should just stop.

I was wrong. The Feynman Lectures hasn't been unglued.

Update, March 3: Verma made a one-line change to the script to un-break it. But it's not a polite script, so don't all go and run it. Better to ask Caltech to use the script to make epubs and mobi's for sale; I would certainly pay for my DRM-free copy!

Update, March 4: Gottlieb e-mailed me to say that Perseus didn't respond to the counter-notice because Github's email notice went to a spam filter, and that more takedowns would be coming. He seemed to think that I am one of the flp.mobi developers and warned that I have put myself "in a precarious legal position". To me clear, I am not involved in the development or publication of flp.mobi. I hope its existence is not used as a pretext to take down or lock down the FLP website. Also, high-quality epub and mobi are on the way!

Update, March 7: Verma e-mailed me to say he is voluntarily taking down his repo:
I'm taking down my copy of the repository on Monday morning, in worry its continued availability will lead Caltech to discontinue free online access to FLP. You're each welcome to adopt maintainership if you prefer, though I would rather if you did not.
Techdirt has a post and commentary.

Update, March 10: Verma's repo is now history, but forks of it remain in 15 places, including, bizarrely, Gottlieb's own Github page. 
Enhanced by Zemanta

Friday, February 28, 2014

Open Access Honesty

I've spent a large part of February becoming acquainted with Open Access ebook publishers. And the one thing that troubles me is that too many of them are not putting honesty first. Because existing distribution channels do not reward forthrightness in Open Access publishers; in fact the channels actively discourage it.

Let's take Amazon, for example. They don't like free ebooks, because there's no money in it for them. If you're a publisher and you want your ebook to be free for people to load onto their kindles, Amazon will charge you for the privilege. They rationalize that they're paying for a separate wireless network, "Whispernet", so it's only fair to assess "delivery charges" to free  ebook publishers. If you use their 70% royalty option, the delivery charge is 15 cents per MB of data, and the minimum price you can set is 99 cents. The only way to get Amazon to deliver your ebook for free is to select their 35% royalty option, and then invoke this "matching Competitor Pricing" clause:
From time to time your book may be made available through other sales channels as part of a free promotion. It is important that Digital Books made available through the Program have promotions that are on par with free promotions of the same book in another sales channel. Therefore, if your Digital Book is available through another sales channel for free, we may also make it available for free. If we match a free promotion of your Digital Book somewhere else, your Royalty during that promotion will be zero. (Unlike under the 70% Royalty Option, if we match a price for your Digital Book that is above zero, it won't change the calculation of your Royalties indicated in C. above.)
Apple, Kobo, and Google are much happier to set prices to zero, because they make some money on hardware sales or advertising, so you can get Amazon to give your ebook away for free by getting people to report your zero price on Apple, Kobo, and Google.

So just to get your ebook to be free on Kindle, you're forced to be incompletely honest with your customers and distributers.

But Amazon creates a great temptation. Why not use the suckers paying on Kindle to subsidize the free availability for those smart users who come to  your website? Isn't it convenience that these people are happily paying for?

And libraries are another temptation. They'll pay for the convenience of getting your ebook though their preferred platform, Overdrive or whatever, even as you offer the book for free to users at all the libraries that don't pay for your ebook. But would they still buy if they knew they could get the ebook for free? Maybe you shouldn't ask questions when you don't want to know the answer.

So here's my simple, unproven postulate: in the long run, full disclosure about pricing and an honest relationship with readers will be in the best, mutual interests of authors, publishers, readers, and libraries. And customers will prefer a distribution channel that enables that honesty.

Stop laughing.

Saturday, February 1, 2014

Crowd-Frauding: Why the Internet is Fake

Power in human societies derives from the ability to get people to act together. Armies, religions, governments, and businesses have dominated societies using weapons, beliefs, laws and money to exert collective effort. In modern societies, mass media have emerged as a similar organizing power.

A new kind of collective organization, mediated by the internet, code and connections, is emerging as another avenue of power. It's no longer ridiculous to think that social networks, crowd-sourcing and crowd-funding could achieve the social consensus, action and compulsion that were the province of governments, armies and religions. As the founder of a crowd funding site for ebooks, I'm naturally optimistic that the crowd, connected by the social internet, will be an immensely powerful force for good.

I'm also continually reminded that the bad guys will use the crowd, too. And it won't be pretty.

In December, our site, Unglue.it, began to get a surge of new users. But there was something fishy. The registrations were all from hotmail, outlook, and various dodgy-sounding email hosts. The names being registered were things like Linette, Ophelia, Rhys, Deanne, Agueda, Harvey, Darcy, Eleanore and Margene. Nothing against the Harveys of the world, but those didn't look like user handles. Of course it was registration bots coming from many different IP addresses.

But they were stupid registration bots- they never complete the registration, so the fake accounts can't leave comments or anything. It wasn't causing us any harm except it was inflating our user numbers. It was mystifying. So I started studying why bots around the world might start making inactive accounts on our site.

And that's how I learned about the dark side of the crowd-force. The best example of this is a program called Jingling, also known as FlowSpirit. It's been around for 5 years or so.

Jingling is an example of a "cooperative traffic generation" tool. It's software-organized crime. Crowd-frauding, if you will.

It works like magic. You download the Jingling software and install it on your computer. You then enter the URLs for four webpages that you want to promote. (or more, if you have a good internet connection.)  Although the user interface is in Chinese, you can get annotated instructions in English on YouTube or websites like theBot.net. Once you've activated Jingling, the webpages you want to promote start getting hundreds of visitors from around the world. The visitors look real, they click around your page, they click on the advertisements, they register accounts on websites, they click "like" buttons and follow you on Twitter.

Meanwhile, your computer starts running a website-visiting, ad-clicking daemon. It visits websites specified by other Jingling users. It clicks ads, registers on sites, watches videos, makes spam comments and plays games. In short, your computer has become part of a botnet. You get paid for your participation with web traffic. What you thought was something innocuous to increase your Alexa- ranking has turned you into a foot-soldier in a software-organized crime syndicate. If you forgot to run it in a sandbox, you might be running other programs as well. And who knows what else.

The thing that makes cooperative traffic generation so difficult to detect is that the advertising is really being advertised. The only problem for advertisers is that they're paying to be advertised to robots, and robots do everything except buy stuff. The internet ad networks work hard to battle this sort of click fraud, but they have incentives to do a middling job of it. Ad networks get a cut of those ad dollars, after all.

The crowd wants to make money and organizes via the internet to shake down the merchants who think they're sponsoring content. Turns out, content isn't king, content is cattle.

Jingling is by no means alone; there are all sorts of bots you can acquire for "free". Diversity of bots is enforced because the click fraud countermeasures only attack the most popular bots; new bots are being constantly developed and improved.

What does this mean for advertising, ad-supported websites, and the internet in general?

It means that the internet rich will get richer and power will concentrate. Let me explain.

I used to run my own mail server. It was a small process on one of my old machines. I was a small independent internet entity. The NSA couldn't scan my emails and I could control my mail service. But as spammers cranked up their assault, it became more and more complicated to run a mail server. At first, I could block some bad ip addresses. But when dictionary attacks could be run by script kiddies, running my own email server got to be a real drag. And because I was nobody, other people running mail servers started blocking the email I tried to send. So I gave up and now I let big brother Google run my email. And Google gets to decide whether email reaches me or gets blocked by spam. That doesn't make me happy, and I still get a fair amount of spam. (Somehow the SEO and traffic generation scammers still get through!)

It's probably the same way that kingdoms and countries arose. Farmers farmed and hunters hunted until some bad guys started making trouble. People accepted these kings and armies because it was too much trouble for farmers and hunters to deal with the bad guys on their own. But sooner or later the bad guys figured out how to be the kings. Power concentrated, the rich got richer.

So with the crowd-frauders attacking advertising, the small advertiser will shy away from most publishers except for the least evil ones- Google or maybe Facebook. Ad networks will become less and less efficient because of the expense of dealing with click-fraud. The rest of the the internet will become fake as collateral damage. Do you think you know how many users you have? Think again, because half of them are already robots, soon it will be 90%. Do you think you know how much visitors you have? Sorry, 60% of it is already robots.

Sooner or later the VCs will catch on to the fact that companies they've funded are chasing after bot-clicks and bot views. They'll start demanding real business models; those of us older than forty may remember those from college. And maybe reality will have a renaissance. But more likely, the absolute power of Google, Amazon, Apple and the rest will corrupt them absolutely and we'll suffer through internet centuries of dark ages (5 solar years at least) before the arrival of an internet enlightenment.

Until then, let's not give in to the dark side of the force, OK?

Notes:
  1. Last year, I wrote about some strange Twitter bots. I now think it's likely that the encoded messages I saw are part of a cooperative traffic generation scheme. If you're trying to orchestrate a vast network of click-bots, what better way to communicate with them than twitter?
  2. There are now disposable email hosts that will autoclick confirmation links. These email hosts are the registration spammer's best friends. A list is here.

Friday, January 17, 2014

EPUB has a Steep Road Ahead (Notes from Open Book Hack 2014)

People I talk to about ebook technology belong to one of two camps.
  1. EPUB3 is the future of ebooks. 
  2. EPUB3 has lost the ebook format war because no one is supporting it. 

Here I'm helping Fortitude fix a bug.
(photo by Ray Schwartz)
So it was very enlightening to join a group of developers at New York Public Library last weekend at Open Book Hack 2014. Sponsored by NYPL Labs and the Readium Foundation, the event was convened to examine the challenges of the "Open Book" on the Web. Since I had written that another publishing hackathon "pretty much ignored ebooks", I felt compelled to attend, and write a report.

Open Book Hack was the first event I've been to where developers were actually working on EPUB, the format that is emerging as the underpinning of the digital book industry. This was both encouraging and discouraging. Encouraging because you could start to glimpse possibilities that EPUB will enable, and discouraging because the road ahead looks so steep.

Open Book Hack attracted a very stimulating mix of developers from around the world (who were in New York for Digital Book World) and local developers looking for interesting problems. Notably, there was a very impressive contingent of students from New York's Flatiron School. The projects are listed on a github wiki.

I was very eager to meet the folks from Berkeley's iSchool that have been working on epub.js, a javascript package that lets you read EPUBs from the web in your browser. One of them, Jake Hartnell, is also on the team building Hypothes.is, a tool designed to let everyone annotate the web (and one of the sponsors of the event). (The others were AJ and Fred.) Sooner or later you'll see these tools added to Unglue.it. But there's still a lot of work left to do.

For example, the project at Open Book Hack with the highest ratio of usefulness to impressive-sounding-ness was the project to add scrolling to epub.js (check it out, it works!). That's right, they added an option (mostly) to make chapter 2 come below chapter 1. You will soon be able to use ebooks on the web using an open-source 9th century reading interface.

There was one prize awarded. The winning project, "See and Read" allowed two people with two screens to interact via an ebook. Very nicely executed and well-deserved, but if I told you that I had invented something that allows two people to interact with a book, you would say "oh?" because you don't really need 100 billion transistors and two glowing screens to interact with a book.

Another project "Breadcrumbs and Beanstalks", extended Harvard's StackLife interface to enable ebook browsing similar to that found in a physical collection of books, with 2 dimensional browsing (the 2 axes being publication date and the general-to-specific axis of subject headings. It looked a bit clunky, but thar's gold in there somewhere.

Hugh MacGuire, Max Fenton, Jean Kaplansky, and Fendi M were doing something brave with linking in  PressBooks, but having already done too much linking in my day, I decided not to understand it. A team from Sophia University in Japan was doing something clever with textbooks and Readium. There was work on converting PDF to EPUB and a project to use Phonegap to make apps from EPUBs.

The project I joined up with focused on making a book club application around a shared EPUB reader. (It works, too!) We based it on epub.js and Hypothes.is. I wasn't very useful to the effort, because the three Flatiron-trained Ruby-on-Rails developers in our group were too awesome for my plodding-but-powerful python to compete with. So I helped with exploration and documentation of the Hypothes.is API, and finding bugs and deployment gotchas in epub.js. I now know how to configure CORS for buckets in S3. Yeah that was my weekend. I also fixed a bug by staring, in an intimidating way, over someone's shoulder. Ah, good times.

What became apparent to me in working with these tools was that the freshly trained developers got everything to work by un-EPUB-ing everything. The web platform just works, with the one exception being that centering text blocks in CSS just doesn't, unless you look away from the screen. The EPUB platform always throws something in your way, for reasons that even StackOverflow doesn't explain. Ruby Zips won't unzip. Cross Sites won't request. Java Scripts won't bind.

EPUB's competition isn't Amazon and KF8 fixed layout, it's the web and HTML5 and its huge gravitational pull. For 90% of ebooks, the benefits of EPUB over HTML are scant (because EPUB is based on HTML!) and the development barriers are significant. It's been years, and still EPUB authoring tools aren't mature or mass market. Deployment tools are barebones.

Don't get me wrong. I'm still betting big on EPUB, but dammit, Publishing Industry, for an $80 billion pillar of modern society, you're investing a nanoscopic amount on your basic infrastructure (i.e. EPUB), despite the herculean efforts of the people I met last weekend.

Notes:
  1. I was really impressed with the Flatiron School students I worked with. If the rest of the students are anything like Edina, Tiff, and Dan (Ivan helped a bit, too), they are going to have a huge impact on the New York area economy. Maybe I should learn RoR.
  2. Bill McCoy has done an amazing job bringing people together under the IDPF and Readium umbrellas. Imagine what he could do with financial support commensurate to his task. Perhaps he should take up bootlegging.
  3. Jake wrote up his impressions, too.
  4. As have Virginie Clayssen and Camille Pène, in French.
  5. Would have posted sooner, but MILESTONE IN UNLUE.IT.

Enhanced by Zemanta

Wednesday, January 1, 2014

In 2013, eBook Sales Collapsed... in My Household.

2013 was not the year the ebook industry was expecting. We hoped that ebooks would continue their explosive year-on-year revenue growth, and that the replacement of print by digital would proceed apace. We suspected that the growth of ebook sales might moderate, because, as HarperCollins CEO Brian Murray told Publisher's Weekly, "Nothing grows by triple digits for too long." But just as CDs replaced vinyl and digital downloads replaced CDs, it seemed obvious that the the age of the printed book was nearing its end; the century of the ebook was dawning.

We got a few things right. Internationally, ebook sales growth was strong. Print continued its slow decline. Bookstores continued to close. But for some reason, ebook sales in the US stopped increasing. And even started declining!

There are many possible explanations for this turn of events. There are technicalities with the data collection, particularly with publishers such as Amazon's imprints that don't report their sales numbers. Young Adult sales dropped steeply, as there was no smash hit to follow on the huge success of Hunger Games. 50 Shades of Gray didn't turn out to be a lasting relationship. And there's been a downward trend on prices, particularly as publishers start to use dynamic pricing to stimulate sales. But it seems to me that something in the environment is changing, more than just a market maturation.

Amazon probably has enough data to understand what's happening, but they're notoriously opaque about reporting numbers. On the other hand, they're quite good about reporting to customers what they've bought. So I decided to analyze my own household's Amazon data. I had the impression that my family was spending less on ebooks, but I wasn't sure, because they still seem to spend all hours of the day reading. The results were kind of shocking.

The graph shows my household Kindle ebook purchases from 2009-2013. As you can see, 2013 marked a steep drop from the 2009-2011 peak years of about $1000 per year.

I don't buy Kindle ebooks myself (I buy ePub only, so I can hack on them) but other members of my household have bought quite a lot. The average price paid is about $7, and this has held quite steady. But in 2013, Kindle purchases stopped almost completely, and they were not replaced by purchases on other platforms.

Based on detailed "interviews" with the subject ebook purchasers, here are some non-factors in this collapse:
  1. "Netflix-for-Books" services. Nobody but me has heard of them.
  2. Kindle Owner's Lending Library. Despite an well-used Amazon Prime subscription, they haven't figured out how to use it for ebooks.
  3. Our public library. Nobody but me has used it for ebooks.
  4. Piracy. As if!
The two main reasons for this spending collapse turn out to be:
  1. The Kindle acquired in early 2009 reached end-of-life due to a cheaply made power cord, and was replaced by an iPad. The lack of in-app purchase for the Kindle App has resulted in a significant impediment to Kindle purchases. The iBookStore has not attracted a single ebook purchase.
  2. The iPad owner now spends the vast majority of her reading time in fan-fiction websites, mostly fanfiction.net and ArchiveOfOurOwn.org. Same for the iPad borrower, but a different mix of websites.
I find it worrying that the Justice Department pursued an big antitrust suit against Apple and 5 of the big 6 publishers, won, and despite making an issue of Apple's in-app purchase ban in iOS, it seems to have lost the argument with Judge Cote. We'll see how that turns out.

It's worth paying close attention to the fan fiction sites. After all, 2012's biggest revenue engine for the book industry, 50 Shades, was a repackaged fanfic. On an iPad with a decent internet connection, the fanfic sites work better than ePubs. They link and they script. Just try making a link from one ePub to another and you'll get my point.  They deliver content in smaller, more addictive chunks, and they integrate popular culture MUCH more effectively than books do, for reasons relating primarily to copyright. The authors are responsive and deeply connected to readers; they often ARE the readers!

There's a fanfic site to appeal to every reader; I highlighted Wattpad earlier this year. ArchiveOfOurOwn.org ("AO3"), a project of the Organization for Transformative Works, a non-profit, experienced the growth in 2013 that was missing from the ebook sector. The number of works hosted by AO3 doubled to just under a million works, covering almost 14,000 "fandoms". (A good example of a fandom is the "Dragonriders of Pern" fandom, which currently hosts 534 works). Fanfiction.net, an advertising supported site, hosts almost 2000 fandoms and over 1.3 million works, more than half of which are in the Harry Potter or Twilight fandoms. Game oriented discussion forums also engage in fanfiction. (Popular in my house is spacebattles.com)

My anecdata might be completely anomalous, although Amazon, a very data-driven company, seems to be aware of the same phenomena. They've been making the Kindle into a full featured tablet to go head-to-head with the iPad. They've also launched a fanfic site called Kindle Worlds, which has 15 worlds and 341 works.

Early stage venture capitalist Josh Kopelman says that many of the best opportunities for startups are not those in expanding markets. "We love investing in technologies and business models that are able to shrink existing markets. If your company can take $5 of revenue from a competitor for every $1 you earn – let's talk!",  he has written on his firm's website. Kopelman founded Half.com in the early days of the internet, a company which shrank the book market by getting people to resell the books they had just bought for a fraction of the price of a new book. Microsoft's Encarta shrank the Encyclopedia business from $1.2B to $600M before Wikipedia shrank the business by another 90%.

In 2014, I'm guessing it's the book publishing industry's time to shrink. A convergence of tech startups, tech monsters, and tech non profits seems to be ready for the assault. The fanfic sites, the Wattpads, the Project Gutenbergs and the Manybooks, the Readmills, the Leanpubs and the Smashwords (and I hope the Unglue.its); these are people building the foundations of a creative industry that will flourish even if the ebook sales collapse that I see around me spreads to your house as well.

Happy New Year!
Enhanced by Zemanta