Showing posts with label wikipedia. Show all posts
Showing posts with label wikipedia. Show all posts

Monday, December 31, 2018

On the Surveillance Techno-state

I used to run my own mail server. But then came the spammers. And  dictionary attacks. All sorts of other nasty things. I finally gave up and turned to Gmail to maintain my online identities. Recently, one of my web servers has been attacked by a bot from a Russian IP address which will eventually force me to deploy sophisticated bot-detection. I'll probably have to turn to Google's recaptcha service, which watches users to check that they're not robots.

Isn't this how governments and nations formed? You don't need a police force if there aren't any criminals. You don't need an army until there's a threat from somewhere else. But because of threats near and far, we turn to civil governments for protection. The same happens on the web. Web services may thrive and grow because of economies of scale, but just as often it's because only the powerful can stand up to storms.  Facebook and Google become more powerful, even as civil government power seems to wane.

When a company or institution is successful by virtue of its power, it needs governance, lest that power go astray. History is filled with examples of power gone sour, so it's fun to draw parallels. Wikipedia, for example, seems to be governed like the Roman Catholic Church, with a hierarchical priesthood, canon law, and sacred texts. Twitter seems to be a failed state with a weak government populated by rival factions demonstrating against the other factions. Apple is some sort of Buddhist monastery.

This year it became apparent to me that Facebook is becoming the internet version of a totalitarian state. It's become so ... needy. Especially the app. It's constantly inventing new ways to hoard my attention. It won't let me follow links to the internet. It wants to track me at all times. It asks me to send messages to my friends. It wants to remind me what I did 5 years ago and to celebrate how long I've been "friends" with friends. My social life is dominated by Facebook to the extent that I can't delete my account.

That's no different from the years before, I suppose, but what we saw this year is that Facebook's governance is unthinking. They've built a machine that optimizes everything for engagement and it's been so successful that they they don't know how to re-optimize it for humanity. They can't figure out how to avoid being a tool of oppression and propaganda. Their response to criticism is to fill everyone's feed with messages about how they're making things better. It's terrifying, but it could be so much worse.

I get the impression that Amazon is governed by an optimization for efficiency.

How is Google governed? There has never existed a more totalitarian entity, in terms of how much it knows about every aspect of our lives. Does it have a governing philosophy? What does it optimize for?

In a lot of countries, it seems that the civil governments are becoming a threat to our online lives. Will we turn to Wikipedia, Apple, or Google for protection? Or will we turn to civil governments to protect us from Twitter, Amazon and Facebook. Will democracy ever govern the Internet?

Happy 2019!

Thursday, January 18, 2018

GitHub Giveth; Wikipedia Taketh Away


One of the joys of administering Free-Programming-Books, the second most popular repo on GitHub, has been accepting pull requests (edits) from new contributors, including contributors who have never contributed to an open source project before. I always say thank you. I imagine that these contributors might go on to use what they've learned to contribute to other projects, and perhaps to start their own projects. We have some hoops to jump through- there's a linter run by Travis CI that demands alphabetical order, even for cyrillic and CJK names that I'm not super positive as to how they get "alphabetized". But I imagine that new and old contributors get some satisfaction when their contribution gets "merged into master", no matter how much that sounds like yielding to the hierarchy.

Contributing to Wikipedia is a different experience. Wikipedia accepts whatever edits you push to it, unless the topic has been locked down. No one says thank you. It's a rush to see your edit live on the most consulted and trusted site on the internet. But then someone comes and reverts or edits your edit. And instantly the emotional state of a new Wikipedia editor changes from enthusiasm  to bitter disappointment and annoyance at the legalistic (and typically white male) Wikipedian.

Psychologists know that that rewards are more effective motivations than punishments so maybe the workflow used on GitHub is kinder than that used on Wikipedia. Vandalism and spam are a difficult problem for truly open systems, and contention is even harder. Wikipedia wastes a lot of energy on contentious issues. The GitHub workflow simplifies the avoidance of contention and vandalism but sacrifices a bit of openness by depending a lot on the humans with merge privileges. There are still problems - every programmer has had the horrible experience of a harsh or petty code review, but at least there are tools that facilitate and document discussion.

The saving grace of GitHub workflow is that if the maintainers of a repo are mean or incompetent, you can just fork the repo and try to do better. In Wikipedia, controversy gets pushed up a hierarchy of privileged clerics. The Wikipedia clergy does an amazingly good job, considering what they're up against, and their workings are in the open for the most part, but the lowly wiki-parishioner rarely experiences joy when they get involved. In principle, you can fork wikipedia, but what good would it do you?

The miracle of Wikipedia has taught us a lot; as we struggle to modernize our society's methods of establishing truth, we need to also learn from GitHub.

Update 1/19: It seems this got picked up by Hacker News. The comment by @avian is worth noting. The flip side of my post is that Wikipedia offers immediate gratification, while a poorly administered GitHub repo can let contributions languish forever, resulting in frustration and disappointment. That's something repo admins need to learn from Wikipedia!

Sunday, September 30, 2012

CC BY and the Truth-Printing Business

Why are dollars worth anything? Why are digits on a bank statement worth anything? When my server tells our payments provider to move bits from your credit card, why does it matter to you?

In practical terms, dollars are valuable because other people will give you stuff or do things for you in exchange. Or at least they will if you can convince their bank to change the digits in their bank account. Their bank has to trust your bank which has to trust you. It all works because we all trust it will work. And why do we trust that it will work?

There are governments and laws to back them up. Why do we trust the government and laws? In practical terms we trust the government and laws because... well... they have ballot boxes. And judges and police forces. But mostly we trust the government and legal system because it sort of works and is often not abusive. At the bottom, it's because there's this web of trust which collectively holds everything together. Until of course, it doesn't. Because there isn't a bottom, it's turtles all the way down.

If you haven't heard of Bitcoin, let me give you this non-technical summary. Bitcoin is a recent implementation of the idea that money based on a web of cryptographically secured assertions is sounder than money based on a web of governmentally secured assertions. If as many people believed in cryptography as believe in astrology, we'd be using Bitcoin today.

The magic result is that an entity that gets society to trust its currency can then print money.

When the currency is truth rather than coin, judges and guns don't work so well. Traditional hierarchical authority systems are breaking down. What's replacing them is open authority systems. Systems such as wikipedia which allow everyone to participate in the construction of truth, not by being correct, but by being fixable. And to the frustration of many, Wikipedia delegates all its authority to things that are "citeable".

So how do you get to be an authority that Wikipedia believes? The two criteria that seem to matter most are
  1. Openness. If wikipedians can't read you, you don't exist. 
  2. Authority. People need to believe you. 
If you notice the circularity here, you'll see that printing truth and printing money are not so different.

As usual, I take a long time getting around to my point. Which is this: If you want to be in the business of printing truth, the best license to choose for your business is the Creative Commons Attribution License (CC BY). For now. And if you're printing science, medicine, technology or even philosophy, I really hope you want to print truth.

The Creative Commons part speaks to the need to be open. In the age of the internet, you can't print truth and keep it secret. No one will believe you.

The Attribution part builds your most valuable asset, your reputation. No one believes anonymous assertions.

You might ask about other options, for example, Non-Commercial (NC), No Derivatives(ND), Share-Alike (SA).

I've written about reasons to use NC and ND. Those reasons don't apply to the truth-printing business.

Can you imagine if your dollar bill said "This note is legal tender for all non-commercial debts public or private". That would be silly. The whole point of money is that it doesn't change depending on its use. And its the same with truth. There ain't no such thing as non-commercial truth. You can't control the uses of the truth you print. You can't even demand that people who consume your truth share that truth the same as you do..

A lot of people get confused about using no-derivative licenses. They think that if you print that the sky is blue, your credibility will be hurt if someone reprints a derivative of your truth and says the sky is black. But that's exactly what the attribution requirements prevent. But more than that, if you print your truth as chiseled in stone, then no one will believe it in a few years or so, because we all know that the truth hasn't been chiseled in stone for at least two thousand years. Nowadays we can make cryptographically strong proofs that assertions aren't being fiddled with and were made by the entities they're attributed. We can track the trail of assertions through history. And the provider of that chain of provenance is you, the truth printing proprietor. The longer the trail of conflicting assertions, the more crucial your authority as a truth printer becomes.

The problem of turning the currency of truth into harder currency is left as an exercise for the reader.
Enhanced by Zemanta

Thursday, October 20, 2011

Creative Commons - NC (Non-Commercial)

wormsby Wahj  (CC BY-NC-ND)Real worms don't come in cans. The last time I saw worms offered for sale, they came in paper buckets, the kind that usually hold Chinese take-out. You open these up, and the worms don't jump out at you. Maybe if you left the bucket open for a day or two, the worms would eventually find their way out, but any resulting problem is manageable. So when we say that doing something "opens up a can of worms", the main thing to think about is not about the calamities that will emerge from the bucket, it's whether or not you want to go fishing today.

The "Non-Commercial" attribute of Creative Commons "NC" licenses is definitely a can of worms. "Non-commercial" is subject to varied interpretation, and the license is not entirely successful at removing ambiguities. What uses are commercial? For example, is a blog that attracts advertising revenue allowed to post a CC NC licensed ebook? Is a for-profit distributor of ebooks allowed to distribute an NC ebook? Is a non-profit charity allowed to print paper copies of an NC ebook and sell the copies? Is a for-profit copy shop allowed to charge a school to make copies of an NC textbook? Is a for-profit company allowed to use an NC ebook about widget manufacturing to improve its factory? Is it ok for me to show you this picture of worms? Et cetera.

In building Unglue.it, we hope that readers and institutions will financially support Creative Commons relicensing of books that are important to them. In doing so, we need to make sure that the licenses we choose enable the things that the supporters want to do. We need to balance these uses against the rights and concerns of the authors and other rights holders who would be granting these licenses.

To get a better feel for what's allowed and what isn't under NC, we have to look at the "legal code" of the license. Here's what it says in the Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported (CC BY-NC-ND) License:
You may not exercise any of the rights granted to You [...] in any manner that is primarily intended for or directed toward commercial advantage or private monetary compensation. The exchange of the Work for other copyrighted works by means of digital file-sharing or otherwise shall not be considered to be intended for or directed toward commercial advantage or private monetary compensation, provided there is no payment of any monetary compensation in connection with the exchange of copyrighted works.
Let's examine some of our wriggling worms against this "code". Remember that I'm not a lawyer, and you should not rely on my scribblings for legal advice of any kind.
  1. Go ahead and improve your lucrative widget factory. The rights restricted by the NC clause are the rights to reproduce, distribute and publicly perform the work. The Creative Commons licenses do not restrict other uses of the work. If there are million-dollar ideas in the book, your ability to exploit them for commercial gain is not restricted by a CC license.
  2. If your blog is your business, it's not a good idea to build it on NC licensed photos from flickr, even if you don't charge for access. But if you're a book blogger and you make money with advertising, is posting a free ebook "primarily directed towards commercial advantage"? This worm is jiggling a bit! If you're a potential supporter of a book, this is the sort of use you probably want to support. It's not really clear how to apply the NC clause. Similarly, Apple, Amazon and Google are big companies that make a lot of money in the course of distributing ebooks. Distribution of some ebooks for free gives them indirect commercial advantages. To the extent that the uncertainly in the NC provision prevents seamless distribution of the works to their users, it goes counter to what most book lovers would want.
  3. Even if you're a non-profit, you can't print and sell copies of an NC e-book to raise money for starving orphans with cancer.
The areas of uncertainty includes some use cases that we think are non-commercial uses. To make it clear that we consider the distribution of unglued ebooks for free to be an allowed activity under NC licenses, rights holders who offer works to the public through Unglue.it will agree to the following:
For purposes of interpreting the CC License, Rights Holder agrees that "non-commercial" use shall include, without limitation, distribution by a commercial entity without charge for access to the Work.
We may also require a statement to the same effect in the front matter of the released ebook; we're still working out the file format details.

Given this clarification, why not go all the way, and require that Unglue.it rights holders agree to commercial distribution of works that get unglued?

 It turns out that the alternative to our can of worms harbors some poisonous snakes. Let me introduce you to one of these. Look at the Amazon page for Dance Dance Revolution (Wii Video Game) a 140 page paperback supposedly edited by Lambert M. Surhone, Mariam T. Tennoe, and Susan F. Henssonow. The so-called publisher, "Betascript Publishing" takes Wikipedia articles and turns them into books. So far, so good. Perfectly legal within the scope of Wikipedia's Creative Commons License (BY-SA). But how would you feel if you found a Wikipedia article that you wrote (with minor edits from others) on sale at Amazon for $57.47? When this happened to my brother he was mostly amused at the audacity of it all. But I think that if the same thing happend to a book I had worked on for a year of my life, it would seriously piss me off. If I had contributed money to "give the book to the world" I would be similarly aggravated. You could argue that Betascript is providing a valuable service by providing attractive formatting and improving the discovery of the article, but please don't.

chinese takeout boxby gabrielsaldana  (CC BY-SA)Retention of commercial rights is potentially of significant value to authors, and can reduce their asking price for ungluing books. I've previously written about Cory Doctorow's experience with selling "deluxe bound" versions of his Creative Commons Licensed books. There's also the possibility that authors' prior publishing contracts preclude them from offering commercial Creative Commons licenses (I'll write more about that soon). Since most of the uses we imagine for unglued ebooks, including the uses most important to libraries, are not affected by our use of the NC-flavored licensed, we've decided to open this "can of worms" in hopes of catching more "fish". We'll allow rights holders to offer non-NC licenses, but we won't expect them to do so.

Notes: I posted yesterday on the "Attribution" in Creative Commons Licenses. Here are some links on Betascript, which has over 350,000 "books" listed in some book directories:
Enhanced by Zemanta

Monday, May 2, 2011

Open Access eBooks, Part 2. What does Open Access mean for e-books?

No Shelf Required: E-books in LibrariesHere's the second section of my draft of a book chapter for a book edited by No Shelf Required's Sue Polanka. I previously posted the introduction; subsequent posts will include sections on Business Models for Open Access E-Books, and Open Access E-Books in Libraries. Note that while the blog always uses "ebook" as one word, the book will use the hyphenated form, "e-book". The comments on the first section have been really good; please don't stop!

What does Open Access mean for e-books?


There are varying definitions for the term “open access”, even for journal articles. For the moment, I will use this as a lower-case term broadly to mean any arrangement that allows for people to read a book without paying someone for the privilege. At the end of the section, I’ll capitalize the term. Although many e-books are available for free in violation of copyright laws, I’m excluding them from this discussion.

Public Domain

The most important category of open access for books is work that has entered the public domain. In the US, all works published before 1923 have entered the public domain, along with works from later years whose registration was not renewed. Works published in the US from 1923-1963 entered the public domain 28 years after publication unless the copyright registration was renewed. Public domain status depends on national law, and a work may be in the public domain in some countries but not in others. The rules of what is in and out of copyright can be confusing and sometimes almost impossible to determine correctly.

In addition to public domain books that are made available by Project Gutenberg, works digitized by other efforts may be available on an open access basis. It’s not true, however, that any digitized public domain book is also open access. That’s because the digitizer can restrict access to the works using license agreements. For example, JSTOR has many digitized public domain works included in its subscription products, but the terms of the subscription prevent republication of their scans. Similarly, Google puts restrictions on the public domain books from partner libraries that it has scanned, digitized and included in Google Books. While they’re available for free, there are limits on what you can do with them.

The public domain is more than just free; it belongs to everyone. Public domain works can be copied, remixed, altered or extended. A book publisher can take a public domain text, print up bound volumes, and sell them in bookstores. A movie producer can create a cinematic dramatization of the public domain work; derivative works such as the movie acquire copyrights of their own and are not in the public domain.

Free Copyrighted Content

Laypeople often confuse public domain for “free”, and vice versa. Most content available for free on the web is copyrighted, which restricts what people can do with it. Often, the content is made available using an advertising model, trading the opportunity to read and interact with content for the user’s attention to ads or links to e-commerce websites. But website users are usually not free to republish content or email the content to friends beyond the bounds of fair use. They’re bound by whatever terms and condition the website chooses to employ; if there are no explicit terms and conditions, they still can’t copy the website’s content for other uses.

Even professional publishers are sometimes confused by copyright on the web. In 2010, the editor of “Cooks Source”, a Massachusetts magazine got into hot water for republishing a blogger’s work without permission. The publisher’s response to the blogger, on being asked for restitution, made the rounds of the Internet, and is striking for the bellicose ignorance it betrays:
Yes Monica, I have been doing this for 3 decades, having been an editor at The Voice, Housitonic Home and Connecticut Woman Magazine. I do know about copyright laws. It was “my bad” indeed, and, as the magazine is put together in long sessions, tired eyes and minds somethings forget to do these things. But honestly Monica, the web is considered “public domain” and you should be happy we just didn’t “lift” your whole article and put someone else’s name on it! It happens a lot, clearly more than you are aware of, especially on college campuses, and the workplace. If you took offence and are unhappy, I am sorry, but you as a professional should know that the article we used written by you was in very bad need of editing, and is much better now than was originally. Now it will work well for your portfolio. For that reason, I have a bit of a difficult time with your requests for monetary gain, albeit for such a fine (and very wealthy!) institution. We put some time into rewrites, you should compensate me! I never charge young writers for advice or rewriting poorly written pieces, and have many who write for me… ALWAYS for free!
Many free e-books are available on a similar basis as free websites. They may include advertising or advocacy. Promotional literature and instruction manuals often fall into this category. Many publishers make free e-books available for limited periods of time as a means of marketing them; that doesn’t make them free to redistribute, though it happens.

Creative Commons Licensing

Creative Commons licensing arose to expand the range of creative works available for others to build upon legally and to share. Many authors really want their works to be redistributed for free in venues such as Cooks Source, but they want to make sure attribution is given, and often want to prevent their work from being altered or chopped into pieces. Others want to make sure that if their work is altered or somehow improved, the altered or improved version will also be available for free. Sometimes, authors are happy to have their works reused non-commercially, but want to keep their works from being commercially exploited without permission. Creative Commons licenses give authors the tools they need to accomplish these goals.

CC BY-SA mark
The different licenses available from Creative Commons are designated with a special mark, with added code letters that indicate the features invoked by the rights holder. For example, the “Attribution-ShareAlike” license is denoted by the letters “CC BY-SA” and the mark shown. This license requires attribution as to the author of the work, and the ShareAlike features bind the licensee to share any modifications or improvements.

It’s important to note that in the Creative Commons licenses, the owner of the copyright does not give up ownership of the work. The owner is free to re-license the work under any terms they desire, and can still sue people who infringe on the copyrights. The owner licenses the work to the user, who accepts the license as a condition of use. The user can in turn distribute the work along with a copy of the license to other users, who accept the terms of the same license from the copyright owner as a condition of their use.

Creative Commons licensing is now widely used for free e-books distributed on the web. Perhaps the best known e-books using CC are the works of Cory Doctorow, a blogger, science fiction author and advocate for copyright law reform. It’s also used for Wikipedia contributions, and is supported by Flickr for use in photos.

Copyleft

While Creative Commons licenses are the most frequently used for e-books, other licenses can be used to allow for the free reading of books. Noteworthy among these is the GNU Free Documentation License (FDL), created by the Free Software Foundation to allow software documentation, manuals and other text to be distributed with strong “copyleft” provisions compatible with the GPL software they’re meant to accompany. The GNU FDL can easily be applied to e-books; many ebooks have been released with this license and with other Free Software Foundation licenses.

The idea of copyleft is that licenses can be used to prevent someone from taking from the commons without also giving back. For example, when a book publisher adds commentary and illustrations to the text of a Shakespeare play, the resulting book is covered under copyright and permission must be given for redistribution even though the underlying work is in the public domain. This would not be allowed by a copyleft license. The Creative Commons SA licenses have weak copyleft; the GNU FDL is stronger, and even forbids the use of DRM. It’s not clear whether it would be legal to distribute a GNU FDL e-book to a Kindle e-reading device without permission from the author.

Open Access vs. open access

How Wikipedia Works: And How You Can Be a Part of ItConsider the book How Wikipedia Works by Phoebe Ayers, Charles Matthews, and Ben Yates. Is it an open access e-book? Based on the page at the Free Software Foundation, you might assume the answer is an easy yes, because it comes with a GNU FDL license. If you search for this book on Google, however, you’ll have to dig quite a bit to get a free e-book. Amazon will sell you the Kindle version for $21.64. You can buy it in three different formats from O’Reilly or from No Starch Press, the publisher, for $23.95. Google books has it through their publisher program; it appears to fully available and Google doesn’t try to sell it to you. You can find the e-book in a library through Worldcat, but the libraries that hold it restrict access to their own users. Wikipedia itself has a page for it, but no download link; for that you need to look on the talk page.

The intent of the publisher of this book doesn’t seem to be to make the e-book available openly, even though it uses a “free” license. The free distribution of the e-book is not effective. There are a lot of ways to license content, but at the end of the day, it’s the intent of the rights holders and the effectiveness of the free distribution that makes an e-book “Open Access” with capital OA.

Notes:
  1. How Wikipedia Works: is available (GNU FDL license) as PDF (here (15 MB)). The Google books version is here. It's listed on a GNU web page.
<- previous post in series    next post in series ->

    Sunday, August 15, 2010

    Charlie Chan Actor Warner Oland Not Mongolian, Say Wikipedia

    When my mom was pregnant with her third child, my dad loved it when people asked if they were expecting a boy or a girl. "Well" he'd answer with a twinkle in his eye. "They say one of every 3 children born in the world are Chinese, so for our third child, that's what we're expecting!"

    My parents were Swedish. My father was born in Gary, Indiana, but grew up in northern Sweden; my mother was born in Sweden and her mother was a Lapp, or Saami. After his retirement, my father became very interested in genealogy, and he traced his ancestors and relatives, almost 10,000 of them. Since about the year 1400 Sweden has done a very good job of recording births and deaths in church records, and since people didn't move around much, it's not hard for us to trace people. In the farming villages where my parents came from, everybody is related to everybody else.

    I've inherited my dad's database and I've put it online. Doing so has has put me in touch with a fascinating variety of distant cousins. Among my distant relatives was the actor Warner Oland, who became famous for portraying Charlie Chan in Hollywood movies. Warner Oland, whose real name was Johan Verner Ölund, was a third cousin to my father's mother. My father noted in his database that he remembered when Warner Oland came to their village by car and met my grandparents. It must have been the same year Warner Oland died, 1938.

    Naturally, I pay attention whenever Oland in mentioned in the media. Over the last week, I've read articles in the New Yorker and in the New York Times about a new book by UCSB English Professor Yunte Huang. The book is entitled Charlie Chan: The Untold Story of the Honorable Detective and His Rendezvous with American History; it tells the story of the "real" Charlie Chan, a detective in Honolulu, Hollywood's portrayal of Charlie Chan, and Huang's own story as a chinese immigrant in America. A significant part of the book recounts the odd story of how a Swedish actor came to portray the quintessential Chinese detective.

    When I read the New Yorker article, I immediately put the book on my "must read" list. (Unfortunately, it's not available as an ebook, and is sold out of my local bookstores!) But one sentence of the New Yorker review, written by Harvard history professor Jill Lapore, stuck out for me:
    Oland, born in Sweden in 1880, had, beginning in 1917, specialized in playing Oriental villains, including Dr. Fu Manchu. (Oland's mother was Russian, and he had slavic features.)
    Oland's mother was NOT Russian. Oland's mother was my grandmother's third cousin. His father was a 5th cousin to my grandmother. The Swedish genealogist Sven-Erik Johansson has specialized in the digitization of the church records in the region of northern Sweden where Oland and my grandparents came from and has published an ancestor chart for Warner Oland going back 5 generations. None of those ancestors come from Russia. To top it off, Warner Oland was born in 1879, not 1880 as reported in the New Yorker.

    So where did the idea that Oland had a Russian mother come from? Doesn't the New Yorker have fact checkers? I went to Wikipedia to find out. The Wikipedia article said that "His mother was Russian of Mongolian descent.", referencing a "page not found" Internet Movie Database (IMDB) article. I refound that article, which says:
    He didn't need make-up when he played Charlie Chan; all he would do is curl down his moustache and curl up his eyebrows. In fact, the Chinese often mistook him for one of their own countrymen. He attributed this to the fact that his Russian grandmother was of Mongolian descent.
    So IMDB says it's his grandmother who's Russian and of "Mongolian descent"; the key thing to note is the attribution. I immediately edited the Wikipedia article to omit to spurious information. A day later, a wikipedian had put back the Mongolian bit, but more accurately worded as being something Oland said. A proper reference, to a book by Ken Hanke, Charlie Chan at the Movies: History, Filmography, and Criticism (Google Books, Amazon) had been added. That book says:
    "Even before the role of Charlie Chan came his way, Oland was a frequent onscreen Oriental, despite the fact that he was born in Sweden to a mixture of Swedish and Russian Parents. Physically, he had an exotic look to begin with, and the addition of an Oriental-style mustache and beard made the transformation complete. "I owe my Chinese appearance to the Mongol invasion," he once told Keye Luke. "That's true," Luke agrees, "because the Mongols did get up there around Sweden and Finland and naturally sired some children, and so, he said, 'I come by it naturally.' And, his whole family looked like that." There was never any need for elaborate make-up. "All he did," explains Luke. "was put that little goatee on his chin. Otherwise, he had his own mustache. Everything was just like that. No make-up. It's just amazing."
    At this point, I must make an observation. Please look at the photo and decide for yourself. As far as I can judge, Warner Oland didn't look the least bit Oriental. He looked like most everyone else living in that area of northern Sweden would look if they put on a smudge of eyebrow makeup. But the resemblance to that Chinese detective in the movies is uncanny!

    There is, however, a story I remember my dad telling about a deserter from the Russian army. (The Russians burned down the closest city, Umeå, in a war in 1720.) It was said that this deserter hid in the woods or disguised himself as one of the locals. The way my dad told it, it was quite a scandal, even 200 years later. So maybe Warner Oland was joking when he said his mother was "Russian". In any case, the mysterious Russian in my family does not appear in the church records!

    If Oland really had exotic features, it's much more likely he got them from a source other than a stray Mongolian. The closest the Mongols got to northern Sweden was Lithuania. In the area where Oland's family originated, the ethnic mix was dominated by Finns, Swedes, and Saami.

    Take a look at a photo of my mother's cousin (unrelated to Oland), a pure Saami. With a bit of make-up (and some acting talent), she would have easily been able to play a Chinese woman. The Saami look quite different from the Finns and the Swedes. They are an indigenous people of Scandinavia, and no one really knows where they came from. Though their language is related to Finnish, they are not genetically related to the Finns. A recent DNA study (PDF, 399KB) published in the American Journal of Human Genetics suggests that they are related to the Berbers of northern Africa. It may well be that Oland thought they might be related to the Mongols.

    There's another interpretation of Oland's references to his "Mongolian" blood. In his time, children with Down's syndrome were referred to as "Mongoloid". In the 19th century, Down's syndrome was regarded as an expression of genetic "degeneration" toward the inferior "Mongoloid" races. It could well be that jokes about Mongolian ancestry reflected a belief that cases of Down's Syndrome were a result of racial contamination. My father's database shows many examples of women with large families bearing children into their 40's; his own familiy of 11 included one Down's child.

    So it seems likely that Warner Oland's statements about his ancestry were either inventions or jests. What's interesting to me is how this truth is constructed. It's not hard for people to look at the evidence now available and decide that a genealogist working with church records is probably more reliable than a co-star's recollection of an actor's constructed persona with regard to Oland's ancestry. Yunte Huang, the author of the new book, emailed me to say he agreed that it was a jest of Oland, who was known to be "quite a wisecracker". Now THAT sounds like my Dad's family!

    At first glance you might say Wikipedia is totally unreliable, because anyone can change it. But compared to the New Yorker, IMDB, and a book published in 2004, Wikipedia is more reliable because it CAN be changed, and because it supports a version history and a culture of citation and transparency for any information that might be disputed. While I'm optimistic about Wikipedia's ability to construct truth, I'm worried about systems that extract facts from Wikipedia articles and feed them in to the semantic web. While editing the article on Warner Oland, I deleted the assertion that he was a "Swedish Person Of Russian Descent". I wonder about the lifespan of this assertion as it has been copied and distributed throughout the world. There are really no good mechanisms to de-sert this sort of assertion. It's only with context that assertions can build truth.

    As it happens, I married into a family that really IS Chinese. I remember showing my mother-in-law old pictures of Saami ancestors in their traditional dress. "Those look like Manchu people!" she exclaimed. It's true. If you put aside the lens of race, we all look more or less alike, and we all look a bit exotic.

    Update: The author of the New Yorker article, Jill Lapore, got back to me to report that her article relied on the entry for Warner Oland in American National Biography (Oxford:  Oxford University Press, 2000) for the assertion about Oland's mother's ancestry.

    Update, August 23: Some additional research shows that Oland is also a third cousin of my grandmother.

    Tuesday, June 2, 2009

    Who's the Boss, Steinbrenner or Springsteen?

    As I started playing with Mathematica when it first came out (too late for me to use it for the yucky path integrals in my dissertation), I just had to try Wolfram|Alpha. The vanity search didn't work; assuming that's what most people find, its probably the death knell for W|A as a search engine. Starting with something more appropriately nerdy, I asked W|A about "Star Trek"; it responded with facts about the new movie, and suggested some other movies I might mean, apparently unaware that there was a television show that preceded it. Looking for some subtlety with a deliberately ambiguous query, I asked about "House" and it responded "Assuming "House" is a unit | Use as a surname or a character or a book or a movie instead". My whole family is a big fan of Hugh Laurie, so I clicked on "character" and was very amused to see that to Wolfram|Alpha, the character "House" is Unicode character x2302, "⌂". Finally, not really expecting very much, I asked it about the Boss.

    In New Jersey, where I live, there's only one person who is "The Boss", and that's Bruce Springsteen. If you leave off the "The", and you're also a Yankees fan, then maybe George Steinbrenner could be considered a possible answer, and Wolfram|Alpha gets it exactly right. Which is impressive, considering that somewhere inside Wolfram|Alpha is Mathematica crunching data. The hype around Wolfram|Alpha is that it runs on a huge set of "curated data", so this got me wondering what sort of curated dataset knows who "The Boss" really is. To me, "curated" implies that someone has studied and evaluated each component item and somehow I doubt that anyone at Wolfram has thought about the boss question

    The Semantic Web community has been justifiably gushing about "Linked Data", and the linked datasets available are getting to be sizable. One of the biggest datasets is "DBpedia". DBpedia is a community effort to extract structured information from Wikipedia and to make this information available on the Web. According to its "about" page, the dataset describes 2.6 million "things", and is currently comprised of 274 million RDF triples. It may well be that Wolfram Alpha has consumed this dataset and entered facts about Bruce Springsteen into its "curated data" set. (The Wikimedia Foundation is listed as a reference on its "the boss" page.) If you look at Bruce's Wikipedia page, you'll see that "The Boss" is included as the "Alias" entry in the structured information block that you see if you pull up the "edit this page" tab, so the scenario seems plausible.

    Still, you have to wonder how any machine can consume lots of data and make good judgments about who is "The Boss". Wikipedia's "Boss" disambiguation page lists 74 different interpretations of "The Boss". Open Data's Uriburner has 1327 records for "the boss" (1676 triples, 1660 properties), but I can't find the Alias relationship to Bruce Springsteen. How can Wolfram|Alpha, or indeed any agent trying to make sense of the web of Linked Data, deal with this ever-increasing flood of data?

    Two weeks ago, I had the fortune to spend some time with Atanas Kiryakov, the CEO of Ontotext, a Bulgarian company that is a leading developer of core semantic technology. Their product OWLIM is claimed to be "the fastest and most scalable RDF database with OWL inference", and I don't doubt it, considering the depth of understanding that Mr. Kiryakov displayed. I'll write more about what I learned from him, but for the moment I'll just focus on a few bits I learned about how semantic databases work. The core of any RDF- based database is a triple-store; this might be implemented as a single huge 3 column data table in a conventional database management software; I'm not sure exactly what OWLIM does, but it can handle a billion triples without much fuss. When a new triple is added to the triple store, the semantic database also does "inference". In other words, it looks at all the data schemas related to the new triple, and from them, it tries to infer all the additional triples implied by the new triple. So if you were to add a triples ("I", "am a fan of", "my dog") and ("my dog", "is also known as", "the Boss"), then a semantic database will add these triples, and depending on the knowledge model used, it might also add a triple for ("I", "am a fan of", "the Boss"). If the data base has also consumed "is a fan of" data for millions of other people, then it might be able to figure out with a single query that Bruce Springsteen, with a million fans, is a better answer to the question "Who is known as 'the Boss'" than your dog, who, though very friendly, has only one fan.

    As you can imagine, a poorly designed data schema can result in explosions of data triples. For example, you would not want your knowledge model to support a property such as "likes the same music" because then the semantic database would have to add a triple for every pair of persons that like the same music- if a million people liked Bruce Springsteen's music, you would need a trillion triples to support the "likes the same music" property. So part of the answer to my question about how software agents can make sense of linked data floods is that they need to have well thought-out knowledge models. Perhaps that's what Wolfram means when they talk about "curated datasets".

    Tuesday, May 26, 2009

    There is no truth on the internet

    In his retirement, my father took up genealogy as a hobby, and after he died, his database of thousands of ancestors (most of them in northern Sweden) passed to me. If you're interested, you can browse through them on the hellman.net website. Having all this data up on the web has been rather entertaining. Every month or so, I get an e-mail from some sixth cousin or such who has discovered a common ancestor through a google search, and the resulting exchanges of data allow me to make occasional corrections and additions.

    Since I've taken the database on, huge amounts of genealogic information has become available on the internet. When I first started finding this information, I made the mistake of trying to suck it into my database, since I had become more a less a professional data sucker and spewer in my work life. Once I had spent hour after hour pulling data in, I started to wonder what the point of it all was. Could I relly determine, and did I really care whether Erik Eriksson, born 1837 in Backfors, was really my fourth cousin thrice removed or not? What is the relationship between the data I sucked in and the truth about all the real people listed in the database? I quickly regretted my data gluttony.

    Traditional genealogists focusing on Sweden use a variety of material as primary sources of information. Baptismal records typically give a childs name and birthdate along with the names of their parents; burial and marriage records similarly give names and dates. The genealogist's job is to connect names on different records to construct a family tree. But things are not always simple. Probably 20% of males in the Backfors region were named Erik, and since patronymics were used, 20% of those males were also named Eriksson, though the name might be abbreviated in the records as "Ersson". To judge whether a girl named Hanna listed on a birth record from 1877 which lists "Erik Eriksson" as the father is really the daughter of the Erik Eriksson born in 1837 in Backfors, the genealogist must consider all the information available together with conditional probabilities.

    The internet genealogist (e.g., me) has a different task. Rather than looking at the birth records and assessing the likelihood of name coincidences, the internet genealogist looking at the same question searches the internet and finds that the web site "sikhallan.se" lists Hanna as Erik's daughter. The internet genealogist then makes a judgement about the reliability of the Sikhallan website. For example, how do we know that Sikhallan's source for Erik's birthdate isn't just the hellman.net website? If the two databases disagree, who should be believed? In my case, I just look at my father's meticulous notes about where his information comes from and if he noted some uncertainty, then I'm much more likely to believe the other sources available to me. Unless of course my data has come from one of my data sucking binges, in which case the source of my data has been lost and I can no longer judge its reliability.

    In my last two posts on reification (Part 1, Part 2), I promised that I would have a third post evaluating whether the reification machinery in RDF was worth the trouble. This is not that third post, this is more of a philosophical interlude. You see, another way to look at genealogic information on the internet is to think of it as a web of RDF triples. For example, imagine if Sikhallan made its data available as a set of triples, e.g. (subject: Erik Eriksson; predicate: had daughter; object Hanna). Then we could load up all the triples into an RDF-enabled genealogy database, and all our problems would be solved, right? Well, yes, unless of course we wanted to retain all the supporting information behind the data, the data provenance, all the extra care in citation of source taken by my Father and and ignored by me in my data-sucking orgies. In reality, the triple itself is worthless, devoid of assessable truth. If the triple were associated with provenance information its truth would become assessable, and thus valuable. The mechanism that RDF provides for doing things like this is... reification.

    Wikipedia is the most successful knowledge aggregation on the internet today and is also, not coincidentally, the best example of the value of comprehensive retention of provenance and attribution. Wikipedia keeps track of the data and author of every change in its database, and relentlessly purges anything which is not properly cited. Wikipedia is, in my opinion the best embodiment of my view that there is no truth on the internet- there are only reified assertions.