Showing posts with label languages. Show all posts
Showing posts with label languages. Show all posts

Saturday, March 9, 2013

African Drummers Invented an Internet

Maybe in 50 years we'll reminisce how there used to be one internet that covered the globe. But even before the telephone was invented, there were internets of a sort that covered regions in Africa. Probably there were others all around the globe, maybe even now the dolphins have their own version of an internet, disconnected from ours.

An internet, for the purposes of this article, is "a digital communications network that connects intelligent nodes distributed thoughout a region".

Even civilizations without written languages needed to communicate with each other. If you lived in a rain forest where villages were separated by miles of bush, the best way of communication with a neighboring village was to use drums. The "talking drums" of Africa used digital codes that could be understood from distances of as much as 5-10 kilometers. The codes weren't at all like Morse code, but were based on the tones of spoken languages.

A "talking drum" has two tones, so the signal is essentially binary. To overcome the lack of consonants, drum languages would add habitual phrases to words to disambiguate one word for another, resulting in an error-correcting code. Ruth Finnegan's chapter on "Drum Language and Literature" in Oral Literature in Africa gives some wonderful examples.
In the Kele language the words meaning, for example, ‘manioc’, ‘plantain’, ‘above’, and ‘forest’ all have identical tonal and rhythmic patterns. By the addition of other words, however, a stereotyped drum phrase is made up through which complete tonal and rhythmic differentiation is achieved and the meaning transmitted without ambiguity. Thus ‘manioc’ is always represented on the drums with the tonal pattern of ‘the manioc which remains in the fallow ground’, ‘plantain’ with ‘plantain to be propped up’, and so on. Among the Kele there are a great number of these ‘proverb-like phrases’ to refer to nouns. ‘Money’, for instance, is conventionally drummed as ‘the pieces of metal which arrange palavers’, ‘rain’ as ‘the bad spirit son of spitting cobra and sunshine’, ‘moon’ or ‘month’ as ‘the moon looks down at the earth’, ‘a white man’ as ‘red as copper, spirit from the forest’ or ‘he enslaves the people, he enslaves the people who remain in the land’, while ‘war’ always appears as ‘war watches for opportunities’. Verbs are similarly represented in long stereotyped phrases. 
So that's how information was transmitted digitally, but it's not an internet yet. For that, you need a network. It turns out that drummed messages of note would be retransmitted to the next village. In modern terminology, packet switching. I imagine the drummers used a protocol similar to that used by Ethernet to ensure a clear channel for retransmission. Thus announcements, warnings, poetry (maybe even advertisements!) were packet-switched between nodes based on topic and relevancy.

A lot of expressive power of the drum language was used for names. From Oral Literature in Africa:
Personal drum names are usually long and elaborate. In the Benue-Cross River area of Nigeria, for instance, they are compounded of references to a man’s father’s lineage, events in his personal life, and his own personal name . Similarly among the Tumba of the Congo, all-important men in the village (and sometimes others as well) have drum names: these are usually made up of a motto emphasizing some individual characteristic, then the ordinary spoken name; thus a Belgian government official can be alluded to on the drums as ‘A stinging caterpillar is not good disturbed’. Carrington describes the Kele drum names in some detail. Each man has a drum name given him by his father, made up of three parts: first the individual’s own name; then a portion of his father’s name; and finally the name of his mother’s village. Thus the full name of one man runs ‘The spitting cobra whose virulence never abates, son of the bad spirit with the spear, Yangonde’. Other drum names (i.e. the individual’s portion) include such comments as ‘The proud man will never listen to advice’, ‘Owner of the town with the sheathed knife’, ‘The moon looks down at the earth / son of the younger member of the family’, and, from the nearby Mba people, ‘You remain in the village, you are ignorant of affairs’. (citations omitted)
So the drum languages seem to put importance on uniquely identifying individuals, something that our Internet is just starting to figure out. (See ORCID.) Reputation of individuals was important; I wonder if creators of particularly compelling drum poems were identified by custom, as we're starting to learn how to do with Attribution licenses.

It goes without saying that the literary forms transmitted by drumming were not copyrighted and the there was no notion of paying a creator for "copies" of a drummed message. But certainly the practitioners of this early digital literature were valued by their societies.
drumming tends to be a specialized and often hereditary activity, and expert drummers with a mastery of the accepted vocabulary of drum language and literature were often attached to a king’s court. 
Masters of unwritten literatures found many ways of making a living. The "court poet" is a familiar role to us; modern writers often find wealthy patrons. In addition, Finnegan relays another way that creators of oral literature earned their livings:

The singer arrives at a village and finds out the names of the important and wealthy individuals in the area. Then he takes up his stand in public and calls out the name of the individual he has decided to apostrophize. He proceeds to his praise songs, punctuated by frequent and increasingly direct demands for gifts. If they are forthcoming in sufficient quantity he announces the amount and sings his thanks in further praise. If not, his innuendo becomes gradually sharper, his delivery harsher and more staccato. This is practically always effective—all the more so as the experienced singer knows the utility of choosing a time when all the local people are likely to be within hearing, in the evening, the early morning before they have left for the farm, or on the occasion of a market which leaves no escape for the unfortunate object singled out for these ‘praises’. The result of this public scorn is normally the victim’s surrender. He attempts to silence the singer with gifts of money or, if he has no ready cash, with clothes or a saleable object like a new hoe.
So even "astroturfing" is not an exclusively modern phenomenon.

I learned all this reading Ruth Finnegan's Oral Literature in Africa which was the first book made free to the world by Unglue.it, working with Open Book Publishers. Download and enjoy. Open Book Publishers have just launched an ungluing campaign for a second book, called Feeding the City, a translation of a seminal work from the original Italian, about the dabawallahs of Mumbai, a subject Internet entrepreneurs could learn a lot from. Support the campaign to make it free to the world!
Enhanced by Zemanta

Sunday, September 5, 2010

What "Ping" Really Means

When I was little, some Swedish was spoken in my house. At some point, I realized that our bathroom words were different from those used by my Ohio playmates. In my house, we didn't do "poop" or "poo" and definitely not "crap" and most certainly not "shit". We did "bice". It may be a complete coincidence, but I have never dined at the Italian restaurant named "Bicé".

The Story about Ping (Reading Railroad Books)I know what you're thinking. Ok, that's number two, so what did you call number one? No, not "pee" or "wee", and definitely not "piss". But I remember exactly when my mom explained to us that the word we used was not the one used by most English speakers. It was when my mom read us the book The Story about Ping by Marjorie Flack. My little sister roared with laughter, because "ping" was our word for urine.

I have no idea whether "ping" is a widely used bathroom word, here, in Sweden or anywhere; I don't go around talking much about ping. But I can tell you that Ping, Apple's new iTunes feature, is a piss-poor excuse for a social network.

I read Dave Winer's post describing the lameness of Ping, but I was still eager to try for myself. Apple's ability to create things that "just work" is justifiably reknowned. But having put my toes in the water, my reaction was more along the lines of Swizec's scatologically titled post.

I was stunned that Apple had not implemented the obvious functionality. When you listen to a song, you should be able to push a comment for it to your followers. In Ping, you can't. iTunes knows the songs I've rated most highly. Inexplicably, these are not the songs it suggests for my profile. Ping seems only interested in things I've bought recently in the store. But it even appears to be inept at using my iTunes-store sanctioned activity in my profile. It's hard to believe that something so poorly executed could get released by Apple.

The Napster corporate logoAfter thinking about it for a while, I realized what had happened. Imagine a system that can tell people what songs you have on your computer, and can connect you directly with the people interested in those songs. You want to share information about the songs and connect to people. Does that sound vaguely familiar? Do you remember Napster? The only difference between a well implemented Ping and the legally challenged Napster is a way to push files around.

I think that Apple showed Ping to some music publishers, who flipped out at the possibility that it would be used for file sharing and forced Apple to cripple Ping. Or maybe Apple saw the file-sharing potential itself and worried that Ping could kill off its music based revenue stream. Or could it possibly be that Apple is being incredibly devious, and is expecting that someone, somewhere will see how to add file sharing to Ping, resulting in huge popularity of Ping even while some elusive third party assumes all the Napster liability? We shall see.

The only thing I'm sure of is that Apple isn't likely to discuss what "Ping" really means – outside of its own bathroom.
Enhanced by Zemanta

Saturday, July 24, 2010

The Curious Case of eCapitalization

An unresolved problem faced by all technology writers is what to do with creative capitalization. When you want to lead off a sentence with a word like "iPad" or "eBook", how do you capitalize it? Do you go with "Ipad" and "Ebook"? Or perhaps "IPad" and "EBook"? Do you stay with "iPad" and "eBook" and consider them to be capitalized versions of "ipad" and "ebook"? Horror of horrors – you could put in a dash. Maybe you just finesse the issue by changing your sentence around to avoid having a camelCased word leading off the sentence. Even then, you have the problem of what to do if the word is in the title of your article, for which you probably use Title Case unless you're a cataloging librarian, in which case you use Sentence case, not that the problem goes away! If you're using an iPhone, you know it has a mind of its own about the first letter of your email address being capitalized.

The practice of capitalizing titles presents issues particularly when the titles are transported into new contexts, for example via an RSS feed or search engine harvest. ALL CAPS MIGHT LOOK OK AS A <TITLE> ON YOUR WEB PAGE, but a search engine might hesitate to scream at people

This is not by any means a new problem, but it's one that changes from era to era because of the symbiotic relationship between language and printing technology. Here's what Charles Coffin Jewett wrote in 1853 when discussing how libraries should record book titles:
The use of both upper-case and lower-case letters
in a title-page, is for the most part a matter of the printer's taste,
and does not generally indicate the author's purpose.  To copy them in a
catalogue with literal exactness would be exceedingly difficult, and of
no practical benefit.  In those parts of the title-page which are
printed wholly in capitals, initials are undistinguished.  It would be
unsightly and undesirable to distinguish the initials where the
printer had done so, and omit them where he had used a form of letter
which prohibited his distinguishing them.  It would teach nothing to
copy from the book the initial capitals in one part of the title, and
allow the cataloguer to supply them in other parts.
The standard practice of libraries in English speaking countries has been to record book titles in Sentence case, in which the first word of the title is capitalized and the rest of the words are capitalized it only if the language demands it (unless the first word is an article like "A", then the second word is also capitalized). The argument for this is that this capitalization style allows for the most meaning to be transmitted; a reader can tell which words of a title are proper names or other words that are capitalized. Which begs two questions: Why are libraries alone in presenting titles this way? Why do libraries persist in this practice when no one in recorded history has ever asked for sentence case titles?

In German and other languages, nouns are capitalized; this used to be true of English (take a look at the US Constitution).  In German, it's easy to tell nouns from verbs, which might be very useful if we still had it in English. Still, I enjoy being able to write that something is A Good Thing. It gives me a way to intone my text with an extra bit of information.

The rules for how English should be capitalized have become quite complicated. Here and here are two web pages I found devoted to collecting capitalization rules. Some of them are pretty arcane.

It's fun to speculate on the future of capitalization. In the late 19th century, there was a fashion to simplify spelling, grammar and capitalization, led by people like Melvil Dewey. I'm guessing part of the reason was the annoyance of needing to press a shift key on those newfangled typewriters. But spelling and capitalization reform didn't get very far. Perhaps they tried to publish articles and got stopped in their tracks by a unified front of copy editors.

If anything, the current trend is in the direction of making capitalization even more idiosyncratic. In addition to a proliferation of Product names like iPod and eBay that have crossed over into the language mainstream,  the shift from print to electronic distribution of text does a better job of preserving the capitalization chosen by the author, thus allowing it to better transmit additional meaning.

The ability to increase the information density in text is useful in a wide range of situations, for example, when you have only 140 characters to work with, or when you want a meaningful function name, like toUpperCase(). If your family name is McDonald, you probably have strong feelings on the issue.

My guess is that life will become increasingly case sensitive. You may already be aware that it takes 8 seconds, not one, to transmit a 1 GB file over a 1 Gb/s link. And that SI unit Mg is a billion times the mass of a mg. If you are a Java programmer, If you know the difference between an integer and an Integer, you'll quickly learn about NullPointerExceptions.

The shift from ascii to Unicode has made it much easier to cling to language specific capitalization rules. Did you know that there are a small number of characters that are different in "upper case" than in title case? They are: Letter DZ, LETTER DZ WITH CARON, LETTER LJ, LETTER NJ, and LETTER DZ. The lower case versions are dz, dž, lj, and nj; the upper case versions are DZ, DŽ, LJ, NJ, and the title case versions are Dz, Dž, Lj, and Njfi, fl, ffi, ffl, ſt, st. And don't forget your Armenian ligatures, ﬓ, ﬔ, ﬕ, ﬖ, ﬗ. For this reason, being "case insensitive" is poorly defined- two strings that are equal when you've changed both to uppercase are not necessarily equal after you've changed them to lower case!

So what do I do when I write about ebooks I don't use a dash? When the word appears in a title, I capitalize the "B". I can't wait till they translate this rule into Armenian.

Friday, November 20, 2009

Putting Linked Data Boilerplate in a Box

Humans have always been digital creatures, and not just because we have fingers. We like to put things in boxes, in clearly defined categories. Our brains so dislike ambiguity that when musical tones are too close in pitch, the dissonance almost hurts.

The aesthetics of technical design frequently ask us to separate one thing from another. It's often said that software should separate code from content and that web-page mark-up should separate presentation from content. XML allows us to separate element content from attribute data; well designed XML schemas make clear and consistent decisions about what should go where.

In ontology design, the study of description logics has given us boxes for two types of information, which have been not-so-helpfully named the "A-Box" and the "T-Box". The T-Box is for terminology and the A-Box is for assertions. When you're designing an ontology, an important decision is how much information should be built into your terminology and how much should be left for users of the terminology to assert.

It's not always easy to decide where to draw the terminology vs. assertion line. For example, if you're building a dog ontology, you might want to have a BlackDog class for dogs that are black. Users of your ontology could then make a single assertion that Fido is a BlackDog, saving them the trouble of making the pair of assertions that Fido is a Dog and Fido is colored black. The audience, on the other hand, would have to understand the added terminology to be able to understand what you've said. In one case, the binding of color to dogs is done in the T-Box, in the second, the A-Box. The A/B box choice boils down to a question of whether users would rather have a concise assertion box and a complex terminology box, or a verbose assertion box and a simple terminology.

Although I designed my first RDF Schema over ten years ago, I had not had a chance to try out OWL for ontology design. Since OWL 2 has just just become a W3C Recommendation, I figured it was about time for me to dive in. I was also curious to find out what kind of ontology designs are preferred for linked data deployment, and I'd never even heard of description logic boxes.

Since I gave the New York Times an unfairly hard time for the mistakes it made in its initial Linked Data release, I felt somewhat obligated to do what I could to participate helpfully in their Linked Open Data Community. (Good stuff is going on there- if you're interested, go have a look!) The licensing and attribution metadata in the Times' Linked Data struck me as highly repetitive, and I wondered if this boilerplate metadata could be cleaned up by moving it into an OWL ontology. It could; if you're interested in details, go to the Times Data Community site and see.

It's not obvious which box this boilerplate information should be in. It's really context information, or assertions about other assertions. The Times wants people to know that it has licensed the data under a creative commons license, and that it wants attribution. If it's really the same set of assertions for everything the Times wants to express (i.e. it's boilerplate) then one would think there would be a better way than mindless repetition.

My ontology for New York Times assertion and licensing boilerplate had the effect of compacting the A-Box at the cost of making the T-Box more complex. I asked if that was a desirable thing or not, and the answer from the community was a uniform NOT. The problem is that there are many consumers of linked data who are reluctant to do the OWL reasoning necessary to unveil the boilerplate assertions embedded in the ontology. Since a business objective for the Times is to enable as many users as possible to make use of its data and ultimately to drive traffic to its topic pages, it makes sense to keep technical barriers as low as possible. Mindlessness is a feature.

I could only think of one reason that a real business would want to use my boilerplate-in-ontology scheme. Since handling an ontology may require some human intervention, the use of a custom ontology could be a mechanism to enforce downstream consideration of and assent to license terms, analogous to "click-wrap" licensing. Yuck!

The conclusion, at least for now, is that for most linked data publishing it is desirable to keep the terminology as simple as possible. Linked Data Pidgin is better than Linked Data Creole.

Saturday, July 4, 2009

How Semantic Technology Unified China in the Qin Dynasty

The first Swedish Rap recording was made by the great troubadour Evert Taube in 1960. It's called "Muren och böckerna", and here's a YouTube video for your listening pleasure:

I became aware of this recording from another song called "Evert berättar" by Peter Carlsson and the Blå Grodorna (Blue Frogs). My Swedish isn't that good, but one day a few months ago the song came up on my iPod Shuffle while I was running, and I suddenly realized that the song had something to do with burning books and the Great Wall of China. As soon as I got back home I started researching Evert Taube and Qin Shi Huangdi, the subject of the original song (whose title translates as "The Wall and the Books").

Shi Huangdi (pinyin: Shǐ Huáng Dì, Chinese: 始皇帝 ) means literally, "first emperor". Just as Julius Caesar's name became synonymous with Emperor continuing to the present in titles such as "Kaiser", "Czar" and "Shah", Huangdi was the term used for Chinese emperors for over two thousand years. Shi Huangdi was the king of the Qin state from 246 BCE to 221 BCE, when he became the first emperor of a unified China. Even the word "China" comes from his "Qin" state (pronounced “chin”), even though most Chinese people are really "Han" rather than "Qin".

Shi Huangdi's unification of China put an end to what historians call "the Warring States Period. Under his leadership, the Qin state defeated one rival state after the other. The Warring States period, though politically chaotic, saw a great deal of economic, cultural and technological growth. Iron replaced bronze, and both Confucianism and Taoism (the Hundred Schools of Thought) developed in this period. The Qin state, however, grew strong because of the adoption of a competing philosophy, called Legalism, which emphasized the rule of law in a totalitarian state. Like Caesar, Shi Huangdi extended his dominion by improving communications and implementing standards. He build roads and canals to link the different parts of China. He standardized the length of axles of carts, the units of weights and measures, and the coinage. His most important acheivement, however, may have been the standardization of the Chinese script. For the first time, the machinery of local governments could communicate with functionaries of government throughout the realm. You might say that this was the first semantic web.

Shi Huangdi's innovations were not achieved by gentle persuasion or community consensus, but rather by imperial edict and brutal force. In order to stifle dissent (not to mention the outlawed non-official scripts), he ordered the destruction of all books other than a few in subjects he deemed to be useful: agriculture, medicine and alchemy, and in particular, he outlawed the works of the competing Hundred Schools of Thought. Those caught possessing any of the illegal works were to be conscripted and sent to work on the public works project now known as the Great Wall of China. In many classical accounts, Shi Huangdi ordered 460 scholars to be buried alive, then beheaded.

Although the Qin dynasty of the first Emperor failed to last more than a decade after his death, the non-political aspects of the unification of China through communication, trade, laws, administration and a standard script have lasted more than 2200 years.

Why would a Swedish troubador be interested in Shi Huangdi? Why would he invent a form similar to modern HipHip to sing it in? Evert Taube seems to be most interested in Shi Huangdi's act of burning "all the books in China", so that "history could begin with him". Shi Huangdi exiled his mother because of some "court intrigue" and Taube thinks that burning the books was an act of destroying history, forgetting the his unhappy past, and thinking only of what can be accomplished for the future. It often strikes me that today we're in another period of forgetting the past- because the internet dates back a relatively short time, modern students often behave as if anything that isn't on the internet doesn't exist, and never has existed. There are ongoing monumental efforts to digitize books and bring them back into view; small wars are being fought over how this will occur and all of the combatants claim the banner of preserving history for eternity. Eternal life was also one of Shi Huangdi's obsessions- the famed army of terra cotta warriors he had made was a product of this obsession.

I think that the musical form chosen by Evert Taube is not an anticipation of HipHop, but rather an evocation of the history that society is so eager to forget. Taube had been a sailor and an adventurer, and no doubt had been exposed to the traditional musical forms of both the Far East and of Africa. I think his intent was to evoke the forgotten primitive past with rhythms that speak across the ages.

We live in a time when the language and mechanisms of human interaction are undergoing great change. We are entering an era in which machines are learning to participate in our conversation. Efforts are under way to standardize and unify notations for the real world concepts and entities that underlie our communication. Success in these endeavors may result in the creation of great wealth and power, and new projections of existing wealth and power. It is possible that we are living in a Warring States/Hundred Schools of Thought period, and standardization of our notations will lend itself to a totalitarian communications regime with global extent such as Shi Huangdi's or Julius Caesar's. Another possibility is that our intercourse will become governed by something like a theocracy, in which texts are governed by a priesthood and preserved by monks. Or perhaps information and its underpinnings will devolve to a dictatorship of the proletariat.

On this 233rd anniversary of the Declaration of Independence, I'd like to suggest that a democratically derived and governed semantic machinery for the internet should also be possible. Humans who interact in large groups, such as they are doing in places like Facebook, Twitter and the like, naturally develop languages and syntax on their own, and machines should bow to our will if they are to participate helpfully in our conversations. We need not only a common language and script to be able to communicate with each other, we need liberty to say what we want to say.

Happy Fourth of July!

Wednesday, June 10, 2009

Conference Hashtags Don't Evolve

Leigh Dodds suggested yesterday that someone should do an "evolutionary study" on successful and failed conference hashtags. Since I've been interested in the way that vocabulary propagates on networks, I decided to take up the idea and do a bit of data collection.

If you're not on Twitter, or haven't been to a conference recently, you may not have encountered the practice of hashtagging a conference. The hashtag is just a string that allows search engines to group together posts on a single topic. It's become popular to use twitter as a back-channel to discuss conference presentations while they happen (does anyone still use IRC for this?), or to report on things being discussed at a meeting. The hashtags can also be used on services such as Flickr. I'll be attending the Semantic Technology conference in San Jose next week, and there was a bit of back and forth about what the right hashtag should be, semtech09 or semtech2009. Someone connected to the conference asserted that the longer version was preferred, sparking Leigh's remark.

To collect data, I search Twitter for "conference" AND "hashtag" and compiled all the results from 7 and 8 days ago. That gave me a list of 30 conferences, and I searched for all tweets which included the hashtags for all these conferences. "MediaBistro Circus" (#mbcircus) was the most referenced of any of these, with 1500 tweets in all.

There was no evidence at all of anything resembling evolution. Almost all the hashtags appeared spontaneously and without controversy, each one apparently via the agency of an intelligent designer. I did not find a single example of multiple hashtags competing with each other in a "survival of the fittest" sort of way. I found two instances of dead-on-arrival hashtags, which were proposed once and never repeated. In only one case was there significant usage of an alternate hashtag- the America's Future Now conference appeared in 22 tweets as #afn09, compared to 949 tweets for #afn. Even in this case, there appeared to be no competition, as the #afn09 tweeters stuck with their hashtag.

One question posed by the initial query was whether "09" or "2009" was the correct convention. Although "09" was 2 to 3 times more popular than "2009" in hashtags, having no year indicator at all was twice as frequent as having a year. What was very clear, however, was that avoidance of unrelated hashtags was the clear preference of hashtag selectors. The most interesting examples of this were the "#cw2009" and "#cw09" hashtags used for ComplianceWeek Conference and CodeWorks Conference, respectively.

None of this is surprising in retrospect. It's quite easy to see if a hashtag is the right one- you just enter in a search to see what comes up before you put it in your tweet. If you can't find one that works as you want it to, most likely you will not use a hashtag at all. Conferences tend to be meetings of people who are connected to each other via common interests, and a small number of people tend to update frequently and be followed by large numbers of people with common interests. Conference goers also seem to be very motivated to propagate and adopt hashtags- hashtag announcements for conferences are quite frequently retweeted.

The behavior of people selecting hashtags is quite uniform. However, I've previously noted another meeting of semantic web folk that had trouble getting their hashtag straight. Perhaps people who develop taxonomies for a living are less likely to adopt other peoples' suggested vocabulary based on a feeling that their way is the best way, and thus are more likely to silo themselves. I've seen this phenomenon before- the phones at Bell Labs never seemed to work very well, and my wife's experience with computers at IBM was not exactly problem-free. The Librarian Conferences I've been to seem to have horrible classification systems. Perhaps one way to improve vocabulary propagation on the semantic web is to get rid of the ontologists!

Wednesday, May 27, 2009

Twitterdata and How Chinese Could Be the Future of Tweeting

True confession time: I love Unicode. I think that Unicode was one of the most important achievements of 20th century civilization. I am so much of a Unicode wonk that one of my first thoughts when I heard about Twitter was "I wonder whether it's 140 characters or 140 bytes?" If I were a true Unicode geek rather than a Unicode wonk, it wouldn't have taken until today for me to do the test to see for sure. In case you're wondering, it really is 140 characters; if it were bytes, you'd only get to send 70 Chinese characters in a tweet. But there's a catch- The SMS network restricts SMS messages to 70 characters if any Unicode characters past the first 256 are used. So somehow international tweets are sent as two SMS messages if Chinese characters are used. There's also the catch that tweet recipients may not be equipped to handle the full suite of Unicode characters that you might want to send. I had to change a setting on Tweetdeck before I could see my Chinese tweet; I was unsuccessful in sending a legible Chinese SMS from Skype to my iPhone- not sure where that problem comes from!

I've been spurred into Unicode tweeting because of a recent proposal called Twitterdata. If you've been reading my blog for a while, which I know you haven't been, you'll know that I've been interested in the way that Twittering seems to be developing in ways that resemble the development of human languages. I'm certainly not the first one to notice that Twitter has many semantic-web-like features, and there has been discussion about ways to add semantics to the Twitter stream. The Twitterdata people have made a very interesting proposal: they suggest some very simple additions to tweet grammar that would make tweets more meaningful to machines. They suggest to use the "$" character to denote the name in a name-value pair of meaningfulness. I think this proposal is brilliant, but my thoughts on the matter are entirely irrelevant, because the twitterdata proposal has an approximately zero chance of being widely adopted. My prediction is that one year from now, there will have been more human-generated tweets in Klingon than in Twitterdataese. Here are the reasons I think that:
  1. Twitterdataese is ugly. Example: "@bdelacretaz: #wmodata $id DW1428 $temp 69F $wangle 232 $wspeed 4.0mph $rh 50% $dew 49F $press 1015.2mb http://bit.ly/lxvlh #twitterdata". I rest my case.
  2. Twitterdataese doesn't lend itself well to imitation. In a previous post, I discussed the importance of imitation on the establishement of languages. Without reading the twitterdata documentation, can you figure out what the "$" does in the tweet "@toddfast $likes movies $likes Twitter"? I don't think I would have been able to.
  3. Twitterdata doesn't relieve pain. When I started my first company, a more seasoned entrepreneur gave me some great advice. "People will spend a lot of money to relieve a toothache- but they're much more reluctant to spend money on toothache prevention. Make sure the product you're selling relieves someone's pain." Somehow I doubt that @toddfast would be suffering much if he just liked movies as opposed to $linking them.
So what is causing Twitterers pain? Or to cast the question in terms of language evolution, what are the competitive pressures on Tweet vocabulary and syntax? So far I have been able to discern two strong competitive pressures.
  1. findability- Twitterers want their tweets to be found and read. This pressure is addressed by hashtags.
  2. terseness- Twitterers want to say more in one tweet than permitted by 140 characters. This pressure has led to the proliferation of URL shorteners (another thing I would like to write about) and to innumerable ROTFL and LOL inventions.
Which brings me back to Unicode tweeting. Chinese characters are much terser in terms of character count than any nonideogrammatic language. So maybe, just maybe, we'll start seeing Chinese characters creep into our tweeting to help us say more with fewer characters- first the really easy characters like 中 for china or 山 for mountain or 水 for water. 好吗?

Monday, May 11, 2009

Dancing Parrots and how the Semantic Web will Happen

Last week there was a story on NPR about a dancing parrot. A neuroscientist in San Diego discovered a sulfur-crested cockatoo named Snowball dancing to the Backstreet Boys on a YouTube video. There were a number of interesting aspects to the story, including the fact that YouTube is now being used as a research corpus for animal behavior. What caught my ear however was the mention of follow-on YouTube research by a graduate student in the psychology department at Harvard, which found that the only animals which exhibited dancing skills on YouTube were 14 species of parrot and an elephant. The graduate student notes that like humans — and unlike dogs or cats — parrots and elephants are both known to be vocal mimics. They can imitate sounds. The hypothesis, then, is that our ability to dance is a byproduct of our ability to vocally mimic others.

In the development of language, mimicry is crucial. We learn to speak by repeating what others our saying. We acquire vocabulary by hearing the words that others use. We almost never acquire vocabulary by looking for words in a dictionary.

Last week, I attended a "Semantic Web Meet-Up" in NYC. One of the speakers was describing Knoodl.com, which is described as facilitating "community-oriented development of OWL based ontologies and RDF knowledgebases." I think its fair to say that Knoodl would like to be a sort of dictionary for the semantic web. To my mind, the semantic web just hasn't happened yet because its been very hard to connect data from different knowledge silos. Vocabulary used in one silo tends not to get used in other silos. Ironically, the twitterers attending the meeting couldn't even arrive at a common hashtag for the meeting- at least three tags for the meeting were used. I'm not sure if that fact says much about Twitter hashtags or about the people attending the Meetup. One of the most intriguing things to me about Twitter has been to observe how hashtags are propagated. I find myself mimicking others as I slowly learn the vocabulary and grammar of the new environment. It struck me that it is this quality of Twitter that makes me want to anoint it as a substrate for semantic web actualization.

Maybe the semantic web needs more than just dictionaries and registries and authorities and linked data to become the next big thing. Maybe what it really needs is some dancing parrots. Software agents that have the capacity to mimic the semantics of the other software agents in a global environment.