Showing posts with label language. Show all posts
Showing posts with label language. Show all posts

July 08, 2014

A Quick Look at Global Language Patterns on Twitter

Today's post is derived from some testing we were doing within our data on language and since the results were interesting, we thought we'd share. This is a first step of a longer process of comparing language use at the global scale so much remains to be done.

Starting from a 10% sample of all global geotagged tweets from the calendar year 2013, we collected tweets that used a variety of non-Latin characters as a proxy for linguistic prevalence (see the map titles below for the list of characters searched). Using composite counts of what we found to be the five most commonly used characters in each of the given languages, we mapped normalized values at the country level in order to understand where these languages are most dominant. In other words, these maps represent the relative level of tweets containing non-Latin characters compared to all tweets; the US has plenty of tweets with Arabic, Chinese and Korean characters but these numbers are small compared to the overall number of tweets within the country.  

There are some issues with the data we collected -- for instance, we relied on non-definitive sources for our list of the most commonly used characters, and the constraints of the way we've structured our data makes (how we treat boolean queries and computing constraints) make our data somewhat incomplete. But still the initial results provide a reasonable snapshot of where Twitter is being used by people who don't speak languages which can be easily expressed in Latin characters. 

Arabic Characters:   ل   ن   م   ي   ا      

The spatial pattern of Arabic-language tweeting is interesting in that it seems to mimic a conventional distance decay effect. Saudi Arabia is the undoubted center of Arabic tweeting, with its immediate neighbors having relatively lower amounts, with their immediate neighbors having even lower concentrations, with practically no discernible differences once you reach Sub-Saharan Africa to the south, India to the east, or Europe to the north and west.

Chinese Characters:   的   一   是   不   了

While Japan has the highest absolute number of tweets containing Chinese characters, due to the fact that the Japanese language relies on written Chinese characters, the relative measure shows China to, quite unsurprisingly, be the center of Chinese-language tweeting. The territory of Greenland shows up as well, mainly because of the relatively low number of total tweets making the few tweets with Chinese characters relatively more frequent. We could, of course, account for this by requiring certain thresholds but for this initial look, we left it in. Given the increasing dominance of China within the global economy, it's somewhat interesting to see that there is very little Chinese-language tweeting happening in other parts of the world.

Korean Characters:   뭐   그   안   근데   거

The final language we explored was Korean and while it is not surprising that South Korea has by far the most Korean tweeting, it is interesting to note that North Korea, despite its almost complete disconnection to the global system, also appears on the map. Again, it seems that the scattering of relatively high scores for places such as Greenland and Somalia has more to do with the relatively low level of overall tweeting in these places than with some previously unknown concentration of Korean-speakers.

While there's not much definitive here, we believe this to be a useful, if incredibly brief, look at how online spaces such as Twitter remain connected to conventional, offline geographies, such as those of language and culture. And given the recent emergence of domain names in non-Latin characters, these maps might offer clues into the evolving geography of domain names, while also offering some potential for future research using such data.

May 27, 2014

Hey Y'all! Geographies of a Colloquialism

Here at Floatingsheep, we've spent the last several years trying to demonstrate the potentials, as well as the pitfalls, of using user-generated internet content for geographic research [1]. A key focus has been how the online world of social media at times reflects, and at other times distorts, our understandings of the offline, material world. 


With all the recent hoopla around the geographies of language, we wanted to return to this topic, using a relatively straightforward example: the geography of y'all. No, not the geography of each and every one of you, the geography of the word "y'all" (see definition below). 


Rather than conducting a survey to measure the term's usage, we decided (after careful thought and rigorous debate) to do something new and use geotagged tweets [2]. Searching all of the geotagged tweets in the United States from July 2012 through March 2014 for variations of "yall" (the most commonly used y'all, as well as ya'll and yall to capture typos or alternative spellings), we found a total of 1,870,687 tweets using this folksy second-person plural pronoun, more than enough to make some definitive conclusions (or at least some maps).

Using only the absolute number of tweets with references to y'all to begin, this is clearly a geographically-specific phenomenon. While some places are extremely saturated with references (we'll get to these in just a sec!), there are 250 counties in the United States with no y'all tweets whatsoever, and approximately 60% of the country's 3,143 counties had fewer than 100 y'all tweets in the nineteen month period from which our data originates.

Still using only these absolute numbers of tweets referencing y'all, Texas, Georgia, Florida, North Carolina and California make up the Top 5 states, while the cities of Dallas, Houston, Chicago, Philadelphia and Los Angeles make up the Top 5 metro areas. And while not exactly mimicking population distribution, there is something clearly suspect about believing that folks in Los Angeles, Chicago and Philadelphia say y'all more than good old fashioned Southerners do. So, to make the map below, we instead normalized the county-level data by the total number of tweets originating in those counties during the same time period.

Geotagged Tweets Referencing Y'all, July 2012 - March 2014

On the broadest level, all suspicions and previous research on the matter is confirmed using our normalized tweet dataset: y'all is much more likely to be uttered (or tweeted) in the South than in any other part of the United States... or even the world, for that matter, as there are approximately sixteen times more references to the term in the USA than in the rest of the world combined [3]! But even still, there are some interesting anomalies worth commenting on...

Using these normalized values, we can see a new hierarchy emerge at the state level, with Louisiana, Alabama and Georgia having the highest relative number of tweets, much more in line with what one would expect. At the county level, 97 of the top 100 normalized values are located within the south (by practically any definition). The only three counties outside of this region in the top 100 are Boundary County, Idaho, Dawson County, Montana and Goshen County, Wyoming, the first two of which surprisingly rank #2 and #3 overall in these normalized rankings, led only by Talbot County, Georgia, the epicenter of y'all-related tweeting [4]. But even the South isn't homogenous when it comes to the usage of y'all, as the central Appalachian region of eastern Kentucky, West Virginia and southwest Virginia remains relatively untouched by Twitter references to y'all, despite being more-or-less surrounded by them. Indeed, Kentucky (spiritual homeland of Floatingsheep) is relatively sparse in references to y'all, despite selling these extremely expensive sweatshirts that attempt to capitalize on the state's southern charm.

Apart from some of these slight anomalies, much of this should come as no surprise to anyone who has spent much time in -- or even knows somebody from -- the South. So we thought it might be interesting to compare our own map to a handful of similar maps that have been circulating around the internet recently.

Some Other Maps of Y'all

The first map shows a stark north/south divide between the places that say "you guys" and those that say "y'all" (and, well, Pennsylvania, the western portion of which is also known for its use of "yinz"). The second map, taken from the New York Times interactive dialect quizdeveloped by Joshua Katz, largely resembles our own map, but seems to place the epicenter of y'all much further west than our own, in southeast Louisiana, bleeding over somewhat into Mississippi. 

So while there is some general agreement that Louisiana, Mississippi, Alabama, Georgia and Texas form the territorial heart of y'all, our work, along with the data from the Times' dialect survey, disputes the cut-and-dry story told by the first map. While it shows significant portions, if not all, of Missouri, Oklahoma, Arkansas, Kentucky, West Virginia and Tennessee, among others, firmly in y'all country, the dividing line appears to be both quite a bit further south, and quite a bit more squiggly [5] in nature. While some conventionally Southern states have only relatively confined pockets of references to y'all in our dataset (as well as in the Times' data), it's equally important to recognize that there are pockets of y'all densely concentrated in some more far flung areas of the country as well.

But ultimately as long as you have a group of friends worth using a second-person plural pronoun -- contracted or otherwise -- in reference to, we imagine you're doing just fine. 

Y'all come back now, y'hear?!
-----
[1] Wait, wait, wait... there are pitfalls to this?!?!?!?!
[2] This was also the most convenient data to use, since we had them lying around.
[3] The Bahamas and South Africa come in at #2 and #3 globally in references to y'all.
[4] We suspect that Talbot County, Georgia is the epicenter of exactly nothing else. Although we fully expected that someone from there will angrily correct us very shortly.
[5] That's a technical cartographic term.

September 05, 2013

Schmos, Schmucks and Schlongs, Oy vey!

Oy vey. It has been a very busy, long summer and due to some glitches we've fallen behind in producing posts for the blog like some kind of nebbish. We could kvetch some more but no one likes a nudnik and beside we know all of our readers are real mensches and won't complain and become pains in our tukhus.

Besides, Rosh Hashanah is upon us and we have just enough time to power up the patented FloatingSheep mapping chutzpah and create a special holiday post... Mazel Tov!

And in case you haven't figured it out, today's theme is Yiddish, that wonderfully expressive language of the Jews of central and eastern Europe and more recently (by which we mean the past century) of New York. Drawing from the DOLLY database, aka the golem of the geoweb, we compiled maps of tweets in the USA for the most common yiddish words used in English. Ok, well, Wikipedia complied the list and we made the maps.

Since it is a holiday, we'll keep things short and simple. A key finding is that Yiddish words are alive and well on Twitter within the US, albeit primarily used as single words rather than in whole phrases or sentences. For example, there is a whole lot of "Oy" and "Oy vey" in the Twitterverse. Likewise, the surprisingly long list of Yiddish terms for penis (putz, schlong, schmuck) are running amuk like some kind of meshuggener, which upon reflection makes sense. Nosh is also very popular relative to other Yiddish terms, such as the delightful zaftig which is not as heavily used.

Below, you'll find a series of maps showing how these various terms are distributed across the U.S. Shalom.

Yiddish words are predominantly used in large cities in the US. The map of Yiddish speakers on Wikipedia suffers from the modifiable area unit problem, so not aggregating to the level of the state is more illustrative here.

chutzpah: nerve, guts, daring, audacity, effrontery (Yiddish חוצפּה khutspe, from Hebrew)

kvetch: to complain habitually, gripe; as a noun, a person who always complains (from Yiddish קװעטשן kvetshn 'press, squeeze', cf. German quetschen 'squeeze')

People in the north east kvetch more on Twitter than in other areas of the country.

mensch: an upright man; a decent human being (from Yiddish מענטש mentsh 'person', cf. German Mensch

nosh: snack (noun or verb) (Yiddish נאַשן nashn, cf. German naschen)


oy or oy vey: interjection of grief, pain, or horror (Yiddish אוי וויי oy vey 'oh, pain!' or "oh, woe"; cf. German oh weh

schlep: to drag or haul (an object); to walk, esp. to make a tedious journey (from Yiddish שלעפּן shlepn; cf. German schleppen)

schlong: (vulgar) penis (from Yiddish שלאַנג shlang 'snake'; cf. German Schlange)

There was more intense discussion of schlongs in smaller cities and suburbs throughout the United States.

schmo: a stupid person. (an alteration of schmuck; see below)

schmuck: (vulgar) a contemptible or foolish person; a jerk; literally means 'penis' (from Yiddish שמאָק shmok 'penis', maybe from Polish smok 'dragon')

schmutz: dirt (from Yiddish שמוץ shmuts or German Schmutz 'dirt')

schnoz or schnozz also schnozzle: a nose, especially a large nose (perhaps from Yiddish שנויץ shnoyts 'snout', cf. German Schnauze)

shtup: vulgar slang, to have intercourse (from Yiddish שטופּ "shtoop" 'push,' 'poke,' or 'intercourse'; cf. German stupsen 'poke')

Shtup is used evenly across the country, perhaps as a misprint for "shut up" in conversations, but then again shtuping is a popular activity across time and space.

spiel or shpiel: a sales pitch or speech intended to persuade (from Yiddish שפּיל shpil 'play' or German Spiel 'play')

Spiel is used more in small cities, such as around Marion, Illinois and Sandusky, Ohio.



tush (also tushy): buttocks, bottom, rear end (from tukhus

yutz: a fool 

zaftig: pleasingly plump, buxom, full-figured, as a woman (from Yiddish זאַפֿטיק zaftik 'juicy'; cf. German saftig 'juicy') 

There is no particular pattern of where 'zaftig' is used more, apparently the pleasingly plump are distributed throughout the continental United States.

July 03, 2013

Welcome to 'Merica (or is it 'Murica?)

"Chicken & waffle flavored lays? #Murica."
On this day (well, technically the day before) in which we celebrate our independence from those limey redcoats and their tea-guzzling ways [1], it's time we take on one of the truly great debates tearing at the fabric of our country... 'Merica? or 'Murica?

When dropping the first letter of America (either sarcastically or to preserve our limited supply of vowels), is it more correct to (a) continue as if it were still there and use the term 'Merica? or (b) produce an altogether different word, 'Murica, to express our facetiousness and/or lack of spelling ability?

For instance, the emotionally incensed Twitter user below makes a compelling argument for 'Merica:
"P.s. please stop spelling it #murica or #mericuh or any other variation. It's #MERICA. #northerngirlprobs"
In contrast, this erudite tweeter prefers the more guttural 'Murica spelling:
"I don't know Harry, I heard the French are assholes" true statement. Elated to be back in 'Murica" 
But sadly, there is no consensus around this important issue, which if left unchecked (or at least unmapped) could threaten to undermine the very foundation of the nation. Even more tragic is that someone [2] was so unthoughtful as to bring up this topic on the day in which all 'Mericans/'Muricans should join together in our hatred of everyone who doesn't acknowledge that we're so totally superior to them. As such, we dutifully bring you an investigation of this debate that you may not have even been aware of. You're welcome.

In this endeavor, we collected all geotagged tweets referencing "murica" or "merica" in the United States from July 1, 2012 to June 30, 2013, producing 12,407 references to "murica" and 80,344 references to "merica". If you believe that absolute numbers solve the debate, read no further, as we should obviously err on the side of 'Merica. But if you believe that, you must also believe that "On dit que Dieu est toujours pour les gros bataillons" [3], which we must point out is in FRENCH, and hence your opinion on this day can easily be ignored. Again, you're welcome.

Seeing as there is such a significant preference for 'Merica, we created a normalized measure at the county level to allow for geographic comparison in spite of the massive difference in usage of the terms.  Thus, the maps below illustrate counties' share of tweets for each of the two terms.

For example, Cook County, Illinois had the absolute most tweets for either term, with 201 for "murica" and 782 for "merica". But because its 201 tweets represented 1.6% of all tweets referencing 'Murica, and its 782 were only 0.97% of the tweets refrencing 'Merica, it was determined to have a relatively greater usage of 'Murica, and is shaded as such on the map. So in this first map, the areas that are the darkest shade of red are those places where that place produces a significantly greater share of the overall number of tweets for 'Murica than it does for tweets referencing 'Merica. Confused? You're welcome.

The Misspellings of America

While it might be remarked that this unusual methodology unfairly tilts the linguistic playing field in favor of the much less used 'Murica, we would respond with: who cares? This is our map and we can do what we want with it. Also, we're academics (aka commies) and are totally OK with doing things like changing the rules to benefit the less well-off. Also, note the holiday appropriate color ramp of blues to white to reds. Clever, yes? You're welcome.

As you can see, use of 'Murica tends to be associated with the east and west coasts, with there being fairly little usage of the term, even by relative measures, in the interior of the United States. So it appears that those living in "flyover country" tend to prefer the more simple 'Merica, the coastal elite like to step up their sarcasm an extra notch by exchanging an 'e' for a 'u'.

While some of the country's biggest cities -- Los Angeles, New York City, Chicago, Boston, Phoenix, Minneapolis, Seattle and D.C. -- have a relatively greater amount of 'Murica-ness (or should that be 'Murica-lity), the divide between the two spellings doesn't break down along clear urban/rural lines. Oklahoma City and Indianapolis are two of the biggest users of 'Merica, while parts of the Charlotte and Atlanta metropolitan regions are also on the list of counties who believe that it's spelled 'Merica, not 'Murica.

Indeed, if you further normalize by creating a location quotient -- in effect controlling for absolute size -- a similar picture emerges, albeit one which tends to emphasize the large urban areas much less, regardless of whether they see themselves (or others) as 'Mericans or 'Muricans.

The Misspellings of America (by Location Quotient)

Unlike in the previous map, the most red end of the spectrum here actually shows the places where there is the most parity between the usage of the two spellings, even if there are still a greater number of absolute references to 'Merica than to 'Murica. So less populous counties, or those with many fewer Twitter users, such as Piscataquis County, Maine, with fewer than five or ten overall references to either term, will generally tend to be more red.

But perhaps the most interesting (and actually rather methodologically valid) ways of examining the data is to simply look at a ranked list of the top ten counties for each term. One sees here that the top ten counties for 'Merica are almost exclusively in the South, while the top ten counties for 'Murica are outside the South and within large metropolitan areas. So, our working hypothesis (which we suggest you discuss over beer and burgers on this fine day), is that 'Murica is likely a derivative of 'Merica, used ironically by slow-pour-coffee-drinking, skinny-jean-wearing hipsters in big cities. Our extensive examination of hipsters (n=1) confirms this hypothesis and places the epicenter of this plague somewhere in the Greater Boston area. But you can probably spell it however you'd like.

Happy 4th of July everyone!
-----
[1] No offense intended. Verily, some of the FloatingSheep collective members are British and have yet to make the move to the promised land of 'Merica/'Murica.
[2] That would be us.
[3] "It is said that God is always on the side of the big battalions." -Voltaire

October 31, 2012

Hurricane Sandy and the Geographies of Flooding on Twitter

With the worst of Hurricane Sandy now past, we wanted to build on our initial map of references to "Frankenstorm" and construct a fuller picture of how the storm was represented and discussed on Twitter. The first alternative representation we offer visualizes how Twitter discussed the most obvious impact of the storm, the massive flooding (felt particularly acutely in New York City) that has not only disrupted the every functioning of the city, but also had likely long-lasting impacts on many individual lives and the way we prepare for and attempt to manage such 'natural' disasters.

To begin, we have been collecting tweets containing the terms "flood" and "flooding" in order to examine how Twitter usage might reflect lived experiences of the storm. By examining the digital data shadows of an intensely material event, we can hope to gain some understanding of how the intertwining and interfacing of virtual and material spaces apart from the immediate consequences of this particular event.
An interactive version of this map is available at:

The map reveals a few important findings. First, like the map of references to Frankenstorm, tweets referencing flooding are almost exactly where you would expect them to be; in other words, the vast majority of tweets were located in the path of the hurricane. Nonetheless, it is interesting that so few people elsewhere in the US are tweeting about the unprecedented flooding and resulting damage taking place on the East Coast. In this sense, the geography of data shadows drawn from Twitter appear to be quite effective at reflecting experiences of the storm. The hurricane, in essence, leaves a digital trail.

Second, we are able to see that these data become significantly less useful if we want to draw insights at a scale finer than the county level. Until noon GMT on Tuesday, October 30th, there were only 5,209 geocoded tweets about flooding, a fairly small number over such a broad area. We even initially intended to map references in both English and Spanish to reflect the potential differences in experience between different linguistic groups affected by the storm, but despite the millions of Spanish-speakers undoubtedly affected, we were only able to collect five Spanish-language tweets!

In other words, it is the absences on this map that are almost more interesting than the mapped results. The lack of published content in Spanish means that we are necessarily only including published content from English speakers in these representations. The absences in the rest of the country are also revealing. Why are so few people in Kentucky, Missouri, Wisconsin, etc. tweeting about East Coast flooding? Is it because the act of tweeting about such an event is only really likely to be performed by people in situ, experiencing the storm? Are people outside the direct path of the hurricane interested in other impacts apart from flooding (for instance, the significant snowfall in parts of central Appalachia)? Are they interested at all? Or does the necessarily limited representation offered by Twitter constrain any possible explanations?

July 24, 2012

SheepCamp 2012: Derek Watkins on the Digital Facets of Place

In his talk, Mapping the Digital Facets of Place, Derek Watkins presents a case study of geoweb representations in English and Spanish along the U.S.-Mexico border and how these digital constructions create and influence perception of places along the border.

SheepCamp 2012, Derek Watkins from UK College of Arts & Sciences on Vimeo.

Derek's website/blog: http://blog.dwtkns.com/
On Twitter: @dwtkns

April 23, 2012

The geolinguistic contours of digital content in Spain

Following up on our post about augmented realities and uneven geographies, we wanted to post a few more maps that came out of the project.

This first one compares content indexed in Spanish (Castilian) to content in Catalan. Throughout much of the Catalonian region in the Northeast coastal areas there is considerably more content in Catalan than in Spanish.

The second compares content containing the word "love" in English and Spanish. The map reveals that while the Spanish term is much more predominant overall, there are clusters of locations along the Mediterranean coast at which there are more references to the English word.

These agglomerations are centered in tourism regions of Costa Brava, Costa Blanca, and the Andalusian coastline and closer inspection reveals that these concentration of hits are tied primarily to tourism related references to hotels, restaurants and other activities that are target to non-Spanish visitors.

One key thing that this map does then is reveal how the audiencing of augmentations can be alternately directed to a range of groups: ranging from the highly local (e.g. interpersonal relationships) to the global (e.g. tourist sites).

You can read more about the methods we used and our full conclusions in our new paper: "Augmented Realities and Uneven Geographies: Exploring the Geo-linguistic Contours of the Web."

March 26, 2012

Augmented Realities and Uneven Geographies: Exploring the Geo-linguistic Contours of the Web

Mark and Matt have just had a paper accepted to Environment and Planning A (Augmented Realities and Uneven Geographies: Exploring the Geo-linguistic Contours of the Web). The paper is concerned with the ways in which augmented inclusions and exclusions, visiblilities and invisibilities will shape the way that places become defined, imagined, and experienced.





The maps above are all taken from an earlier draft of the paper. They visualise the layers of information indexed by Google and segment the data by language in order to map some of the geo-linguistic contours of the Web. Have a glance through the paper, and let us know if you have any comments or questions. The publication date of the full paper should be some time in early 2013.

January 03, 2012

Augmented realities and uneven geographies

In the "better late than never category" we offer the presentation that Mark and I gave last September at the iCS-OII symposium. The paper version is available as well is you email me.

October 25, 2011

The Globalization of Beer in the Eurozone

Beer is no laughing matter. Wars have been started over less...and what about bar fights? Granted we're usually hiding under the table when the glasses start flying but we are certainly not laughing.

And while just saying "beer" to the bartender will likely work in most of the world, you don't want to be stuck asking someone for a beer who only knows آبجو (Persian), or asking for a piwo (Polish) when all they've got is 맥주 (Korean). To aid the faithful followers of the Floating Sheep in their ongoing explorations in landscapes of liquid lubrication, we present the following geolinguistic guide to Europe's landscape of beer.

Because simply mapping references to beer in the world's most spoken languages yielded a relatively homogeneous result due to the significant number of references to "beer" and "ale" in English, we thought a more locally specific analysis would be appropriate. So we instead mapped references to beer in twelve languages spoken primarily in Europe that were not included in our earlier map. And while this map obviously doesn't include all of the many languages spoken on the continent, these languages were chosen because of their relative prominence within a larger sample of languages.

Mapping Beer in Europe's (Relatively) Smaller Languages [1]
As we would expect, many countries are dominated by references in their native languages -- Denmark, Poland, the Czech Republic and Hungary display patterns that closely mirror the political borders of the material world.

However, it is the discrepancies where the digital transcends the material expectations that present the most interesting findings. For example, there is an abundance of references in Romanian, even infringing on the virtual territory of Italy, Spain and England (though Spanish and English aren't included in this comparison). While we can have no certain answer, perhaps this is because the Romanian word for beer is "bere", which could of course be an understandable typo for the English-language word. Similarly, Dutch-language references not only fill the entirety of the Netherlands, but also Germany and a not insignificant portion of France. Even Lithuanian references are prominent throughout the Baltic states, despite Estonia's prominence in the global information economy.

Despite the usefulness of this particular grouping, it remains useful to consider how some of the most spoken languages in the world stack up to these more country-specific languages, so in the map below we reintroduce references in English, as well as references in German, Portuguese, Russian and Spanish, to some of Europe's more widely spoken tongues.

The Globalization of Beer in Europe
While this graphic complicates the picture provided by our first map -- there continues to be a significant amount of content in the expected, native languages of each country -- English remains prominent throughout Europe, especially in reference to beer. This could potentially have a number of causes:
  1. Use of English as a second language by many native Europeans in creating user-generated placemarks, signaling the increasing use of English as a global language.
  2. Creation of content in English by native English-speakers traveling throughout other parts of Europe.
  3. Concerted efforts by beer-serving establishments throughout the continent to present English-language content online, so as to attract more English-speaking tourists as patrons.
While these are not testable hypotheses with our current dataset, the results strongly support the idea that the cyberscape of beer is impacted by the forces of globalization, especially in the creep of geolinguistic uniformity. We can only hope that this creep is limited to to linguistics, rather than beer making techniques. At the same time, however, it is also evident that there remains a considerable amount of content in the local languages of many countries across Europe, and it is unlikely that such ties to local language will disappear. This should be good news for the Trappist and Lambic beers of Belgium and the Světlé and Černé beers of the Czech Republic!

So while it is always good to learn the local term for beer, the English word seems likely to get you what you were looking for. Or you can try our technique which is to start a bar fight and while everyone is distracted, grab someone's glass and hide under the table.

---------------
[1] Smaller in that they were not one of the world's ten most spoken languages (by # of native speakers).

October 12, 2011

Wherever You Are, Just Ask for a "Beer"

Now that we've gotten mapping soft drinks out of the way, not to mention other mind-altering substances, it's time we get back to good ole fashioned beer. But rather than mapping different colloquial terms for beer, as we did with pop/soda/coke, we return to our long-standing interest in investigating how different socio-linguistic groups are represented in the geoweb. Only this time, we do it through the proverbial lens of a pint glass (which happens to resemble the geoweb in its distortive capabilities).

The below map shows the relative prevalence of the word for beer in the world's ten most spoken languages (by # of native speakers). However, because of the fact that there were no points at which the number of references in the world's sixth most-spoken language, Bengali, were greater than references to each of the other nine languages, we have excluded Bengali in this particular case. So while we're sad to see Bengali left off the map, the fact that a language with 181 million native speakers has so few references to "beer" is telling of either vast inequalities in the way Bengalis are represented within the geoweb, or perhaps just their general distaste for beer.

B-E-E-R M-A-P!*
While many of our maps are extremely clear in showing that the content within the geoweb reflects traditional state borders, mapping references to beer leaves a much hazier picture. So while most of the content in Russia is in Russian, China in Mandarin, Japan in Japanese, Germany in German and Portugal and Brazil in Portuguese, the cases presented by English and Spanish references are much less clear.

Spanish, the world's second most represented native language, spoken throughout nearly the entirety of the Americas south of the U.S.-Mexico border (not to mention the actual country of Spain), has relatively few references compared to the land mass of countries in which Spanish is spoken. Indeed, outside of Spain, Mexico, Chile and Argentina, there are relatively few points in the world at which there are a significant number of references to "cerveza" compared to references in other languages.

In fact, many places in Latin America have more references in English than in Spanish (or any other language). This appears to be reflected significantly in Europe, as well, where English appears to be the default second language of the geoweb in many of these countries. As a follow-up map will show, when including more nationally-defined languages such as Italian, Polish, Dutch, Danish and Hungarian, these respective countries show a significant preference for their own terms, as one would expect.

When in Europe, Just Ask for "Beer"
Zooming in to Europe only further accentuates the relative dominance of English among these languages, with significant portions of Portugal, Spain, and Germany even showing more references to beer than in their native languages. Interesting, however, that much of France is a mixture of English and German references, even in the much more southern portions of the country.

Whether you want something cheap, something a bit fruity, something hoppy or something to just make the pain go away... our results clearly show that no matter where you may be in the world, it's a safe bet that if you just say "beer", the bartender will probably know what you mean.

Happy drinking, friends!

* Sung to the tune of "Beer Run" by Todd Snider...

June 13, 2011

Distribution of References to Food in Arabic and Hebrew

Continuing our look at the distribution of language in the geoweb, the map below shows the pattern of references to the word food in Arabic and Hebrew. The locations marked in gray are places in which neither language had more references which usually meant both languages had zero. Locations in white are either indicative of water (e.g., the Black Sea) or are places without any placemark references.

The dominance of Hebrew in the Israel/Palestine area corresponds to some of our earlier findings. Continental Europe shows some interesting clusters with much of Italy, Belgium, the Netherlands and parts of Germany contain more Arabic references, while Switzerland, Austria and parts of Germany have more references to food in Hebrew.

References to the word "food" in Arabic and Hebrew, Data from 2010
(Green=more references in Arabic; Red=more references in Hebrew)

June 09, 2011

Geography of Beer by Language

With the summer months upon us, the FloatingSheep Collective is busy with travel and paper-writing and as a result, we've not been posting as much.

This will be changing over the next weeks as we are working on topics ranging from zombies to augmented reality to marijuana pricing to the interaction between material and virtual flows in the economy. We'll be pushing some of this material out later in June and July.

We're also continuing to work on the languages of the geoweb with specific case studies in a range of locations such as Belgium, the corridor between Toronto and Quebec, Kenya, the UAE, France, and Spain. This will likely start coming out in August and September. But to give an initial sense of what we're finding, we offer the following look at languages in Europe...

We searched for the term "beer" in about 70 different languages -- some native to Europe, others from around the world -- to see what kind of patterns we could see. The map below shows the distribution of six languages that we selected to highlight the tight ties between online use of language and offline patterns.

The clustering of references corresponds very closely with the distribution of the speakers of each language, even languages that exist within a state with another dominant language. For example, Welsh appears within Wales but in few other places within the United Kingdom and Catalan is concentrated around Barcelona within Spain. The other interesting finding is that most languages have a micro-cluster of references to beer within Brussels. Whether this is due to the high quality of Belgian beer or the fact that the E.U. is headquartered there remains to be seen.

The Geography of Beer References by Language
(Red=Estonian; Orange=Welsh; Purple=Czech; Black=Italian;
Blue=Castillian/Spanish; Yellowish Green = Catalan)
Note: The size of the circles are consistent within a language but should not be compared between languages. For example, there are many fewer references to beer (or anything) in Welsh than in Italian.

Search Terms for Beer Used in the Map Above

August 16, 2010

Whisk(e)y: to "e" or not to to "e"?

Now that we've mapped both the areas that prefer drinking over eating and which parts of the world prefer which types of alcohol, it seemed apropos to introduce our fascination with language into the cybergeography of alcohol.

It all goes back to an age old question: do you spell whisk(e)y with, or without, the "e"? And while the roots of the disagreement come from linguistic differences between the Scottish and Irish, the reverberations of this debate can be felt at pubs all across the world. So who spells it with the "e"? And who spells it without [1]?

Sadly enough, we are unable to come to come to a definitive answer on which the world prefers. As the amber haze which covers the above map indicates, most of the world seems to have a fuzzy grasp of how to spell the term, with references to each spelling being equal across most of the world [2]. Perhaps they are just too intoxicated to care which way it is spelled?

Interestingly, the few places where references to one or another spelling predominate are not located in the actual places associated with those spellings. Instead, references to "whisky" are most prevalent in South America, while an interesting mixture of the two spellings can be seen in Australia, New Zealand and South Africa. Even Ireland and Scotland seem to be a bit mearbhall about which is spelling is correct.

[1] For what it's worth, Willie Nelson spells it with the "e", but Whisky a Go Go leaves the "e" out.
[2] Places where references to both spellings equal zero were excluded.

July 12, 2010

The Floating Sheep translation project

As many of our readers have pointed out, a key issue with the methods that we have been employing for some of our maps is that we usually only conduct searches in English. With that criticism in mind, we have decided to perform a search for eight words (food, mother, house, land, family, love, water, sex) in sixty-six languages that are commonly spoken in Europe (we are limiting the search to Europe for the time being).

But to do this we need your help. We'd like to ask all of our readers to take a quick look at our spreadsheet of languages and suggest appropriate translations for any of these words in any languages that you know.

We are well aware that in some languages the selection/spelling of words depend upon the context and grammar. So there will be comparability issues, but we'd like to give it a try anyway.

You are welcome/encouraged to record your name in the comments section of the spreadsheet, and we may even send a floating sheep t-shirt to some of the most prolific contributors.

April 28, 2010

Football (or is it soccer?) in nine and a half languages

On Monday we created a map illustrating the geography of virtual references to the words "football" and "soccer". In today's post, we've added eight more languages into the mix: German, Portuguese, Dutch, Spanish, Japanese, Korean, Thai and Chinese. The map below visualizes which of these various ways of referring to "football" are most visible at any particular location in the Google Maps database.
What struck us most was how the map reproduces expected patterns (based on language groups) with very few exceptions: most points in Korea reference the Korean word for football more than the same word in any other language. The same thing is true in Japan, Thailand, Brazil/Portugal and every other country associated with the languages that we conducted this batch of searches in.

Ultimately, Australia wins the prize for having the most homogeneous footballing cyberscape. There is only one place in the country with a reference to football in a language other than English: A reference to Fussball (German) somewhere around the vicinity of Alpine National Park in Victoria. Perhaps there is some sort of odd colony of football playing Germans (is there any other kind?) in this National Park (would any Aussie readers mind checking up on this for us?).

Sweden and Poland are interesting cases: a diverse mix of references to the sport in English (both "football" and "soccer"), German and Spanish, with a small smattering of Dutch and Portuguese. Of course, if we had searched in Swedish or Polish the results would likely have been otherwise.

English appears to be the dominant language for references to the sport in most parts of the world with no direct connection to one of the languages in which we conducted the search (e.g. in Iran, Finland and Russia). We should also point out the the French word for football is "football," so it is difficult to distinguish between references made in English and French using this keyword.

This map is about more than just a sport. We are interested in using this method to study and map cyberscapes in a range of languages. This map was just a first step to test some of the boundaries of the method. We will eventually be mapping a range of other terms in a lot more languages in the near future. Suggestions are welcome.

p.s. This may be a dagger in the heart of many calcio loving Italians, but despite having won the World Cup four times we simply forgot to do a search in your language. Ci scusiamo. We don't know what we were thinking.

March 26, 2010

Multi-lingual cyberscapes: the case of Bangkok, Thailand

For the most part, our maps have been based on the results of keywords in English which has complicated our analysis at the global level (see here and here). Although many English words are used in other languages, it remains an issue. As a result, we have recently begun mapping and comparing the results of searches based in a variety of other languages including those written using characters other than the Latin alphabet.

This post outlines four terms we searched for in Bangkok in both English and Thai.

The first map is similar to other city-scale maps we've published in the past. However, this time we specifically decided to search for references to beer. This map shows results of searches for placemarks containing the word "beer" overlaid on satellite imagery of Bangkok (courtesy of Google Earth). The color and size of each circle indicates the number of results at every given point. Big red circles have the most references and small purple circles have few. Locations without any circles have no references.

A few parts of the city stand out with an abundance of references to beer: the backpacker ghetto of Khao San Rd at the northwestern corner of the map, Patpong in the middle of the map, and the infamous Nana Plaza and Soi Cowboy at the far eastern edge of the map. A cluster of references to beer can also be seen slightly to the north of the main Hua Lamphong train station (the area of white blobs on the map), although we're not really sure why.

"Beer" in English
Interestingly when the same search is conducted using Thai charcters instead of English (เบียร์), a relatively similar pattern is evident. The same areas stand out on the map, although this time Patpong is far more visible. Also evident is the fact that references to beer have a far more dispersed geography in Thai than they do in English. This is likely due to the fact that English speakers are far more likely to reference the few tourist hot-spots in the city, while Thai speakers are more likely to be familiar with a much broader range of places . In any case, we are gratified to see that beer seems to international.

"Beer" in Thai (เบียร์)

This difference in the parts of the city understood and mapped by tourists versus locals can also be observed when a search for "silk" is conducted. Silk is one of famous exports of Thailand and this is reflected in Bangkok's cyberscape. When looking at references to silk in English, we see a map not too dissimilar from the map of beer. Khao San Rd. and Patpong stand out again (likely because silk is often sold in the night markets in both places). The map also picks up references to silk stretching from Patpong down Silom Rd. to the Chao Phraya river and stretching from Nana down Sukhumvit road in both directions (both roads are lined with silk shops that are oriented towards tourists).

"Silk" in English
When the Thai word for silk is used (แพร), a very different pattern can be seen. References to silk are scattered throughout the city without the clear clustering seen in the English (and presumably tourist oriented) cyberscapes of silk. Very different geographies and understandings of place are therefore being constructed between English and Thai representations of the city.

"Silk" in Thai (แพร)

These differences between English and Thai cyberspaces are observable in a whole range of terms. Mapping references to "temple" in English, unsurprisingly highlights the Grand Palace Complex and the many temples in the Phra Nakhon District of the city.

"Temple" in English
A search for the Thai word for temple (วัด) highlights entirely different parts of the city. Instead of the Grand Palace area, we see a focus on the Temple of Dawn and the many other temples on the eastern bank of the river. There are also clusters of placemarks around the Golden Mount, Chinatown and the temples in Sathon (e.g. Wat Yan Nawa). Again these are locations in the city that are more likely to be known and frequented by Thais than foreigners.

"Temple" in Thai (วัด)
Finally, we wanted to look at the geographies of an industry that is far more visible in Bangkok than in most parts of the world: the sex industry. We started with a search for the term "brothel" in English. We see a few dominant clusters showing up on the map: most notably the famous red-light districts of Nana and Patpong.

"Brothel" in English
But when searching for brothel in Thai (ซ่อง), a radically different geography can be seen. The Chong Nonsi area especially stands out (a part of the city not particuarly famous as a red-light district).

"Brothel" in Thai (ซ่อง)
We recognize that ensuring that our terms are equivalent in both languages is problematic and would welcome thoughts on ways to improve our searches. For example, would English speakers use the term brothel or sex club? Or something else?

Nevertheless what this shows, is that very different representations of a city are being produced and reproduced for different people. Our understandings of the cities we move through are heavily influenced by representations of place and we therefore need to find useful ways to map and understand the fluid and sometimes fleeting representations that exist on the internet.