Showing posts with label digital divide. Show all posts
Showing posts with label digital divide. Show all posts

April 13, 2012

Where do Tweets come from?

Last week, we posted a map of all georeferenced tweets mentioning the #Kony video. The patterns were interesting, but not entirely unexpected.

A more interesting question though, would be to see what percentage of all tweets from each country reference #kony, in order to get a better sense of how focused people were on the event. However, to do that, we need to figure out how much content in Twitter actually comes from each country.

Mark and Devin Gaffney collected all georeferenced tweets sent between March 5 and March 13 (it is important to point out that we are only dealing with a very small percentage of total tweets here [less than 1%], and so there may be significant geographic biases in where/how people georeference their content). We then took a random 20% sample of that dataset: giving us about 4.5 million tweets that we spatially joined to countries. The results are below:


(countries are on the x-axis)



The bar chart shows us the degree of inequality in where this content is coming from: with people in a few countries producing the bulk of content, and then a very long tail of countries from which very little content is produced.

Interestingly though, it is not just the usual suspects that are producing the bulk of content. The top six tweeters are:

(1) USA
(2) Brasil
(3) Indonesia
(4) UK
(5) Mexico
(6) Malaysia

Only two of the countries on that list are in the Global North and traditional hubs of the production of codified knowledge. What does this all tell us then? It is possible that Twitter is truly allowing for a 'democratisation' of information production and sharing because of its low barriers to entry and adaptability to mobile devices.

However, we need to do more work in this area to really figure out where content is coming from in the platform. Our sample in this post was limited, and more importantly we are still only dealing with georeferenced tweets that make up less than 1% of the total content that passes through the platform. An interesting start nonetheless.

**We'll post the Kony map normalised by the number of tweets in each country soon.

March 20, 2012

Geographies of the World's Knowledge E-book Now Available for Tablets

Our booklet, "Geographies of the World's Knowledge", is now available for iPads from Apple's iTunes store. The publication is free and is optimized for tablet viewing (we've included lots of cool interactive features). If you have a tablet, I highly recommend you check it out!


If you don't, you can always download our PDF version in both English and German. Let me know if you have any questions/suggestions.

March 10, 2012

Big Data and the End of Theory?


The Guardian just published a short post by Mark which looks at the discourses surrounding 'big data.'

In it he argues that:

Gender, geography, race, income, and a range of other social and economic factors all play a role in how information is produced and reproduced. People from different places and different backgrounds tend to produce different sorts of information. And so we risk ignoring a lot of important nuance if relying on big data as a social/economic/political mirror.

We can of course account for such bias by segmenting our data. Take the case of using Twitter to gain insights into last summer's London riots. About a third of all UK Internet users have a twitter profile; a subset of that group are the active tweeters who produce the bulk of content; and then a tiny subset of that group (about 1%) geocode their tweets (essential information if you want to know about where your information is coming from).

Despite the fact that we have a database of tens of millions of data points, we are necessarily working with subsets of subsets of subsets. Big data no longer seems so big. Such data thus serves to amplify the information produced by a small minority (a point repeatedly made by UCL's Muki Haklay), and skew, or even render invisible, ideas, trends, people, and patterns that aren't mirrored or represented in the datasets that we work with.

Big data is undoubtedly useful for addressing and overcoming many important issues face by society. But we need to ensure that we aren't seduced by the promises of big data to render theory unnecessary.
We may one day get to the point where sufficient quantities of big data can be harvested to answer all of the social questions that most concern us. I doubt it though. There will always be digital divides; always be uneven data shadows; and always be biases in how information and technology are used and produced.

And so we shouldn't forget the important role of specialists to contextualise and offer insights into what our data do, and maybe more importantly, don't tell us.

You can check out the full piece here.



February 09, 2012

Cape Town Cyberscapes: Khayelitsha and the digital divide

For a recent project with Professor Stan Brunn, we updated and expanded our visualization on the cyberscape of Cape Town, South Africa (the original version from 2009 is here). Again this map was based on the amount of geo-coded material indexed in Google Maps using a fine grid of points approximately 1/10 of a mile apart.

This time around we were particularly interested in Khayelitsha, an informal (and fast growing) township in the Cape Town area. You can see it in the lower right of the map below (it is highlighted and vaguely boomerang shaped). The main take away from the map is the clear difference in amount of geo-coded material in Khayelitsha versus other richer, whiter parts of the region.

Map generated by Jeff Levy

Based on our previous work, this is entirely unsurprising. Nonetheless, it remains useful to visualize these inequalities at the metropolitan level in order to demonstrate that the digital divide in user-generated content operates at a variety of scales, and that even the most populated areas in terms of content are surrounded by areas with very little.

July 05, 2010

Wikipedia and Internet Use

The following map displays the total number of Wikipedia articles normalised by the number of internet users at the country level. The countries with the highest number of articles per 100,000 internet users are Nauru (4667), the Central African Republic (1253) and Myanmar (824). In fact most of the places that score highly by this measure, like the countries listed above, have extremely low levels of internet use per capita.

In contrast, countries with higher level of per-capita internet usage tend to have far lower rates of Wikipedia article per 100,000 internet users (e.g. the United Kingdom (70) and France (67)). While it is entirely possible that the high rates of articles per internet users in some countries is an indication of dedicated Wikipedia editors, it seems instead more likely that Myanmar, the Central African Republic and most other nations with low levels of internet penetration are being represented by editors from outside of their boundaries.

June 09, 2010

2010 Internet Penetration Rates

Today's post comes courtesy of data available from Internet World Stats. The map below presents the most recent statistics on global internet usage. The shading reflects the proportion of the population that uses the internet within each country. The height of each bar indicates the total number of internet users in each country.

Iceland has the world's highest penetration rate: over 93% of the population are internet users. Almost all of Europe and North America also have relatively high rates (at least at the national scale, as there are likely to be significant digital divides in every country). China, interestingly, is already home to the world's largest population of internet users (384 million) despite having a penetration rate of less than 30%. India is another interesting case. 81 million Indians are internet users (there are more Indian internet users than there are people in the UK), yet this represents only 7% of the Indian population.

June 03, 2010

International Internet Bandwith

Today's map displays international internet bandwidth globally. "International bandwidth" is another way of referring to the contracted capacity of international connections between countries for transmitting Internet traffic. These data are kindly made available from the World Bank's new open data initiative.

Like most other geographies of Internet-related data, the patterns in this map are highly uneven. Countries in northern Europe generally have the most available kilobits per person. The Netherlands has 78kb per person, Sweden 50kb, and the UK 40kb. A number of micro-states and small nations also score highly on this measure: Hong Kong (not displayed on the map) has 315kb per person, Singapore has 23kb, Antigua and Barbuda has 17kb and Panama has 16kb. Surprisingly, the United States has fewer available kilobits per person than any of these countries (11kb).

At the other end of the scale, there is a long-tail of countries in Africa, Asia and South America that have less than 1kb per person. Guinea, for instance, has only 0.21 bits (0.00021kb) per person (our next post will focus specifically on bandwidth in Africa).

These data seem to mirror the geographies of content at the global scale, a topic we plan on exploring in much more detail in a future paper.

March 19, 2010

How does the density of placemarks vary across space?

One of the most fundamental questions in our research is also one of the most basic. How does the density of placemarks vary over place? Back in June 2009, we took an initial look at information inequalities but had to rely on keyword searches for "0" and "1" (based on the assumption that there would be no particular spatial bias to these terms) as proxies for the total amount of content produced about a place. It worked fairly well but was less ideal than we hoped.

Recently it became possible to conduct wildcard searches (using the "*" operator) and this post revisits the same question, How does the density of cyberscape vary across locations? We conducted a wildcard search at approximately 260,000 points on the Earth's surface and collected the total number of placemarks indexed there. As always, a direct observation is preferable to a proxy measure so we're quite excited by these maps.

One sees that the United States contains the most placemarks (77 million) with almost twice as many as China which has 43 million. The only other countries that also have over ten million placemarks are the usual suspects when it comes to technology use: Germany, Japan, the UK, France and Italy. However, looking at the raw number of placemarks per country only tells part of the story. So, we decided to normalize these data by population and area. In doing so, some interesting patterns emerge.

Most countries in western Europe have extremely high levels of user-generated content per person despite having fewer placemarks than countries like China or the US. Denmark in particular stands out as having the world's highest ratio of placemarks per person. We're not sure why the Danes are so well represented in cyberscapes. Perhaps Danes have the perfect combination of high levels of disposable time and income to allow them to engage in the construction of user-generated content (the country has the world's highest level of income equality, a large welfare state and one of the highest levels of internet access). An alternate theory (which we're not putting a lot of store in) rests on the well established fact that all things internet-related can usually be explained by pornography. Denmark was the world's first country to legalize pornography and, as such, it stands to reason that they have a head start when it comes to producing content for the internet. We should point out that we haven't yet had a chance to explore the actual content that the Danes are producing.

Moving swiftly on, it is remarkable that China, despite being home to 1.3 billion people, continues to have a relatively high ranking when the data are normalized by population. The finding is a testament to the enormous amount of content being created about China. Interestingly in many of our maps so far, China has not shown up very strongly but this is likely connected to our focus on English search terms. For instance, we're currently searching using the Chinese characters for temple which is producing some interesting patterns that are also much denser than the searches on the English word temple.Finally, we decided to normalize the data by area. Here, very different patterns emerge. Small, densely populated countries like the Maldives and Singapore rise to the top of the list. Much of Europe as well as Japan and South Korea also stand out as having a large number of placemarks per square kilometre.

These maps show that there is no single way to represent the multiplicity of the world's cyberscapes. Depending on the particular way that these cyberscapes are measured and normalized, some quite different results can be found. And yet, irrespective of how the data are measured, a general 'digital divide' can be observed in these virtual representations of place. Western Europe, North America and parts of East Asia are represented by a significant amount of virtual content, while much of the rest of the world (in particular most of Africa and the Middle East) remains, both literally and figuratively, off the map.