Showing posts with label twitter. Show all posts
Showing posts with label twitter. Show all posts

December 18, 2014

Deconstructing the (most detailed tweet) map (ever)

If you’re the kind of person who visits our blog with any regularity, you’re almost certainly also the kind of person who would have seen some version of the map below in the last couple of weeks. Created by Eric Fischer of Mapbox, this map was released along with a blogpost entitled “Making the most detailed tweet map ever”, discussing some of the data cleaning and visualization methods necessary to produce such a striking map. The map is undoubtedly interesting and has sparked a great deal of interest from all corners of the internet, but there’s just something about the framing that rubs us the wrong way. While Eric’s post emphasizes the making part of the equation, the internet hype cycle around it has caused us to read the title a bit more along the lines of:

"Making THE MOST DETAILED tweet map EVARRRR!!!!"

That is to say, for all of the admittedly really great detail about what went into making this map, the framing of this map as not only a detailed map of six billion or so geotagged tweets, but as the most detailed tweet map ever, raises more questions than it answers. For example, what constitutes ‘detail’ in tweet maps? What do competing definitions of ‘detail’ reveal about what we value in this kind of analysis? What do these particular ideas of ‘detail’ foreclose in terms of other possibilities for analysis?

These are important questions, regardless of whether they’re applied to this particular map or any other one. The issue in this case, however, seems to be that the answers to some of these questions conflict with one another, or with the ways the project is itself described. The detail that seems to be valued here is of the “every tweet ever” variety, or, put simply “more = better”, the fetish for bigger data at the expense of all else.

But more data isn’t necessarily better, and it certainly doesn’t mean that there’s more detail, especially when the only bit of detail you're concerned with in each of these six billion points is the latitude and longitude coordinates. Each of these individual tweets contains a wealth of other interesting information, from information about the user and the way they describe themselves, to the time the tweet was created to the text of the tweet itself, which might contain hashtags that link up with bigger conversations, or @-mentions to other Twitter users that might be used to understand social networks and interactions. All of these bits of information represent a kind of detail that is not included in this, the most detailed tweet map ever

As we’ve been arguing for the past two years or so, there are a range of social and spatial processes represented in geotagged tweets that we can’t get at if all we’re concerned with is the latitude and longitude coordinates. So to say that this represents the most detailed tweet map ever serves to reify what we see as two of the most problematic assumptions of contemporary big data/social media research: (1) that more data is equivalent to better data, and (2) that the only important aspect of the data is the geographic coordinates attached to it. There's lots of interesting stuff that can be done with this kind of data, and we can do better than simply plotting points on a map and calling it a day [1].

Even if one were inclined to accept the argument that more tweets equals more detail, how should we interpret the fact that this map only visualizes about 9% of all geotagged tweets, due to the design decisions necessary in order to make the map nice and pretty [2]? Due to the existence of exact or near-duplicate coordinates that would make points indistinguishable from one another, this, the most detailed tweet map ever, actually eliminates about 91% of the detail that it seems to value most (i.e., the presence or absence of points on the map). The Gizmodo headline about the map reads, “The Most Detailed Tweet Map Ever Includes 6,341,973,478 Tweets”... except that, you know, it doesn’t [3].

Of course, there’s also good bit of imprecision in the locational accuracy associated with geotagged tweeting; our iPhones don’t come with military grade GPS units installed in them. So while Mapbox CEO Eric Gunderson was marveling at the detailed micro-geographies of an airport gate seen in the map, he was ignoring both the fact that all of those folks on the jet bridge could just have well been 40 feet away, and that a number of tweets might have been eliminated from the initial dataset due to a lack of precision in the geotagging process. Take all of that together and a lot of the detail that’s being celebrated here starts to give way to fuzziness. This map is more art than science, though the striking visuals and discursive framing give the illusion of precision and absolute insight. 

To be clear, there’s no problem with fuzziness. It’s something we all live with every day, it’s something we academics may embrace from time to time through the use of overly obtuse language. But taking all of this fuzziness and then repackaging it as the most detailed tweet map ever, comes off a bit wrong to us. These initial misgivings were only amplified when brought down to a more local level, when we saw a post from a local urbanist blog in Louisville wondering “What we can learn from where people in Louisville are using Twitter”. While relatively mundane, and certainly not nearly as celebratory, the blog’s ultimate conclusion was that "These locations [with the highest concentrations of tweets] make sense as they are places where people gather and are often held captive by events.”


This, in general, is true, but also a bit… how do we put it? Meh. More fundamentally, people tweet where people are. It comes as no surprise to anyone with even the vaguest familiarity with Louisville that people tweet in larger numbers from downtown (including 4th Street Live!), the University of Louisville campus, Bardstown Road and the St. Matthews / Oxmoor Mall area than anywhere else in the city. These are (some of) the primary gathering points on a day-to-day basis within the city.

But just identifying these locations doesn’t really help us to ‘learn’ anything beyond the fact that those are, indeed, the places with the highest concentrations of geotagged tweets in Louisville [4]. In fact, the map doesn’t even really show us actual concentrations of tweeting activity, but rather concentrations of unique tweeting locations. Take, say, two hypothetical city squares, one of just 50 x 50 meters, and another much larger one of 500 x 500 meters, both the originating point of one million geotagged tweets spread randomly over the squares. In Fischer’s method, these two squares would not 'glow' in equal amounts, but rather the larger square would show up as much more visually prominent because it has many more unique tweeting locations while many of the tweets from the smaller square would be filtered out due to a duplication of coordinates.

Further, from a data collection standpoint, all of these tweets in Louisville reveal little that isn't revealed by mapping a random sample of tweets (say 1% of tweets from 2013, see map below). If all we’re really concerned about is the question of where people are tweeting from, there isn’t much that looking at all the tweets reveals that couldn't also be found from a smaller subset, and it’s much easier to collect or analyze a few hundred thousand tweets than it is to collect 6,341,973,478 of them. But even still, all we can ‘learn’ from these kinds of maps is where people have created geotagged tweets and, to some extent, where they have not [5].


But if that’s all we can learn from this map, again, why call it the most detailed tweet map ever? Again, there are any number of details that are excluded from analysis by only looking at the locations of geotagged tweets. What if we instead took a different approach to this data, such as examining at the use history of individual Twitter users, or even collectives of Twitter users based on some kind of shared experience or identity, such as association with particular neighborhoods or other places?

OK, you're right. This particular question is a bit self-serving, as this is precisely the kind of thing we've been working on for some time now. And so rather than just offering a critique of someone else's work, we really want to see if we can push this kind of analysis in more productive directions. So we offer up the map below, which comes from a paper we currently have under review, that attempts to demonstrate how geotagged tweets can help us to better understand urban socio-spatial inequality beyond simply identifying the presence or absence of tweets in a given area, as is so often done.


Using Louisville and the now-common ‘9th Street Divide’ trope as a starting point, we sought out to understand how people from different parts of the city used and moved around the city in different ways. So in a manner not uncommon to some other things Eric Fischer has done previously, we identified a number of Twitter users as belonging to one of two groups, those with close ties to either the West End (traditionally a poorer and predominantly African-American part of the city) or the East End (a more affluent and largely white part of the city), and collected all of the geotagged tweets from those users [6]. We then compared the spatial footprint of these groups' tweeting activity via an odds-ratio measure. On the map areas in purple represent places with greater-than-usual levels of West End user tweeting activity, while orange hexagons represent places where East End users were relatively more dominant than expected. Those places which demonstrate roughly equivalent or expected levels of tweeting are signified by those hexagons with hashes.

This map, in short, represents those places in the city of Louisville which are more socially heterogeneous and homogeneous, dominated either by West End or East End residents, or characterized by a relative mix of people from parts of the city. Though it’s evident that there is indeed a kind of divide between the West End and the rest of the city, this map also shows that West End residents are incredibly spatially mobile within the city, while East End residents tend to be much more spatially constrained, sticking to their own parts of town.

While there are certainly a lot of underlying factors driving this process, suffice it to say that this map provides an alternative way of understanding socio-spatial inequality than simply identifying those places that do or do not have significant concentrations of geotagged tweets [7]. Through our analysis, we also learned that contrary to the kind of assumptions often made about this kind of informational inequality, West End users actually produce a significantly greater number of geotagged tweets than their East End counterparts, it's just that many of these tweets are created in other parts of the city. This is, of course, an important kind of detail that we can draw from the mapping and analysis of geotagged tweets and one that, in many ways, is more detailed than the most detailed tweet map ever.

There is, of course, a whole lot more detail in the paper that this one map and blog post can’t capture, just as is the case with Eric Fischer’s map. Just to be clear, we think Eric Fischer does some fantastic and beautiful work with geotagged social media data, and commend him for openly discussing and sharing his methods. And yet, we can’t help but feel like the characterization of his map as being the most detailed tweet map ever is at best a half-truth, and helps to reproduce some of the most common problems with the analysis of geotagged social media data. But the more we think about it, we’re not so sure that a single most detailed tweet map could exist, or that it’s even desirable to have such a thing. Instead, we should be striving to create any number of highly-detailed, geographically-situated tweet maps, that collectively contribute to better understandings of the complex social and spatial processes that are represented and reproduced through this kind of data. 

----------------
[1] That’s the royal we. 
[2] Which it most certainly is.
[3] As Fischer notes, there are actually no more than about 590 million dots on the map due to his filtering process. When one zooms all the way out on the map so that the entire globe is represented in a single map tile, there are only 1,586 visible tweets, a far cry from the 6 billion number that seems so, well… big.
[4] #tautology
[5] This is qualified in this way because, as Kenneth Field pointed out in a Twitter exchange with Eric Fischer about these maps, geotagged tweets that he has consciously created from his house do not appear on the map. So while we know that all of the tweets on the map were created in that place, we can't say definitively that tweets were not also created in places where they do not appear on the map.
[6] In order to do this classification, we collected all geotagged tweets created within the defined boundaries of these two areas, and then identified those users with more than 40 tweets within either area, where those 40+ tweets represented greater than 50% of their overall geotagged tweeting activity. This concentration of activity indicates that users had a strong association with, and presence within, either area, while also making sure that no users were identified as belonging to both areas.
[7] We also see this map as complicating the conventional narrative in Louisville of 9th Street as representing a kind of impenetrable barrier within the city. But since this is less directly relevant to our argument here, we'll make you wait to hear more about that particular line of reasoning.

April 02, 2014

New Book Chapter on the Geographies of Beer on Twitter

We're pleased to announce a new publication by members of the Floatingsheep team. Just released is "Offline Brews and Online Views: Exploring the Geography of Beer on Twitter", a new book chapter written by Matt and Ate that analyzes the geographies of beer-related tweeting activity. Published in a new edited collection from Springer appropriately- and straightforwardly-entitled The Geography of Beer, Matt and Ate's paper -- the latest in Floatingsheep's long line of investigations into the geographies of beer -- shows that geotagged tweets about beer, and other alcoholic beverages for that matter, are reflective of people's offline consumption preferences.

Using a database of one million geotagged tweets from June 2012 to May 2013 containing the keywords "wine", "beer" or the names of a range of light or cheaper beers within the continental US, some clear regional variations in alcoholic beverage preference are detected. For instance, when comparing tweets referencing "wine" to those referencing "beer", wine-related tweets tend to be more dominant along both the east and west coasts of the US. But this kind of variation is present even when comparing different brands of light beer. While Bud Light is more popular in the eastern and southeastern US, Coors Light tends to dominate the west coast, with Miller Lite and Busch Light being preferred in the midwest and Great Plains. The dominance of these brands in virtual space is no surprise, as they also dwarf the competition in actual sales.

But these regional variations are even more distinct when one looks at locally- or regionally-specific brands. While some of these cheaper (which is not to say less delicious!) beers have reached a national or even international market, others remain popular in only a very limited region, owing either to local tradition or simply limited distribution outside of their home-markets. Nonetheless, by mapping the concentrations of geotagged tweets referencing each of these brands, we're able to uncover these regional particularities, as is shown in the map below, taken from Matt and Ate's chapter.

Aggregated Geographies of Tweets referencing Regional 'Cheap' Beers

From Sam Adams in New England to Yuengling in Pennsylvania to Grain Belt and Schlitz in the upper Midwest, these beers are quite clearly associated with particular places. Other beers, like Hudepohl and Goose Island are interesting in that they stretch out from their places of origin -- Cincinnati and Chicago, respectively -- to encompass a much broader region where there tend to be fewer regionally-specific competitors, at least historically. On the other hand, beers like Lone Star, Corona and Dos Equis tend to have significant overlap in their regional preferences, with all three having some level of dominance along the US-Mexico border region, but with major competition between these brands in both Arizona and Texas.

Beer, like many other social practices, may be millennia-old, but the socio-spatial practices associated with it – checking into a brewery, posting a review, geotagging a photo – continue to evolve with technological change. As such, this kind of data provides an important way to capture these socio-spatial practices and preferences, while demonstrating how even in an era of supposed globalization and homogenization, regional histories and cultures continue to be reflected online in important ways.

If you don't have access and would like to read more about this, please contact Matt at zook [at] uky [dot] edu for a pre-publication version of the chapter. Bottoms up!

The full citation for Matt and Ate's chapter is below:
Zook, M. and A. Poorthuis. 2014. "Offline Brews and Online Views: Exploring the Geography of Beer Tweets". In The Geography of Beer, eds. M. Patterson and N. Hoalst-Pullen. Springer. pp. 201-209.

September 18, 2013

What do Twerking and Syria have in common? Not much, except for the Twerking Tramp Stamp of America

User-generated data, especially generated via social media, provides a useful look into the day-to-day experiences and conversations that occur around the world. This kind of data provides insights, however partial, into our hopes, our triumphs and our fears. Sometimes it reveals things that we would rather not think about, like hate speech or racially charged discourse. But in all cases, it enlightens us to what matters to people, or at least what matters to the people participating in online discussions.

The past month has seen a sharp increase in two topics of discussion, the first representing an African American cultural meme from the bounce music tradition of New Orleans but more recently (and cynically) appropriated by white pop artists, and the other tied to an ongoing conflict that rapidly garnered calls for international intervention. We speak, of course, of twerking [1] and the ongoing civil war in Syria. While the two topics share little in common apart from recent media attention, the divergences between the space-time patterns of these geo-coded tweets show how online and offline actions and characteristics are intricately and imperfectly connected.

Using DOLLY, we extracted all geocoded tweets from July 1, 2012 to September 11, 2013 from North America that referenced either "twerk*" or "syria*". It quickly became apparent that twerking is a much more popular topic on Twitter, with 775,000 geocoded tweets during this time period, while there were only 75,000 references to Syria [3]. These numbers have changed significantly over the past month as the U.S. has called for military strikes in response to reports of chemical weapon attacks in Syria, but nevertheless there have still been three times as many references to twerking as Syria in August and September, with 133,000 tweets against just 43,000.

Indexed Volumes of Twerk and Syria Tweets, July 2012 to August 2013
The evolution in volume of each kind of tweet overtime is also strikingly different.  Compared to the number of mentions in July 2012, twerking has steadily become a more popular topic of Twitter conversations over the course of the past thirteen months. In contrast, discussion about Syria largely declined over the course of the year, and only in August did it became a hot topic. Since July 2013, the relative amount of conversation about Syria increased almost tenfold, from an index value of 61 to an index value of 526. Though they took much different trajectories, both topics have around five times as much discussion as was the case a year ago.

Looking at the geography of these tweets, just for the month of August 2013, illustrates how tweeting behavior varies across space during a time when there was a lot of national attention to both topics. Using a simple ratio of the # of Syria Tweets / # of Twerking Tweets, which we term the Twerkyria Index, the maps below show this distribution at the state and county levels.

The Twerkyria Index by State, August 2013



One of the most compelling results is the clear difference in ratios for Washington D.C., which has three times as many Syria tweets as twerking tweets, bucking the national average which is three to one in the opposite direction. The next closest areas are Vermont and Alaska, with relatively small African American populations, which have roughly the same number of tweets for each topic. The rest of the country is divided into red states that have more than the national average of twerking and pink states that are slightly less twerking obsessed (at least relative to attention to events in Syria). The concentration of states in the southeast -- from Texas to South Carolina -- forms a Twerking Tramp Stamp across America [4], with a few other concentrations tastefully tattooed across the Great Plains and Midwest.

The pattern suggests a number of possible connections between the Twerkyria index of Twitter activity and offline demographics. Despite Miley Cyrus's "act of cultural appropriation being passed off as a rebellious reclamation of her sexuality after a childhood in the Disneyfied spotlight", twerking's roots are within
African American culture, especially as it relates to southern hip-hop. In this context, sending a tweet containing twerk is likely much more about cultural expression of local and identity politics than the more recent appropriation of the work by the dominant culture to work through "a raft of personal, socioeconomic and third-wave-feminist issues". Likewise, tweets containing Syria represent a wide range of political views and stances towards possible U.S. intervention.

In short, it's complicated. Far from a simple unitary meaning, the use of both twerking and Syria on Twitter are complicated expressions of cultural and political expression in the U.S., relating these locations both to particular regional cultures within the country and particular geopolitical configurations that span the globe.

In an effort to test some of these relationships, we ran a quick (and relatively crude) OLS regression at the state level with the Twerkyria index as the dependent variable and a range of demographic variables including:
  • Population under 18 years, percent, 2012  
  • African American Population, percent, 2012 
  • Median value of owner-occupied housing , 2007-2011
  • Population per square mile, 2010 
The model excludes Washington DC, given that it is a city rather than a state [5] and includes dummy variables for Hawaii and Alaska given their unique spatial position vis-a-vis the lower 48 states.

Modelling the Twerkyria Index

While there are any number of ways to critique/improve upon this basic model, it works well for simple illustration [6]. The model explains about 65% of the variation in the Twerkyria Index. States with younger populations and a higher percentage of African Americans are associated with more twerking tweets. Interest in foreign affairs is more difficult to measure (at least with standard Census data), but a combination of housing price and population density provides a measure of a state's urbanity and presumed interest in international affairs. Locations with more expensive real estate and higher population densities, i.e., more urbanized states, have relatively fewer tweets about twerking and more with Syria. This model shows that the Twerkyria Index correlates fairly well with some reasonable theoretical expectations about the nature of the offline demographics of the points of origin of these tweets.

The Twerkyria Index by County, August 2013

The state level, however, masks many of the subtle distinctions that emerge within these relatively large spatial units.  Examining the Twerkyria index at the county level (see below) shows that the higher number of Syria tweets in Washington DC is also evident in a number of counties, roughly equal in size, to the district. While these counties are scattered across the country, a particularly large concentration is found in the San Francisco Bay with Alameda County, containing Oakland and Berkeley, representing the source of fully 8% of all Syria-related tweets. Given this large concentration, it is not surprising, that the Twerkyria Index diverges greatly from the rest of the country. In contrast, southern California largely conforms with the larger national trend of tweeting considerably more about twerking.

In summary, we have no easy answer to the age old, "To twerk, or not to twerk" [7], but looking at the socio-spatial dimension of online activities provides useful insight on the complicated interconnections between our online and offline activities.

-------------------------
[1] For those few unaware of what twerking is, we refer you to the Wikipedia definition, which defines it as "a type of dancing in which the dancer, usually a woman, shakes her hips in an up-and-down bouncing motion, causing the dancer's buttocks to shake, "wobble" and "jiggle".  If you still have trouble understanding it we suggest this thoughtful overview provided by the New York Times. After all, where else would one go to truly understand an artifact of twenty-century African American urban culture than the grand grey lady of journalism?  Or you could view the three videos at YouTube with the most hits 1) Booty Me Down Song By Kstylis; 2) How to Twerk and 3) Jimmy Kimmel Reveals "Worst Twerk Fail EVER - Girl Catches Fire" Prank".

We're still working on obtaining video of someone from the FloatingSheep collective twerking but an array of technical and legal difficulties have prevented this thus far. [2]

[2] Also the fact that no one has volunteered has been a bit of problem. But we have high hopes that we'll eventually wear Mark's resistance down.

[3] For sake of international comparison, the UK has 26,607 tweets on Syria and 28,342 tweets containing Twerk.  Apparently, twerking still has room for expansion in Britian.

[4] Not to be confused with the Beer Belly of America, which is an entirely different socio-spatial phenomenon that we anthropomorphized into a regional definition.

[5] Although models that include DC actually have a higher r-squared. It just seems better to exclude it from a state level model.

[6] Seriously, this is just a blog post comparing twerking and Syria. For this, you expect peer review?

[7] Except in the case of Mark (see [2]) in which case the answer is yes.

August 14, 2013

Visualizing the Relational Spaces of Hurricane Sandy

Nearly a year ago, Hurricane Sandy made landfall on the eastern seaboard of the US, wreaking havoc on the lives of millions of people in its path. At the time, we threw together some quick maps of where Sandy was being talked about on Twitter, and how the geographies of Sandy-related tweets were both intensely connected to the material impacts of the storm, but also somewhat incongruent.

Since then, we've been putting the finishing touches on a paper that extends our initial interest in the data shadows of Hurricane Sandy to a more comprehensive look at how we can use Sandy-related tweeting to understand the multidimensionality of the geographies of social media activity. All too often, a one-to-one connection is made between the location of a geotagged tweet or other piece of social media content and the content of that tweet [1]. We have instead been attempting to understand how we can think through, and then visualize, how geotagged tweets reflect and produce much more complex socio-spatial relations, which include both intense connections to the places where such content is produced, as well as much more physically distant locations which are brought closer in relational space through such informational flows. The rest of this post is adapted from our paper-in-progress, and outlines how we can map and measure the relational spaces of Hurricane Sandy.

Using T-100 Domestic Market data from the Research and Innovative Technology Administration (RITA) on flights and the number of passengers between city pairs in 2012, we determined the 50 cities that have the most passenger traffic with New York City, ranging from Chicago (3.5 million passengers back and forth) to Kansas City (175,000 passengers). Since operations and activities at some airports close to New York were directly affected by Sandy’s landfall, we exclude any airport within 500 kilometres of Manhattan in this analysis. For the remaining airports we used a buffer of 5km to collect all Hurricane Sandy related tweets and calculated the lower bound of the odds-ratio (or location quotient).  This metric measures the level of Hurricane Sandy tweets relative to overall Twitter activity . If relational networks did not play a significant role in Sandy-related tweeting, one would expect to see a direct distance decay effect: as the distance from New York City increases the odds-ratio should decrease.

Twitter Activity vs. Physical Distance

Our map shows, however, that physical distance has no significant relationship with the relative level of tweeting activity about Hurricane Sandy as is evidenced by both the scatterplot and the map (Spearman’s rho is -0.05). The map uses an azimuthal equidistant projection with New York City as the center, where the size of each airport is proportional to its odds ratio. Airports that are equally distant in physical terms from New York have widely diverging measures of Sandy-related Twitter activity. In addition, the average odds ratio in each 1000km zone does not decrease the further away one travels from New York.

In contrast, a slightly altered version of our map shows that the number of passengers between each city and New York City exhibits a much stronger positive correlation with the odds-ratio metric of Twitter activity (Spearman’s rho is 0.34). This figure preserves the directional bearing of each city with respect to New York City, but instead uses an inverse of the number of passengers to recalculate the relational distance between the cities. Airports are thus no longer displayed according to their physical distance from New York City, but rather based on the intensity of air traffic between the two cities. Since the bearing has remained the same, airports with a higher intensity will move closer to New York along that line, and vice versa. In addition to the correlation coefficient, we can also visually determine that cities with a lower odds-ratio, such as Pittsburgh and Memphis, have a tendency to move towards the outer circles while cities with a higher odds-ratio, such as San Francisco and Los Angeles, move relatively closer.

Twitter Activity vs. Air Traffic Interactivity

In other words, it is the relational connection to New York, measured by number of air travelers, not physical distance, which better explains the level of concern with Hurricane Sandy as expressed via Twitter. This concern, however, can vary within metropolitan territories depending upon the scale of analysis; some parts of an urban area may have much stronger relational ties to distant cities, while other parts are largely disconnected from such global flows.

To test the extent to which the data shadows of Sandy-related tweeting are a localized phenomenon within certain parts of metropolitan areas (rather than a more generalized territorial phenomenon), we increased the initial buffer around each airport from 5km to 25km. Thus, rather than just capturing neighborhoods that are spatially proximate to the airport, this measure captures a much wider swath of each metropolitan area. With this larger buffer, there is a near-reversal of the correlations illustrated in our first map, as Pearson’s rho for total number of passengers is now 0.06 (rather than 0.34), while the distance effect starts to emerge (rho is -0.15). In other words, even though the sociospatiality of a phenomenon like Sandy is expressed partly through a network of connections between territories, these connections are very much bounded by the locally-specific practices of place. This once again highlights the complex ways in which the digital data shadows of a material event are manifest through the intertwinement of different dimensions of social space.

As evidenced by these examples, Sandy’s data shadows are not evenly distributed through the continental United States. They are instead quite intense in some locations, while hardly reaching others at all, demonstrating the multiple spatial dimensions of social processes such as the response to Hurricane Sandy.
---------------------------
[1] We're as guilty of this as anyone.

May 10, 2013

The Geography of Hate

UPDATE (5/13/13 @ 10:45pm): We have written and published a FAQ to respond to some of the questions and concerns raised in the comments here and elsewhere. Please review our comments there before commenting or emailing.

Following the 2012 US Presidential election, we created a map of tweets that referred to President Obama using a variety of racist slurs. In the wake of that map, we received a number of criticisms - some constructive, others not - about how we were measuring what we determined to be racist sentiments. In that work, we showed that the states with the highest relative amount of racist content referencing President Obama - Mississippi and Alabama - were notable not only for being starkly anti-Obama in their voting patterns, but also for their problematic histories of racism. That is, even a fairly crude and cursory analysis can show how contemporary expressions of racism on social media can be tied to any number of contextual factors which explain their persistence.

The prominence of debates around online bullying and the censorship of hate speech prompted us to examine how social media has become an important conduit for hate speech, and how particular terminology used to degrade a given minority group is expressed geographically. As we’ve documented in a variety of cases, the virtual spaces of social media are intensely tied to particular socio-spatial contexts in the offline world, and as this work shows, the geography of online hate speech is no different.

Rather than focusing just on hate directed towards a single individual at a single point in time, we wanted to analyze a broader swath of discriminatory speech in social media, including the usage of racist, homophobic and ableist slurs.

Using DOLLY to search for all geotagged tweets in North America between June 2012 and April 2013, we discovered 41,306 tweets containing the word ‘nigger’, 95,123 referenced ‘homo’, among other terms. In order to address one of the earlier criticisms of our map of racism directed at Obama, students at Humboldt State manually read and coded the sentiment of each tweet to determine if the given word was used in a positive, negative or neutral manner. This allowed us to avoid using any algorithmic sentiment analysis or natural language processing, as many algorithms would have simply classified a tweet as ‘negative’ when the word was used in a neutral or positive way. For example the phrase ‘dyke’, while often negative when referring to an individual person, was also used in positive ways (e.g. “dykes on bikes #SFPride”). The students were able to discern which were negative, neutral, or positive. Only those tweets used in an explicitly negative way are included in the map.

Tweets negatively referring to "Dyke"
All together, the students determined over 150,000 geotagged tweets with a hateful slur to be negative. Hateful tweets were aggregated to the county level and then normalized by the total number of tweets in each county. This then shows a comparison of places with disproportionately high amounts of a particular hate word relative to all tweeting activity. For example, Orange County, California has the highest absolute number of tweets mentioning many of the slurs, but because of its significant overall Twitter activity, such hateful tweets are less prominent and therefore do not appear as prominently on our map. So when viewing the map at a broad scale, it’s best not to be covered with the blue smog of hate, as even the lower end of the scale includes the presence of hateful tweeting activity.

Even when normalized, many of the slurs included in our analysis display little meaningful spatial distribution. For example, tweets referencing ‘nigger’ are not concentrated in any single place or region in the United States; instead, quite depressingly, there are a number of pockets of concentration that demonstrate heavy usage of the word. In addition to looking at the density of hateful words, we also examined how many unique users were tweeting these words. For example in the Quad Cities (East Iowa) 31 unique Twitter users tweeted the word “nigger” in a hateful way 41 times. There are two likely reasons for higher proportion of such slurs in rural areas: demographic differences and differing social practices with regard to the use of Twitter. We will be testing the clusters of hate speech against the demographic composition of an area in a later phase of this project. 

Hotspots for "wetback" Tweets
Perhaps the most interesting concentration comes for references to ‘wetback’, a slur meant to degrade Latino immigrants to the US by tying them to ‘illegal’ immigration. Ultimately, this term is used most in different areas of Texas, showing the state’s centrality to debates about immigration in the US. But the areas with significant concentrations aren’t necessarily that close to the border, and neither do other border states who feature prominently in debates about immigration contain significant concentrations.

Ultimately, some of the slurs included in our analysis might not have particularly revealing spatial distributions. But, unfortunately, they show the significant persistence of hatred in the United States and the ways that the open platforms of social media have been adopted and appropriated to allow for these ideas to be propagated.

Funding for this map was provided by the University Research and Creative Activities Fellowship at HSU. Geography students Amelia Egle, Miles Ross and Matthew Eiben at Humboldt State University coded tweets and created this map.

The full interactive map is available here: http://users.humboldt.edu/mstephens/hate/hate_map.html

February 13, 2013

The Urban Geographies of Tweets in Africa

This is a quick post containing a few visualisations of information densities in a selection of African cities.

Below, you can find maps of all geocoded tweets published in November 2012 in Accra, Cairo, Dar es Salaam, Cape Town, Johannesburg, Lagos, Tunis, Nairobi, Kigali, Mogadishu, and Addis Ababa.

Look for the information presences and absences; groups of people who are and aren't participating in each city. But also look at the significant differences between cities. Cities like Nairobi, Cairo, and Cape Town are swimming in thick clouds of information, whereas in Mogadishu and Addis Ababa we barely find any digital geospatial information at all. 

If you're interested in why these geographies of information might matter, then check out any the articles at the end of this post. Otherwise, enjoy the maps - and please share any insights you have about any of these cities (or let us know if there are other cities that you'd like to see mapped).











Relevant articles:

Graham, M. 2013. Virtual Geographies and Urban Environments: Big data and the ephemeral, augmented city. In Global City Challenges: debating a concept, improving the practice. eds. M. Acuto and W. Steele. London: Palgrave. (in press).

Graham, M and M. Zook. 2013. Augmented Realities and Uneven Geographies: Exploring the Geo-linguistic Contours of the Web. Environment and Planning A 45(1) 77-99.

Graham, M., M. Zook., and A. Boulton. 2012. Augmented Reality in the Urban Environment: contested content and the duplicity of code. Transactions of the Institute of British Geographers. DOI: 10.1111/j.1475-5661.2012.00539.x  

January 29, 2013

New Special Issue of E&PA: Situating Neogeography

The new special issue of Environment and Planning A on neogeography edited by Matthew Wilson and Mark Graham, and featuring a handful of pieces by members of the Floatingsheep team and other friends of the sheep, is now out and available to download. The complete table of contents is below:

Theme issue: Situating neogeography

Guest editors: Matthew W. Wilson, Mark Graham

Guest editorial
Situating neogeography
Matthew W. Wilson, Mark Graham

Neogeography and volunteered geographic information: a conversation with Michael Goodchild and Andrew Turner
Matthew W. Wilson, Mark Graham

Crowdsourced cartography: mapping experience and knowledge
Martin Dodge, Rob Kitchin

Situating performative neogeography: tracing, mapping, and performing “Everyone’s East Lake”
Wen Lin

Neogeography and the delusion of democratisation
Mordechai (Muki) Haklay

Commentary: Political applications of the geoweb: citizen redistricting
Jeremy W. Crampton

Augmented realities and uneven geographies: exploring the geolinguistic contours of the web

Mark Graham, Matthew Zook

Featured graphic: Mapping the geoweb: a geography of Twitter
Mark Graham, Monica Stephens, Scott Hale

p.s. feel free drop Mark a note if you don't have institutional access to journal and would like email copies of any of the articles. 

January 11, 2013

Premier League teams on Twitter (or why Liverpool wins the league and the Queen might support West Ham)

Have you ever wondered where Premier League football teams draw most of their support from? Or what the geography of fandom is? We have too, and set about to better understand how Premiership teams are reflected in Twitter usage across the UK.

The Floatingsheep team, with the help of two researchers from the Oxford Internet Institute - Joshua Melville and Scott A. Hale (both of whom did most of the work) - have created a neat interactive map for you to both explore the geography of Twitter mentions of specific teams, and let you explore the patterns of five key rivalries. Click on the screenshot below to be brought to the full interactive map


The data used include all geotagged tweets mentioning any of the Premiership football teams and their associated hashtags (e.g., #MUFC or #YNWA) that were sent between August 18 and December 19, 2012. We have only included one tweet per user to prevent 'loud' fans from skewing the results. The users were then aggregated to postcode districts in order to see a fairly fine-grained geography of results. The number of tweeters per district is normalized by the total 'population' of Twitter users based on a 0.25% random sample of all tweets within the UK. 

What do the data show us, you ask? In Manchester, for instance, there is the oft-repeated stereotype that Manchester City are the 'real' local team, while Manchester United attract support from further afield. Our map doesn't really support that idea though. There are only a few parts of Greater Manchester in which we see significant more tweets mentioning Manchester City than their local rivals. We also, strangely, see more mentions of Manchester City in Scotland and Merseyside, and more support for Manchester United in Northern Ireland.

The Merseyside rivalry (Liverpool vs. Everton) is another interesting one to map. There we see that Liverpool have the slight edge in the postcode that is home to both team's stadiums. However, there is no clear winner in the rest of the region: with most postcodes having a fairly close split between the two teams. Interestingly, many postcodes in Scotland seem to have more mentions of Everton; while many in Northern Ireland have more mentions of Liverpool.

We can also zoom into particular postcodes and see which teams are most mentioned there. The  academics in Oxford (for some strange reason) mention Manchester City more than any other team. Central Edinburgh (when not focusing on Hearts or Hibs) has more mentions of Everton than any other Club. And the Queen's home of SW1A goes for West Ham.

What about maybe the most important question of all. Who wins the league based on total number of Tweets sent from anywhere in the UK? The answer is Liverpool (a team that hasn't won the actual league since 1990).  Manchester United are a somewhat distant second, joined by Everton and Tottenham in the Champions League spots. We also find out that Fulham, Swansea, and Wigan are the three teams that get relegated due to their quite abysmal scores. Apparently just not that many people want to tweet about Wigan.

There is no doubt that using Tweets as a proxy for fandom is messy and not always reliable. In other words, we are mapping mentions and not measuring sentiment. But, the data do give us a rough sense of who is interested in (or at least talking about what), and where they are doing it from. It allows to begin to counter myths (e.g. that Mancunians don't support Manchester United), develop new insights about places that we don't necessarily have good data about, and most importantly, have some guesses as to which team the Queen might support.

See also:
A broader take on how information augments place (a second paper on the topic can be accessed here)
Other examples of our Twitter mapping (racism, flooding, earthquakes)
The code behind this visualisation (made freely [CC-BY-NC-SA] available on Github)

December 27, 2012

The Bluegrass Basketball Battle

In Kentucky, basketball means everything -- especially college basketball, and especially the intrastate rivalry between the Kentucky Wildcats and the Louisville Cardinals, one of the greatest in all of college sports. Growing up in Louisville, one can't help but choose sides and develop one's debating skills, arguing with classmates, family and friends over whether Patrick Sparks traveled in 2004 or whether Rick Pitino is the modern-day basketball equivalent of Benedict Arnold. But given our connections to the University of Kentucky (and Taylor's fandom), the upcoming game and the tools at our disposal, we thought it might be time to wade in on the age-old debate between the two sides.

A recent public opinion poll of Kentucky by Public Policy Polling piqued our interest, as it found that Kentucky fans outnumber Louisville fans in the state by an overwhelming 66% to 17% margin. But how do the two fanbases stack up on Twitter?

We took to DOLLY to collect references to the two general-purpose hashtags used by fans of each team and promoted by the respective athletics departments -- #BBN (for Big Blue Nation) and #L1C4 (for Louisville First, Cards Forever) -- in geotagged tweets created between June 21, 2012 and December 20, 2012, in order to measure the both the absolute numbers and geographic distribution of UK and UL fans at the national, statewide and local scales as reflected by Twitter.

Number of Tweets referencing #BBN or #L1C4
According to the aforementioned poll's 66-to-17 margin, there are ~3.9x more UK fans than UL fans in Kentucky. This finding is mirrored almost exactly by our measures of tweeting, where the 6,371 geotagged references to #BBN in the state are also 3.9x greater than the 1,628 references to #L1C4. And while the number of tweets for each team are essentially equal within the city of Louisville, UK fandom becomes even more dominant once one moves outside of the Commonwealth, with there being over 10.5x more #BBN tweets than #L1C4 tweets in the US outside of Kentucky, for a total of 4.9x more UK tweets than UL tweets nationwide. So not only does UK hold an ever-so-slight advantage within Louisville's homebase, it shows increasing popularity as one moves to the larger scales of the state and nation.

#BBN vs. #L1C4 Nationwide
But when we visualize these tweets, we get a better idea for just how geographically concentrated these patterns of fandom are. For instance, 599 of the 3,141 US counties had references to either #BBN or #L1C4. But of these, only 35 counties had a greater number of references to #L1C4, with Butler County, KY holding the dubious honor of being the only county in the Commonwealth with more references to #L1C4. Of the remaining counties, 554 had more references to #BBN, and only 10 counties in the country had an equal number of tweets referencing #BBN and #L1C4.

#BBN and #L1C4 in the Commonwealth of Kentucky
Also interesting is that no county in the US apart from Jefferson County, KY (Louisville and Jefferson County have a merged government, and so are coterminous) has more than 100 tweets with references to #L1C4, highlighting the essentially limited spatial distribution of UL fans. And though Jefferson County does have a few more UK tweets than UL tweets, one doesn't have to go far to find the county with the largest margin of UL-related tweets over UK-related tweets; right across the river from Louisville in Clark County, Indiana there are 20 more #L1C4 than #BBN tweets.

Meanwhile, Kentucky holds a decisive advantage in its hometown of Lexington-Fayette County, with 1,588 more #BBN tweets than #L1C4 tweets. But the county with the second-highest margin favoring UK is all the way south in Broward County, Florida (Ft. Lauderdale) with a +299 margin favoring UK.

#BBN vs. #L1C4 in Louisville
Within Louisville, the absolute number of tweets are almost equal, as mentioned previously; but, interestingly enough, the geographies of UK and UL tweeting are quite different. The clustering of #L1C4 tweets tends to be around the UL campus and downtown areas, while UK tweeting tends to be more spatially distributed, with many tweets coming from more suburban, residential areas in the city. So while the vast majority of UL tweets across the country are located in Louisville, a still significant number come from within just a handful of square miles surrounding the UL campus in downtown Louisville, perhaps indicating the limited appeal of a team that's lost four-straight games to the defending-champion Wildcats.

#L1C4? More like #L1C4.9xLessPopularThanUK.

UPDATE: See today's article over at ESPN.com, "The Commonwealth's great divide", which discusses some of the same geographic dimensions of UK and UL fandom we are showing here. It includes this interesting passage:
In 2005, the Courier-Journal polled fans on their sports loyalties and 53.7 percent within the city counted themselves as UL fans compared to just 33 percent who identified themselves as Cats fans. And according to the two schools' alumni associations, Louisville understandably has a far greater base in Jefferson County (54,872 living alumni) than Lexington (16,112). 

But here's the catch: There are just 22,160 living Louisville alumni in the rest of the state and other than Fayette County (where Lexington sits), none of Kentucky's 120 counties boasts more UK grads than Jefferson.
While we weren't aware of these figures at the time of our initial post, they not only tend to confirm some of our findings, but indeed only lend even more credence to our assertion that UK fans seem to be more voracious tweeters than their UL counterparts, as the roughly 50-50 split in tweeting in Louisville is significantly askew from the 54-33 numbers from the Courier-Journal's 2005 survey.

November 28, 2012

Digital Data Trails of the UK Floods

What do data scraped from the Internet tell us about a range of social, economic, political, and even environmental processes and practices? As ever more people take to social media to share and communicate, we are seeing that the data shadows of any particular story or event become increasingly well defined. 

The ongoing UK floods offer a useful example of some of the links between digital data trails and the phenomena they represent. In the graphics below, we mapped every geocoded tweet between Nov 20 and Nov 27, 2012 that mentioned the word "flood" (or variations like "flooded" or "flooding").


Unlike many maps of online phenomena (relevant XKCD),careful analysis and mapping of Twitter data does NOT simply mirror population densities. Instead concentration of twitter activity (in this case tweets containing the keyword flood) seem to closely reflect the actual locations of floods and flood alerts even when simply look at the total counts. This pattern becomes even clearer when we do normalise the map (the second map is a location quotient where everything greater than 1 indicates that there are more tweets related to flooding than one would expect based on normal Twitter usage in that area), the data even more closely mirror the UK Environment Agency's flooding map.

As we demonstrated with our maps of Hurricane Sandy, it is important to approach these sorts of maps with caution. At least in the information-dense Western world, they are often able to reflect the broad contours of large phenomena. But, because we are still necessarily measuring subsets of subsets, our big data shadows start to become quite small and unrepresentative at more local levels. This is particularly an issue when the use of the relevant technology is unevenly distributed across demographic sectors such as was the case in post-Katrina New Orleans

Nonetheless, with every new large event, movement, and phenomena, we are undoubtedly going to see a much more research into both the potentials and limitations of mapping and measuring digital data shadows. This is because physical phenomena like hurricanes and floods don't just leave physical trails, but create digital ones as well. 

November 12, 2012

Mapping the Eastern Kentucky Earthquake

Last week's post on racist tweets in the wake of the US presidential election received much more attention than we ever expected. A number of questions about and critiques of our method were raised, which we attempted to respond to in a special FAQ with the post (first time we had to do that). Nonetheless, we thought it might be useful to demonstrate the utility of our technique on a less controversial subject in order to demonstrate how we can leverage a relatively small number of geocoded tweets in order to understand particular offline phenomena, and maybe even assuage some concerns about such an approach.

The 4.3 magnitude earthquake that occurred on Saturday, November 10th around 12:08pm EST, about eight miles west of Whitesburg, Kentucky, provides just such an example. Given our own connections to Kentucky, and the significant number of our own friends and family who tweeted or updated their statuses about the earthquake, we were naturally interested in what we might be able to bring to such an analysis.

But before showing our own results, it is useful to note that the US Geological Survey also collects user-generated data on earthquakes through their "Did You Feel It?" reporting system in which individuals contribute their location and experience with quake. The USGS then aggregates these reports into a crowd sourced map like the one below in order to visualize an approximation of how the earthquake was experienced in different locations.

Rather than use such a direct system of user-generated data collection, we fired up DOLLY in order to gather geocoded tweets referencing the earthquake in its immediate aftermath. We were able to collect 795 geotagged tweets referencing "earthquake" from 12:08pm -- where the first tweet we uncovered near Hyden in Leslie County, KY simply said "EARTHQUAKE HOLY SHAT" -- until around 4:05pm in an area comprising most of central and eastern Kentucky, southern Ohio, West Virginia, southwest Virginia, western North Carolina and east Tennessee (we limited our query based on a bounding box drawn around the epicenter of the quake).

This area includes several cities such as Louisville and Lexington in Kentucky and Knoxville, TN, as well as many more rural areas. As much of our earlier work has clearly shown, population centers typically possess a greater level of online activity simply by virtue of population size, so it was important to look beyond just the raw numbers of earthquake-related tweeting. Therefore, in order to normalize the data, we also collected a 1% sample of all geotagged tweets from the month of October within in the same area. This totaled 30,699 tweets, which we used to normalize the tweets about the earthquake and construct a location quotient measurement in exactly the same way as with the racist tweet analysis [1]. We again aggregated from individual tweets to a larger areal unit, in this case, counties.


First and foremost, though we did not use an entirely contiguous area, it is easy to notice that our map roughly conforms with the map of crowdsourced reports from the USGS, generally confirming the relevance of a relatively small set of user-generated data to understanding such an event.

Second, by looking at the blue dots representing each individual tweet, we can see concentrations within the counties containing the largest cities in the specified search area. These include Knox Co., TN (Knoxville), Jefferson Co., KY (Louisville), Fayette Co., KY (Lexington), Madison Co., KY (Richmond), and Cabell Co., WV (Huntington). None of these localities are particularly close to the epicenter of the quake in eastern Kentucky, but are more likely is a product of the higher population in these cities (increasing the likelihood that Twitter users would feel the quake and take to Twitter to report it), as well as their importance as regional centers with close social and economic connections to eastern Kentucky.

Third, and interestingly enough, there were only six counties where there were more earthquake tweets than there were tweets within the given 1% sample from October [2]. Leading this group of counties is Letcher County, where the earthquake epicenter was located. Letcher County also has a location quotient of nearly 100, indicating the fact that the earthquake generated a much greater than average number of tweets in Letcher County than one would expect on average. Each of the other counties, though possessing many fewer tweets both in the earthquake and reference datasets, are also located in close proximity to Letcher County and the epicenter of the earthquake. These include Bath Co., KY, Leslie Co., KY, Polk Co., TN, Johnson Co., TN and Rockingham Co., VA.

We can also look at patterns of tweets without aggregating to an administrative unit. In this case, we estimate the intensity of the earthquake tweet pattern (again normalized for what would be expected based on a random sample of tweets) in the region using Gaussian kernel smoothing. Interestingly, the 'epicenter' of earthquake tweets is only 6.7 miles away from the real epicenter of the earthquake (indicated by the red star). Not coincidentally, the center of intensity of our tweet map is located in the nearby town of Hazard, KY, which has a higher population density (resulting in more twitter users) than the more rural town of Whitesburg, the epicenter as measured by the USGS.

Ultimately, these results are not necessarily surprising, as they indicate both the extremely localized nature of a phenomenon like reporting an earthquake as evidenced by the greater location quotient values nearer the epicenter, as well as the essentially networked nature of such phenomena mediated by the internet in the clustering of user-generated internet content in cities quite distant from the earthquake's origin.

From a methodological standpoint, it shows that the fairly simple technique of calculating location quotients, or even the more involved technique of Gaussian kernel smoothing, can provide powerful ways of uncovering the spatial dimensions of online reflections of essentially offline phenomena.

We hope that this example -- which uses about the same number of tweets (particularly relative to the number of administrative units) as our racist tweets map -- will help alleviate some of the methodological concerns raised in our previous post.
---------
[1] The equation used to calculate the location quotient is as follows:

# of tweets referencing "earthquake" per county / total # of tweets referencing "earthquake"
------------------------------------------------
# of reference tweets per county / total # of reference tweets

[2] We should note that this doesn't mean that there were more earthquake-related tweets in the given time period on Saturday than total tweets in the entire month of October. Rather, this simply represents an indicator of how many earthquake-related tweets there were relative to the expected amount of content in that place.

November 05, 2012

Can Twitter Predict the US Presidential Election?

Can Twitter predict the outcome of tomorrow's US presidential election? If the results of our preliminary analysis are anything to go by, then Barack Obama will be easily re-elected. The data presented below, including all geocoded tweets referencing Obama or Romney between October 1st and November 1st, out of a sample of about 30 million, give some insight into the visibility of each of the candidates on Twitter.


We see that if the election were decided purely based on Twitter mentions, then Obama would be re-elected quite handily. In fact, the only states in the electoral college that Romney would win are Maine, Massachusetts, New Mexico, Oregon, Pennsylvania, Utah, and Vermont. Romney also wins in the District of Colombia, and we unfortunately didn't collect data on Alaska or Hawaii. Some of the results seem to be interesting reflections of social and political characteristics of particular places. It makes sense that Romney has captured more of the public imagination in Utah, likely due to the state's considerable conservatism and large Mormon population, and Massachusetts, the state that he governed not all that long ago.

However, this drubbing that Romney receives in the Twitter electoral college belies the close nature of the final popular (Twitter) vote, re-raising the issue of whether the electoral college is the most suitable means of deciding the country's political future. There are a total of 132,771 tweets mentioning Obama and 120,637 mentioning Romney, giving Obama only 52.4% of the total and Romney 47.6%, a breakdown that is remarkably similar to current opinion polls, though not reflected when looking at the state-level aggregations in absolute terms. If you want to explore the data in more detail, please play around with the interactive map below:


We can also visualize the data using a sliding scale, so as to see how close the margin of victory is for each candidate in a given state.


Romney's largest margins of victory are in Pennsylvania and Massachusetts, while Obama's largest victories are in California and, strangely, Texas. The cases of Massachusetts and Texas, not to mention large portions of the south and plain states, likely point to the fact that many references on Twitter would tend to be negative.

It is also worth noting that we compared Twitter mentions of both Vice-Presidential candidates: Biden and Ryan. Ryan, interestingly, wins the head-to-head competition in every single state. This makes for a rather boring map, so we decided to instead compare references to Ryan and Romney in the map below (Romney shaded in grey for his ebullient personality, and Ryan in pink as a result of his staunch support for gay rights).


As might be expected, there are more references to Romney in most states (Kansas, Michigan, North Dakota, Rhode Island, South Dakota, and Vermont being the exceptions here). However, when looking at total references, we again don't see a large gap between the two men. Ryan has 94,707 tweets compared to Romney's 120,637.

What do these data really tell us? Ultimately, I doubt that they will accurately predict the election, as Obama's seeming victory in Texas or Romney's in Massachusetts will almost certainly not come to pass. But they do certainly reveal that many internet users in California, Texas, and much of the rest of the country for that matter, tend to talk more about Obama than Romney. And, of course, in order to truly equate tweets with votes, we would need to employ sentiment analysis or manually read a large number of the election-related tweets in order to figure out whether we are seeing messages of support or more critical posts, as has been done in a couple of interesting projects by Twitter available here and here and another project by Esri available here.

Maybe the most revealing aspect of these data is that the 'popular vote' is split between the two candidates. While the social and political data shadows that we are picking up may not accurately tell us much about the electoral college results, when aggregated across the country they may be a rough indicator of tomorrow's outcome, pointing to the more-or-less equal and evenly divided nature of the American two-party political system. While this work may seem like a contemporary attempt at soothsaying, something we tend to shy away from, the data more appropriately serve as a useful benchmark in order to allow us to analyze what social media data shadows might actually reflect, as no matter the level of participation, they remain distorted mirrors on the offline material world.