Showing posts with label Population. Show all posts
Showing posts with label Population. Show all posts

Tuesday, February 18, 2014

In the Era of Big Data - Sampling VS Collecting Everything

The Conundrum - Data Everywhere


We live in a world where much of what we do is tracked by both governments and companies. When online, as you browse the web, various metrics are tracked for every web page you view (including this one). In the outside world cell phones can provide companies your location; credit cards, bank accounts, and frequent shopper cards show them your spending habits. 

With enough data it is possible find correlation trends between seemingly unrelated topics, such as web searches and illness. Below is a youtube video showing how google uses search data to predict flu outbreaks in near real-time.

In the past while it was possible to gather huge amounts of data, analysis was difficult and generally required sampling. Basically, a selection of the available data would be collected and analyzed to form assumptions about the overall population. Now there are new tools for analysis that allow for true population analysis. So should we look to abandon sampling in favor of simply collecting and analyzing everything? What does using more data gain us, and at what cost?

 Collecting Everything - The NSA


The NSA provides a very high profile example of an organization compiling all available data (though there are unknowns as to how exactly it analyzes what it gets).  In a recent Forbes article, Gil Press compiles data from various sources in an attempt to decipher whether or not this results in better results than simply sampling from the population.

Beyond the question of if the NSA should collect the data, is the question of whether the glut of data is unmanageable.
NSA Datacenter Image - AP PHOTO/GOOGLE, CONNIE ZHOU
"The unspoken assumption here is that possessing massive quantities of data guarantees that the government will be able to find criminals, and find them quickly, by tracing their electronic tracks. That assumption is unrealistic. Massive quantities of data add cost and complexity to every kind of analysis, often with no meaningful improvement in the results. Indeed, data quality problems and slow data processing are almost certain to arise, actually hindering the work of data analysts. It is far more productive to invest resources into thoughtful analysis of modest quantities of good quality, relevant data.”

On the other hand, there is a concern that sampling will miss outliers, which is really what you want in this sort of an analysis. In a discussion post regarding big data and sampling, Paige Roberts had this to say:

"When finding outliers is the goal, sampling is counter-productive. When finding trends in the overall data is the goal, then sampling is a shortcut that may or may not do the job. But one that has become standard practice because up until now, we didn't have the data crunching power to do anything else in a sensible time frame."

So, Should We Sample or Not!?!?


It may seem unsatisfying, but in the end the answer to the question of whether or not to sample appears to be "it depends". If the sampling is done in a statistically valid way, way sampling, especially for common trends, seems to be a very cost effective method to get what we want.  If we're looking for specific outliers (as opposed to just finding about how many there would be in a given sample) however, there might not be a good alternative to combing through much more data. Even in cases where outliers are being sought out it is important to note that the amount of the data in the population and the amount of processing power and tools being used to analyze have to align. A thorough reading of the NSA case shows the organization might have bit off more than it can properly chew.

So, should you sample? I'll leave that decision up to you.

__________________________

References


In the order they appear in-article:

- http://www.forbes.com/sites/gilpress/2013/06/12/the-effectiveness-of-small-vs-big-data-is-where-the-nsa-debate-should-start/

- http://global.fncstatic.com/static/managed/img/Scitech/NSA%20Phone%20Records%202.jpg / http://www.foxnews.com/tech/2013/06/11/inside-nsas-secret-utah-data-center/

- http://www.techrepublic.com/blog/big-data-analytics/why-samples-sizes-are-key-to-predictive-data-analytics/

- http://smartdatacollective.com/users/paige-roberts

Monday, February 4, 2013

Normalizing Geographic Data

Geographical data has always been fascinating to me.  Probably because maps are fascinating.  Overlaying a map with data to create a heat map provides an interesting perspective to behavior and markets based on location.  There is a beautiful map of Walmart’s growth over time on FlowingData.  This graphic provides a lot of insight into Walmart's growth, but this information can be misleading.  The concept of normalization can provide some clarity.

Normalization is a concept that allows two metrics with different dimensions to be compared using a common base.  While shopping we may have the choice between 100 paper plates for 2.99, or 250 plates for $5.99.  Simple division tells us that the first option is $.03 per plate and the second option is $.024 per plate.  We normalized the price to the unit (plates) in order to compare the two offerings using a common unit.  

The issue with geographic based data is that heat maps are usually created using a standard map of the United States (or the world, specific state,...) but these maps depict area.  Since populations are not distributed uniformly across areas, it rarely makes sense to compare metrics using acreage as the common unit.  

A common example of this is with election data.  We all see the results roll in and the data is presented with red and blue shadings of states to indicate which way the voters in that state voted.  



The 2012 election results are pictured above.  At first glance the heat map would indicate the red candidate is the favorite.  But unless votes are assigned to acres, this doesn’t make since.  Populations are often congregated near coasts, but in this map Wyoming with less than 800k people appears to carry more weight than say Maryland with just under 6MM people.  Normalizing the data by the underlying population can provide more insight into the actual results.



Here the map (actually a cartogram) is morphed so that the area encompassed by each state boundary reflects the population of each state, not the geographical area.  This modified map indicates the blue candidate won a large percentage of the vote.  Of course we know this to be true, but this graphic displays the data in a more appropriate context.  Similarly we could modify the map to reflect electoral college votes, but population is a proxy for that.

So what does this have to do with web analytics?  Web Analytics is often about sales and markets.  Google Analytics, Adobe, Tableau and the host of other platforms make it very easy to display your data on a geographic heat map.  But don't let the temptation to make something eye catching cause you to create something that can mislead your intended audience. 
Let’s say we have 100k conversions in California, and 100k conversions in Wyoming.  On a standard heat map, these two states would have the same color shading indicating they are equal.  Yet if we normalize over population, we would see that we have a much larger presence in Wyoming, and have really just barely tapped California.  The first approach might cause us to choose the same strategy for California and Wyoming, but the second would cause us to choose very different strategies.  This would hopefully be obvious when we look at these two states, but if we expand the segmentation to zip code this might not be as clear.

Even if we do not have the software necessary to morph maps by population (or the number of customers in our target segment) we can normalize the metrics we display by population.  Population data is readily available on the Census Bureau website.  Using percentages of the total population, instead of measured values (sales, views,...), is one easy way to accomplish this.

Other applications of this concept might be more product specific.  For instance, if you are selling skis, on a heat map of sales you might notice western mountain states carry a greater load of the total sales.  This makes sense intuitively.  But if you were to acquire data that depicted the populations of avid skiers by zip code, you might find significant marketing opportunities outside of what was expected.  For instance Texas might be under represented compared to Colorado, but the skiing population is smaller.  Normalizing might allow you to see that conversions could actually be better on Texas based consumers than Colorado based consumers.  This could allow you to target this demographic with greater precision.

Getting back to the Walmart graphic in the first paragraph, you might have noticed a lot of under representation in the west.  Take another look at Nevada.  There are no stores anywhere in the middle of the state.  Hopefully by now you realize that normalizing this data would show that Walmart strategically positioned their stores in population centers, and there is very little population in the middle of the state.  

Resources -
http://www.esri.com/news/arcuser/1000/files/normalize.pdf
http://www-personal.umich.edu/~mejn/election/2012/
http://adam.webanalyticsdemystified.com/2010/04/13/comparison-reports/
http://semphonic.blogs.com/semangel/2010/01/tactics-in-web-analytics-visitor-segmentation.html