Showing posts with label Just Plain Bad. Show all posts
Showing posts with label Just Plain Bad. Show all posts

Monday, April 28, 2014

why I disdain most infographics

Too many offenses to sensible data visualization to list. It's unfortunate, too, because there are some compelling stats lost in the cartoony graphics.

Gates Foundation Inventions
Source: MPHOnline.org

Thursday, April 10, 2014

just because you have numbers doesn't mean you need a graph

I subscribe to updates from the Pew Research Center. They arrive in my inbox with subject lines like "Future of Internet, News Engagement, God and Morality" (yes, this was an actual title from their March 13th update - quite a span of topics!) and probably 90% of the time get moved to my trash without a second thought. But in a fraction of cases, something in that subject line catches my eye and I open the email to read more. Sometimes, this even prompts me to click further to the full article.

The snippet that caught my attention this time was "Stay-at-Home Mothers on the Rise." The link I clicked on within my email brings you here.

A quick scan through and I found that I was hardly able to focus on the article because of the issues plaguing the visuals that accompany it. There are many. But I'll focus on just a single one today and keep this rant very short and sweet:

Just because you have numbers doesn't mean you need a graph!

The following graph prompted this adage:


That's a whole lot of text and space for a grand total of two numbers. The graph does nothing to aid in the interpretation of numbers here! Even the fact that 20 is less than half of 41 doesn't really come across clearly here visually (perhaps because of the way the numbers are place above the bars?).

Rather, the above can be conveyed in a single sentence:
20% of children had a "traditional" stay-at-home mom in 2012 (compared to 41% in 1970). 

Just because you have numbers doesn't mean you need a graph!

For a less ranting delivery of a similar lesson, check out my post the power of simple text.

Wednesday, April 2, 2014

US prison population revisualized

The following graph caught my eye recently in my Twitter feed:
Source: https://plot.ly/~Dreamshot/361/ 

I've been debating whether to post about it (and finally decided that I couldn't resist).

I don't want to rip it apart.

Well, that's not entirely true.

I do want to rip it apart, but it's not in an effort to be mean. The above visual breaks pretty much every best practice out there when it comes to effective graph design. It's simple data. Probably not so much is being lost in terms of being able to interpret the data through this less-than-stellar data viz. But the specifics of the design choices (or lack thereof) drive me batty. To the extent that I can't help but comment and resolve to show what it has the potential to be.

First, let me list the main components that get under my skin (and I should note that it's possible some of these are constraints in the Plot.ly tool through which the above visual was published, which I have not used directly) :
  • No meaningful ordering to the data (rather, the categories are shown in reverse alphabetical order... not so helpful);
  • Lack of axis labels (sure, we can infer, but why should we have to?);
  • Diagonal text on x-axis (avoid, avoid, avoid!); and
  • Grey background, white vertical gridlines, and black bar outlines add unnecessary clutter (eliminate!).

I think the only positives I have to say about the original visual are: 1) a horizontal bar chart is a good choice here because we're dealing with categorical data with long category names, 2) good descriptive graph title, and 3) it makes me happy to see the data source listed (both as general good graph hygiene, as well as because it allows me to get to the source data to remake the visual).

Speaking of remaking the visual, here's what it could look like when we tackle the above issues:


If we just want to show the data, we could proceed with the above. Taking a cue from the original visual - a single point is labeled with its corresponding value: Drug Offenses - perhaps there is a story here worth highlighting. If that's the case, our visual might look something like the following:


Meta-lesson: if you're going to go through the effort of visualizing data, take the time to be thoughtful about your design choices!

If you're interested in the Excel version of the above makeovers, you can download it here.

Thursday, September 20, 2012

bar charts must have a zero baseline

This is one rule of data visualization that I see broken too often: when it comes to bar charts, the y-axis must begin at zero.

When our eyes interpret bar charts, we are comparing the relative heights of the bars. When we cut the height off at something greater than zero, it skews this visual comparison, over-emphasizing the difference between the bars in a way that simply isn't honest. Most recently, I saw this in a visual that was forwarded by a friend of a colleague. The offender: Fox News.


There are a number of things that bother me about this visual. Beyond the unnecessary visual clutter of tiny gridlines and strange chart borders, the y-axis isn't labeled (I think it's Top Tax Rate, as noted by the subtitle, but this would be a lot clearer if the axis itself were labeled) and it is placed on the right-hand size of the visual, so it's the last thing I see as my eyes scan across from left to right, making it even less likely that I see the biggest issue with the graphic, the fact that the y-axis starts at 34%. This makes the difference between Now (35%) and Jan 1, 2013 (39.6%) appear to be way bigger than it actually is.

How big of an issue is this? Let's do some math to find out. The way it's graphed, the height of the bars are 1 (35-34) and 5.6 (39.6-34). This represents a visual increase of 460% ((5.6-1)/1). If we graph the bars with a zero baseline so that the heights are accurately represented - 35 and 39.6 we get a visual (and actual) increase of 13% ((39.6-35)/35). Perhaps that is still significant and that is the point that Fox News was attempting to make. That's fine, but I wish they would have done it without this visual misrepresentation of the truth.

A couple related things to consider (and I have my own opinion on each of these that I'll of course make clear):
  • I've heard the argument that if you're graphing something that has a sort of "natural" baseline of something greater than zero, then it might be appropriate to start with that. For example, if we consider the baseline unemployment rate to be 5%, then the argument goes that you could use this 5% as the baseline. I don't like it. For me, it isn't a valid visual comparison, so if that were the case, I'd use a different way to show it (perhaps plot the entirety of the bars but then also highlight 5% horizontal line and label it in a way that makes it clear how to use it for comparison).
  • When it comes to line graphs, the zero baseline rule does not hold. In other words, you can get away with a non-zero baseline in a line graph. With line graphs, we compare the lines to each other more than their height from the x-axis. Still, you need to be careful. I would advise to make it clear to your audience that you're using a non-zero baseline so they interpret the information correctly (one approach: label the y-axis and highlight the minimum value in bold so attention is drawn that it's something other than zero). And you need to be careful about zooming in too much and making a change that is minor look big - this gets you back into the visual misrepresentation place that we want to avoid.
My advice to Fox News (and to those communicating with data in general) would be to first determine the story you want to tell. Then determine what data will best support this story. Don't compel your audience with visual misrepresentations; rather, convince them with accurately displayed data that backs up the point you are trying to make.

Related note: there are a number of posts by others on this and related topics. In case you're interested in reading more, here are a few I'm aware of (not an exhaustive list):

Thursday, September 6, 2012

color me bad(ly)

Recently, a contact shared the following image with me, along with his thoughts. I found both amusing, so thought I'd share with you here, along with some of my own thoughts and a makeover:

From: http://www.consultingmag-digital.com/consultingmag/201207?pg=6&pm=2&fs=1#pg26

Commentary accompanying the visual:
This seems like some seriously simple data to present, but SOOO poorly executed. Looking at it hurts my head and leaves me with nothing but questions:
  • How much time does it take others to figure out the color pattern(s)?
  • Is there really even a pattern?
  • Why are the two legends/color-schemes different? Don't make me work so hard!
  • Why use donuts/pies instead of some simple paired bars/columns, or even just a pair of lines (i.e., a simple histogram)?

No matter what your content, this is the sort of reaction we should work to avoid in our data visualizations. In this case, it seems the color and donut form is meant to make the data more visually interesting, but it hinders our ability to understand the data.

There are a number of lessons we can employ here to make this data easier to comprehend:
  • If there is an intrinsic order in your categories, leverage it. In this case, the 2011 data has categories in order of increasing days away from home (starting at the lower middle left of graph with the light green segment and working clockwise around), but somehow neither this ordering or the colors of the categories carried over to the 2012 graph; rather, this graph appears to be sorted numerically by category. This makes comparing the segments of the pies even more difficult than it would otherwise be. Speaking of which...
  • Don't make people compare segments of two different pies (or donuts, in this case - substitute your fave dessert dataviz). Our eyes have a hard time measuring angles and areas: this difficulty is amplified when we're meant to do it across different pies/donuts, where the pieces are in slightly different places and there is no consistent baseline.
  • Put the things you want to compare close together. Physical distance between the things we're meant to compare makes comparing those things more difficult. In this case, a bar graph would allow us to put 2011 and 2012 right next to each other so we can get an easy visual comparison.
  • Use color strategically. Don't use color to make something colorful; rather, use color sparingly and strategically to draw your audience's attention to where you want it.
  • Tell a story with your data! Don't assume your audience will want to look at the data and make up their own story. If you look at the full article, the point they are trying to make is that consultants are traveling less in 2012 than prior years. I'm not actually sure this data shows that (it could be that the other groups surveyed are traveling less but the consultants are traveling just as much - we don't have that breakdown of the data to see). At any rate, I'd suggest making the point more clearly with the data and actually calling out the takeaway within the data visualization to help your audience know where to look for the evidence of what you're telling them.
Here's an alternative view of the same data, employing the lessons I've outlined above:


Thanks, Andy, for passing this less-than-stellar viz along and for your thoughts!

For those interested, you can download my Excel file with the above visual here.

Friday, August 10, 2012

evaluating word clouds

Word clouds created a bit of buzz when they first became popular a couple of years ago (or at least that's when I encountered them for the first time). Like the infographic, they have a bit of sex appeal that draws you in. As in the case of infographics, however, I often find that upon further evaluation they tend to be a letdown - full of fluff without so much informative value.

While facilitating a workshop recently, I heard a horror story about someone who had tried to create a word cloud by hand (perhaps the scariest part of the story involved scaling text boxes one at a time). Lesson: in data viz (and in life), if you find yourself doing something tedious and repetitive like that, stop to reevaluate. At minimum, do a Google search. Even better if you can find a blog post or related article on the topic from someone who has encountered the same challenge before and identified an eloquent solution.

In the case of word clouds, there are a number of applications you can use to generate them. Wordle is a popular free product (created by Jonathan Feinberg of IBM, note that if you upload your Wordle to the gallery, the data goes with it, though you can also opt for local-only word cloud generation) that allows for quite a bit of customization of color, size, font, etc. Google docs has a word cloud gadget within spreadsheets. There are a number of others, easily located via a Google search.

But before you start thinking about generating word clouds, let's continue our discussion on their efficacy. Their sexiness can draw you in. But is there value beyond that? I think it comes down to the use case. I've got one example for the negative and one for the affirmative.

Poor use of word clouds
First, let's take a look at an example from a Community Health Center. My understanding is that they employed a consultant to analyze some survey data from their clients. The consultant put together a report filled with pretty word clouds like this one:


Good service is... minutes? Part of the challenge in this case is that the connotation has been completely stripped away from the nouns, removing the sentiment behind the comments. Which is kind of the important part of the comments, in my opinion. But in reading the report, buried near the end of it, I found the following:

The consultants took the time to content-code the comments. These categories and their descriptions are much more useful for understanding what people value than the word cloud. With this info, we can direct action: we get an understanding of what's going well that we want to maintain, as well as potential areas for improvement. We could take this a step further of making the data visual like this:


In this case, I think the simple bar chart is much more useful (in terms of both understanding the information and determining how to act on it) than the word cloud. Now let's look at a better use of word clouds.

Thoughtful use of word clouds
Caveat: this example came to me by way of the telephone game (I heard it from someone who heard it from someone), which means it's guaranteed that I don't have the details totally right. But I think this still serves well as an example of a good use of word clouds. The story goes: Apple stores obviously really value customer service. They use surveys to collect info about each store. Each day, they create a word cloud for each store based on customer comments. What they are looking for are 5 (I'm making that number up, I don't know what the real number is) specific words - things that are considered must-haves when it comes to customer service in their stores. It's when these [5] words don't show up prominently on the word cloud for a given store that a red flag is raised and some sort of action is taken.

This is what I would consider a thoughtful and actionable use of word clouds. If the required word doesn't appear, some sort of intervention happens.

We can generalize this to the following: when you're considering using a word cloud, think about what you want your audience to know and what you want your audience to do. Then ask yourself if a word cloud will enable them to know and do those things.

And for goodness sake, if you do use a word cloud - leverage some of the tools that exist - don't try to create it by hand!

Monday, January 16, 2012

wisdom of crowds: how can we improve this viz?

I am in the exciting process of selling a house. In case you're unsure, that sentence is dripping with sarcasm. While I tend to have fun on the buying side, selling (at least in my recent experience) seems to be all about fixing things and losing money.

Ok, that's enough ranting (almost)...on to the data viz. The broker I'm using apparently has a kiosk in the local mall that is meant to drive traffic to their office. My mild curiosity (and disbelief that this is an effective way to find buyers) prompted me to ask my agent what percentage of sales can be traced back to the kiosk. She sent me the following graphs:



While I can answer my question based on the visuals (4% of transactions appear to be driven by the kiosk - not much, but more than I would have guessed), these graphs seem to be poster children for how not to present data. In my current ranting mood, I could go on and on about what I'd change, but instead I'll try to get over myself and let you join in on the fun: what about the above visuals bothers you the most? Leave a comment. Bonus points for discussing what you'd do differently if you were presenting this information.

Anyone looking to buy a cute rambler in Poulsbo, WA? :-)

Tuesday, December 27, 2011

don't fall victim to this

I came across this graph recently when catching up on some reading over the holiday. My question to you is simple: can you read it?


The website where this interactive visual resides is called worldshapin, and it implores you to "compare countries through their shape." It visualizes data from the Human Development Report 2011 as a "star plot" along the six dimensions of education, population, health, workplace equality, carbon footprint, and living standards. As shown above, you can look at this data between countries and as it compares to continents and the world (when the world isn't obscured by the countries and continents you've chosen, as it is above).

Before I get to the don't fall victim portion of this blog post, let me first say that I do think this helps make the data in the report more accessible by making it visual. You can get a quick idea of how one part of the world stacks up to another across these dimensions that you wouldn't get with a table of data, for example. This is fine for information discovery. This assumes you are making it available for an audience who will have an appetite to "play" with the data.

This visual is not fine, however, if you have a specific story that you want to tell through data.

To convince you of this, I'm going to take one of my own failed data visualizations from my past and remake it into something that works. First, a bit of history:

I used to make charts like this. I called them "spider graphs." In a prior life, I worked in banking, managing home equity fraud. When it comes to fraud, the ways you can impact it can be classified into 8 categories (where each category is a piece of the fraud management lifecycle): deterrence, prevention, detection, mitigation, analysis, policy, investigation, and prosecution (Wes Wilhelm, The Fraud Management Lifecycle Theory). So if we were to look at our efforts in each of these areas and rate the activities along a scale from 0 (we have nothing in place) to, say, 10 (the unattainable utopia of fraud management - we've solved every problem), we could show how well we're doing on a relative basis in each area, with the goal of maximizing our coverage and balancing activity across the different parts of the lifecycle. The spider graph was perfect for this!

I was able to locate an old annual review on the topic of home equity fraud that I put together that highlighted progress to date and introduced forward-looking plans. I'm going to assume it's ok to share an excerpt here, given that the financial institution I did this work for is now defunct (due to much bigger issues than my poor data viz). Here's what it looked like:


The visual starts off with an explanation, shows an example of how to read the graphs on the right, followed by the real-data-graphs across the bottom (the titles across the very bottom are the 5 different types of home equity fraud that we were tracking).

Lesson 1 (foreshadowing): if you have to have a graph to show how to read your graph, your visual may be too complicated.

When it comes to the visual at the bottom ("FML for Home Equity"), let's try to look past the black background and meaningless colors (while annoying, we have bigger fish to fry here) to the actual data. Same question as I led this post with: can you read it?

Before I answer that question with my current data viz lens on, let's back up the better part of a decade to take a look at what I thought of these visuals when I created them. I thought they looked really cool. Sexy, even. I also thought they clearly showed what I wanted to show: mainly, that we had a lot of work to do - we were failing in a lot of places and needed to make some changes.

But people found them really hard to read. I found myself explaining, repeatedly (to the same people even!) how to read them. At the time, I thought this was an issue with my audience.

When I look at the graphs through today's lens, I recognize that the issue was not with my audience, but rather with me. It was a visual design failure. I stubbornly persisted to show data in a way that wasn't straightforward for my audience to consume (even when it became obvious through their questions that it wasn't clear!). When information isn't straightforward, it's hard to look at. For an audience, this feels uncomfortable. Most people don't want to spend a lot of time with things that make them feel uncomfortable. Even when you try to convince them to. Can you blame them?

Let's talk about some other ways to visualize this same data. The sort of data we have lends itself easily to a matrix structure, with fraud management lifecycle stage across one axis and fraud type across the other. When I see the data organized this way, I think heatmap. But the main drawback to a heatmap in this scenario is that, while it gives us a decent visual comparison of how we're doing across the different buckets (both by fraud management lifecycle stage and by fraud type), we don't get a visual comparison of where we are vs. where we'd like to be, which I think is the most important piece here.

Instead, I'll leverage one of my best friends: the bar chart. Bar charts are great because people already know how to read them. This means there's no learning curve for your audience to face to get to the information you want to provide. Rather than spending their time deciphering how to read the graph, they can spend it understanding the information it shows. There also more likely to spend time on a visual that doesn't make them feel uncomfortable. Here's another way to visualize this data:

Note that the actual numbers aren't so important here - they were somewhat subjective to begin with - so I opted not to show a numerical scale at all. What is important is the relative distance from where we consider ourselves to be currently and where we want to be (as close to "we've solved every problem" as possible). I've drawn attention to this gap by showing the opportunity that remains outlined in blue.

The overarching lesson is this: don't fall victim to choosing sexy over utility when it comes to data viz for telling a story. When your audience tells you something is hard to read, or you find yourself explaining the visual more than discussing the information it shows, listen and adjust!

If interested, my Excel file is here. Leave a comment to let me know what you think!

Wednesday, July 20, 2011

death to pie charts

I hate pie charts. 

I mean, really hate them.

Those who have heard me speak on data visualization will have learned that the only thing I hate more than a pie chart is a 3D, exploding pie chart - they are the absolute worst - but the plain vanilla pie charts are pretty bad, too. Here's a recent one from TechCrunch, which is intended to show how much they cover start-ups versus big companies (full article):


I'll start with the lesser evil of the above visual: meaningless color. The pie above is what happens if you put the data in Excel and say "chart data". I've said this before and I'll say it again: graphing your data with a tool like Excel should be the first step in your design process, not your last! In TechCrunch's pie, the color itself doesn't represent anything, it's simply used as a categorical differentiator. One unintended side effect is the optical illusion you get with a darker colored slice appearing larger than a same-size slice of a ligher color.

My strong opinion is that color should always be an explicit choice and should be used strategically to draw the audience's eye. This preattentive power is being wasted here. If you must use a pie chart, at least make the slices the same color and highlight only the one or two you want to draw attention to. Or if you don't want to highlight a particular slice, but rather are intending the visual to aid in information discovery, you may consider something like the following:

Hopefully you can see that this still isn't a very good visualization. The labels are messy. Only a few things are immediately apparent: General Consumer Web is the biggest piece, there are a lot of small slices.

My main beef with pie charts like the one above (and in general) is this: our eyes aren't good at attributing quantitative value to two dimensional spaces. In English: pie charts are really hard for people to read! When segments are close in size, it'd difficult (if not impossible) to tell which is bigger. When they aren't close in size, the best you can do is determine that one is bigger than the other, but you can't judge by how much. To get over this, you can add data labels, as they've done in the TechCrunch version. But I'd still argue the visual isn't worth the space it takes up.

What should you do instead? My typical advice would be to replace a pie chart with a horizontal bar chart, organized from greatest to least or vice versa (unless there is some intrinsic value in the categories, in which case that should be followed). With bar charts, our eyes compare the end points. Because they are aligned at a common baseline, it’s very easy to assess relative size. This makes it easy to see not only which segment is the largest (for example), but also how incrementally larger it is than the other segments. Here's what this looks like with the TechCrunch data:


One might argue that you lose something in the transition from pie to bar. The unique thing you get with a pie chart that is absent in a bar chart is the concept of there being a whole, and thus, parts of a whole. But if the visual is difficult to read, is it worth it? Ultimately, it's up to the designer of the visual. My advice is as follows:
  1. Don't use pie charts.
  2. If you find yourself unable to follow #1, keep in mind the challenges with pie charts: if relative sizes are important, you'll need to include data labels. Also be aware of impact of color on 2D space (darker looks larger); don't let your tool decide your color scheme. 
Personally, I will continue to avoid pie charts.

Wednesday, June 15, 2011

learning from bad graphics

They say when it comes to understanding how to be a good people manager, you can learn perhaps as much from a bad boss (what not to do) as you can from a great manager (what to emulate). I have to think the same is probably true when it comes to honing one's data visualization skills.

A good visual display invites you in. It's straightforward to read and understand. It reveals something interesting. A bad visuals turn you off. It may look too complicated or messy. It's hard to figure out what's going on. It may also misrepresent the data.

A colleague forwarded an article from TechCrunch recently, Look At Who's Winning The Talent Wars in Tech, 6/7/11. (Perhaps the lack of concision in the title serves as an early warning of bad things to come.) The article discusses the flow of talent between tech companies. It includes the following visual:


To put it bluntly: I am not a fan of this visual. Beyond looking clunky, it's wasting precious space that could be visually informing but is not. How could we make better use of the space this graphic occupies to tell a visual story?

I think the biggest opportunity lost here is having a visual sense of the relative magnitudes of talent flow. Currently, all the arrows convey is directionality; we actually have to read the numbers to get a sense of relative magnitudes, which would probably be much more straightforward in a table than a visual (at least then all of the text would be uniformly horizontal so we could read it!). Scaling the arrows to reflect the relative magnitude of talent flow would solve this.

It would also be great if the size of the circles could represent something meaningful to add context and not just take up space (e.g. number of employees in the given organization). 

If the arrows and circles were scaled, it would give the audience a quick visual sense of what's going on, helping to focus attention to interesting pieces and make information easier to pull out of the visual. A program like Circos could work well for this (view related blog post).

I actually had been thinking of recreating the visual using Circos, but a quick calculation (from info I have exposure to in my day job) made me super suspect of the figures provided, so I decided not to spend my time graphing bad data. The information used in the TechCrunch visual was mined from a social networking job seeker website, so the underlying dataset is not ideal for getting to the information they are after. As one of my colleagues pointed out: this is like using Facebook data to figure out how many people delete MySpace profiles to join Facebook! 

Ah, what we can learn from things done poorly...