Data cleaning tips

Posted to Statistics  |  Tags:  |  Nathan Yau

When you first learn statistics, visualization, or any data-related subject, the data usually is given to you in a ready-to-use format. This is so that you can spend most of your time on the topic of interest. But once you step outside the learning bubble, data rarely comes in the format you want.

Marc Bellemare, an associate professor in the Department of Applied Economics at the University of Minnesota, provides some practical tips on how to deal with this. Bellemare’s parting advice:

Really, there is no big secret to cleaning data other than “Document everything” and to save everything in different files and in different locations (i.e., your computer, Dropbox, Google Drive), and there is no other way to learn data cleaning than by doing it.

Yep.

Some of the tips are in the context of specific software environment, but you can easily apply them to more general situations.

Favorites

Divorce Rates for Different Groups

We know when people usually get married. We know who never marries. Finally, it’s time to look at the other side: divorce and remarriage.

Causes of Death

There are many ways to die. Cancer. Infection. Mental. External. This is how different groups of people died over the past 10 years, visualized by age.

The Changing American Diet

See what we ate on an average day, for the past several decades.

Jobs Charted by State and Salary

Jobs and pay can vary a lot depending on where you live, based on 2013 data from the Bureau of Labor Statistics. Here’s an interactive to look.