How to ask for datasets

Posted to Data Sources  |  Nathan Yau

There’s a lot of data readily available online. We know this. However, there’s also a lot of data available that’s offline, sitting on people’s hard drives. You just have to ask for it — in the right way. Christian Kreibich, a researcher for the International Computer Science Institute, provides a guide.

Make sure you’ve done your homework. This, too, is key. You need to demonstrate that you know your stuff and have good reasons for the inquiry. To give a negative example, I frequently receive requests for “botnet data”. That could mean anything. Are you interested in malware binaries, traffic captures, NetFlow data? Why, and why would you need mine? Understand the meaning and potential of the data you’re asking for, and be concrete. Understand the implications of obtaining certain datasets, such as privacy concerns, risks to others, or repeatability of the experiment.

Other tips include don’t be a jerk, make your affiliation clear, and make your purpose clear. But the first tip — don’t be shy — is the best. Just because the data isn’t online doesn’t mean the source doesn’t want to share, and it never hurts to ask.

This tweet from Kevin Quealy came to mind:


Where People Run in Major Cities

There are many exercise apps that allow you to keep track of your running, riding, and other activities. Record speed, …

How You Will Die

So far we’ve seen when you will die and how other people tend to die. Now let’s put the two together to see how and when you will die, given your sex, race, and age.

Think Like a Statistician – Without the Math

I call myself a statistician, because, well, I’m a statistics graduate student. However, the most important things I’ve learned are less formal, but have proven extremely useful when working/playing with data.

The Best Data Visualization Projects of 2011

I almost didn’t make a best-of list this year, but as I clicked through the year’s post, it was hard …