Extract CSV data from PDF files with Tabula

Posted to Software  |  Tags: ,  |  Nathan Yau

Tabula, by Manuel Aristarán, came out months ago, but I’ve been poking at government data recently and came back to this useful piece of free software to get the data tables out of countless free-floating PDF files.

If you’ve ever tried to do anything with data provided to you in PDFs, you know how painful this is — you can’t easily copy-and-paste rows of data out of PDF files. Tabula allows you to extract that data in CSV format, through a simple interface.

It’s not the fastest software in the world, but it really is simple to use and it sure beats manual entry. You just load a PDF file into Tabula, which runs on your computer, highlight the table to extract, and the program does the rest. Save as a CSV and do what you want with it.

Download Tabula here. Find out a little more about it on Source.

Favorites

This is an American Workday, By Occupation

I simulated a day for employed Americans to see when and where they work.

Jobs Charted by State and Salary

Jobs and pay can vary a lot depending on where you live, based on 2013 data from the Bureau of Labor Statistics. Here’s an interactive to look.

Years You Have Left to Live, Probably

The individual data points of life are much less predictable than the average. Here’s a simulation that shows you how much time is left on the clock.

Reviving the Statistical Atlas of the United States with New Data

Due to budget cuts, there is no plan for an updated atlas. So I recreated the original 1870 Atlas using today’s publicly available data.