Saturday, April 12, 2014

Working with Data in R -- Import, Manipulate, Export -- Part 1

Goal

Goal

The goals of this post are to demonstrate how to (1) import data; (2) manipulate vectors and data.frames to create and store new variables (from existing data); (3) calculate groupwise summary statistics; and (4) export data.

This post will use a non-trivial example for demonstration purposes. Let's say you are interested in determine the extent to which unemployment rates in Califoria counties are related to the national unemployment rate. Historical unemployment data can be downloaded for California and sub-state geographies here. Using this tool, I downloaded all available unemployment data for all California counties. This will be the primary data used for this post – data for the U.S. writ large will be addressed at a later point in time (trust me, it's much easier to obtain!).

Importing data

The first step is to import data – to tell R to read a structured dataset. This can actually be quite tricky because idiosyncracies of operating systems can get in the way. I exclusively use Windows machines to run R so some of the tricks I show in terms of navigating a system may only apply to Windows. But there are plenty of sources online that can help with specific problems – Google search is your friend! That being said, these problems should only apply to locating files in a system – they will not affect how to program with R.

In order to tell R to read in a file, one must tell the machine where to look for said file. Programmers refer to the location of a file on a computer as the “file path”. Folders are called “directories”. While one can read in files by using a specific file path each time, I find it more efficient to set working directories instead. So instead of telling R to “import file 'A' from path 'x/y/z/½/3', I prefer so tell R that "I'm going to be working in location 'x/y/z/½/3', now import file 'A' or 'B' or 'C'”.

How does one locate and refer to a directory? To locate a directory in windows, simply keep clicking until you find the file that you downloaded. Yes, that is a cheeky answer but this post assumes one can perform basic operations on Windows.

This is what you should see once you're done clicking. See the row at the top of the box? Click that row. This is what you should see now. Copy the highlighted text (CTRL + C or right-click) – that is your working directory. However, note that Windows directories use backslashes ("\") to separate folders, whereas R uses forward slashes ("/), so just change all backslashes to forward slashes and you will be all set. Now we can do some programming!

A side note – most of the learning experience is doing so I will not make life easy and provide the file that I downloaded. Follow the link I provided, use the download tool, download the file, create a well-organized location for your project(s), and move the file there.

# Set working directory
setwd("C:/Users/taylor/Dropbox/democratic-Data/Data/Unemployment/California/County")

# print the working directory to confirm the change
getwd()
[1] "C:/Users/taylor/Dropbox/democratic-Data/Data/Unemployment/California/County"

# Examine the contents of the directory
list.files() # this gives the names of files and folders in the given location. The name of the file that you downloaded will show up here
[1] "DA2014150.txt"

## Read in the data
# The website exports comma delimited .txt files, which are values separated by commas. Comma separated values saved as .txt or .csv are the most common type of files used with R and many other applications. Familiarize yourself with how they look.

# One simply has to tell R where the data is and how the data is formatted in order to read it in
args(read.table)
function (file, header = FALSE, sep = "", quote = "\"'", dec = ".", 
    row.names, col.names, as.is = !stringsAsFactors, na.strings = "NA", 
    colClasses = NA, nrows = -1, skip = 0, check.names = TRUE, 
    fill = !blank.lines.skip, strip.white = FALSE, blank.lines.skip = TRUE, 
    comment.char = "#", allowEscapes = FALSE, flush = FALSE, 
    stringsAsFactors = default.stringsAsFactors(), fileEncoding = "", 
    encoding = "unknown", text) 
NULL

CA_unemployment = read.table("DA2014150.txt", # make sure to include the extension (the .txt part)
                             header = T, # this says that there is a series of column names for the data
                             sep = ",", # this says that the separator for the values is a comma
                             stringsAsFactors = F) # just do this, it will make your life easier. The curious can Google "factors in R" to understand why. 

# What type of object is the data stored as?
class(CA_unemployment)
[1] "data.frame"

# Make sure the data seems to have been read in properly -- compare what R "sees" and what you see when you open the file in windows
head(CA_unemployment, 3)
  Year Period           Area Adjusted Preliminary Labor.Force Employment
1 2014    Jan Alameda County  Not Adj  Not Prelim     780,600    727,600
2 2014    Feb Alameda County  Not Adj      Prelim     782,000    729,900
3 2013 Annual Alameda County  Not Adj  Not Prelim     783,100    725,000
  Unemployment Unemployment.Rate
1       53,000               6.8
2       52,000               6.7
3       58,000               7.4
tail(CA_unemployment, 3)
      Year Period        Area Adjusted Preliminary Labor.Force Employment
16168 1990    Oct Yuba County  Not Adj  Not Prelim      21,700     19,700
16169 1990    Nov Yuba County  Not Adj  Not Prelim      21,600     19,200
16170 1990    Dec Yuba County  Not Adj  Not Prelim      21,500     18,800
      Unemployment Unemployment.Rate
16168        2,000               9.2
16169        2,400              11.3
16170        2,700              12.6

# Everything looks good -- data has been successfully imported!

Summary

That is all for today – part 2 will demonstrate how to manipulate data and how to create/store new variables. As always, please feel free to email me at democraticdata@gmail.com if you have any questions.

No comments:

Post a Comment