Goal
The goals of this post are to (1) introduce basic R objects and functions and (2) to show how useful operations can be performed.
vectors, matrices, and data.frames
Below, I will define three types of objects, examine them, and perform operations on them. Note that “#”, without the parentheses, is R's comment character. That means, any text after # on a given line will NOT be interpreted by R. It is important to use comments to explain what the code is doing, both for the programmer's benefit and for anyone that might use the code later.
### Defining a vector
vector_1 = c(1:100)
### How many elements are in the vector?
print(length(vector_1))
[1] 100
### What are the first 5 elements of a vector? There are multiple ways to
### extract this information, but I will show two methods:
# Method 1
print(vector_1[1:5])
[1] 1 2 3 4 5
# Method 2
print(head(vector_1, 5))
[1] 1 2 3 4 5
### Define objects that equal the mean, median, minimum, maximum, 25th
### percentile, 75th percentile, and sum of the vector
mean_vector_1 = mean(vector_1)
median_vector_1 = median(vector_1)
max_vector_1 = max(vector_1)
min_vector_1 = min(vector_1)
sum_vector_1 = sum(vector_1)
q_25_vector_1 = quantile(vector_1, probs = 0.25)
q_75_vector_1 = quantile(vector_1, probs = 0.75)
### How can this information be extracted more efficiently? Many R objects
### have summary functions that are quite useful -- it doesn't hurt to try the
### summary() function on an object to see what values are returned.
summary_vector_1 = summary(vector_1)
print(summary_vector_1)
Min. 1st Qu. Median Mean 3rd Qu. Max.
1.0 25.8 50.5 50.5 75.2 100.0
### What if I wanted to store this information for later use or for export? I
### will define a data.frame for storage
summary_df_vector_1 = data.frame(Mean = mean_vector_1, Median = median_vector_1,
Max = max_vector_1, Min = min_vector_1, Sum = sum_vector_1, Q25 = q_25_vector_1,
Q75 = q_75_vector_1, stringsAsFactors = F)
row.names(summary_df_vector_1) = NULL
### Let's take a look at the object I just defined (you may have guessed by
### now that the equals sign in 'new thing = some other thing') is used for
### object assignment/definition
print(summary_df_vector_1)
Mean Median Max Min Sum Q25 Q75
1 50.5 50.5 100 1 5050 25.75 75.25
# Object class
print(class(summary_df_vector_1))
[1] "data.frame"
# What are the objects dimensions?
print(dim(summary_df_vector_1))
[1] 1 7
# How many rows?
print(nrow(summary_df_vector_1))
[1] 1
# Columns?
print(ncol(summary_df_vector_1))
[1] 7
### Note that I didn't examine the dimensions of the original vector_1 object
### -- I only examined the length. This is because, vectors DO NOT HAVE
### dimensions == they only have length. Observe:
print(dim(vector_1))
NULL
# NULL' in R literally means NOTHING or NO VALUE so there are no dimensions
### What values does the Sum variable in the summary_df_vector_1 have? There
### are multiple ways to extract the values of a vector that is held within a
### data.frame
# Method 1
print(summary_df_vector_1$Sum)
[1] 5050
# Method 2
print(summary_df_vector_1[, "Sum"])
[1] 5050
### I can also define a new object using data from summary_df_vector_1
extracted_sum = summary_df_vector_1$Sum
### Is this new vector the same as the original vector?
# One test:
print(identical(extracted_sum, sum_vector_1))
[1] TRUE
# Another:
print(extracted_sum == sum_vector_1)
[1] TRUE
### Note that I used ' == ' and not ' = ' in the instance above. The ' == '
### is a LOGICAL OPERATOR and tests the TRUTH of some statement, whereas' = '
### assigns a value to some object.
How do functions work?
The first section used several functions but without much explanation. This section demonstrates how one uses R functions. R Functions have names and they have arguments. Some arguments are defined by default and others require additional input from the user. Users can also program their own functions to perform a series of tasks as well as to return output from the work of the function.
### Virtually all R functions have some sort of documentation in help files or
### online that demonstrate how they work, along with examples. To query the
### help files on one's computer, simply type '?' along with the function name
### into the command line and run the code.
# Example: ?data.frame
### SINCE FUNCTION NAMES AND OBJECT NAMES ARE CASE SENSITIVE IN R, SOMETIMES
### IT IS NECESSARY TO SEARCH THE WEB FOR INFORMATION REGARDING A PARTICULAR
### FUNCTION google: 'data.frame in R' and you'll get some helpful pages as
### the top results. These help pages show what arguments the function has,
### how they change things, (sometimes) some background on the function, and
### (sometimes) helpful examples the user can run on their own.
### Here is how one can extract the arguments from a function.
print(args(data.frame))
function (..., row.names = NULL, check.rows = FALSE, check.names = TRUE,
stringsAsFactors = default.stringsAsFactors())
NULL
### Arguments can be 'passed' to a particular function 'positionally' or by
### directly assigning values to various arguments.
# Example: rnorm() -- the rnorm() function is one of the most commonly used
# functions in R (and in R help literature!) as it performs random draws
# from a normal distribution with mean and standard deviation parameters
# that are specified by the user. Let's see how this function works!
print(args(rnorm))
function (n, mean = 0, sd = 1)
NULL
### The function has arguments of n, mean, and sd. I will define an object
### using different methods of defining object parameters.
# Positional
normal_draw_1 = rnorm(1000, 10, 1)
# Defining each argument
normal_draw_2 = rnorm(n = 1000, mean = 10, sd = 1)
# Both positional and defined
normal_draw_3 = rnorm(1000, mean = 10, sd = 1)
### All three of those objects are the same (bonus points if you test my
### assertion yourself from functions I used in the first code chunk!). Note
### that when the arguments were printed, the mean and sd parameters were
### assigned values of 0 and 1 respectively. Those are the function defaults,
### which means that, as long as n is assigned, the function can operate
### properly and it will use the default values of mean and sd if none are
### supplied.
Summary
That is all for today. I will end the lesson with a little advanced coding just to give folks something to look forward to (and perhaps cheat ahead a little!). Feel free to email me at democraticdata@gmail.com if you have any questions – I will do my best to respond within a reasonable amount of time. Also, note that there is an excellent listserve that handles everything related to R. Here is the link to sign up to the R help listserve: https://stat.ethz.ch/mailman/listinfo/r-help.
standard_normal_distribution = rnorm(1000)
### Now, I'm going to cut ahead a little and show you some cool things R can
### do -- don't worry about understanding what's going on yet but bonus points
### if you can figure it out! Loading libraries
require(ggplot2)
## Loading required package: ggplot2
require(reshape2)
## Loading required package: reshape2
# define new data.frame for plotting
combined_normal_draws = data.frame(Index = c(1:length(normal_draw_3)), random_draw = normal_draw_3,
standard_normal = standard_normal_distribution)
# take a look
print(head(combined_normal_draws))
## Index random_draw standard_normal
## 1 1 8.733 0.2096
## 2 2 10.816 0.7517
## 3 3 9.509 -0.7431
## 4 4 10.857 -0.3762
## 5 5 10.560 0.9273
## 6 6 9.657 -0.7291
# Create 'Long' format for ggplot2
long_combined_normal_draws = melt(combined_normal_draws, id = "Index", measure = c("random_draw",
"standard_normal"))
# Change variable names for plotting
long_combined_normal_draws$variable = as.character(long_combined_normal_draws$variable)
long_combined_normal_draws$variable[long_combined_normal_draws$variable == "standard_normal"] = "Standard Normal\n(Mean 0, SD 1)"
long_combined_normal_draws$variable[long_combined_normal_draws$variable == "random_draw"] = "Random Draw\n(Mean 10, SD 1)"
# take a look
print(head(long_combined_normal_draws))
## Index variable value
## 1 1 Random Draw\n(Mean 10, SD 1) 8.733
## 2 2 Random Draw\n(Mean 10, SD 1) 10.816
## 3 3 Random Draw\n(Mean 10, SD 1) 9.509
## 4 4 Random Draw\n(Mean 10, SD 1) 10.857
## 5 5 Random Draw\n(Mean 10, SD 1) 10.560
## 6 6 Random Draw\n(Mean 10, SD 1) 9.657
print(tail(long_combined_normal_draws))
## Index variable value
## 1995 995 Standard Normal\n(Mean 0, SD 1) 1.1738
## 1996 996 Standard Normal\n(Mean 0, SD 1) 0.4880
## 1997 997 Standard Normal\n(Mean 0, SD 1) -0.8823
## 1998 998 Standard Normal\n(Mean 0, SD 1) -0.6630
## 1999 999 Standard Normal\n(Mean 0, SD 1) 1.1272
## 2000 1000 Standard Normal\n(Mean 0, SD 1) 1.9558
# use a scale parameter
scale = 2
ggplot(long_combined_normal_draws, aes(value)) + facet_wrap(~variable, scales = "free") +
theme_bw() + labs(title = "Density plots of random draws from \nthe normal distribution\n",
x = "\n", y = "Density\n") + theme(axis.text = element_text(size = 12 *
scale, face = "bold"), strip.text = element_text(size = 15, face = "bold"),
strip.background = element_rect(fill = "lightblue"), axis.title = element_text(size = 11 *
scale, face = "bold"), title = element_text(size = 13 * scale, face = "bold")) +
stat_density()
No comments:
Post a Comment