Wednesday, March 19, 2014

Education vs. Infant Mortality, Part 1

Goal

Goal

The goal of this post will be to demonstrate how to use R to analyze an open data source and to produce high quality output in a short amount of time. Specifically, this will be a multi-part post that seeks to examine the relationship between education of females and infant mortality on a global scale. I will start with directly with analysis and output and and the end of the series, I will make all of my code available for review.

The Data

Data were obtained from the World Bank, which offers an API or application programmers interface, which greatly reduces the time and effort required to locate and download data series of interest. The WDI package in R, authored by Vincent Arel-Bundockin, R provides a useful interface with World Bank Data. In this case, I simply started with two search terms, “education” and “mortality” and I was able to find the unique identifiers and descriptions for the data I was interested in.

Data were extracted from all available countries from 1980 to 2014. However, it is often the case that global macro data has a significant amount of missingness and this is no exception. Missing data doesn't mean that it has been lost per se – it means that there was no official record of a particular data point for a particular country in a particular year. Missing data can occur for any number of reasons. Some causes of missingness can actually become a problem when trying to estimate relationships between variables if the cause of the missingness is related to the outcome of interest.

In this case, the question to ask is: are countries/years with more missing data more likely to have higher or lower rates of infant mortality? Without having done any analysis to investigate this question, I would assume that the answer is yes. The reasons: (a) data is more likely to be missing in earlier time periods because many countries in the sample were less likely to track and record macro statistics and (b) there is reason to suspect (putting aside my own prior knowledge of the subject) that infant mortality rates have decreased over time. At the very least, this means that estimates of the trend in infant mortality rates over time will be understated.

Comparing Female Eduation to Infant Mortality

Graphical Analysis

A direct comparison between the primary outcome of interest (infant mortality) and the primary predictor (female education) is always a good place to start. A visual inspection of the data allows one to infer a great deal about subsequent analysis. The first things to look for are the type, strength, and direction of a relationship. That being said, I will produce another post on the topic of assessing functional form, strength, and direction of relationships so this post can be more focused on the analysis at hand.

First, note that each plot on the graph represents a country-year, i.e. “the United States in 2000” has one data point on the graph. In addition, the purple line is the “least squares regression line,” which is a common method of estimating and summarizing the relationship between variables. In this case, the relationship between (a lack of) female education and infant mortality is linear, relatively strong, and postive. That is to say, as the proportion of females aged fifteen and over without education increases, so too does the infant mortality rate and the mortality rate of children five and younger. There is some clustering at the bottom left at the plot, which is probably the domain of wealthier, more “developed” countries. Subsequent investigations should determine the extent to which the relationship between female education and mortality differs between developed and developing nations.

plot of chunk main_plot

Estimate the Relationship

After taking an initial look at primary variables of interest, it is helpful to estimate the strength of the relationship. In this case, I want to estimate the slope of the purple line that appears in the plot above. The questions is, for every percent increase in females without education, what is the associated increase in the number of infant deaths per 1,000 births? First note that a causal relationship is not assumed here – this pattern can be caused by a number of factors. In order to make causal inferences, one must first consider the plausible causal processes that link x to y and one must do at least a reasonable job of adjusting for other factors that also have an impact on infant mortality. For instance, it may be that countries that are more likely to have a high proportion of uneducated females are also more likely to have poor access to adequate nutrition and health care, which should also have an impact on mortality rates. In that case, it may be that education of males may have a similar relationship to child mortality

That being said, regressions can still be useful investigative tools, even if causal inferences cannot necessarily be made from the results.

Below is a table of regression results:

————————————————–

Outcome Variable Estimate Std. Error t value Pr(>|t|)
Mortality rate, infant (per 1,000 live births) (Intercept) 12.318 1.17286 10.503 1.255e-22
Mortality rate, infant (per 1,000 live births) df$Female_Education 1.064 0.04126 25.792 4.035e-83
Mortality rate, under-5 (per 1,000 live births) (Intercept) 13.469 1.95098 6.904 2.373e-11
Mortality rate, under-5 (per 1,000 live births) df$Female_Education 1.786 0.06863 26.029 4.855e-84

————————————————–

The “Outcome” column indicates which outcome was used in the regression. The “Variable” column indicates which variable was estimated. The “Estimate” column includes the estimates of the slope for a particular variable. Ignore the other columns for now. The “(Intercept)” variables indicate what value of y (when x is zero) best aids in producing the best fit to the data – this doesn't help in interpreting results so it is usually just ignored. The “df$Female_Education” variable is the main predictor, the proportion of females aged fifteen and over with no education. In order to aid in interpretation of estimates, I multiplied this proportion by 100 (to convert the proportion to percent).

The estimate for the first outcome is 1.064. This means, for a 1 unit (percent) increase in x (prorportion of uneducated females), one should expect an increase of 1.064 y (infant mortality rate per 1,000 live births). Therefore, for every 10% increase/decrease in the proportion of uneducated females in a given country-year, one should expect an increase/decrease of 10.64 infant deaths per 1,000 live births. The estimate for the under five mortality rate is 1.786 – this value has the same interpretation.

The “Pr(>|t|)” column indicates the probability that the relationship implied by the estimates reported here are due to chance. In this case, the “P values” are 4 x 10-83 and 4 x 10-84, which are incredibly small numbers. Depending on the subject at hand, the conventional P value for which something is deemed “statistically significant” is 0.05, and these are well below that. This simply means that the association we have observed is unlikely due to chance alone.

Summary

That is all for Part 1. The next step will be to produce summary statistics and associated visualizations to examine underlying characteristics of the data. As always, feel free to email me at democraticdata@gmail.com if you have any questions.

————————————————–

No comments:

Post a Comment