name age sex loc
1 Raju 11 <NA> Urban
2 Raj 9 boy Rural
3 Raba NA girl Urban
4 Rahul 10 boy Urban
5 Rimi 5 girl <NA>
2 Some useful functions
# Variable names of the data framenames(df)
[1] "name" "age" "sex" "loc"
# Dimension of the data framedim(df)
[1] 5 4
# Details of a dfstr(df)
'data.frame': 5 obs. of 4 variables:
$ name: chr "Raju" "Raj" "Raba" "Rahul" ...
$ age : num 11 9 NA 10 5
$ sex : Factor w/ 2 levels "boy","girl": NA 1 2 1 2
$ loc : Factor w/ 2 levels "Urban","Rural": 1 2 1 1 NA
# Summary of the data framesummary(df)
name age sex loc
Length :5 Min. : 5.00 boy :2 Urban:3
N.unique :5 1st Qu.: 8.00 girl:2 Rural:1
N.blank :0 Median : 9.50 NAs :1 NAs :1
Min.nchar:3 Mean : 8.75
Max.nchar:5 3rd Qu.:10.25
Max. :11.00
NAs :1
# Summary of a specific variablesummary(df$age)
Min. 1st Qu. Median Mean 3rd Qu. Max. NAs
5.00 8.00 9.50 8.75 10.25 11.00 1
# Frequency table of a variabletable(df$sex)
boy girl
2 2
3 Ordering data frames
We want to reorder the observations of the data df by the variable age.
Recall:order() is used to order an atomic vector by its value. Remember the following example?
age <-c(11, 9, 8, 10, 5)sort(age)
[1] 5 8 9 10 11
order(age)
[1] 5 3 2 4 1
age[order(age)] #equivalent to sort()
[1] 5 8 9 10 11
# Original datadf
name age sex loc
1 Raju 11 <NA> Urban
2 Raj 9 boy Rural
3 Raba NA girl Urban
4 Rahul 10 boy Urban
5 Rimi 5 girl <NA>
# Ordering the data by `age`df[order(df$age), ]
name age sex loc
5 Rimi 5 girl <NA>
2 Raj 9 boy Rural
4 Rahul 10 boy Urban
1 Raju 11 <NA> Urban
3 Raba NA girl Urban
4 Handling missing data
The NA (Not Applicable) character is used as a placeholder of missing observation in R
Most of the R functions have an argument na.rm, which takes a logical value to exclude the missing value from the calculation
mean(c(1:10, NA, 14:16), na.rm =TRUE)
[1] 7.692308
na.omit() is used to exclude all rows of a data frame that include a missing observation
df
name age sex loc
1 Raju 11 <NA> Urban
2 Raj 9 boy Rural
3 Raba NA girl Urban
4 Rahul 10 boy Urban
5 Rimi 5 girl <NA>
na.omit(df)
name age sex loc
2 Raj 9 boy Rural
4 Rahul 10 boy Urban
5 Adding new column or rows
Adding a new variable using $
df$place <-c("UK", "BN", "PK", "IN", "BN")df
name age sex loc place
1 Raju 11 <NA> Urban UK
2 Raj 9 boy Rural BN
3 Raba NA girl Urban PK
4 Rahul 10 boy Urban IN
5 Rimi 5 girl <NA> BN
# removing locdf$loc <-NULLdf
name age sex place
1 Raju 11 <NA> UK
2 Raj 9 boy BN
3 Raba NA girl PK
4 Rahul 10 boy IN
5 Rimi 5 girl BN
name age sex place
1 Raju 11 <NA> UK
2 Raj 9 boy BN
3 Raba NA girl PK
4 Rahul 10 boy IN
5 Rimi 5 girl BN
# A subset of boy's datadf_boy <- df[df$sex =="boy", ]df_boy
name age sex place
NA <NA> NA <NA> <NA>
2 Raj 9 boy BN
4 Rahul 10 boy IN
# Mean age of boysmean(df_boy$age)
[1] NA
7 Exercise 8.1
The data mtcars comprises fuel consumption and 10 aspects of automobile design and performance for 32 automobiles. Load the data by running data(mtcars)
Obtain the variable list of the data frame mtcars
How many observations and variables do the mtcars data have?
Check the types of the variables of mtcars
Rename the variable hp to horsepower
Order the dataset in ascending order of the variable mpg (miles per gallon)
Convert the variable cyl (number of cylinders) to factor type variable
Create a subset of the mtcars dataset where mpg is less than 30, retaining only the first five variables. Save the resulting dataset as mtcars_subset.
8 Frequency table
A frequency table (known as frequency distribution) is a tabular format of summarizing data where frequency corresponding each data point is presented
Frequency table of ungrouped data is often not so useful.