-
Notifications
You must be signed in to change notification settings - Fork 0
Home
Welcome to the GettingAndCleaningData wiki!
Choose a column Zip from a variable dat dat$Zip
Setting up a directory
setwd("")
Display the names of the column names()
Find the reccurence of a number in a dataframe: Save the variable in a variable and use length zipcode <- xpathSApply(rootNode,"//zipcode",xmlValue) length(zipcode[zipcode==21231])
Load the csv data into a dataframe library(data.table) DT <- fread(input="fsspid.csv", sep=",")
**Calculate the mean of a variable system.time(DT[,mean(pwgtp15),by=SEX]) system.time(tapply(DT$pwgtp15,DT$SEX,mean)) system.time(mean(DT$pwgtp15,by=DT$SEX))
Sending a url file in a variable
fileURL <- "..."
Download URL file on internet
download.file(url=fileUrl,destfile="idaho_housing.csv",mode="w",method="curl")
Loading packages require(package)
work with xml files fileUrl2 <- "http://d396qusza40orc.cloudfront.net/getdata/data/restaurants.xml"
Parse data and save it to a variable doc <- xmlTreeParse(file=fileUrl2,useInternal=TRUE) rootNode <- xmlRoot(doc) xmlName(rootNode)
List Files
list.files(".")
Send csv file into variable mydata <- read.csv("idaho_housing.csv") head(mydata)
Find data above one value = Length of a file length(mydata$VAL[!is.na(mydata$VAL) & mydata$VAL==24])
Install packages install.packages("....")
**Tidy data Principles ** Tidy data has one observation per row. Each tidy data table contains information about only one type of observation. Each variable in a tidy data set has been transformed to be interpretable. Tidy data has one variable per column.
Load xlxs library library(xlxsjars)
install xlxs library install.packages("xlsxjars")
Read excel file and send it to a variable dat
read.xlsx(file="gov_NGAP.xlsx",sheetIndex=1,colIndex=colIndx,startRow=18,endRow=23, header=TRUE)
Send selected row and columns to a variable rowIndex <- 18:23 colIndx <- 7:15
extract data, download data from html and copy them to a file
f <- file.path(getwd(), "ss06hid.csv") download.file(url, f) dt <- data.table(read.csv(f))
Extract data from a file and assign them to a variable agricultureLogical <- dt$ACR == 3 & dt$AGS == 6
download jpeg image f <- file.path(getwd(), "jeff.jpg") download.file(url, f, mode = "wb") img <- readJPEG(f, native = TRUE)
Load data dtGDP <- dtGDP[X != ""] dtGDP <- dtGDP[, list(X, X.1, X.3, X.4)] setnames(dtGDP, c("X", "X.1", "X.3", "X.4"), c("CountryCode", "rankingGDP", "Long.Name", "gdp"))
match data in fonction of data
dt <- merge(dtGDP, dtEd, all = TRUE, by = c("CountryCode")) sum(!is.na(unique(dt$rankingGDP)))
Sort the data frame in descending order by GDP rank (so United States is last) dt[order(rankingGDP, decreasing = TRUE), list(CountryCode, Long.Name.x, Long.Name.y, rankingGDP, gdp)][13]
What is the average GDP ranking for the "High income: OECD" and "High income: nonOECD" group? dt[, mean(rankingGDP, na.rm = TRUE), by = Income.Group]
Cut the GDP ranking into 5 separate quantile groups. Make a table versus Income.Group. How many countries are Lower middle income but among the 38 nations with highest GDP? breaks <- quantile(dt$rankingGDP, probs = seq(0, 1, 0.2), na.rm = TRUE) dt$quantileGDP <- cut(dt$rankingGDP, breaks = breaks) dt[Income.Group == "Lower middle income", .N, by = c("Income.Group", "quantileGDP")]