Sitelet https://github.com/NhaTrang/GettingAndCleaningData/wiki
Skip to content
NhaTrang edited this page Feb 26, 2016 · 11 revisions

Welcome to the GettingAndCleaningData wiki!

Choose a column Zip from a variable dat dat$Zip

Setting up a directory setwd("")

Display the names of the column names()

Find the reccurence of a number in a dataframe: Save the variable in a variable and use length zipcode <- xpathSApply(rootNode,"//zipcode",xmlValue) length(zipcode[zipcode==21231])

Load the csv data into a dataframe library(data.table) DT <- fread(input="fsspid.csv", sep=",")

**Calculate the mean of a variable system.time(DT[,mean(pwgtp15),by=SEX]) system.time(tapply(DT$pwgtp15,DT$SEX,mean)) system.time(mean(DT$pwgtp15,by=DT$SEX))

Sending a url file in a variable fileURL <- "..."

Download URL file on internet download.file(url=fileUrl,destfile="idaho_housing.csv",mode="w",method="curl")

Loading packages require(package)

work with xml files fileUrl2 <- "http://d396qusza40orc.cloudfront.net/getdata/data/restaurants.xml"

Parse data and save it to a variable doc <- xmlTreeParse(file=fileUrl2,useInternal=TRUE) rootNode <- xmlRoot(doc) xmlName(rootNode)

List Files list.files(".")

Send csv file into variable mydata <- read.csv("idaho_housing.csv") head(mydata)

Find data above one value = Length of a file length(mydata$VAL[!is.na(mydata$VAL) & mydata$VAL==24])

Install packages install.packages("....")

**Tidy data Principles ** Tidy data has one observation per row. Each tidy data table contains information about only one type of observation. Each variable in a tidy data set has been transformed to be interpretable. Tidy data has one variable per column.

Load xlxs library library(xlxsjars)

install xlxs library install.packages("xlsxjars")

Read excel file and send it to a variable dat read.xlsx(file="gov_NGAP.xlsx",sheetIndex=1,colIndex=colIndx,startRow=18,endRow=23, header=TRUE)

Send selected row and columns to a variable rowIndex <- 18:23 colIndx <- 7:15

extract data, download data from html and copy them to a file

f <- file.path(getwd(), "ss06hid.csv") download.file(url, f) dt <- data.table(read.csv(f))

Extract data from a file and assign them to a variable agricultureLogical <- dt$ACR == 3 & dt$AGS == 6

download jpeg image f <- file.path(getwd(), "jeff.jpg") download.file(url, f, mode = "wb") img <- readJPEG(f, native = TRUE)

Load data dtGDP <- dtGDP[X != ""] dtGDP <- dtGDP[, list(X, X.1, X.3, X.4)] setnames(dtGDP, c("X", "X.1", "X.3", "X.4"), c("CountryCode", "rankingGDP", "Long.Name", "gdp"))

match data in fonction of data

dt <- merge(dtGDP, dtEd, all = TRUE, by = c("CountryCode")) sum(!is.na(unique(dt$rankingGDP)))

Sort the data frame in descending order by GDP rank (so United States is last) dt[order(rankingGDP, decreasing = TRUE), list(CountryCode, Long.Name.x, Long.Name.y, rankingGDP, gdp)][13]

What is the average GDP ranking for the "High income: OECD" and "High income: nonOECD" group? dt[, mean(rankingGDP, na.rm = TRUE), by = Income.Group]

Cut the GDP ranking into 5 separate quantile groups. Make a table versus Income.Group. How many countries are Lower middle income but among the 38 nations with highest GDP? breaks <- quantile(dt$rankingGDP, probs = seq(0, 1, 0.2), na.rm = TRUE) dt$quantileGDP <- cut(dt$rankingGDP, breaks = breaks) dt[Income.Group == "Lower middle income", .N, by = c("Income.Group", "quantileGDP")]

Clone this wiki locally