Skip to main content

Data Import in R

Importing data means choosing a parser, preserving the intended column types, and checking the resulting object before analysis. This guide covers local delimited files, Excel, JSON, RData, and RDS that fit in memory; database queries and streaming pipelines are outside its scope. Start with R vectors, data frames, and missing values if those operations are unfamiliar.

The inline-text and temporary-file examples run without project data. Paths under data/data-import/ are illustrative inputs, not bundled datasets: replace them with your files and ensure the parent directory exists before writing. Relative paths are resolved from getwd(). Base R readers need no extra packages; examples using readr, data.table, readxl, or jsonlite require those packages already installed.

The R for Data Science import chapter provides a small students dataset containing missing markers and an age written as text. Work through its column-type controls and inspect parsing problems; a successful read should not be your only check.

Symbol-Separated Files​

Delimited-text files represent tabular data by separating fields with delimiters such as commas or tabs. Knowing how to import them is essential. Here, "symbol" refers to any delimiter used to separate data, commonly commas (,), and tab characters (\t), known as CSV and TSV files, respectively.

CSV​

CSV files typically have the .csv extension. The extension does not affect the file's content but helps quickly identify the format and aids automatic interpretation by some software. For example, a CSV file representing student grades might look like this:

student,chinese,math,english
stu1,99,100,98
stu2,60,50,88

R provides read.table() to import delimited files. Here’s how to use it to import the data directly from text:

stu <- read.table(text = "
student,chinese,math,english
stu1,99,100,98
stu2,60,50,88
", header = TRUE, sep = ",")
stu
class(stu)

Usually, data is stored in files on the computer. To import a CSV file using read.table():

cars <- read.table(file = "data/data-import/mtcars.csv", header = TRUE, sep = ",")

Check the first few rows:

head(cars)

Alternatively, use read.csv() for CSV files, which simplifies the process:

cars2 <- read.csv(file = "data/data-import/mtcars.csv")
head(cars2)

Preserve meaning before calculating​

Type inference can turn an identifier like 001 into the number 1. A quoted comma belongs inside a field, while an unquoted comma separates fields; splitting lines with a plain string split would lose that distinction. Declare types and missing-value markers when the format is known:

csv <- 'id,score,note
001,10,"good, complete"
002,,pending'
scores <- read.csv(
text = csv,
colClasses = c("character", "numeric", "character"),
na.strings = c("", "NA"),
check.names = FALSE
)
stopifnot(
identical(names(scores), c("id", "score", "note")),
nrow(scores) == 2L,
identical(scores$id, c("001", "002")),
is.na(scores$score[2]),
!anyDuplicated(scores$id),
all(is.na(scores$score) | (scores$score >= 0 & scores$score <= 100))
)
str(scores)
colSums(is.na(scores))

The result has two rows and three columns; the second score is missing, not zero. The uniqueness and range checks express this example's schema, not universal CSV rules. head() alone cannot detect a malformed value near the end. For real files, also check expected row counts, required columns, units, and date interpretation. fileEncoding = "UTF-8" requests input decoding when that is the file's actual encoding; guessing it does not repair damaged text. read.csv2() handles the common semicolon-separated, decimal-comma variant. read.table() defaults to comment.char = "#", whereas read.csv() disables comments; use comment.char = "" when a literal # is data.

Efficient Data Import with readr and data.table​

readr and data.table::fread() provide alternative readers. Compare parsing controls, returned classes, memory use, and measured runtime on your own input rather than assuming one reader is always faster.

Create one temporary CSV so each reader sees the same input:

temp_csv <- tempfile(fileext = ".csv")
writeLines(c("id,value", "1,10", "2,20"), temp_csv)

z1 <- read.csv(temp_csv)

Using readr​

z2 <- readr::read_csv(temp_csv, show_col_types = FALSE)

Using data.table​

z3 <- data.table::fread(temp_csv)

The three functions return related but different object classes:

class(z1)
class(z2)
class(z3)

unlink(temp_csv)

With readr, specify col_types when guessing is unsafe and inspect readr::problems(z2) for parsing failures. show_col_types = FALSE hides the type message, not validation work. A successful read can still have wrong inferred types, repaired names, or unexpected missing values. All examples here load data into memory; a faster parser does not remove that capacity limit.

TSV and Other CSV Variants​

TSV files use the tab character as a delimiter and can be imported similarly by specifying sep = "\t":

mt <- read.table("data/data-import/mtcars.tsv", sep = "\t", header = TRUE)
mt

Using readr:

mt2 <- readr::read_tsv("data/data-import/mtcars.tsv")
mt2

Using data.table:

mt3 <- data.table::fread("data/data-import/mtcars.tsv")
mt3

Excel​

Excel files are widely used for data storage and processing. The readxl package imports Excel data:

library(readxl)
mt_excel <- read_excel("data/data-import/mtcars.xlsx")
head(mt_excel)

To read a specific sheet, define the workbook path first:

excel_path <- "data/data-import/example.xlsx"
readxl::excel_sheets(excel_path)
iris <- readxl::read_excel(excel_path, sheet = "iris")
head(iris)

The workbook path and sheet name are project inputs; replace them with an existing workbook path and the name of a sheet in that workbook.

Check the selected rectangle, header row, and column types, not just the sheet name. readxl::read_excel() accepts range, skip, col_types, and na; for example, force an identifier column to "text" when inference would lose meaning. Blank cells are missing by default. If Excel already stored 001 as the numeric value 1 with display formatting, reading it as text cannot reconstruct the original intended identifier automatically.

JSON​

JSON is a lightweight data-interchange format. The jsonlite package converts between JSON and R objects:

jsonlite::toJSON(letters)
jsonlite::toJSON(c(a = 1L, b = 2.0))
jsonlite::toJSON(data.frame(a = 1:3, b = 2:4))
jsonlite::toJSON(list(a = 1L, b = 2:5, c = c(TRUE, FALSE), d = NULL))

Save JSON data to a file:

jsonlite::write_json(list(a = 1L, b = 2:5, c = c(TRUE, FALSE), d = NULL), path = "data/data-import/example.json")

Read JSON data:

jsonlite::read_json("data/data-import/example.json")
jsonlite::read_json("data/data-import/example.json", simplifyVector = TRUE)

read_json() returns nested lists by default; simplifyVector = TRUE attempts to turn compatible structures into vectors or data frames. Inspect str() before assuming tabular rows. A missing field, JSON null, and an empty array carry different meanings; decide how each should map to missing values in your analysis. JSON does not preserve every R class or attribute, so it is not a general replacement for RDS when exact R object structure matters.

R Data Files​

Using R's native data storage formats, RData and RDS, is efficient and common for saving and loading R objects.

RData​

RData files can save multiple named objects. Because load() writes them into an environment, use a dedicated environment when names should not leak into the global workspace:

d1 <- head(mtcars)
d2 <- head(iris)
save(d1, d2, file = "data/data-import/example.RData")

loaded <- new.env(parent = emptyenv())
load("data/data-import/example.RData", envir = loaded)
ls(loaded)

RDS​

RDS files are for single objects and allow renaming upon loading:

saveRDS(mtcars, file = "data/data-import/mtcars.rds")
mtcars_rename <- readRDS("data/data-import/mtcars.rds")
head(mtcars_rename)

RDS stores one object, which may itself be a list of many objects. readRDS() returns it for assignment; load() restores stored names into an environment and returns those names. The separate environment above prevents accidental name replacement, but it is not a security sandbox. Only deserialize R files from sources you trust.

Common Issues and Solutions​

Loading Data from Clipboard​

This "clipboard" connection is a Windows pattern, not a portable file path; clipboard access on other platforms depends on the session and available facilities. See R connections. For a repeatable analysis, save the input as a file instead.

data <- read.table('clipboard', header=TRUE)

Reading Line by Line​

Use readLines() to read file content line by line:

fil <- tempfile(fileext = ".data")
cat("TITLE extra line", "2 3 5 7", "", "11 13 17", file = fil, sep = "\n")
readLines(fil, n = -1)
unlink(fil) # Clean up

Fixed-Width File Format​

Use read.fwf() or readr::read_fwf() for fixed-width format files.

Explore connectionsOpen network