Real-world datasets often come with messy text fields and inconsistent date–time formats. In R, the stringr package delivers a consistent, easy-to-use API for text manipulation, while lubridate simplifies parsing, formatting, and arithmetic on date–time objects. This post covers:
Pattern detection with
str_detect()Text replacement using
str_replace()and friendsParsing dates with
ymd(), parsing times withhms()Converting between time zones
Calculating durations and intervals
By the end, you’ll be equipped to clean textual data and handle complex date–time workflows with confidence.
1. String Manipulation with stringr
The stringr package (part of the tidyverse) wraps base R regex functions in a uniform interface. It ensures consistent behavior and clearer code.
install.packages("stringr")
library(stringr)
1.1 Detecting Patterns: str_detect()
Use str_detect() to flag which strings match a pattern. It returns a logical vector of the same length as your input.
emails <- c("alice@example.com", "bob@site", "carol@domain.org")
valid <- str_detect(emails, "^[[:alnum:]._%+-]+@[[:alnum:].-]+\\.[A-Za-z]{2,}$")
# valid: TRUE FALSE TRUE
Common options:
negate = TRUEto return non-matchesignore_case = TRUEfor case-insensitive matching
1.2 Replacing Text: str_replace() & str_replace_all()
Clean or standardize text by replacing patterns:
raw <- c("Price: $100", "price - 200 USD", "Cost: 300$")
cleaned <- str_replace_all(raw, "\\D+", "")
# cleaned: "100", "200", "300"
str_replace()replaces only the first match per string.str_replace_all()replaces every match.
1.3 Other Handy stringr Functions
str_trim()removes leading/trailing whitespace.str_to_lower()/str_to_upper()normalize case.str_split()breaks strings into pieces.str_c()concatenates vectors with a separator.
# Normalize product codes
codes <- c(" Abc-123 ", "XYZ-789")
codes_clean <- codes %>%
str_trim() %>%
str_to_upper()
# "ABC-123", "XYZ-789"
2. Date–Time Parsing with lubridate
Parsing dates and times manually is error-prone. lubridate auto-detects common formats and streamlines conversion.
install.packages("lubridate")
library(lubridate)
2.1 Parsing Dates: ymd(), mdy(), dmy()
Choose the function matching your input order:
dates1 <- ymd(c("2023-01-15", "2023/02/20"))
dates2 <- mdy("03-25-2023")
dates3 <- dmy("31.12.2022")
ymd_hms() handles full timestamps:
timestamps <- ymd_hms("2023-01-15 08:30:00", tz = "UTC")
2.2 Parsing Times: hms()
Isolate time components without dates:
times <- hms(c("12:30:45", "00:05:00"))
# Period object: "12H 30M 45S", "5M 0S"
2.3 Extracting Components
Once parsed, extract parts with accessor functions:
year(dates1) # 2023
month(dates1) # 1, 2
day(dates1) # 15, 20
hour(timestamps) # 8
2.4 Formatting Dates and Times
Convert back to strings in custom formats:
format(dates1, "%d-%b-%Y") # "15-Jan-2023"
format(timestamps, "%H:%M %Z")# "08:30 UTC"
Or use stamp() for reusable formatters:
st <- stamp("Jan 01, 2023")
st(dates1) # "Jan 15, 2023", "Feb 20, 2023"
3. Time-Zone Conversion
Handling global datasets means juggling time zones. lubridate provides two key functions:
force_tz()sets a time-zone without altering the clock time.with_tz()converts a date–time to a new time-zone, adjusting the clock.
utc_time <- ymd_hms("2023-01-15 12:00:00", tz = "UTC")
ny_time <- with_tz(utc_time, tzone = "America/New_York")
# "2023-01-15 07:00:00 EST"
4. Calculating Durations and Intervals
Quantify time differences with durations, periods, and intervals.
4.1 Durations
Durations represent exact seconds:
dur <- duration(3600) # one hour
dur * 5 # 5 hours
4.2 Periods
Periods respect human calendars:
per <- years(1) + months(6) + days(15)
ymd("2022-01-01") + per # "2023-07-16 UTC"
4.3 Intervals
Intervals span between two date–time objects:
start <- ymd_hms("2023-01-01 00:00:00", tz = "UTC")
end <- now(tzone = "UTC")
intv <- interval(start, end)
as.duration(intv) # total seconds
as.period(intv) # years, months, days, etc.
Use int_diff <- intv / dyears(1) to compute fractional years.
5. Practical Example: Cleaning and Computing Durations
Imagine event logs with messy text fields and string timestamps:
library(dplyr)
raw_logs <- tibble(
user = c(" Alice ", "BOB"),
event = c("login", "logout"),
timestamp= c("2023-01-15 08:30:00 UTC", "2023-01-15T10:45:00Z")
)
cleaned <- raw_logs %>%
mutate(
user = str_trim(str_to_title(user)),
timestamp= ymd_hms(timestamp),
tz = tz(timestamp)
) %>%
arrange(timestamp) %>%
mutate(
duration = timestamp - lag(timestamp),
duration = as.period(duration)
)
The result:
user normalized (
"Alice","Bob")timestamp parsed into POSIXct
duration between events as a human-readable period
6. Best Practices
Always trim and normalize case before pattern matching.
Explicitly specify formats when auto-detection fails.
Be cautious with time zones—store and convert times consistently.
Choose periods vs. durations based on calendar vs. exact-second needs.
Validate parsing by checking
class()and summary statistics.
7. Conclusion and Next Steps
You’ve learned how to harness stringr for robust text cleaning and lubridate for flexible date–time parsing, formatting, and arithmetic. These tools unlock deeper insights from messy real-world data.
In the next chapter, we’ll explore Functional Programming and Advanced Tools—using purrr for mapping, writing custom functions, and error handling to elevate your R workflows. If you encountered any parsing challenges or have tips on handling exotic date-time formats, share them in the comments below!

Comments
Post a Comment