Skip to main content

10 String and Date–Time Handling in R: Cleaning Text and Managing Dates

 

"stringr, lubridate, string handling R, date-time manipulation R, data cleaning R, R tutorial

Real-world datasets often come with messy text fields and inconsistent date–time formats. In R, the stringr package delivers a consistent, easy-to-use API for text manipulation, while lubridate simplifies parsing, formatting, and arithmetic on date–time objects. This post covers:

  • Pattern detection with str_detect()

  • Text replacement using str_replace() and friends

  • Parsing dates with ymd(), parsing times with hms()

  • Converting between time zones

  • Calculating durations and intervals

By the end, you’ll be equipped to clean textual data and handle complex date–time workflows with confidence.

1. String Manipulation with stringr

The stringr package (part of the tidyverse) wraps base R regex functions in a uniform interface. It ensures consistent behavior and clearer code.

r
install.packages("stringr")
library(stringr)

1.1 Detecting Patterns: str_detect()

Use str_detect() to flag which strings match a pattern. It returns a logical vector of the same length as your input.

r
emails <- c("alice@example.com", "bob@site", "carol@domain.org")
valid <- str_detect(emails, "^[[:alnum:]._%+-]+@[[:alnum:].-]+\\.[A-Za-z]{2,}$")
# valid: TRUE FALSE TRUE

Common options:

  • negate = TRUE to return non-matches

  • ignore_case = TRUE for case-insensitive matching

1.2 Replacing Text: str_replace() & str_replace_all()

Clean or standardize text by replacing patterns:

r
raw <- c("Price: $100", "price - 200 USD", "Cost: 300$")
cleaned <- str_replace_all(raw, "\\D+", "")
# cleaned: "100", "200", "300"
  • str_replace() replaces only the first match per string.

  • str_replace_all() replaces every match.

1.3 Other Handy stringr Functions

  • str_trim() removes leading/trailing whitespace.

  • str_to_lower() / str_to_upper() normalize case.

  • str_split() breaks strings into pieces.

  • str_c() concatenates vectors with a separator.

r
# Normalize product codes
codes <- c(" Abc-123 ", "XYZ-789")
codes_clean <- codes %>%
  str_trim() %>%
  str_to_upper()
# "ABC-123", "XYZ-789"

2. Date–Time Parsing with lubridate

Parsing dates and times manually is error-prone. lubridate auto-detects common formats and streamlines conversion.

r
install.packages("lubridate")
library(lubridate)

2.1 Parsing Dates: ymd(), mdy(), dmy()

Choose the function matching your input order:

r
dates1 <- ymd(c("2023-01-15", "2023/02/20"))
dates2 <- mdy("03-25-2023")
dates3 <- dmy("31.12.2022")

ymd_hms() handles full timestamps:

r
timestamps <- ymd_hms("2023-01-15 08:30:00", tz = "UTC")

2.2 Parsing Times: hms()

Isolate time components without dates:

r
times <- hms(c("12:30:45", "00:05:00"))
# Period object: "12H 30M 45S", "5M 0S"

2.3 Extracting Components

Once parsed, extract parts with accessor functions:

r
year(dates1)      # 2023
month(dates1)     # 1, 2
day(dates1)       # 15, 20
hour(timestamps)  # 8

2.4 Formatting Dates and Times

Convert back to strings in custom formats:

r
format(dates1, "%d-%b-%Y")    # "15-Jan-2023"
format(timestamps, "%H:%M %Z")# "08:30 UTC"

Or use stamp() for reusable formatters:

r
st <- stamp("Jan 01, 2023")
st(dates1)                     # "Jan 15, 2023", "Feb 20, 2023"

3. Time-Zone Conversion

Handling global datasets means juggling time zones. lubridate provides two key functions:

  • force_tz() sets a time-zone without altering the clock time.

  • with_tz() converts a date–time to a new time-zone, adjusting the clock.

r
utc_time <- ymd_hms("2023-01-15 12:00:00", tz = "UTC")
ny_time  <- with_tz(utc_time, tzone = "America/New_York")  
# "2023-01-15 07:00:00 EST"

4. Calculating Durations and Intervals

Quantify time differences with durations, periods, and intervals.

4.1 Durations

Durations represent exact seconds:

r
dur <- duration(3600)     # one hour
dur * 5                   # 5 hours

4.2 Periods

Periods respect human calendars:

r
per <- years(1) + months(6) + days(15)
ymd("2022-01-01") + per  # "2023-07-16 UTC"

4.3 Intervals

Intervals span between two date–time objects:

r
start <- ymd_hms("2023-01-01 00:00:00", tz = "UTC")
end   <- now(tzone = "UTC")
intv  <- interval(start, end)
as.duration(intv)         # total seconds
as.period(intv)           # years, months, days, etc.

Use int_diff <- intv / dyears(1) to compute fractional years.

5. Practical Example: Cleaning and Computing Durations

Imagine event logs with messy text fields and string timestamps:

r
library(dplyr)
raw_logs <- tibble(
  user     = c(" Alice ", "BOB"),
  event    = c("login", "logout"),
  timestamp= c("2023-01-15 08:30:00 UTC", "2023-01-15T10:45:00Z")
)

cleaned <- raw_logs %>%
  mutate(
    user     = str_trim(str_to_title(user)),
    timestamp= ymd_hms(timestamp),
    tz       = tz(timestamp)
  ) %>%
  arrange(timestamp) %>%
  mutate(
    duration = timestamp - lag(timestamp),
    duration = as.period(duration)
  )

The result:

  • user normalized ("Alice", "Bob")

  • timestamp parsed into POSIXct

  • duration between events as a human-readable period

6. Best Practices

  • Always trim and normalize case before pattern matching.

  • Explicitly specify formats when auto-detection fails.

  • Be cautious with time zones—store and convert times consistently.

  • Choose periods vs. durations based on calendar vs. exact-second needs.

  • Validate parsing by checking class() and summary statistics.

7. Conclusion and Next Steps

You’ve learned how to harness stringr for robust text cleaning and lubridate for flexible date–time parsing, formatting, and arithmetic. These tools unlock deeper insights from messy real-world data.

In the next chapter, we’ll explore Functional Programming and Advanced Tools—using purrr for mapping, writing custom functions, and error handling to elevate your R workflows. If you encountered any parsing challenges or have tips on handling exotic date-time formats, share them in the comments below!

Comments

Popular posts from this blog

Alfred Marshall – The Father of Modern Microeconomics

  Welcome back to the blog! Today we explore the life and legacy of Alfred Marshall (1842–1924) , the British economist who laid the foundations of modern microeconomics . His landmark book, Principles of Economics (1890), introduced core concepts like supply and demand , elasticity , and market equilibrium — ideas that continue to shape how we understand economics today. Who Was Alfred Marshall? Alfred Marshall was a professor at the University of Cambridge and a key figure in the development of neoclassical economics . He believed economics should be rigorous, mathematical, and practical , focusing on real-world issues like prices, wages, and consumer behavior. Marshall also emphasized that economics is ultimately about improving human well-being. Key Contributions 1. Supply and Demand Analysis Marshall was the first to clearly present supply and demand as intersecting curves on a graph. He showed how prices are determined by both what consumers are willing to pay (dem...

Fundamental Analysis Case Study NVIDIA

  Executive summary NVIDIA is analyzed here using the full fundamental framework: balance sheet, income statement, cash flow statement, valuation multiples, sector comparison, sensitivity scenarios, and investment checklist. The company shows exceptional profitability, strong cash generation, conservative liquidity and net cash, and premium valuation multiples justified only if high growth and margin profiles persist. Key investment considerations are growth sustainability in data center and AI, margin durability, geopolitical and supply risks, and valuation sensitivity to execution. The detailed numerical work below uses the exact metrics you provided. Company profile and market context Business model and market position Company NVIDIA Corporation, leader in GPUs, AI accelerators, and related software platforms. Core revenue streams : data center GPUs and systems, gaming GPUs, professional visualization, automotive, software and services. Strategic advantage : GPU architecture, C...

“This Sentence Is False”: The Liar Paradox, from Ancient Crete to Modern Code

 “All Cretans are liars,” said the Cretan Epimenides.  “This sentence is false,” echoes every logic textbook.  We’re still arguing 2,600 years later—and the paradox is winning.   _____________________________  /                             \ |   “THIS SENTENCE IS FALSE.”  |  \_____________________________/               |               |  self-reference               v    +---------------------------+    |  Truth flips back on     |    |  itself — paradox loop!  |    +---------------------------+ 1. Meet the Liar The classic one-liner: L: “This sentence is false.” If L is true, then what it asserts—its own falsity—must hold, so L is false. If L is false, then what it asserts isn’t the ca...