Skip to main content

Individual Course

Data Science: Wrangling

Course Length

8 weeks

1-2 hours a week

Featuring faculty from:

Harvard T.H. Chan School of Public Health LogoHarvard T.H. Chan School of Public Health

Enroll as Individual

Certificate Price:

$ 149

Enroll as Individual

Certificate Price:

$ 149

In this online course taught by Harvard Professor Rafael Irizarry, learn to process and convert raw data into formats needed for analysis.

As part of our Professional Certificate Program in Data Science, we cover several standard steps of the data wrangling process like importing data into R, tidying data, string processing, HTML parsing, working with dates and times, and text mining. Rarely are all these wrangling steps necessary in a single analysis, but a data scientist will likely face them all at some point.

Very rarely is data easily accessible in a data science project. It's more likely for the data to be in a file, a database, or extracted from documents such as web pages, tweets, or PDFs. In these cases, the first step is to import the data into R and tidy the data, using the tidyverse package. The steps that convert data from its raw form to the tidy form is called data wrangling.

This process is a critical step for any data scientist. Knowing how to wrangle and clean data will enable you to make critical insights that would otherwise be hidden.

Self-Guided

EDX

Learning Outcome

Importing data into R from different file formats

Learning Outcome

How to tidy data using the tidyverse to better facilitate analysis

Learning Outcome

String processing with regular expressions (regex) and wrangling data using dplyr

  • Learn from Harvard faculty
  • Do it on your own time
  • Get a certificate, add it to your resume
  • Be part of the Harvard Community
Data Science for Business values

Faculty

Rafael Irizarry

Your Instructor

Rafael Irizarry

Professor of Biostatistics, Harvard T.H. Chan School of Public Health

Rafael Irizarry is a Professor of Biostatistics at the Harvard T.H. Chan School of Public Health and a Professor of Biostatistics and Computational Biology at the Dana Farber Cancer Institute. For the past 15 years, Dr. Irizarry’s research has focused on the analysis of genomics data. During this time, he has also taught several classes, all related to applied statistics. Dr. Irizarry is one of the founders of the Bioconductor Project, an open source and open development software project for the analysis of genomic data. His publications related to these topics have been highly cited and his software implementations widely downloaded.

Read full bio.

Complete your journey with this Professional Certificate Series

These courses can be bundled together to receive a professional certificate at a discounted price.

Learn More
  • 9 Courses
  • 1 Year & 5 Months
  • Earn Your Certificate
An example HarvardX certificate

Ways to take this course

Audit or Pursue a Verified Certificate

A Verified Certificate costs $149 and provides unlimited access to full course materials, activities, tests, and forums. At the end of the course, learners who earn a passing grade can receive a certificate.

⁠Alternatively, learners can Audit the course for free and have access to select course material, activities, tests, and forums. Please note that this track does not offer a certificate for learners who earn a passing grade.

Stay tuned for more

Don’t miss a thing. Subscribe to our newsletter and get updates on exclusive content for Harvard Online learners.

FAQs