Data Science: Capstone
15-20 hours a week • Start today
Individual Course
Course Length
8 weeks
1-2 hours a week
Featuring faculty from:
Harvard T.H. Chan School of Public Health
Enroll as Individual
Certificate Price:
$ 149
Enroll as Individual
Certificate Price:
$ 149
Learn how to use R to implement linear regression, one of the most common statistical modeling approaches in data science.
Linear regression is commonly used to quantify the relationship between two or more variables. It is also used to adjust for confounding. This course, part of our Professional Certificate Program in Data Science, covers how to implement linear regression and adjust for confounding in practice using R.
In data science applications, it is very common to be interested in the relationship between two or more variables. The motivating case study we examine in this course relates to the data-driven approach used to construct baseball teams described in Moneyball. We will try to determine which measured outcomes best predict baseball runs by using linear regression.
We will also examine confounding, where extraneous variables affect the relationship between two or more other variables, leading to spurious associations. Linear regression is a powerful technique for removing confounders, but it is not a magical process. It is essential to understand when it is appropriate to use, and this course will teach you when to apply this technique.
Self-Guided
edX
Understand how linear regression was originally developed by Galton
Learn about confounding and how to detect it
Learn how to examine the relationships between variables by implementing linear regression in R
Your Instructor
Professor of Biostatistics, Harvard T.H. Chan School of Public Health
Rafael Irizarry is a Professor of Biostatistics at the Harvard T.H. Chan School of Public Health and a Professor of Biostatistics and Computational Biology at the Dana Farber Cancer Institute. For the past 15 years, Dr. Irizarry’s research has focused on the analysis of genomics data. During this time, he has also taught several classes, all related to applied statistics. Dr. Irizarry is one of the founders of the Bioconductor Project, an open source and open development software project for the analysis of genomic data. His publications related to these topics have been highly cited and his software implementations widely downloaded.
Read full bio.
These courses can be bundled together to receive a professional certificate at a discounted price.
Learn More15-20 hours a week • Start today
1-2 hours a week • Start today
1-2 hours per week • Start today
1-2 hours a week • Start today
2-4 hours a week • Start today
1-2 hours a week • Start today
1-2 hours a week • Start today
1-2 hours a week • Start today
Ways to take this course
A Verified Certificate costs $149 and provides unlimited access to full course materials, activities, tests, and forums. At the end of the course, learners who earn a passing grade can receive a certificate.
Alternatively, learners can Audit the course for free and have access to select course material, activities, tests, and forums. Please note that this track does not offer a certificate for learners who earn a passing grade.
Don’t miss a thing. Subscribe to our newsletter and get updates on exclusive content for Harvard Online learners.