Ir al contenido principalSkip to Xpert Chatbot

HarvardX: Data Science: Wrangling

4.4 stars
28 ratings

Learn to process and convert raw data into formats needed for analysis.

Data Science: Wrangling
8 semanas
1–2 horas por semana
A tu ritmo
Avanza a tu ritmo
Gratis
Verificación opcional disponible

Elige tu sesión:

¡Ya se inscribieron 98,213! Una vez finalizada la sesión del curso, será archivadoAbre en una pestaña nueva.
Comienza el 22 nov
Termina el 18 dic
Comienza el 22 nov
Comienza el 16 abr 2025

Sobre este curso

Omitir Sobre este curso

In this course, part of our Professional Certificate Program in Data Science,we cover several standard steps of the data wrangling process like importing data into R, tidying data, string processing, HTML parsing, working with dates and times, and text mining. Rarely are all these wrangling steps necessary in a single analysis, but a data scientist will likely face them all at some point.

Very rarely is data easily accessible in a data science project. It's more likely for the data to be in a file, a database, or extracted from documents such as web pages, tweets, or PDFs. In these cases, the first step is to import the data into R and tidy the data, using the tidyverse package. The steps that convert data from its raw form to the tidy form is called data wrangling.

This process is a critical step for any data scientist. Knowing how to wrangle and clean data will enable you to make critical insights that would otherwise be hidden.

De un vistazo

  • Language English
  • Video Transcript English
  • Associated programs
  • Associated skillsHyperText Markup Language (HTML), Data Wrangling, Parsing, Data Science, Text Mining, Web Pages

Lo que aprenderás

Omitir Lo que aprenderás
  • Importing data into R fromdifferent file formats
  • Web scraping
  • How to tidy data using the tidyverse tobetter facilitateanalysis
  • String processing with regular expressions (regex)
  • Wrangling data using dplyr
  • How to workwith dates and times as file formats
  • Text mining

Preguntas frecuentes

Omitir Preguntas frecuentes

Honor code statement
HarvardX requires individuals who enroll in its courses on edX to abide by the terms of the edX honor code. HarvardX will take appropriate corrective action in response to violations of the edX honor code, which may include dismissal from the HarvardX course; revocation of any certificates received for the HarvardX course; or other remedies as circumstances warrant. No refunds will be issued in the case of corrective action for such violations. Enrollees who are taking HarvardX courses as part of another program will also be governed by the academic policies of those programs.

Research statement
By registering as an online learner in our open online courses, you are also participating in research intended to enhance HarvardX's instructional offerings as well as the quality of learning and related sciences worldwide. In the interest of research, you may be exposed to some variations in the course materials. HarvardX does not use learner data for any purpose beyond the University's stated missions of education and research. For purposes of research, we may share information we collect from online learning activities, including Personally Identifiable Information, with researchers beyond Harvard. However, your Personally Identifiable Information will only be shared as permitted by applicable law, will be limited to what is necessary to perform the research, and will be subject to an agreement to protect the data. We may also share with the public or third parties aggregated information that does not personally identify you. Similarly, any research findings will be reported at the aggregate level and will not expose your personal identity.

Please read the edX Privacy Policy for more information regarding the processing, transmission, and use of data collected through the edX platform.

Nondiscrimination/anti-harassment statement
Harvard University and HarvardX are committed to maintaining a safe and healthy educational and work environment in which no member of the community is excluded from participation in, denied the benefits of, or subjected to discrimination or harassment in our program. All members of the HarvardX community are expected to abide by Harvard policies on nondiscrimination, including sexual harassment, and the edX Terms of Service. If you have any questions or concerns, please contact harvardx@harvard.edu and/or report your experience through the edX contact form.

Este curso es parte del programa Data Science Professional Certificate

Más información 
Instrucción por expertos
9 cursos de capacitación
A tu ritmo
Avanza a tu ritmo
1 año 5 meses
2 - 3 horas semanales

¿Te interesa este curso para tu negocio o equipo?

Capacita a tus empleados en los temas más solicitados con edX para Negocios.