1. Getting started with R and RStudio. Organizing the workspace: creating projects, scripts, reports in RMarkdown, working with variables. The main data type — vector, indexing. Ways to create vectors, slicing. Data types. Matrices. Basic descriptive statistics (min, max, mean, median). Indices and values. Accessing help resources.
2. Rectangular data — tables. Creating tables from the command line. Structure and features of tabular data. Reading tables from files in various formats. Writing tables to files. Data manipulation using base R. Which. Handling missing data. Lists. Loops. Apply family functions.
3. Installing packages and loading libraries. Tidyverse. The logic of tidyverse packages, differences from base R. tibble vs data.frame. Creating tibbles. Basic tabular data manipulation with dplyr. Working with groups. Reading and writing tibbles to files in various formats. Using the pipeline in base R and in tidyverse.
4. Working with strings — stringr. Regular expressions. Working with factors — forcats. ggplot2 — the logic of layered graph construction. Scatter plot, bar chart. Graph parameter settings. Wide and long data formats. Working with color, shape, transparency. Saving graphs to files in various formats.
5. ggplot2 (continued). Graph types: line chart, histogram, pie chart, bubble chart, density plot, box plot, violin plot, raincloud plot. Adding additional data to graphs. Combining multiple tables. Using metadata. Advanced tabular data manipulation with dplyr and tidyr.
6. Properties of the normal distribution and the central limit theorem. What statistics is and why it is needed. Population and sample. The difference between parameter estimation and its true value. Sample representativeness. What exploratory data analysis is and how to perform it. Hypothesis H0 and the alternative. Differences. How to assess hypothesis validity using computational simulations. Type I and Type II errors. P‑value and significance level. z‑test and Student’s t‑test. Conditions for application. One‑sample and two‑sample Student’s t‑tests.
7. Difference between paired and two‑sample Student’s t‑tests. Chi‑square test, Fisher’s exact test. Nonparametric tests.
8. Writing custom functions. Building functions, function parameters, default values. Functional programming (map family of functions). Handling exceptions/errors. Importing functions from a file. Creating R scripts that accept multiple input variables.
9. ggplot2 and beyond. Histogram with multiple axes. Grid of graphs. Creating panels of graphs. Simple and complex heatmaps. Visualizing flows. Visualizing networks. Specific considerations for preparing figures for publication or presentation.
10. Basics and logic of Quarto. Creating interactive visualizations and customizable reports.
11. Creating and using interactive dashboards.
12. The problem of multiple testing. Methods for addressing multiple testing. FDR and FWER. What correlation is and what it is not. Correlation.
13. ANOVA. Introduction to regression analysis.