[R] Why data manipulation is crucial and sensitive?

What does a data scientist really do?

Identifying the pattern in cultural consumption, making fancy graph, engage a dialogue between data and the existing literature, refining hypothesis....(done within one months with three to four online meetings with partners = no more than 35 hours to agree on the main assertions)

|---------------------------------------------------------------------------------|-----|------|
| Litteratue review | 60 | 20% |
| Primary definition of the hypothesis | 5 | 2% |
| Getting familiar with the codebook and the survey | 10 | 3% |
| Explore the potential variable of interest | 20 | 7% |
| Rename the variable of interest | 15 | 5% |
| Recode the variable of interest and translate in English | 70 | 23% |
| Non answer cleaning | 5 | 2% |
| Rename the labels (levels) | 10 | 3% |
| Primary analysis of the outputs (inspect the recoded variable and bivariate an) | 20 | 7% |
| Reformulation of some hypothesis | 5 | 2% |
| Plotting the first MCA and analyze them | 15 | 5% |
| Compare model strength and understand the primary outputs | 5 | 2% |
| Refining hypothesis and assertions | 15 | 5% |
| Writing the article | 50 | 16% |
| | 305 | 100% |

What is "cleaning and organizing data"?

Definition:

  • Cleaning and organizing data refer to the process of preparing raw data for analysis by identifying and correcting errors, inconsistencies, and inaccuracies, and structuring it in a way that facilitates effective analysis.

Steps Involved:

  • Data Cleaning:
    • Handling missing values.
    • Removing duplicates.
    • Correcting errors and inconsistencies.
  • Data Organization:
    • Structuring data in a readable format.
    • Categorizing and labeling data.
    • Creating meaningful variables.

Removing Duplicates:

R 复制代码
# Removing duplicate rows
unique_data <- unique(raw_data)

Correcting Errors and Inconsistencies

R 复制代码
# Replacing incorrect values
corrected_data <- replace(raw_data, incorrect_condition, replacement_value)

Structuring Data:

R 复制代码
# Creating a data frame
structured_data <- data.frame(variable1 = vector1, variable2 = vector2, ...)

Categorizing and Labeling Data:

R 复制代码
# Creating factors for categorical variables
categorized_data <- factor(raw_data$variable, levels = c("Category1", "Category2", ...))

Creating Meaningful Variables:

R 复制代码
# Creating a new variable based on existing ones
raw_data$new_variable <- raw_data$variable1 + raw_data$variable2

Common issue with online survey

Data not writen in the good format: the typical issue with year of birth

R 复制代码
家庭状况与教育经历
47、您的出生年份是?(请填写整数,例如:1984) (填空题 *必答)
________________________

Section Familial Situation and Education background
47. Which year are you born (Please write number such as 1984)

In the raw data, we have 2 people born in 1898 = 120 years old

Given the average age of the population they are likely to be born in 1998

25 respondents answered using the format Year/Month/Birth

Ex: 19940105

2 respondents answered using very original format

Ex: 930524 / 197674

1 respondent just answer 1

How to clean efficiently with R?

tidyR

= it is a very important package to transform a long table from a wide table

Will not be covered, but basic operation using tidyr are explained in this website: https://mgimond.github.io/ES218/Week03b.html

dplyr

dplyr is a very important package that enables you to select specific variable and data, and to transform them

dplyr Package in R:
  1. Selection of Specific Variables:

    • select() function: It allows you to choose specific columns (variables) from a data frame.

      R 复制代码
      # Example: Selecting columns "variable1" and "variable2"
      selected_data <- select(your_data_frame, variable1, variable2)
  2. Filtering Data:

    • filter() function: Enables you to subset your data based on specific conditions
    R 复制代码
    # Example: Filtering data where "variable1" is greater than 10
    filtered_data <- filter(your_data_frame, variable1 > 10)
  3. Transformation (Mutating) Data:

    • mutate() function: Allows you to create new variables or modify existing ones.
    R 复制代码
    # Example: Creating a new variable "new_variable" as a transformation of existing variables
    mutated_data <- mutate(your_data_frame, new_variable = variable1 + variable2)
  4. Arranging Data:

    • arrange() function: Sorts rows based on specified variables.
    R 复制代码
    # Example: Sorting data based on "variable1" in ascending order
    sorted_data <- arrange(your_data_frame, variable1)
  5. Summarizing Data:

    • summarize() function: Aggregates data, often used with functions like mean, sum, etc.
    R 复制代码
    # Example: Calculating the mean of "variable1"
    summary_stats <- summarize(your_data_frame, mean_variable1 = mean(variable1))

The magrittr package

The magrittr package offers a set of operators which make your code more readable by:

structuring sequences of data operations left-to-right, avoiding nested function calls, minimizing the need for local variables and function definitions, and making it easy to add steps anywhere in the sequence of operations.

The operators pipe their left-hand side values forward into expressions that appear on the right-hand side, i.e. one can replace f(x) with x %>% f(), where %>% is the (main) pipe-operator.

https://magrittr.tidyverse.org/

相关推荐
2301_803875614 分钟前
CSS如何制作导航栏平滑移动_使用transition与left属性
jvm·数据库·python
2501_933329554 小时前
媒介宣发技术实践:Infoseek舆情系统的AI中台架构与应用解析
开发语言·人工智能·架构·数据库开发
zxrhhm5 小时前
MySQL 8.4 LTS 数据库巡检脚本
数据库·mysql
AI木马人5 小时前
9.【AI任务队列实战】如何在高并发下保证系统不崩?(Redis + Celery完整方案)
数据库·人工智能·redis·神经网络·缓存
[J] 一坚5 小时前
嵌入式高手C
c语言·开发语言·stm32·单片机·mcu·51单片机·iot
odoo中国5 小时前
Odoo 19技术教程 : 如何在 Odoo 19 中创建 Many2one 组件
开发语言·odoo·odoo19·odoo技术·many2one
逻辑驱动的ken6 小时前
Java高频面试考点场景题14
java·开发语言·深度学习·面试·职场和发展·求职招聘·春招
2401_883600256 小时前
golang如何理解weak pointer弱引用_golang weak pointer弱引用总结
jvm·数据库·python
aLTttY6 小时前
【Redis实战】分布式锁的N种实现方案对比与避坑指南
数据库·redis·分布式
2301_773553626 小时前
mysql如何评估SQL语句的索引开销_mysql性能追踪与分析
jvm·数据库·python