where \(\overline{y}_{R_t}\) is the mean of the variables in partition \(R_t\)
Algorithm I
How are the partitions \(R_t\) determined?
Beginning from the first node, each split is determined by the variable and value which minimizes the RSS (regression problem) or impurity (classification problem)
End up with the most complex tree explaining every data point
Pruning the tree
Pre-pruning: additional splits are created until cost of depth is too high
Post-pruning: complex tree is created and cut down to a size where cost is met
Post-pruning is preferred because stopping too early bears the danger of missing some split down the line
Algorithm II
How is the tuning parameter \(\lambda\) determined?
Employing cross-validation, we choose the parameter value which minimizes some prediction error in the test/validation data (more on that later)
Exact implementation of the algorithm differs
Different packages may use different impurity/information measures (e.g., entropy) and different cost functions/tuning parameters
e.g., rpart uses post-pruning where the complexity parameter cp regulates the minimum gain in \(R^2\) or minimum decrease of Gini impurity to create another node
Growing the Tree
Advantages and Drawbacks
Trees have many advantages
Easy to understand and very flexible
No theoretical assumptions needed
Computationally cheap
Automatic feature selection
No limitation in number of features
Do not necessarily discard observations
Drawbacks include
No immediate (causal) interpretation of decision
Algorithm is “greedy”
The curse of dimensionality
The Curse of Dimensionality
Volume of the support space increases exponentially with dimensions
Data points become dispersed in a high-dimensional space
Splits become very volatile
To mitigate the problem, random forest is used in practice
A Perfect Summary I
Synthetic Unemployment Data
synthetic_unemployment_data <-read_parquet("data/synthetic_unemployment_data.parquet")set.seed(123)data <- synthetic_unemployment_data |>mutate(train_index =sample(c("train", "test"),nrow(synthetic_unemployment_data),replace=TRUE,prob=c(0.75, 0.25) ) )train <- data |>filter(train_index=="train")test <- data |>filter(train_index=="test")
target_high
target_low
region_1
promised_employment
benefits
benefits_amount
social_security
sex
age
education
family_situation
nationality
disability
nace1
job_sector
asylum
children
age_youngest_child
migrational_background
employment_1m
employment_unsubsidized_1m
unemployment_1m
out_of_labor_force_1m
employment_3m
unemployment_3m
out_of_labor_force_3m
employment_6m
unemployment_6m
out_of_labor_force_6m
employment_1y
unemployment_1y
out_of_labor_force_1y
employment_2y
unemployment_2y
out_of_labor_force_2y
days_employment_unsubsidized_10j
days_unemployment_10j
days_out_of_labor_force_insured_10j
days_employment_unsubsidized_5j
days_unemployment_5j
days_out_of_labor_force_insured_5j
days_employment_unsubsidized_2j
days_unemployment_2j
days_out_of_labor_force_insured_2j
days_to_last_job
income_last_job
employment_subsidy_1j
qualification_subsidy_1j
support_subsidy_1j
employment_subsidy_4j
qualification_subsidy_4j
support_subsidy_4j
contact_pes_6m
contact_pes_2j
job_mediation_6m
job_mediation_2j
regional_unemployment
regional_long_time_joblessness
regional_seasonal_unemployment
regional_promise_employment
regional_gdp
regional_job_openings
train_index
unsuccessful
successful
409-Linz
0
Keine
NA
0
M
17
Pflichtschulausbildung
Ledig
Russland
-
Öffentl. Verwaltung, Verteidigung, SV
Produktionsberufe
Konventionsflüchtling
0
NA
1. Generation
0
0
0
1
0
1
0
0
0
1
0
0
1
0
0
1
34
18
3602
34
18
1775
0
0
731
762
774
0
0
0
0
0
0
0
0
0
0
0.0786348
0.3696403
0.0657192
0.0704333
52400
0.2278938
train
unsuccessful
successful
Andere
1
NH
23.20
0
M
29
Lehrausbildung
Ledig
Österreich
A-Laut AMS
Bau
Produktionsberufe
ohne ASYL
0
NA
Kein Migrationshintergrund
0
0
1
0
1
0
0
0
1
0
0
1
0
0
1
0
335
2257
386
0
1463
336
0
671
0
2331
280
1
0
0
1
1
0
1
6
1
1
0.0600450
0.2943468
0.2819971
0.3153480
29400
0.8180097
train
unsuccessful
successful
Andere
0
NH
33.26
0
M
26
Lehrausbildung
Verheiratet
Österreich
-
Beherbergung und Gastronomie
Dienstleistungsberufe
ohne ASYL
0
NA
Kein Migrationshintergrund
0
0
1
0
0
1
0
1
0
0
1
0
0
1
0
0
2192
1243
46
1127
701
5
359
374
3
139
2592
0
0
0
0
0
0
4
9
1
1
0.0525828
0.1586341
0.2133195
0.2889534
46000
0.4919413
train
unsuccessful
unsuccessful
900-Wien
0
NH
11.69
0
M
33
Hoehere Ausbildung
Ledig
Deutschland
-
Gesundheits- und Sozialwesen
Produktionsberufe
ohne ASYL
0
NA
1. Generation
0
0
1
0
0
1
0
0
1
0
0
1
0
0
1
0
396
1159
1630
382
982
168
0
731
0
940
712
0
1
0
0
1
0
3
12
0
1
0.1495503
0.4561777
0.0597223
0.0440014
49600
0.3659622
train
unsuccessful
unsuccessful
900-Wien
0
NH
23.31
0
M
32
Pflichtschulausbildung
Ledig
Österreich
-
Sonst. wirtschaftliche DL
Dienstleistungsberufe
ohne ASYL
0
NA
Kein Migrationshintergrund
0
0
1
0
0
1
0
0
1
0
0
1
0
1
0
0
975
1524
1179
366
889
804
82
639
0
527
1094
0
1
1
1
1
1
2
8
0
0
0.1495503
0.4561777
0.0597223
0.0440014
49600
0.3659622
train
unsuccessful
unsuccessful
900-Wien
1
ALG
44.65
0
M
55
Lehrausbildung
Ledig
Österreich
-
Beherbergung und Gastronomie
Saisonberufe
ohne ASYL
0
NA
Kein Migrationshintergrund
0
0
1
0
1
0
0
1
0
0
1
0
0
1
0
0
3623
35
22
1796
33
22
697
33
22
31
3641
0
0
0
0
0
0
1
1
0
0
0.1495503
0.4561777
0.0597223
0.0440014
49600
0.3659622
train
unsuccessful
unsuccessful
Andere
0
NH
28.89
0
M
58
Pflichtschulausbildung
Ledig
Polen
A-Laut AMS
Gesundheits- und Sozialwesen
Produktionsberufe
ohne ASYL
0
NA
1. Generation
0
0
1
0
0
1
0
0
1
0
0
1
0
0
1
0
322
2874
351
0
1461
350
0
588
145
2053
1269
0
0
0
0
0
0
3
8
0
0
0.0659340
0.3625010
0.1338427
0.1111248
32600
0.3064870
train
Fitting Trees with RPART
tree <-rpart( target_low ~ days_unemployment_2j + age + days_to_last_job,data = train |>select(-train_index, -target_high),cp =0.007 )tree