Groups the score into a small number of contiguous tiers (typically 3, 5
or 7, labeled from "very low" to "very high") fitted on the reference
sample and read on the reference and the study samples. Tiers are a
policy and communication device; they are not a rating scale in the
sense of the internal ratings-based approach (see scr_grades() for
that).
Usage
scr_tiers(x, ...)
# S3 method for class 'scr_scorecard'
scr_tiers(
x,
n_tiers = 5L,
method = c("optimal", "anchored", "quantile"),
criterion = c("deviance", "iv"),
anchors = NULL,
conservative = FALSE,
level = NULL,
min_pct = NULL,
min_events = NULL,
alpha = 0.05,
max_bins = NULL,
labels = NULL,
round_to = NULL,
n_boot = 0L,
seed = NULL,
sample = "holdout",
reference = "train",
max_cells = 1e+05,
...
)
# S3 method for class 'data.frame'
scr_tiers(
x,
score = "score",
y = "y",
objective = "risk",
direction = NULL,
weight = NULL,
sample = NULL,
reference = NULL,
study = NULL,
counts = FALSE,
n = "n",
events = "events",
n_tiers = 5L,
method = c("optimal", "anchored", "quantile"),
criterion = c("deviance", "iv"),
anchors = NULL,
conservative = FALSE,
level = 0.95,
min_pct = 0.05,
min_events = 20L,
alpha = 0.05,
max_bins = 100L,
labels = NULL,
round_to = NULL,
n_boot = 0L,
seed = NULL,
max_cells = 1e+05,
...
)Arguments
- x
An object from
scr_scorecard(), or adata.framewith one row per scored case (or one row per score value withcounts = TRUE).- ...
Passed on to the methods; an unknown argument is an error.
- n_tiers
Number of tiers requested (2 to 9).
- method
"optimal","anchored"or"quantile".- criterion
For
"optimal":"deviance"(binomial log-likelihood) or"iv"(information value).- anchors
For
"anchored": event-rate thresholds, as numbers in (0, 1) or the string"overall"(the reference event rate); a character vector may mix both.- conservative
For
"anchored": compare the anchors with the lower Jeffreys bound of the block rate instead of the rate.- level
Confidence level of the Jeffreys intervals. For a scorecard,
NULLusesconfig$study_level(0.95).- min_pct
Smallest volume share of a tier. For a scorecard,
NULLusesconfig$tier_min_pct(0.05).- min_events
Fewest events, and fewest non-events, of a tier. For a scorecard,
NULLusesconfig$tier_min_events(20).- alpha
Significance level of the adjacency test (divided by
n_tiers - 1) and ofall_distinct.- max_bins
Pre-bins of the search (2 to 500). For a scorecard,
NULLusesconfig$tier_max_bins(100).- labels
Optional labels, one per tier achieved, in ascending order of the event rate. They get the order prefix of
tier_labeltoo.- round_to
Optional positive number: cuts are rounded to its multiples.
- n_boot
Bootstrap resamples of the stability study (
0, the default, skips it).- seed
Seed of the stability bootstrap. A number is local to the call (the user's random stream is restored on exit);
NULLdraws from the user's stream and advances it. For a scorecard,NULLusesconfig$seed.- sample
For a scorecard: the study sample(s),
"holdout"(default) and/or"train". For a data.frame: the name of a column with sample labels, orNULL(all rows are one sample, reference and study at once).- reference
For a scorecard: the sample the bands are frozen on (
"train"). For a data.frame: the label of the reference sample;NULLtakes the first level of thesamplecolumn (the levels of a factor in their order, numbers in numeric order, text sorted).- max_cells
Largest number of distinct score values kept exactly.
- score, y
Column names of the score and of the 0/1 outcome (
NAallowed).- objective
"risk"(the event is the bad case) or"propensity"(the event is the good case).- direction
"higher_is_safer"or"higher_is_riskier";NULLderives it fromobjective.- weight
Optional column of non-negative case weights.
- study
Labels of the study samples;
NULLtakes every level other than the reference.- counts
TRUEwhenxis pre-aggregated: one row per score value with the columnsscore,nandevents(and, optionally,valueandvalue_events).- n, events
Column names of the counts when
counts = TRUE.
Value
An object of class c("scr_study_tiers", "scr_study", "list"):
tableOne row per sample and tier, event-richest tier first:
sample,tier,label,tier_label(the label with its order in front,"01."for the event-richest tier; see the section Labels),score_lo,score_hi,n,pct,events,rate,rate_lo,rate_hi,lift,pct_event,pct_nonevent,woe,p_adjacent(one-sided Fisher exact test that the tier has a higher event rate than the next lower tier) andp_adjacent_adj(Holm).summaryOne row per sample:
sample,n,events,rate,n_tiers,monotone(the rates rise with the tier),all_distinct(everyp_adjacent_adjbelowalpha),ivandpsi(of the tier mix against the reference).cuts,cuts_rawThe cuts in use (rounded when
round_tois given) and the fitted ones.labels,tier_labelsThe tier labels, plain and numbered, lowest rate first.
codes,code_labelsTier number and plain label of every interval in ascending score order, used by
scr_apply()andscr_sql().method,criterion,measure,n_tiers_requested,n_tiersThe fit.
ledgerOne row per step:
step,n_tiers,statusanddetail.stabilityNULL, or withn_boot > 0a list:cuts(per cut:score,median,q25,q75,iqrandn_same, the resamples that gave the same number of tiers),agreement,same_countandn_boot.objective,direction,level,alpha,min_pct,min_events,reference,samples,target,hist,callThe settings and the count table.
Method
The reference scores are cut into at most max_bins pre-bins of equal
share (tie-safe, as in scr_bands()); adjacent pre-bins are then pooled
(pool adjacent violators) until the event rate rises monotonically
toward the event-rich side of the score. Every tier is a run of these
blocks, so the tier rates are monotone on the reference.
method = "optimal" searches every segmentation of the blocks into
n_tiers contiguous tiers with an exact dynamic program and keeps the
one that maximizes the binomial log-likelihood
$$\sum_t \left[e_t \ln p_t + (n_t - e_t) \ln(1 - p_t)\right],
\qquad p_t = e_t / n_t$$
(criterion = "deviance", the same as minimizing the deviance) or the
information value (criterion = "iv"), subject to: every tier holds at
least min_pct of the volume; at least min_events events and
min_events non-events; and every pair of adjacent tiers is distinct by
a one-sided Fisher exact test at alpha / (L - 1), where \(L\) is the
tier count being tried. The rates and the objective use the weighted
counts; the event constraints and the Fisher test use the unweighted
counts. When no segmentation into n_tiers tiers meets the constraints,
n_tiers - 2, n_tiers - 4, ... down to 2 are tried (an odd count stays
odd while possible); when even 2 tiers are infeasible, the last resort is
a single tier. Every attempt is recorded in ledger and a warning is
raised. The tiers are never relabeled silently: the labels follow the
number of tiers achieved.
The previous tier enters the state of the program, so its cost is
\(O(L M^3)\) time and \(L M^2\) memory for \(L\) tiers and \(M\)
blocks (compiled code); the Fisher tests are cached when \(M \le 200\)
and recomputed above. n_tiers is capped at 9 and max_bins (hence
\(M\)) at 500, and the stability study refits once per resample, so
n_boot multiplies the cost.
method = "anchored" places one cut per value of anchors (event-rate
thresholds; "overall" stands for the reference event rate): the cut
sits before the first block, from the low-rate side, whose event rate
(its lower Jeffreys bound with conservative = TRUE) reaches the anchor.
n_tiers is then length(anchors) + 1 and the argument is ignored.
method = "quantile" cuts the reference into n_tiers tiers of equal
share, the baseline.
round_to rounds every cut to the nearest multiple (a policy-friendly
cut-off, such as 500 or 520 points); the tiers are then re-evaluated on
every sample with the rounded cuts, and both sets of cuts are kept. When
the scores were pooled into cells (max_cells), the count table is
rebuilt with the rounded cuts as forced cell edges, so the reported
tiers agree exactly with scr_apply() and scr_sql().
Tiers are left-closed: score >= cut is the upper side, the convention
of scr_cutoff().
With n_boot > 0, the stability of the fit is measured by redrawing the
event and non-event counts of the reference pre-bins from a multinomial
law and refitting: per cut, the median and interquartile range of the
refitted cut in score units, and the agreement rate (the share of the
reference volume that keeps its tier).
Labels
Tiers are numbered by event rate, lowest first: 2 tiers are labeled
"low", "high"; 3 tiers "low", "medium", "high"; 4 tiers
"low", "medium low", "medium high", "high"; 5 tiers "very low",
"low", "medium", "high", "very high"; 6 tiers "very low",
"low", "medium low", "medium high", "high", "very high"; 7
tiers add "extremely low" and "extremely high" to the five; 1, 8 and 9
tiers are numbered "T1", "T2", ... The labels
describe the event rate, so under objective = "risk" they read as
risk and under "propensity" as propensity (measure). The table lists
the event-richest tier first.
For production, every label also exists with its order in front
(tier_label): "01." for the event-richest tier, the first row of the
table, then "02.", ... down to the tier with the lowest event rate,
under every objective and direction, and for labels given in labels
too. With five tiers of a credit score, tier 5 is "01.very high" and
tier 1 is "05.very low": the number tier rises with the event rate,
the prefix sorts from the highest rate down. The prefix is zero-padded to
two digits. scr_apply() and scr_sql() return these numbered labels
(numbered = FALSE gives the plain ones), so their output joins to the
table by tier_label.
References
Brown, L. D., Cai, T. T. and DasGupta, A. (2001). Interval estimation for a binomial proportion. Statistical Science, 16(2), 101-133. doi:10.1214/ss/1009213286
Siddiqi, N. (2006). Credit Risk Scorecards: Developing and Implementing Intelligent Credit Scoring. Wiley.
Yurdakul, B. and Naranjo, J. (2020). Statistical properties of the population stability index. Journal of Risk Model Validation, 14(4), 89-100.
See also
scr_bands(), scr_rag(); scr_apply() and scr_sql() assign
the tiers in production.
Other score-studies:
scr_bands(),
scr_claims(),
scr_detection(),
scr_maturity(),
scr_mix_shift(),
scr_operating(),
scr_overlap(),
scr_rag(),
scr_rag_plan(),
scr_score_cross(),
scr_segments(),
scr_uplift()
Examples
cfg <- scr_config(verbose = FALSE, nthread = 1, use_ranger = FALSE,
use_lightgbm = FALSE, xgb_rounds = 40, n_boot = 20)
res <- scr_select(scr_demo, "default", config = cfg, drop = c("id", "churn"),
date_col = "ref_date")
sc <- scr_scorecard(res)
tr <- scr_tiers(sc, n_tiers = 5)
tr
#> <scr_study_tiers> target "default" | measure risk | higher_is_safer
#> optimal (deviance) | 5 tiers requested, 5 achieved | fitted on 'train'
#> cuts: 502.3609, 516.0750, 538.0089, 555.0927
#> sample n events rate IV PSI tiers monotone distinct
#> train 2,800 399 14.25% 1.162 - 5 yes yes
#> holdout 1,400 203 14.50% 0.744 0.0019 5 yes no
#>
#> Tiers on 'holdout' (event-richest first)
#> tier label score pct rate [95% CI] p_adj
#> 5 01.very high [-Inf, 502.3609) 6.3% 36.36% [26.88%, 46.72%] 0.487
#> 4 02.high [502.3609, 516.075) 6.9% 35.05% [26.11%, 44.87%] 0.033
#> 3 03.medium [516.075, 538.0089) 20.6% 23.18% [18.60%, 28.30%] 0.001
#> 2 04.low [538.0089, 555.0927) 22.4% 12.46% [9.15%, 16.46%] 0.000
#> 1 05.very low [555.0927, Inf) 43.8% 5.06% [3.53%, 7.01%] -
tr$table[sample == "holdout", .(tier, label, score_lo, score_hi, pct, rate)]
#> tier label score_lo score_hi pct rate
#> <int> <char> <num> <num> <num> <num>
#> 1: 5 very high -Inf 502.3609 0.06285714 0.36363636
#> 2: 4 high 502.3609 516.0750 0.06928571 0.35051546
#> 3: 3 medium 516.0750 538.0089 0.20642857 0.23183391
#> 4: 2 low 538.0089 555.0927 0.22357143 0.12460064
#> 5: 1 very low 555.0927 Inf 0.43785714 0.05057096
# policy cuts on round numbers, and the tiers assigned to new scores
tr10 <- scr_tiers(sc, n_tiers = 3, round_to = 10)
tr10$cuts
#> [1] 520 550
head(scr_apply(tr10, c(480, 530, 600)))
#> score tier tier_label
#> <num> <int> <char>
#> 1: 480 3 01.high
#> 2: 530 2 02.medium
#> 3: 600 1 03.low
# anchors on the event rate: below, around and above the overall rate
scr_tiers(sc, method = "anchored", anchors = c(0.08, "overall", 0.25))$summary
#> sample n events rate n_tiers monotone all_distinct iv
#> <char> <num> <num> <num> <int> <lgcl> <lgcl> <num>
#> 1: train 2800 399 0.1425 4 TRUE TRUE 1.120941
#> 2: holdout 1400 203 0.1450 4 TRUE TRUE 0.732374
#> psi
#> <num>
#> 1: NA
#> 2: 0.0006075206
