Reads one score on the segments of a population (a product, a channel, a region) and says, for each, whether the score ranks and calibrates there as it does on the whole: discrimination with its standard error, the observed events against those expected from the pooled bands, the offset on the log-odds scale, the slope of the score relative to the pooled one and a suggested action.
Usage
scr_segments(x, ...)
# S3 method for class 'data.frame'
scr_segments(
x,
segment,
by = NULL,
n_bands = 10L,
level = 0.95,
min_events = 20L,
n_boot = 0L,
auc_tol = 0.03,
offset_tol = 0.25,
slope_tol = 0.25,
score = "score",
y = "y",
objective = "risk",
direction = NULL,
weight = NULL,
seed = NULL,
max_cells = 1e+05,
...
)
# S3 method for class 'scr_scorecard'
scr_segments(
x,
newdata,
segment,
by = NULL,
n_bands = 10L,
level = NULL,
min_events = 20L,
n_boot = 0L,
auc_tol = 0.03,
offset_tol = 0.25,
slope_tol = 0.25,
target = NULL,
seed = NULL,
max_cells = 1e+05,
...
)Arguments
- x
A
data.framewith one row per scored case, or an object fromscr_scorecard().- ...
Passed on to the methods; an unknown argument is an error.
- segment
Name of the segment column. A missing segment is the segment
"(missing)". The segments of a factor column are listed in the order of its levels.- by
Optional name of a column of periods or groups.
- n_bands
Equal-share bands frozen on the pooled rows, for the expected events and the PSI.
- level
Confidence level of the intervals;
1 - levelis the significance of the tests. For a scorecard,NULLusesconfig$study_level(0.95).- min_events
Fewest events, and non-events, for an action other than
"too few events".- n_boot
Bootstrap resamples of the AUC interval;
0(default) keeps the DeLong interval.- auc_tol, offset_tol, slope_tol
Tolerances of the actions: on the difference of the AUC from the weighted mean, on the absolute offset and on the distance of the slope ratio from 1.
- score, y
Column names of the score and of the 0/1 outcome (
NAallowed).- objective
"risk"(the event is the bad case) or"propensity"(the event is the good case).- direction
"higher_is_safer"or"higher_is_riskier";NULLderives it fromobjective.- weight
Optional column of non-negative case weights.
- seed
Seed of the bootstrap. A number is local to the call (the user's random stream is restored on exit);
NULLdraws from the user's stream and advances it. For a scorecard,NULLusesconfig$seed.- max_cells
Largest number of distinct score values kept exactly.
- newdata
For a scorecard: the rows to score with
scr_apply(), holding the candidate variables, the target and the segment column.- target
For a scorecard: the outcome column of
newdata;NULLuses the target of the scorecard.
Value
An object of class c("scr_segments", "list"):
tableOne row per group and segment:
group,segment,n,events,rate,auc,auc_se,auc_lo,auc_hi,gini,ks,psi,psi_critical,expected,oe_ratio,oe_lo,oe_hi,offset,slope,slope_ratio,auc_diff,p_auc,p_auc_adj,p_slope,p_slope_adjandaction.testOne row per group:
group,segments(those in the test),auc_w,statistic,dfandp_value.pooledOne row per group, the pooled rows:
group,n,events,rate,auc,auc_se,auc_lo,auc_hi,gini,ksandslope.cuts,segment,by,level,min_events,auc_tol,offset_tol,slope_tol,objective,direction,target,callThe pooled cuts and the settings.
Statistics
The rows are counted once per segment and distinct score. Per segment:
auc,giniandks, with the DeLong standard errorauc_secomputed from the counts and the intervalauc +/- z * auc_se(cut to [0, 1]); withn_boot > 0, the interval is the percentile interval of the count bootstrap ofscr_bands(). A segment without a positive standard error (fewer than two events or non-events, or every score tied) has no interval.psi: the population stability index of the score distribution of the segment against the pooled one, overn_bandsequal-share bands frozen on the pooled rows (seescr_psi()). The segment is part of the pooled rows, so the two distributions are not independent:psi_critical, the n-adjusted critical value at1 - level, uses \((1/n_s - 1/N)\,\chi^2_{B-1}\) in place of \((1/n_s + 1/N)\,\chi^2_{B-1}\), with \(n_s\) and \(N\) the sizes of the segment and of the pooled rows (under weights, the Kish effective sizes, with the covariance term scaled by the weight share of the segment). It isNAwhen the segment is the whole pool.expected: the events expected by indirect standardization, \(\sum_b n_{s,b} R_b\), with \(n_{s,b}\) the rows of the segment in the pooled band \(b\) and \(R_b\) the pooled event rate of the band;oe_ratio= events /expected, with the Jeffreys interval of the observed rate (on the Kish effective size under weights) divided by the expected rate.offset= logit(observed rate) - logit(expected rate): what a segment intercept would add to the log-odds.NAwhen either rate is 0 or 1.slope: the coefficient of the score in a logistic regression of the outcome within the segment, fitted on the count table with the counts as weights (equal to the fit on the rows);slope_ratioisslopeover the pooled slope.NAwhen the segment has a single score value or its classes are separated by the score.p_slopeis the two-sided Wald test of the slope of the segment against the slope fitted on the rest of its group (the pooled rows without the segment, an independent sample), \(z = (b_s - b_r) / \sqrt{v_s + v_r}\) with the model-based variances of the two fits (scaled to the Kish effective size under weights);p_slope_adjis its Holm adjustment across the segments.
Rows with a missing outcome count in the volume and in the PSI only.
With more than max_cells distinct scores the scores are pooled into
cells, and the slope is fitted on the cell midpoints.
Test of equal discrimination
With \(w_s = 1 / se_s^2\), the inverse-variance weighted mean is
\(AUC_w = \sum_s w_s AUC_s / \sum_s w_s\), and
$$Q = \sum_s (AUC_s - AUC_w)^2 / se_s^2$$
is compared with a chi-square law with \(S - 1\) degrees of freedom
(test: statistic, df, p_value). The segments are disjoint, so
their AUCs are independent. Per segment, auc_diff =
\(AUC_s - AUC_w\) is tested by a two-sided z test with the variance
\(se_s^2 - 1 / \sum_s w_s\) (the mean contains the segment); p_auc is
its p-value and p_auc_adj the Holm adjustment across the segments.
Segments without a positive standard error (fewer than two events or
non-events, or every score tied) are left out of the test. With exactly
two segments the two tests are one and the same, so no adjustment is
made; the same holds for the slope tests.
Action
A convention of this package, read in this order:
"too few events"Fewer than
min_eventsevents, or non-events, in the segment (unweighted counts)."separate model"The score ranks differently:
abs(auc_diff) > auc_tolwithp_auc_adj < 1 - level, orabs(slope_ratio - 1) > slope_tolwithp_slope_adj < 1 - level. A difference must be both material and significant: on a small segment, a slope ratio far from 1 is often noise."offset"The ranking holds but the level does not:
abs(offset) > offset_toland the interval ofoe_ratioexcludes 1."shared"None of the above: the pooled score serves the segment.
The tolerances are starting points, not rules; set them to the policy in force.
Groups
With by (a period, a sample label), the analysis is repeated within
each group: the pooled reference of a segment is the whole of its group.
The bands are frozen once, on all rows. Groups and segments are listed in
the order of their labels; numeric columns in numeric order, and a factor
segment column in the order of its levels.
References
Brown, L. D., Cai, T. T. and DasGupta, A. (2001). Interval estimation for a binomial proportion. Statistical Science, 16(2), 101-133. doi:10.1214/ss/1009213286
DeLong, E. R., DeLong, D. M. and Clarke-Pearson, D. L. (1988). Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach. Biometrics, 44(3), 837-845.
Thomas, L. C., Crook, J. and Edelman, D. (2017). Credit Scoring and Its Applications, 2nd edition. SIAM. doi:10.1137/1.9781611974560
See also
scr_mix_shift() for the change of the event rate between
samples, scr_rag() for lights per segment against a reference.
Other score-studies:
scr_bands(),
scr_claims(),
scr_detection(),
scr_maturity(),
scr_mix_shift(),
scr_operating(),
scr_overlap(),
scr_rag(),
scr_rag_plan(),
scr_score_cross(),
scr_tiers(),
scr_uplift()
Examples
local({
set.seed(1)
n <- 6000
seg <- sample(c("app", "store", "web"), n, TRUE, c(0.5, 0.3, 0.2))
x <- rnorm(n)
# the score ranks the same everywhere; "web" defaults more at every score
d <- data.frame(channel = seg, score = round(600 + 50 * x),
y = rbinom(n, 1, plogis(-2 - x + 0.6 * (seg == "web"))))
sg <- scr_segments(d, segment = "channel")
print(sg)
sg$test
})
#> <scr_segments> target "y" | objective risk | higher_is_safer | segments of 'channel'
#> tolerances: AUC 0.03, offset 0.25, slope 0.25 | fewest events 20 | level 95%
#>
#> Pooled: n 6,000 | rate 17.7% | AUC 0.7375 | equal AUC across 3 segments: chi-square 0.04 (df 2), p 0.9819
#> segment n rate AUC [lo, hi] PSI O/E [lo, hi] offset slope ratio action
#> app 3,042 15.6% 0.7406 [0.717, 0.764] 0.0007 0.89 [0.81, 0.96] -0.14 1.03 shared
#> store 1,711 15.2% 0.7367 [0.705, 0.768] 0.0029 0.87 [0.77, 0.97] -0.17 0.97 shared
#> web 1,247 26.3% 0.7395 [0.708, 0.771] 0.0054 1.45 [1.32, 1.59] +0.47 1.02 offset
#> group segments auc_w statistic df p_value
#> <char> <int> <num> <num> <int> <num>
#> 1: all 3 0.7392968 0.03660684 2 0.9818631
