Follows the units of each score band over time and estimates, at each horizon, the cumulative share that has had the event, allowing for units whose follow-up ends before the horizon (censoring). It shows how fast each band matures, where the curve flattens (the performance window) and how the discrimination of the score changes with the horizon.
Usage
scr_maturity(x, ...)
# S3 method for class 'data.frame'
scr_maturity(
x,
score = "score",
time = "time",
event = "event",
horizons,
objective = "risk",
direction = NULL,
n_bands = 5L,
cuts = NULL,
level = 0.95,
weight = NULL,
max_cells = 1e+05,
...
)Arguments
- x
A
data.framewith one row per unit.- ...
Passed on to the methods; an unknown argument is an error.
- score, time, event
Column names of the score, of the time to the event or to the end of the follow-up, and of the 0/1 event indicator.
- horizons
Times at which the incidence is read, in the unit of
time(positive numbers).- objective
"risk"(the event is the bad case) or"propensity"(the event is the good case).- direction
"higher_is_safer"or"higher_is_riskier";NULLderives it fromobjective.- n_bands
Equal-share bands of the score (tie-safe, the event-richest first) when
cutsis not given.- cuts
Optional ascending cuts of the score, or an object from
scr_bands()orscr_tiers(), whose cuts, numbers and labels are then used, with its objective and direction.- level
Confidence level of the intervals.
- weight
Optional column of non-negative case weights.
- max_cells
Largest number of distinct score values kept exactly.
Value
An object of class c("scr_maturity", "list"):
tableOne row per band and horizon, event-richest band first, then all units (
band = NA,label = "all"):band,label,n,horizon,at_risk,events,censored,incidence,se,lo,hiandpct_of_final.discriminationOne row per horizon:
horizon,n(units with a complete window),events,aucandgini.cuts,codes,labelsThe ascending cuts, and the band number and label of every interval in ascending score order.
horizons,level,objective,direction,score,time,event,n,n_rows,n_dropped,weighted,callThe settings:
nis the volume used (the sum of the weights),n_rowsthe rows used andn_droppedthose left out.
Data
One row per unit (a loan, a customer), with the score at the origin,
time and event:
timeA non-negative number of periods (days, months) from the origin to the event, or to the end of the follow-up.
event1 when the event was observed at
time, 0 when the follow-up ended there without it (censored).
From a monthly panel, take per unit the months from its opening date to
its first month in default (event = 1), or to its last observed month
when it never defaulted (event = 0); the example does it for
scr_demo_panel.
Rows with a missing or infinite score, a missing time or event, or a zero
weight are left out (n_dropped).
Incidence
Per band and for all units, the Kaplan-Meier estimate of the survival is
\(S(t) = \prod_{t_j \le t} (1 - d_j / n_j)\), with \(d_j\) the events
at the distinct time \(t_j\) and \(n_j\) the units still followed
just before it (units censored at \(t_j\) are at risk at \(t_j\)).
incidence is \(1 - S(h)\) at the horizon \(h\), and se its
standard error by Greenwood's formula,
\(S(h) \sqrt{\sum_{t_j \le h} d_j / (n_j (n_j - d_j))}\). The interval
lo, hi is the complementary log-log interval of \(S(h)\),
\(S^{\exp(\pm z \sigma)}\) with
\(\sigma = \sqrt{\sum d_j / (n_j (n_j - d_j))} / |\ln S|\), which stays
inside [0, 1]. It is undefined when no event has happened by the
horizon (lo = 0, hi = NA) and when every unit has had the event
(NA).
Under weights, \(d_j\) and \(n_j\) are weighted sums and each term of Greenwood's sum is computed on the Kish effective size of the risk set, \(d_j / ((n_j - d_j)\, n^{eff}_j)\) with \(n^{eff}_j = n_j^2 / \sum w^2\); equal weights give the unweighted result.
Per band and horizon the table also counts events (events up to the
horizon), censored (units censored before it) and at_risk (the other
units: still followed at the horizon without the event), which add up to
n. pct_of_final is the incidence over the incidence at the largest
horizon: the share of the final events already seen.
Past the last follow-up time of a band nothing more is observed: the
curve is carried flat, and a horizon beyond it repeats the last estimate
with at_risk = 0. Read such a row as "no information", not as a
plateau of the event rate.
Discrimination
At each horizon, the units with a complete window are those with the
event by the horizon and those followed for at least the horizon without
it; units censored earlier are left out. auc and gini are those of
the score for "event by the horizon" among them, from the counts per
score value. The complete window ignores the censored units, so it is
unbiased only when censoring does not depend on the score.
References
Greenwood, M. (1926). The natural duration of cancer. Reports on Public Health and Medical Subjects, 33, 1-26. HMSO.
Kalbfleisch, J. D. and Prentice, R. L. (2002). The Statistical Analysis of Failure Time Data, 2nd edition. Wiley. doi:10.1002/9781118032985
Kaplan, E. L. and Meier, P. (1958). Nonparametric estimation from incomplete observations. Journal of the American Statistical Association, 53(282), 457-481. doi:10.1080/01621459.1958.10501452
See also
scr_bands() and scr_tiers() for the bands, scr_default()
and scr_default_rate() for the default flag and the cohort rates of a
monthly panel.
Other score-studies:
scr_bands(),
scr_claims(),
scr_detection(),
scr_mix_shift(),
scr_operating(),
scr_overlap(),
scr_rag(),
scr_rag_plan(),
scr_score_cross(),
scr_segments(),
scr_tiers(),
scr_uplift()
Examples
# time and event from a monthly panel: months from the opening date to
# the first default (90 days past due), or to the last observed month
p <- scr_demo_panel[order(scr_demo_panel$id, scr_demo_panel$ref_date), ]
month <- 12 * as.integer(format(p$ref_date, "%Y")) + as.integer(format(p$ref_date, "%m"))
first <- !duplicated(p$id)
bad <- which(p$dpd >= 90)
bad <- bad[!duplicated(p$id[bad])]
u <- data.frame(id = p$id[first], score = p$score[first], open = month[first])
u$event <- as.integer(u$id %in% p$id[bad])
u$end <- as.vector(tapply(month, p$id, max)[as.character(u$id)])
u$end[u$event == 1] <- month[bad][match(u$id[u$event == 1], p$id[bad])]
u$time <- u$end - u$open
mt <- scr_maturity(u, horizons = c(6, 12, 24, 35), n_bands = 4)
mt
#> <scr_maturity> event "event" over "time" | objective risk | higher_is_safer
#> 600 units | 4 bands (equal shares) | horizons 6, 12, 24, 35 | 95% complementary log-log intervals
#>
#> Cumulative incidence (Kaplan-Meier), event-richest band first
#> band score n t=6 t=12 t=24 t=35
#> 1 [-Inf, 556.5) 149 12.1% 24.2% 47.7% 58.4%
#> 2 [556.5, 603.5) 150 7.3% 12.0% 27.3% 34.7%
#> 3 [603.5, 643.5) 150 3.3% 4.0% 8.7% 14.7%
#> 4 [643.5, Inf) 151 0.662% 3.3% 5.3% 11.3%
#> all 600 5.8% 10.8% 22.2% 29.7%
#> share of the final (all) 20% 37% 75% 100%
#>
#> Discrimination at each horizon (units with a complete window)
#> horizon n events AUC Gini
#> 6 600 35 0.7492 0.4984
#> 12 600 65 0.7601 0.5202
#> 24 600 133 0.7965 0.5930
#> 35 600 178 0.7692 0.5384
mt$discrimination
#> horizon n events auc gini
#> <num> <num> <num> <num> <num>
#> 1: 6 600 35 0.7492035 0.4984071
#> 2: 12 600 65 0.7601006 0.5202013
#> 3: 24 600 133 0.7965014 0.5930028
#> 4: 35 600 178 0.7692103 0.5384206
