A score is validated with an AUC, but it is used through cuts: who is
approved, who is reviewed, who gets the retention offer. The score
studies of scorecraft answer the questions that come with
the cuts. How does the event rate move along the score? Which few groups
can be named and defended? Is the score still fit on new data? What can
be promised about a group, and with what confidence? Where should the
cut be under a budget or a team’s capacity? How do two scores read on
the same customers relate? When the event rate moves, is it the
population or the score bands? Does one score serve every segment?
Every study aggregates the scored rows once into a table of counts
and works on that table afterwards, so a study of millions of rows costs
one grouped pass plus work proportional to the number of distinct
scores. The exceptions are the rank association of
scr_score_cross(), which sorts the rows of the two scores
once more, and scr_detection(), which sorts the rows of the
entities with an event.
1. Three scores on the demo data
The demo table carries a risk target (default) and a
propensity target (churn) on the same customers. A light
configuration fits a credit scorecard, a mirrored copy of it read as a
fraud score (a higher score means more risk), and a churn scorecard. The
split is out-of-time on ref_date, so both targets share the
same hold-out rows.
library(scorecraft)
library(data.table)
cfg <- scr_config(verbose = FALSE, nthread = 1, use_glmnet = TRUE, use_ranger = FALSE,
use_lightgbm = FALSE, xgb_rounds = 60, n_boot = 20)
res <- scr_select(scr_demo, "default", config = cfg, drop = c("id", "churn"),
date_col = "ref_date")
credit <- scr_scorecard(res)
fraud <- scr_scorecard(res, direction = "higher_is_riskier")
cfg_churn <- scr_config(objective = "propensity", verbose = FALSE, nthread = 1, use_glmnet = TRUE,
use_ranger = FALSE, use_lightgbm = FALSE, xgb_rounds = 60, n_boot = 20)
res_churn <- scr_select(scr_demo, "churn", config = cfg_churn, drop = c("id", "default"),
date_col = "ref_date")
churn <- scr_scorecard(res_churn)2. Credit: bands, tiers and lights
scr_bands() cuts the score into bands of equal share
frozen on the training sample and reads them on the
hold-out. Each band reports its event rate with a Jeffreys interval, the
lift over the overall rate, the cumulative capture of events and the KS
at its lower edge; the last column is the Holm-adjusted p-value of a
one-sided Fisher test that the band is riskier than the band before it
(a rank-order reversal).
b <- scr_bands(credit, n_bands = 10, n_boot = 50, seed = 1)
b
#> <scr_study_bands> target "default" | objective risk | higher_is_safer
#> bands frozen on 'train', read on 'holdout' | 10 requested, 10 effective (uniform)
#> sample n events rate AUC [95% CI] Gini KS IV PSI reversals
#> train 2,800 399 14.25% 0.7856 [0.763, 0.805] 0.571 0.441 1.136 - 0
#> holdout 1,400 203 14.50% 0.7394 [0.710, 0.763] 0.479 0.389 0.804 0.0069 0
#>
#> Bands on 'holdout' (event-richest first; 95% Jeffreys interval of the rate)
#> band score pct rate [lo, hi] lift capture KS p_rev_adj
#> 1 [-Inf, 509.8922) 9.1% 35.43% [27.52%, 44.00%] 2.44 22.2% 0.153 -
#> 2 [509.8922, 523.5026) 9.1% 27.34% [20.19%, 35.51%] 1.89 39.4% 0.248 1.000
#> 3 [523.5026, 533.3684) 9.9% 26.62% [19.81%, 34.39%] 1.84 57.6% 0.345 1.000
#> 4 [533.3684, 542.0923) 10.6% 17.45% [12.01%, 24.14%] 1.20 70.4% 0.370 1.000
#> 5 [542.0923, 550.3612) 10.8% 11.26% [6.96%, 17.03%] 0.78 78.8% 0.342 1.000
#> 6 [550.3612, 557.7648) 10.6% 12.16% [7.64%, 18.15%] 0.84 87.7% 0.322 1.000
#> 7 [557.7648, 566.5601) 10.6% 7.38% [3.99%, 12.41%] 0.51 93.1% 0.261 1.000
#> 8 [566.5601, 576.5187) 9.1% 3.91% [1.51%, 8.35%] 0.27 95.6% 0.183 1.000
#> 9 [576.5187, 590.2783) 8.9% 3.23% [1.10%, 7.49%] 0.22 97.5% 0.102 1.000
#> 10 [590.2783, Inf) 11.2% 3.18% [1.23%, 6.84%] 0.22 100.0% 0.000 1.000The hold-out shows 0 significant reversal(s) of the rank order, and the summary compares the AUC on train and on the hold-out with bootstrap intervals. A band with a wide interval is a band with few events; the interval, not the point rate, is what a policy should be read against.
Ten bands are too many to name in a credit policy.
scr_tiers() groups the score into a few tiers by an exact
dynamic program: every tier holds at least 5% of the volume and 20
events and non-events, adjacent tiers differ by a one-sided Fisher test,
and among the segmentations that meet those constraints the one with the
best binomial likelihood wins. round_to moves the cuts to
round numbers, and n_boot refits on resampled counts to
show how stable each cut is.
tr <- scr_tiers(credit, n_tiers = 5, round_to = 5, n_boot = 30, seed = 1)
tr
#> <scr_study_tiers> target "default" | measure risk | higher_is_safer
#> optimal (deviance) | 5 tiers requested, 5 achieved | fitted on 'train' | cuts rounded to 5
#> cuts: 500, 515, 540, 555
#> sample n events rate IV PSI tiers monotone distinct
#> train 2,800 399 14.25% 1.143 - 5 yes yes
#> holdout 1,400 203 14.50% 0.752 0.0014 5 no no
#>
#> Tiers on 'holdout' (event-richest first)
#> tier label score pct rate [95% CI] p_adj
#> 5 01.very high [-Inf, 500) 5.4% 33.33% [23.45%, 44.47%] 0.753
#> 4 02.high [500, 515) 7.5% 37.14% [28.35%, 46.63%] 0.007
#> 3 03.medium [515, 540) 23.6% 23.03% [18.74%, 27.80%] 0.001
#> 2 04.low [540, 555) 19.6% 11.64% [8.25%, 15.82%] 0.001
#> 1 05.very low [555, Inf) 43.9% 5.04% [3.52%, 6.98%] -
#>
#> Stability (30 resamples): tier agreement 84.5%, same tier count 96.7%
tr$stability$cuts
#> cut score median q25 q75 iqr n_same
#> <int> <num> <num> <num> <num> <num> <num>
#> 1: 1 502.3609 504.4274 502.3609 506.1428 3.781886 29
#> 2: 2 516.0750 516.0750 516.0750 526.5455 10.470557 29
#> 3: 3 538.0089 539.6230 538.0089 542.0923 4.083412 29
#> 4: 4 555.0927 555.0927 552.8810 556.8089 3.927859 29The tiers are fitted on train, so the hold-out tells whether they
hold. Its summary line reads monotone “no”: the tiers are
monotone on train by construction, but on the hold-out at least one pair
of adjacent tiers comes out in the wrong order. The p_adj column and the
intervals of the table say which pair, and whether the two tiers can be
told apart at all on new data. The interquartile range of each refitted
cut, in score points, shows which boundaries the data pins down and
which ones a different sample would move.
scr_rag() turns the comparison of the hold-out with
train into red, amber and green lights for discrimination, calibration,
stability and the variables of the scorecard, each lit only when its
confidence interval or test shows a deviation.
rg <- scr_rag(credit, n_boot = 50, seed = 1)
rg$summary
#> sample group discrimination calibration stability variables overall reason
#> <char> <char> <char> <char> <char> <char> <char> <char>
#> 1: holdout all amber green green amber amber3. Fraud: tail bands, action tiers and a daily capacity
A fraud score is read at its tail: the riskiest 2%, 5% or 10% of the
traffic, not its deciles. spacing = "tail" places the cuts
at cumulative shares counted from the event-rich end, here the high
scores.
bf <- scr_bands(fraud, spacing = "tail", tail_probs = c(0.02, 0.05, 0.10, 0.20, 0.50), n_boot = 0)
bf$table[sample == "holdout", .(band, label, pct, rate, lift, capture)]
#> band label pct rate lift capture
#> <int> <char> <num> <num> <num> <num>
#> 1: 1 [490.1607, Inf) 0.01428571 0.35000000 2.4137931 0.03448276
#> 2: 2 [474.5723, 490.1607) 0.03785714 0.33962264 2.3422251 0.12315271
#> 3: 3 [464.3536, 474.5723) 0.03857143 0.37037037 2.5542784 0.22167488
#> 4: 4 [450.7431, 464.3536) 0.09142857 0.27343750 1.8857759 0.39408867
#> 5: 5 [423.8846, 450.7431) 0.31357143 0.18223235 1.2567748 0.78817734
#> 6: 6 [-Inf, 423.8846) 0.50428571 0.06090652 0.4200449 1.00000000Three actions (pass, review, block) are three tiers; the labels follow the event rate, lowest first.
tf <- scr_tiers(fraud, n_tiers = 3, labels = c("pass", "review", "block"))
tf$table[sample == "holdout", .(tier, label, score_lo, score_hi, pct, rate)]
#> tier label score_lo score_hi pct rate
#> <int> <char> <num> <num> <num> <num>
#> 1: 3 block 458.1708 Inf 0.1321429 0.35675676
#> 2: 2 review 423.2734 458.1708 0.3728571 0.18390805
#> 3: 1 pass -Inf 423.2734 0.4950000 0.05916306How deep can the review queue go? A review team has a capacity per
day. scr_operating() accumulates the score from its
event-rich end and, with a date column, reads the selected volume of
every date at each candidate cut. In the demo the date is monthly, so
the capacity is read per month: here at most 80 cases on nine months in
ten (day_quantile = 0.9), over the six months of the whole
table. The rates of this curve include the training months and are
therefore optimistic; the volumes, which is what a capacity is about,
are not.
d_fraud <- data.frame(score = scr_apply(fraud, scr_demo)$score, y = scr_demo$default,
month = scr_demo$ref_date)
op_f <- scr_operating(d_fraud, objective = "risk", direction = "higher_is_riskier",
max_per_day = 80, date = "month")
op_f
#> <scr_operating> target "y" | objective risk | higher_is_riskier | side event (from the high scores, score >= cut)
#> sample 'all' | 4188 score values | 6 dates | economics: none
#> constraints: max_per_day 80
#>
#> Optimum: cut 463.0245 | depth 10.4% (n 438) | rate 42.9% [38.3%, 47.6%] | capture 31.2% | value -
#> binding: max_per_day | shadow price - per additional case | next marginal rate 33.6%
#>
#> Curve (selected rows)
#> depth cut n_sel rate [lo, hi] capture lift marginal value day_q feasible
#> 1.0% 497.4886 42 66.7% [51.7%, 79.4%] 4.7% 4.65 45.6% - 8.5 yes
#> 5.0% 475.1601 210 47.6% [40.9%, 54.4%] 16.6% 3.32 42.1% - 39.0 yes
#> 10.0% 463.7915 420 42.9% [38.2%, 47.6%] 29.9% 2.99 33.6% - 78.0 yes
#> 10.4% 463.0245 438 42.9% [38.3%, 47.6%] 31.2% 2.99 33.6% - 80.0 yes <- optimum
#> 20.0% 450.096 840 34.4% [31.3%, 37.7%] 48.0% 2.40 22.6% - 152.0 no
#> 50.0% 423.8132 2,100 23.8% [22.0%, 25.7%] 83.1% 1.66 10.3% - 358.5 no
#> 100.0% -Inf 4,200 14.3% [13.3%, 15.4%] 100.0% 1.00 0.0% - 700.0 noWithout economics the optimum is the deepest cut that fits the
capacity, and binding names the constraint that stops it.
The curve shows what the capacity costs: the optimum captures 31.2% of
the defaults, and the next score value beyond it still has a smoothed
default rate of 33.6%, against an overall rate of 14.3%.
4. Churn: anchored tiers, claims and a budget
Under propensity the event is the outcome sought, and tiers can be anchored on event rates that mean something to the business: below 15%, around the overall churn rate and above 50%.
ta <- scr_tiers(churn, method = "anchored", anchors = c(0.15, "overall", 0.5))
ta
#> <scr_study_tiers> target "churn" | measure propensity | higher_is_riskier
#> anchored | 4 tiers requested, 4 achieved | fitted on 'train'
#> cuts: 432.3465, 456.9944, 486.6599
#> sample n events rate IV PSI tiers monotone distinct
#> train 2,800 797 28.46% 0.761 - 4 yes yes
#> holdout 1,400 424 30.29% 0.692 0.0008 4 yes yes
#>
#> Tiers on 'holdout' (event-richest first)
#> tier label score pct rate [95% CI] p_adj
#> 4 01.high [486.6599, Inf) 13.9% 65.46% [58.58%, 71.89%] 0.000
#> 3 02.medium high [456.9944, 486.6599) 34.1% 34.38% [30.22%, 38.73%] 0.001
#> 2 03.medium low [432.3465, 456.9944) 31.7% 24.32% [20.51%, 28.47%] 0.000
#> 1 04.low [-Inf, 432.3465) 20.4% 8.77% [5.90%, 12.47%] -scr_claims() turns such statements into tests. Each
claim names a group (a tier, a band or a score range), a direction and a
rate; it is tested on the hold-out with an exact one-sided binomial
test, Holm-adjusted across the claims, and reported with its one-sided
Jeffreys bound. A claim is “supported” when the data show it, “refuted”
when they show the opposite and “not proven” otherwise.
claims <- data.frame(
name = c("high tier churns", "top of the score", "low tier is quiet"),
label = c("high", NA, "low"), score_lo = c(NA, 480, NA),
op = c(">=", ">=", "<="), rate = c(0.50, 0.60, 0.15))
cl <- scr_claims(ta, claims)
cl
#> <scr_claims> target "churn" | objective propensity | sample 'holdout' | level 95% (one-sided) | adjustment holm | average
#> claim group n rate bound claimed p_adj verdict
#> high tier churns high 194 65.5% 59.7% >= 50.0% 0.0000 supported
#> top of the score score >= 480 278 60.4% 55.5% >= 60.0% 0.4675 not proven
#> low tier is quiet low 285 8.8% 11.8% <= 15.0% 0.0024 supported
#>
#> On 'holdout' (n = 194), rows in tier 'high' had an event rate of 65.5% (95% one-sided lower bound
#> 59.7%); the claim 'high tier churns' (rate >= 50%) is supported.
#> On 'holdout' (n = 278), rows with score >= 480 had an event rate of 60.4% (95% one-sided lower
#> bound 55.5%); the claim 'top of the score' (rate >= 60%) is not proven.
#> On 'holdout' (n = 285), rows in tier 'low' had an event rate of 8.8% (95% one-sided upper bound
#> 11.8%); the claim 'low tier is quiet' (rate <= 15%) is supported.The claim on the top of the score shows the difference between a point rate and a statement: the observed rate is 60.4% against a claim of 60%, so the one-sided lower bound sits below the claim and the claim is not proven.
A rate of a group is an average. type = "floor" tests
the claim on the weakest end of the group instead: the reference rows of
the group are cut into ten pre-bins (floor_bins), a
monotone fit of their rates along the score (pool adjacent violators)
finds the end where the claim is hardest to meet (the lowest rates for a
“rate >= r” claim, the highest for “rate <= r”), and the claim is
tested on the hold-out rows of that end. A dip inside the group is
averaged with its neighbors by the fit, and the pre-bins set the
resolution, so the floor speaks for the weakest tenth or more of the
group, not for each customer.
fl <- scr_claims(ta, claims, type = "floor")
fl$table[, .(name, group, score_lo, score_hi, n, rate, bound, verdict)]
#> name group score_lo score_hi n rate bound
#> <char> <char> <num> <num> <num> <num> <num>
#> 1: high tier churns high 486.6599 496.8521 79 0.5569620 0.4645748
#> 2: top of the score score >= 480 480.0000 483.7317 46 0.4782609 0.3604868
#> 3: low tier is quiet low 430.0975 432.3465 18 0.2222222 0.4082341
#> verdict
#> <char>
#> 1: not proven
#> 2: not proven
#> 3: not provenAn average can be supported while its floor is not: on the tier “high”, the average claim is supported but its floor is not proven, because the customers at the low end of the tier churn less than the tier as a whole.
A retention campaign has a gain per event reached (a churner who gets
the offer; how many of them stay is part of that gain), a cost per
contact and a budget. scr_operating() maximizes the value
of the campaign along the score within the budget and reports the shadow
price, what one more contact beyond the optimum would be worth.
op_c <- scr_operating(churn, gain_event = 100, cost_select = 20, budget = 3000)
op_c$optimum[, .(cut, depth, n_sel, rate_sel, value, binding, shadow_price)]
#> cut depth n_sel rate_sel value binding shadow_price
#> <num> <num> <num> <num> <num> <char> <num>
#> 1: 492.3036 0.1042857 146 0.6986301 7280 budget 37.14286The binding constraint is the budget, and the shadow price of 37.1 per contact is positive: the next contacts would still pay, so the budget, not the score, limits the campaign.
5. Credit and churn on the same customers
The two scorecards score the same hold-out rows.
scr_score_cross() crosses their tiers, reports the event
rate of both targets in every cell, the rank association of the scores,
and the overlap of the customers each score puts first.
ho <- scr_demo[res$split$holdout_idx, ]
d_x <- data.frame(credit = scr_apply(credit, ho)$score, churn = scr_apply(churn, ho)$score,
default = ho$default, churned = ho$churn)
cx <- scr_score_cross(d_x, "credit", "churn", y_a = "default", y_b = "churned",
objective_b = "propensity", cuts_a = tr, cuts_b = ta)
cx
#> <scr_score_cross> A "credit" (risk, higher_is_safer) | B "churn" (propensity, higher_is_riskier)
#> 1,400 rows | outcomes: default, churned | bands: A tiers, B tiers
#> association (raw scores): Spearman -0.432, Kendall tau-b -0.299 | oriented to the event-rich ends: 0.432, 0.299
#>
#> Share of rows
#> A \ B B1 B2 B3 B4 total
#> A1 very low 13.4% 17.1% 11.1% 2.4% 43.9%
#> A2 low 3.2% 7.1% 7.6% 1.8% 19.6%
#> A3 medium 2.9% 6.4% 9.7% 4.6% 23.6%
#> A4 high 0.8% 0.6% 4.0% 2.1% 7.5%
#> A5 very high 0.1% 0.4% 1.7% 3.1% 5.4%
#> total 20.4% 31.7% 34.1% 13.9% 100.0%
#>
#> Event rate of "default"
#> A \ B B1 B2 B3 B4 total
#> A1 very low 2.7% 5.0% 7.1% 9.1% 5.0%
#> A2 low 8.9% 12.1% 9.4% 24.0% 11.6%
#> A3 medium 10.0% 18.9% 25.0% 32.8% 23.0%
#> A4 high 18.2% 33.3% 42.9% 34.5% 37.1%
#> A5 very high 50.0% 33.3% 37.5% 30.2% 33.3%
#> total 5.6% 10.4% 18.4% 27.3% 14.5%
#>
#> Event rate of "churned"
#> A \ B B1 B2 B3 B4 total
#> A1 very low 6.4% 18.8% 19.4% 60.6% 17.4%
#> A2 low 15.6% 27.3% 39.6% 56.0% 32.7%
#> A3 medium 12.5% 30.0% 37.5% 60.9% 37.0%
#> A4 high 0.0% 44.4% 42.9% 79.3% 48.6%
#> A5 very high 50.0% 83.3% 70.8% 72.1% 72.0%
#> total 8.8% 24.3% 34.4% 65.5% 30.3%
#> B bands: B1 low, B2 medium low, B3 medium high, B4 high
#>
#> Overlap (each score selects from its event-rich end)
#> depth share_a share_b n_a n_b both A only B only Jaccard
#> 5.0% 5.0% 5.0% 70 70 17 53 53 0.138
#> 10.0% 10.0% 10.0% 140 140 45 95 95 0.191
#> 20.0% 20.0% 20.0% 280 280 125 155 155 0.287
#>
#> Event rate of each set (B only = swap-in, A only = swap-out)
#> depth outcome A B both A only B only
#> 5.0% churned 71.4% 75.7% 76.5% 69.8% 75.5%
#> 10.0% churned 64.3% 69.3% 77.8% 57.9% 65.3%
#> 20.0% churned 52.9% 60.4% 68.0% 40.6% 54.2%
#> 5.0% default 34.3% 30.0% 35.3% 34.0% 28.3%
#> 10.0% default 35.7% 27.9% 31.1% 37.9% 26.3%
#> 20.0% default 31.4% 25.7% 32.8% 30.3% 20.0%The oriented association (Spearman 0.43) is positive: the customers the credit score calls risky tend to be those the churn score calls likely to leave. The overlap table says how far that goes at the top of each score, and the rates of the “A only” and “B only” sets say what changes when one list is replaced by the other.
6. Why the default rate moved
The hold-out defaults at a different rate than the training sample.
Two things can move a rate: the population shifted along the score (the
mix), or the same score now carries a different risk (the rates).
scr_mix_shift() splits the change between the two, band by
band, on bands frozen on the base. The two effects add up to the change
exactly.
ms <- scr_mix_shift(credit, n_bands = 5)
ms
#> <scr_mix_shift> target "default" | objective risk | higher_is_safer
#> base 'train' | 5 bands frozen on the base | 1 comparison
#> holdout rate 14.2% -> 14.5% (+0.25 pp): mix -0.35 pp, rate +0.60 pp | PSI 0.0027 (critical 0.0102)
#>
#> Largest band effects on 'holdout' (p_adj: Holm-adjusted test of the band rate)
#> band score share rate mix rate total p_adj
#> 1 [-Inf, 523.5026) 20.0% -> 18.2% 36.1% -> 31.4% -0.60 pp -0.90 pp -1.50 pp 0.955
#> 2 [523.5026, 542.0923) 20.0% -> 20.6% 18.8% -> 21.9% +0.12 pp +0.63 pp +0.75 pp 0.955
#> 3 [542.0923, 557.7648) 20.0% -> 21.4% 9.1% -> 11.7% +0.14 pp +0.54 pp +0.68 pp 0.955
#> 4 [557.7648, 576.5187) 20.0% -> 19.8% 4.8% -> 5.8% -0.01 pp +0.19 pp +0.18 pp 1.000
#> 5 [576.5187, Inf) 20.0% -> 20.1% 2.5% -> 3.2% +0.00 pp +0.14 pp +0.14 pp 1.000Of the change of +0.25 percentage points, -0.35 come from the mix and
+0.60 from the band rates; 0 band(s) changed their rate significantly
after the Holm adjustment, and the PSI of the band shares sits below its
critical value. With by, every period is compared with the
base, here the first month of the scored rows:
scr_mix_shift(credit, by = "date", n_bands = 5)$summary[
, .(group, rate_base, rate_cmp, delta, mix_total, rate_total, psi)]
#> group rate_base rate_cmp delta mix_total rate_total
#> <char> <num> <num> <num> <num> <num>
#> 1: 2026-02-01 0.1414286 0.1457143 0.004285714 -2.762334e-04 0.0045619477
#> 2: 2026-03-01 0.1414286 0.1371429 -0.004285714 1.721942e-05 -0.0043029337
#> 3: 2026-04-01 0.1414286 0.1457143 0.004285714 -9.430080e-03 0.0137157940
#> 4: 2026-05-01 0.1414286 0.1400000 -0.001428571 -1.060643e-03 -0.0003679283
#> 5: 2026-06-01 0.1414286 0.1500000 0.008571429 -9.716728e-03 0.0182881565
#> psi
#> <num>
#> 1: 0.002097308
#> 2: 0.005328688
#> 3: 0.013627112
#> 4: 0.006771972
#> 5: 0.0096411917. One score, many segments
A single scorecard is applied to every channel.
scr_segments() reads the score within each segment against
the pooled rows: the AUC with its DeLong standard error and a chi-square
test of equal AUC across the segments, the observed events against those
expected from the pooled bands (indirect standardization), the offset on
the log-odds scale and the slope of the score relative to the pooled
slope. The action column summarizes them under stated
tolerances: a shared score, an intercept offset, a separate model, or
too few events to say.
sg <- scr_segments(credit, ho, segment = "ds_channel")
sg
#> <scr_segments> target "default" | objective risk | higher_is_safer | segments of 'ds_channel'
#> tolerances: AUC 0.03, offset 0.25, slope 0.25 | fewest events 20 | level 95%
#>
#> Pooled: n 1,400 | rate 14.5% | AUC 0.7394 | equal AUC across 3 segments: chi-square 1.17 (df 2), p 0.5577
#> segment n rate AUC [lo, hi] PSI O/E [lo, hi] offset slope ratio action
#> APP 701 17.1% 0.7184 [0.671, 0.766] 0.0357 1.05 [0.89, 1.23] +0.06 0.92 shared
#> STORE 267 12.0% 0.7383 [0.647, 0.829] 0.0955 1.03 [0.73, 1.40] +0.03 1.00 shared
#> WEB 432 11.8% 0.7636 [0.697, 0.831] 0.0304 0.89 [0.68, 1.14] -0.13 1.14 sharedThe test of equal AUC has a p-value of 0.56, and the O/E intervals all cover 1: the hold-out gives no reason to treat a channel apart.
8. Time, treatment, rules and episodes
Four more studies need columns the demo table does not have; each help page has a worked example.
-
scr_maturity()follows each band over time with censoring (a Kaplan-Meier curve per band) and shows where the event rate flattens: the performance window. -
scr_uplift()reads a score on a treated and a control group: the uplift per band with its interval, the Qini coefficient and a check of the randomization. -
scr_overlap()compares expert rules with the alerts of a score: what each catches alone, and which rules the score already covers. -
scr_detection()measures, per alert threshold, how many fraud episodes are detected, after how many fraudulent transactions and with how much of the loss prevented.
9. Production
Bands and tiers are assigned to new scores in R and in SQL from the
same frozen cuts, and every study writes a workbook. A tier label
carries its order in front, 01 for the tier with the
highest event rate, so that the labels sort from the event-richest tier
in a report or an ORDER BY; the tiers table has the same
value in tier_label, and numbered = FALSE
returns the plain labels.
head(scr_apply(tr, c(495, 520, 560)))
#> score tier tier_label
#> <num> <int> <char>
#> 1: 495 5 01.very high
#> 2: 520 3 03.medium
#> 3: 560 1 05.very low
sql <- scr_sql(tr, table = "scored", dialect = "postgres")
substr(sql[9], 1, 80)
#> [1] " CASE WHEN s.score IS NULL THEN NULL WHEN s.score < 500 THEN 5 WHEN s.score <"
substr(sql[10], 1, 80)
#> [1] " CASE WHEN s.score IS NULL THEN NULL WHEN s.score < 500 THEN '01.very high' W"scr_export() writes
study_tiers_<target>.xlsx,
claims_<target>.xlsx,
operating_<target>.xlsx,
score_cross_<a>_<b>.xlsx,
mix_shift_<target>.xlsx and
segments_<target>.xlsx for the objects above.
