Accumulates the score from one end, one score value at a time, and finds the cut that maximizes the value of the decision under volume, budget, daily capacity and event-rate constraints: how many customers to target, how many alerts to raise, or how many applicants to approve.
Usage
scr_operating(x, ...)
# S3 method for class 'scr_scorecard'
scr_operating(
x,
side = NULL,
gain_event = NULL,
cost_select = 0,
revenue_good = NULL,
loss_bad = NULL,
max_n = NULL,
max_share = NULL,
budget = NULL,
max_per_day = NULL,
day_quantile = 0.9,
date = NULL,
min_rate = NULL,
max_rate = NULL,
sample = "holdout",
n_points = 200L,
level = NULL,
max_cells = 1e+05,
...
)
# S3 method for class 'data.frame'
scr_operating(
x,
side = NULL,
gain_event = NULL,
cost_select = 0,
revenue_good = NULL,
loss_bad = NULL,
max_n = NULL,
max_share = NULL,
budget = NULL,
max_per_day = NULL,
day_quantile = 0.9,
date = NULL,
min_rate = NULL,
max_rate = NULL,
sample = NULL,
n_points = 200L,
score = "score",
y = "y",
objective = "risk",
direction = NULL,
weight = NULL,
value = NULL,
study = NULL,
counts = FALSE,
n = "n",
events = "events",
value_events = NULL,
level = 0.95,
max_cells = 1e+05,
...
)Arguments
- x
An object from
scr_scorecard(), or adata.framewith one row per scored case (or one row per score value withcounts = TRUE).- ...
Passed on to the methods; an unknown argument is an error.
- side
"event"or"safe";NULLfollows the objective and the direction (see the section Side).- gain_event
Gain per selected event (side
"event").- cost_select
Cost per selected case (default 0).
- revenue_good, loss_bad
Revenue per accepted non-event and loss per accepted event (side
"safe").- max_n
Largest volume selected.
Largest share of the volume selected, in (0, 1].
- budget
Largest cost,
cost_select * n_sel; needs a positivecost_select.- max_per_day
Largest selected volume per day, read at the
day_quantilequantile of the days; needs a date.- day_quantile
Quantile of the daily volume compared with
max_per_day(0.9: nine days in ten within capacity).- date
For a data.frame: name of a date column, for the daily capacity. For a scorecard: a column of the scored sample;
NULLuses itsdatecolumn when present.- min_rate
Smallest event rate among the selected (side
"event").- max_rate
Largest event rate among the accepted (side
"safe").- sample
For a scorecard: the sample the curve is read on (
"holdout"). For a data.frame: the name of a column with sample labels, as inscr_bands(); the curve is read onstudy.- n_points
About how many rows of the curve to keep.
- level
Confidence level of the Jeffreys intervals. For a scorecard,
NULLusesconfig$study_level(0.95).- max_cells
Largest number of distinct score values kept exactly.
- score, y
Column names of the score and of the 0/1 outcome (
NAallowed).- objective
"risk"(the event is the bad case) or"propensity"(the event is the good case).- direction
"higher_is_safer"or"higher_is_riskier";NULLderives it fromobjective.- weight
Optional column of non-negative case weights.
- value
Optional column of the value of every case: with no
gain_event, the value of the selected events is the gain.- study
For a data.frame: the label of the sample the curve is read on;
NULLtakes the first label other than the first level (the reference ofscr_bands()), or the only one.- counts
TRUEwhenxis pre-aggregated: one row per score value with the columnsscore,nandevents(and, optionally,valueandvalue_events).- n, events
Column names of the counts when
counts = TRUE.- value_events
With
counts = TRUE: the column of the value of the events per score cell.
Value
An object of class c("scr_operating", "list"):
curveThe thinned curve (see the section Curve).
optimumOne row: the columns of the curve at the optimum,
binding,shadow_priceandnext_rate(the marginal rate of the next score value).constraintsOne row per constraint given:
constraint,limit,at_optimum(the constrained quantity at the optimum) andbinding.side,objective,direction,target,sample,level,gain_event,cost_select,revenue_good,loss_bad,day_quantile,callThe settings.
economics,value_columnWhether the curve has a value, and whether it comes from the
valuecolumn instead ofgain_event.select_highTRUEwhen the selection starts at the high scores (score >= cut),FALSEat the low ones (score < cut).n_cells,n_days,quantized,weightedThe number of candidate cuts (score values or pooled cells), the number of distinct dates (
NAwithout a date), whether the scores were pooled intomax_cellscells, and whether weights were used.messageNA, or the explanation of an infeasible or a loss-making optimum.
Side
side = "event" selects from the event-rich end of the score (targeting
under propensity, alerting under fraud, collections under credit);
side = "safe" accepts from the safe end (approval under credit). The
default follows the objective and the direction: propensity selects from
the event-rich end ("event"); risk with higher_is_riskier (fraud)
alerts from the event-rich end ("event"); risk with higher_is_safer
(credit) approves from the safe end ("safe").
The selected rows are score >= cut when the selection starts at the high
scores and score < cut when it starts at the low ones, the convention of
scr_cutoff(). Every cut sits between two adjacent distinct scores (or on
a bucket edge when the scores were pooled into max_cells cells), as in
scr_bands(); the last row of the curve selects every row (cut is
-Inf or Inf).
Curve
One row per candidate cut, in increasing depth: cut, depth (share of
the volume selected), n_sel, events_sel, rate_sel with its Jeffreys
interval rate_lo, rate_hi (on the Kish effective size under weights),
capture (share of all events selected), lift (rate_sel over the
overall rate), marginal_rate (the event rate of the score value just
added, smoothed by pool adjacent violators toward the event-rich end of
the score), cost (cost_select * n_sel), value and feasible. With
a date, day_q (the day_quantile quantile of the selected volume per
day) and pct_days_over (share of days above max_per_day).
The economics:
side
"event":value = gain_event * events_sel - cost_select * n_sel; with avaluecolumn and nogain_event, the sum of the value of the selected events replacesgain_event * events_sel.side
"safe":value = revenue_good * nonevents_sel - loss_bad * events_sel - cost_select * n_sel, the cumulative profit ofscr_strategy()whencost_select = 0(a missing one ofrevenue_goodandloss_badcounts as 0).
Without economics (gain_event, a value column, revenue_good or
loss_bad), value is NA. Non-events are rows with a known outcome
that are not events; rows with a missing outcome count in the volume and
the cost only.
The curve is thinned to about n_points rows evenly spread in depth; the
rows of the optimum, the last row meeting each constraint, the deepest
row and the rows nearest 1%, 5%, 10%, 20% and 50% are always kept. The
optimum and the constraints are evaluated on every cell boundary.
Constraints and optimum
max_n (n_sel <= max_n), max_share (depth <= max_share), budget
(cost_select * n_sel <= budget), max_per_day (day_q <= max_per_day), min_rate (rate_sel >= min_rate, side "event") and
max_rate (rate_sel <= max_rate, the event rate among the accepted,
side "safe"). A row is feasible when it meets every constraint given.
The optimum is the feasible row with the highest value (the smallest
depth on a tie); without economics, the deepest feasible row. A
constraint is binding when dropping it alone, the others kept,
improves the optimum: a higher value, or a greater depth when there are
no economics. A constraint that is slack at the optimum is therefore
never named, and a constraint that stops the curve at the row that is
the best anyway is not binding either. Constraints that stop the optimum
at the same row bind jointly (none improves it alone) and are named
together. When no constraint binds, binding is "value" with
economics (no row is worth more than the optimum) and "end of the curve" without (every row is selected).
shadow_price is the marginal value of the next score value beyond the
optimum, per additional selected case: gain_event * marginal_rate - cost_select on side "event" (with a value column, the smoothed event
value per case of that score value), revenue_good * (1 - marginal_rate) - loss_bad * marginal_rate - cost_select on side
"safe". It is what one more selected case is worth when a constraint
binds; divide it by cost_select for the value of one more unit of
budget. When no row is feasible, the optimum is NA and a warning names
the constraints that the first row already breaks.
Daily capacity
A day is a distinct value of the date column: with a monthly date, read
the capacity per month. For every candidate cut, the selected volume of
each day is a cumulative sum over a table of counts per day and score
value; day_q is its quantile across the days (type 7 of
stats::quantile()). The quantile grows with the depth, so the deepest
cut within max_per_day is found by bisection. Rows with a missing date
count in the curve but not in the daily volumes. A scorecard uses the
dates of its scored sample when they exist. A date-time column counts
every distinct time as a day: convert it with as.Date() first.
References
Brown, L. D., Cai, T. T. and DasGupta, A. (2001). Interval estimation for a binomial proportion. Statistical Science, 16(2), 101-133. doi:10.1214/ss/1009213286
Thomas, L. C., Crook, J. and Edelman, D. (2017). Credit Scoring and Its Applications, 2nd edition. SIAM. doi:10.1137/1.9781611974560
See also
scr_cutoff() and scr_strategy() for the cut-off sweep and the
strategy table of a scorecard, scr_claims() to test statements about
the selected rates.
Other score-studies:
scr_bands(),
scr_claims(),
scr_detection(),
scr_maturity(),
scr_mix_shift(),
scr_overlap(),
scr_rag(),
scr_rag_plan(),
scr_score_cross(),
scr_segments(),
scr_tiers(),
scr_uplift()
Examples
set.seed(1)
x <- rnorm(5000)
d <- data.frame(score = round(500 + 50 * x),
y = rbinom(5000, 1, plogis(-1.5 + 1.2 * x)),
day = as.Date("2026-01-01") + sample(0:29, 5000, TRUE))
# targeting under propensity: a gain per responder, a cost per contact, a budget
op <- scr_operating(d, objective = "propensity", gain_event = 40, cost_select = 6,
budget = 6000)
op
#> <scr_operating> target "y" | objective propensity | higher_is_riskier | side event (from the high scores, score >= cut)
#> sample 'all' | 288 score values | economics: gain 40 per event, cost 6
#> constraints: budget 6,000
#>
#> Optimum: cut 544.5 | depth 19.8% (n 989) | rate 55.5% [52.4%, 58.6%] | capture 47.0% | value 16,026
#> binding: budget | shadow price 9.693 per additional case | next marginal rate 39.2%
#>
#> Curve (selected rows)
#> depth cut n_sel rate [lo, hi] capture lift marginal value feasible
#> 1.0% 615.5 50 86.0% [74.5%, 93.5%] 3.7% 3.68 84.8% 1,420 yes
#> 5.0% 584.5 252 71.8% [66.0%, 77.1%] 15.5% 3.07 61.0% 5,728 yes
#> 10.2% 565.5 510 63.7% [59.5%, 67.8%] 27.8% 2.73 51.9% 9,940 yes
#> 19.8% 544.5 989 55.5% [52.4%, 58.6%] 47.0% 2.37 39.2% 16,026 yes <- optimum
#> 20.2% 543.5 1,011 55.1% [52.0%, 58.1%] 47.6% 2.36 39.2% 16,214 no
#> 49.7% 499.5 2,484 38.6% [36.7%, 40.5%] 82.0% 1.65 18.6% 23,456 no
#> 100.0% -Inf 5,000 23.4% [22.2%, 24.6%] 100.0% 1.00 0.0% 16,760 no
op$optimum[, c("cut", "depth", "n_sel", "rate_sel", "value", "binding", "shadow_price")]
#> cut depth n_sel rate_sel value binding shadow_price
#> <num> <num> <num> <num> <num> <char> <num>
#> 1: 544.5 0.1978 989 0.5551062 16026 budget 9.693215
# alerting under a daily capacity: at most 40 alerts on nine days in ten
scr_operating(d, objective = "propensity", max_per_day = 40, date = "day")$optimum
#> cut depth n_sel events_sel rate_sel rate_lo rate_hi capture lift
#> <num> <num> <num> <num> <num> <num> <num> <num> <num>
#> 1: 546.5 0.1854 927 526 0.5674218 0.53536 0.5990633 0.4499572 2.426954
#> marginal_rate cost value feasible day_q pct_days_over binding
#> <num> <num> <num> <lgcl> <num> <num> <char>
#> 1: 0.4318182 0 NA TRUE 38.2 0.03333333 max_per_day
#> shadow_price next_rate
#> <num> <num>
#> 1: NA 0.3923304
