Tests claims such as "the customers scoring 625 or more respond at a rate of at least 60%" or "tier 'low' defaults at no more than 2%" on a study sample, and writes each verdict as one English sentence with the observed rate and its one-sided bound.
Usage
scr_claims(x, claims, ...)
# S3 method for class 'scr_study'
scr_claims(
x,
claims,
level = NULL,
adjust = c("holm", "none"),
type = c("average", "floor"),
sample = NULL,
floor_bins = 10L,
...
)
# S3 method for class 'scr_scorecard'
scr_claims(
x,
claims,
level = NULL,
adjust = c("holm", "none"),
type = c("average", "floor"),
sample = "holdout",
reference = "train",
n_bands = NULL,
max_cells = 1e+05,
floor_bins = 10L,
...
)
# S3 method for class 'data.frame'
scr_claims(
x,
claims,
level = 0.95,
adjust = c("holm", "none"),
type = c("average", "floor"),
sample = NULL,
score = "score",
y = "y",
objective = "risk",
direction = NULL,
weight = NULL,
reference = NULL,
study = NULL,
counts = FALSE,
n = "n",
events = "events",
n_bands = 20L,
max_cells = 1e+05,
floor_bins = 10L,
...
)Arguments
- x
An object from
scr_bands()orscr_tiers(), an object fromscr_scorecard(), or adata.framewith one row per scored case (or one row per score value withcounts = TRUE).- claims
A
data.frameof claims; see the section Claims.- ...
Passed on to the methods; an unknown argument is an error.
- level
Confidence level of the tests and of the one-sided bounds. For a scorecard,
NULLusesconfig$study_level(0.95); for a score study,NULLuses the level of the study.- adjust
"holm"(default) or"none": multiplicity adjustment across the claims.- type
"average"(the rate of the whole group) or"floor"(the rate at the weakest end of the group under a monotone fit of the rate); see the section Average and floor.- sample
For a score study: the sample the claims are read on,
NULLfor its first study sample. For a scorecard: that sample ("holdout"). For a data.frame: the name of a column with sample labels, as inscr_bands(); the claims are read onstudy.- floor_bins
Pre-bins of equal share of a group for
type = "floor": the resolution of the floor (default 10; 1 makes the floor the average).- reference
For a scorecard: the sample the bands are frozen on (
"train"). For a data.frame: the label of the reference sample;NULLtakes the first level of thesamplecolumn (the levels of a factor in their order, numbers in numeric order, text sorted).- n_bands
Bands of the internal percentile study (their labels and numbers can be used in
claims$label). For a scorecard,NULLusesconfig$study_bands(20).- max_cells
Largest number of distinct score values kept exactly.
- score, y
Column names of the score and of the 0/1 outcome (
NAallowed).- objective
"risk"(the event is the bad case) or"propensity"(the event is the good case).- direction
"higher_is_safer"or"higher_is_riskier";NULLderives it fromobjective.- weight
Optional column of non-negative case weights.
- study
For a data.frame: the label of the sample the claims are read on;
NULLtakes the first label other than the reference.- counts
TRUEwhenxis pre-aggregated: one row per score value with the columnsscore,nandevents(and, optionally,valueandvalue_events).- n, events
Column names of the counts when
counts = TRUE.
Value
An object of class c("scr_claims", "list"):
tableOne row per claim:
name,group(the label or the score range),type,score_loandscore_hi(the range tested),n(rows with a known outcome; their weighted volume under weights),n_eff(Kish effective size, thenof the test; equal tonwithout weights),events,rate,bound(one-sided Jeffreys bound),op,rate_claimed,p_valueandp_adj(test of the claim),p_refuteandp_refute_adj(the opposite test),verdictandstatement(under weights, the sentence quotes the effective n of the test next to the weighted volume).sample,reference,level,adjust,type,floor_bins,objective,direction,target,callThe settings (
floor_binsis used bytype = "floor"only).
Claims
claims is a data.frame with one row per claim:
op">="(the event rate is at leastrate) or"<="(at mostrate).rateThe claimed event rate, in (0, 1).
labelA band or tier of the study: its label (the
labelcolumn of the study table, or the numberedtier_labelof a tiers study) or its number (thebandortiercolumn).score_lo,score_hiInstead of a label, a score range
[score_lo, score_hi); a missing end is open. A claim with neither a label nor a score range is on the whole sample.nameOptional: a name for the printed statement.
Test
Each claim is read on the rows of the group with a known outcome in the
study sample. With weights, the counts are taken to the Kish effective
size \(n = (\sum w)^2 / \sum w^2\) and \(x = \hat p\, n\) effective
events (without weights, the counts themselves). The claim
"rate >= r" is the alternative to \(H_0: p \le r\), tested by the exact
one-sided binomial p-value \(P(X \ge x)\), \(X \sim
\mathrm{Binomial}(n, r)\); \(x\) and \(n\) are rounded to whole numbers
for the binomial only. The refutation is the opposite test,
\(P(X \le x)\). For "rate <= r" the two tests swap. With
adjust = "holm", the claim p-values and the refutation p-values are each
Holm-adjusted across the claims. They are two Holm families, the tests of
the claims and the tests of the refutations, each controlled at
1 - level; a claim and its refutation can never both be significant.
The verdict is "supported" when the adjusted p-value of the claim is
below 1 - level, "refuted" when the adjusted p-value of the
refutation is, and "not proven" otherwise (also for a group without a
row with a known outcome). bound is the one-sided Jeffreys bound at
level in the direction of the claim, on the unrounded effective counts:
the lower bound, the Beta(\(x + 1/2, n - x + 1/2\)) quantile
1 - level, for "rate >= r" (0 when \(x = 0\)), and the upper bound,
its quantile level, for "rate <= r" (1 when \(x = n\)). The bound
describes the uncertainty; the verdict comes from the tests.
Average and floor
type = "average" reads the claim on the event rate of the whole group.
type = "floor" reads it on the weakest end of the group under a
monotone fit of the rate: a claim that holds there holds for the part of
the group where it is hardest to meet, not only for the average.
The reference rows of the group are cut into floor_bins pre-bins of
equal share (tie-safe, as the bands of scr_bands()), and their event
rates are fitted by pool adjacent violators: the rate is made monotone
along the score, rising toward its event-rich end. The weakest end is the
run of pre-bins with the lowest fitted rate of the group for
"rate >= r", with the highest for "rate <= r". Its score range is found
on the reference sample and the claim is tested on the rows of the study
sample in that range, so the end is not chosen on the rows it is tested
on (unless the study sample is the reference). When the fitted rate is
flat over the group, the weakest end is the whole group.
Two limits follow from the fit. A dip of the rate inside the group is
pooled with its neighbors by the monotone fit, so the floor does not see
a weak pocket in the middle of the group, only the end the fit calls
weakest. And floor_bins sets the resolution: the floor speaks for runs
of pre-bins, about one part in floor_bins of the group or more, never
for a single customer (a fit on single score values would spike at its
ends and read the floor on a handful of rows).
Input
An object from scr_bands() or scr_tiers() is read as it is: labels
refer to its bands or tiers, and sample picks one of its samples (by
default its first study sample). A scorecard or a data.frame is first
summarized by a percentile study (scr_bands(), n_bands bands frozen
on the reference, no bootstrap) whose count table keeps every distinct
score up to max_cells; the finite ends of the score ranges of claims
are forced as cell edges, so a score range is exact even when more
distinct scores are pooled into cells. A score range read on a study
whose cells were pooled must not cut through a cell, or the call is an
error.
References
Brown, L. D., Cai, T. T. and DasGupta, A. (2001). Interval estimation for a binomial proportion. Statistical Science, 16(2), 101-133. doi:10.1214/ss/1009213286
Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65-70.
Kish, L. (1965). Survey Sampling. Wiley.
See also
scr_bands() and scr_tiers() for the groups, scr_operating()
to choose a cut under constraints.
Other score-studies:
scr_bands(),
scr_detection(),
scr_maturity(),
scr_mix_shift(),
scr_operating(),
scr_overlap(),
scr_rag(),
scr_rag_plan(),
scr_score_cross(),
scr_segments(),
scr_tiers(),
scr_uplift()
Examples
set.seed(1)
x <- rnorm(4000)
d <- data.frame(score = round(500 + 50 * x),
y = rbinom(4000, 1, plogis(-0.5 + 1.2 * x)))
claims <- data.frame(op = c(">=", ">=", "<="), rate = c(0.60, 0.80, 0.25),
score_lo = c(550, 550, NA), score_hi = c(NA, NA, 450),
name = c("top converts", "top converts well", "bottom is cold"))
cl <- scr_claims(d, claims, objective = "propensity")
cl
#> <scr_claims> target "y" | objective propensity | sample 'all' | level 95% (one-sided) | adjustment holm | average
#> claim group n rate bound claimed p_adj verdict
#> top converts score >= 550 692 78.3% 75.7% >= 60.0% 0.0000 supported
#> top converts well score >= 550 692 78.3% 75.7% >= 80.0% 0.8743 not proven
#> bottom is cold score < 450 645 9.5% 11.5% <= 25.0% 0.0000 supported
#>
#> On 'all' (n = 692), rows with score >= 550 had an event rate of 78.3% (95% one-sided lower bound
#> 75.7%); the claim 'top converts' (rate >= 60%) is supported.
#> On 'all' (n = 692), rows with score >= 550 had an event rate of 78.3% (95% one-sided lower bound
#> 75.7%); the claim 'top converts well' (rate >= 80%) is not proven.
#> On 'all' (n = 645), rows with score < 450 had an event rate of 9.5% (95% one-sided upper bound
#> 11.5%); the claim 'bottom is cold' (rate <= 25%) is supported.
cl$table[, c("name", "n", "rate", "bound", "p_adj", "verdict")]
#> name n rate bound p_adj verdict
#> <char> <num> <num> <num> <num> <char>
#> 1: top converts 692 0.78323699 0.7566252 2.834300e-24 supported
#> 2: top converts well 692 0.78323699 0.7566252 8.743393e-01 not proven
#> 3: bottom is cold 645 0.09457364 0.1148683 1.504897e-23 supported
# the same claims at the weakest end of each group, not only on its average
scr_claims(d, claims, objective = "propensity", type = "floor")$table$verdict
#> [1] "not proven" "refuted" "supported"
# claims on the tiers of a study, by label
tr <- scr_tiers(d, objective = "propensity", n_tiers = 3)
scr_claims(tr, data.frame(label = "high", op = ">=", rate = 0.5))
#> <scr_claims> target "y" | objective propensity | sample 'all' | level 95% (one-sided) | adjustment holm | average
#> claim group n rate bound claimed p_adj verdict
#> rate >= 50% high 1,199 71.5% 69.3% >= 50.0% 0.0000 supported
#>
#> On 'all' (n = 1,199), rows in tier 'high' had an event rate of 71.5% (95% one-sided lower bound
#> 69.3%); the claim 'rate >= 50%' is supported.
