Optimal binning and Weight of Evidence for credit scoring and risk modelling: 37 C++ binning implementations behind one R interface, from a raw feature store to a deployed scorecard.
install.packages("OptimalBinningWoE")
library(OptimalBinningWoE)
CASE, for 14 dialects.
Binning replaces a raw column by its bins; WoE replaces each bin by the log-ratio of the event and non-event distributions it holds.
pi = share of the events that fall in bin i; qi = share of the non-events. A positive WoE is a riskier bin.
An IV at or above 0.50 is far more often leakage than signal, which
is why obwoe_select() rejects it by default. Medium is where the workhorses
of a scorecard sit.
obwoe(data, target, feature = NULL,
min_bins = 2, max_bins = 7,
algorithm = "auto",
control = control.obwoe())
NULL bins every column but the target."auto": jedi
for a binary target, jedi_mwoe for a multinomial one.control.obwoe() on sheet 2.m <- obwoe(german, target = "default", max_bins = 6)
m$summary # feature, type, algorithm, n_bins, total_iv
m$results$age # bin, woe, iv, count, count_pos, count_neg,
# cutpoints, converged, iterations
m$target_type # "binary" or "multinomial"
sort_by = "iv", decreasing."iv" ranking bar chart ·
"woe" profile of one feature · "bins" counts and
event rate. Pass feature=, top_n = 15.Two criteria govern admission: Information Value strength and guaranteed rank ordering. Every candidate returns a row, so an automatic verdict stays reviewable.
sel <- obwoe_select(m,
detail = "summary", # or "full": one row per bin
iv_min = 0.02, iv_max = 0.5,
require_monotonic = "numeric", # "all" | "none"
monotonicity = "weak", # or "strict"
min_bins = 2, max_bins = Inf,
min_bin_pct = 0, allow_degenerate = FALSE,
top_n = NULL, sort_by = "iv", decreasing = TRUE)
monotonic_direction, n_violations, spearman.reason and reason_desc say why.keep <- sel$feature[sel$selected]
obwoe_apply(data, obj,
suffix_bin = "_bin", suffix_woe = "_woe",
keep_original = TRUE, na_woe = 0)
Adds <feature>_bin and <feature>_woe per binned
column. na_woe is the fallback for an unseen category or an unmodelled
missing value. Numerical intervals are half-open on the right, (a, b].
missing_values = c(-999).missing_values = c("NA","Missing","").g <- obwoe_gains(m, feature = "duration") g <- obwoe_gains(scored, target = "default", feature = "score_decile", use_column = "direct", n_groups = 10)
"auto" · "bin" ·
"woe" · "direct"."id" (the algorithm's own order, default) ·
"woe" · "event_rate" · "bin".g$metrics # ks, gini, auc, total_iv, ks_bin plot(g, type = "cumulative") # "ks" | "lift" | "woe_iv"
Related helpers: obwoe_gains_score() for the bin-level
engine behind a fitted result, obwoe_gains_variable() for any binned
data frame plus a grouping column.
ob_preprocess(feature, target, num_miss_value = -999, char_miss_value = "N/A", outlier_method = "iqr", # "zscore" | "grubbs" outlier_process = FALSE, preprocess = "both", iqr_k = 1.5, zscore_threshold = 3, grubbs_alpha = 0.05)
Missing values become a value of their own rather than a dropped row, so the bin that holds them carries its own WoE and stays visible in the model document.
ob_check_distincts(x, target) # cardinality and
# separation warnings
The Statlog (German Credit) benchmark ships with the package.
german <- read.csv(gzfile(system.file( "extdata", "germancredit.csv.gz", package = "OptimalBinningWoE")), stringsAsFactors = FALSE) german$default <- 1L - german$credit_risk german$credit_risk <- NULL m <- obwoe(german, "default", max_bins = 6) sel <- obwoe_select(m) out <- obwoe_apply(german, m)
28 algorithm names cover 37 implementations — 21 accept a numerical feature, 16 a categorical one, 9 both. Every one is callable on its own, and the whole pipeline is callable as a single function.
obwoe_algorithms()
vignette("introduction")
| Family | Name | N C | Optimises |
|---|---|---|---|
| Information- theoretic |
jedi | NC | Joint entropy-driven intervals (default) |
| jedi_mwoe | NC | JEDI for a multinomial target | |
| mdlp | N· | Fayyad–Irani MDL stopping rule | |
| fast_mdlp | N· | MDLP with a monotonicity constraint | |
| dmiv | NC | Divergence measures (Zeng, 2013) | |
| ivb | ·C | IV by dynamic programming | |
| Statistical merging |
cm | NC | Enhanced ChiMerge, χ² on neighbours |
| fetb | NC | Fisher's exact test on neighbours | |
| mob | NC | Monotonic optimal binning | |
| Shape- constrained |
ir | N· | Isotonic regression (PAVA) |
| mrblp | N· | Monotonic risk + likelihood-ratio pre-bins | |
| mblp | N· | Monotonic binning by linear programming | |
| oslp | N· | Optimal supervised partitioning | |
| ldb | N· | Local density binning | |
| lpdb | N· | Local polynomial density binning | |
| gmb | ·C | Greedy merge under monotonicity | |
| Exact optimisation |
dp | NC | Dynamic programming, global optimum |
| bb | N· | Branch and bound | |
| milp | ·C | Mixed-integer formulation | |
| sblp | ·C | Sequential bounded partitioning | |
| Search & metaheuristic |
udt | NC | Entropy-based tree partitioning |
| sab | ·C | Simulated annealing | |
| mba | ·C | Agglomerative monotonic merging | |
| swb | ·C | Sliding window over ordered levels | |
| Unsupervised & streaming |
sketch | NC | Streaming quantile sketch, sketch_k |
| ewb | N· | Equal width, then IV refinement | |
| kmb | N· | K-means initialisation | |
| ubsd | N· | Standard-deviation cut points |
ir, mrblp,
mblp, mob. Categorical: gmb,
mob, mba.sketch, then ewb
or kmb.dp and bb for numerical,
milp and sblp for categorical.sketch, mba,
swb.algorithm = "auto".sc <- obwoe_scale(pdo = 20, score_ref = 600, odds_ref = 50, direction = "higher_is_safer") s <- obwoe_score(link, sc, round = TRUE)
The models predict the log-odds of the event, so the sign is negative: a high
WoE is a risky bin and must score fewer points. Two identities pin the scale
— a case at odds_ref scores exactly score_ref, and
doubling the good-to-bad odds adds exactly pdo points, everywhere.
control.obwoe( bin_cutoff = 0.05, # min share per bin max_n_prebins = 20, # pre-merge ceiling convergence_threshold = 1e-6, max_iterations = 1000, bin_separator = "%;%", # joins merged levels verbose = FALSE)
max_n_prebins is a modelling decision, not a detail. Pre-binning runs before the optimiser sees the data, so a heavy tail can be smeared into one quantile cell and lost with no warning. Leave the default, and tune it against held-out IV for long-tailed continuous predictors only.
r <- ob_numerical_ir(german$duration, y, min_bins = 3, max_bins = 5, bin_cutoff = 0.05, max_n_prebins = 20, auto_monotonicity = TRUE) r <- ob_categorical_jedi(german$purpose, y, bin_separator = "%;%")
Every name in the table is exported as ob_numerical_<name>(),
ob_categorical_<name>(), or both. They return the same bin, woe, iv
and count vectors that obwoe() stores per feature.
card <- obwoe_scorecard(german,
target = "default", split = 0.7,
validation = NULL, # an out-of-time frame
exclude = NULL, feature = NULL,
binning = list(max_bins = 6),
screening = list(iv_max = 0.5),
engine = "glm", # "obwoe" | "glmnet" | list()
control = control.obwoe_scorecard(),
file = "scorecard.xlsx", seed = 42)
Split, bin, screen by IV and correlation, fit, scale to points, and write the model document. Three properties are enforced rather than assumed:
obwoe_select() plus a stage column:
in_model, sign_rejected, corr_pruned,
constant_woe, screened_out.predict(card, new_data, type = "score") # "card" | "link" | "prob" | "woe"
obcorr(df, method = "all", threads = 0)
p <- obwoe_prune(woe_df, ranking = keep,
cutoff = 0.7, method = "pearson")
p$keep; p$dropped # variable, correlated_with, correlation
Prune in the WoE space — that is the space the model actually sees.
s <- obwoe_psi(base, compare, n_groups = 10)
s$psi; s$table; s$flag
| PSI | Flag | Reading |
|---|---|---|
| < 0.10 | stable | No action |
| 0.10 – 0.25 | watch | Monitor the shift |
| > 0.25 | act | Population has moved |
A band populated in one vintage and empty in the other reports
Inf rather than being smoothed away — a segment that has vanished is
exactly what monitoring exists to catch.
rec <- recipe(default ~ ., data = german) |>
step_obwoe(all_predictors(), outcome = "default",
algorithm = "auto",
min_bins = 2, max_bins = 10,
bin_cutoff = 0.05, output = "woe",
na_woe = 0)
prep(), bake(), tidy() and
required_pkgs() are all implemented. Four parameters are tunable, each with
a dials constructor:
| tune() | dials | Range |
|---|---|---|
| algorithm | obwoe_algorithm() | 28 names |
| min_bins | obwoe_min_bins() | 2 – 5 |
| max_bins | obwoe_max_bins() | 5 – 20 |
| bin_cutoff | obwoe_bin_cutoff() | 0.01 – 0.10 |
obwoe_sql(m, table = "risk.applications", features = keep, output = "woe", # "bin" | "index" | "both" style = "select", # "case" | "cte" | "view" dialect = "postgres", na_value = 0, explicit_bounds = TRUE, quote_identifiers = "auto", file = NULL)
CASE WHEN duration IS NULL THEN 0
WHEN duration <= 7 THEN -1.31218638896617
WHEN duration > 7 AND duration <= 10 THEN -0.4519851
... ELSE 0 END AS duration_woe
Every expression opens with an explicit IS NULL branch,
because NULL <= 5 is NULL in SQL, not FALSE.