Skip to contents

First stage of the pipeline, callable on its own. Converts the target to 0/1 (resolving event_level), types the candidates (numerics become double, everything else becomes character) and splits train and hold-out before any supervised fit.

Usage

scr_split(
  data,
  target,
  date_col = NULL,
  ratio = 0.3,
  seed = NULL,
  event_level = NULL,
  drop = character(),
  copy = TRUE
)

Arguments

data

A data.frame or data.table with the target and the candidates.

target

Name of the target column. Binary: 0/1, logical, or a two-level factor/character.

date_col

Date column of the out-of-time cut. NULL uses a random stratified split.

ratio

Target hold-out fraction.

seed

Seed of the random split. NULL draws from the session's random stream (reproducible only through a set.seed() of your own; scr_select() passes the seed of scr_config()). A seed is applied locally: the session's random stream is restored on exit.

event_level

Which target value counts as the event. NULL uses the convention (1, or the second alphabetical level).

drop

Columns that are never candidates (identifiers, sibling targets, free text). They stay in the funnel as 00.config.

copy

If TRUE (default), works on a copy of data. FALSE modifies a data.table by reference (target and typing), saving memory; a data.frame is always converted, hence copied.

Value

An scr_split object with data (typed), target, train_idx, holdout_idx, method, cutoff, date_col and cols (features, var_num, var_cat, dropped, event).

Details

The split prefers out-of-time by date_col: it is the only one that tests generalization to a future period. The cut is made on the distinct date values, not by row quantile: it picks the smallest set of most recent periods that already reaches ratio of the population. Without a date column, or with a single period, it falls back to random stratified by the target. The date column is never a candidate: it is the key of the split and leaves the contest. A text date column is read as an ISO date (YYYY-MM-DD, YYYY/MM/DD, YYYY-MM) or as all-digit periods (YYYYMM); rows with a missing date belong to no period and are left out of both train and hold-out, with a warning in the log.

Event orientation

event_level changes what is modeled. Passing 0 makes class 0 the event: the sign of every WOE flips, the emitted SQL changes, the points change. For a text target, the second level in alphabetical order is the event by default, and the choice is always reported. Not to be confused with config$objective, which only orients the reading and the points scale.

Examples

sp <- scr_split(scr_demo, "default", date_col = "ref_date", drop = "id")
#>   OOT: 4 period(s) in train, 2 in hold-out (hold-out starts at 2026-05-01, 33.3% of rows)
sp
#> <scr_split> target "default" | 4,200 rows: train 2,800, hold-out 1,400
#>   method: out-of-time (hold-out from 2026-05-01)
#>   candidates: 38 (32 numeric, 6 categorical) | dropped: 2
#>   event: class '1'
length(sp$train_idx); length(sp$holdout_idx)
#> [1] 2800
#> [1] 1400