Skip to contents

Draws trials from an evidence accumulation core and replaces a known subset with contaminants generated by one of three psychologically motivated processes. Which trials were replaced, and by which process, comes back with the data. That is what a real data set cannot give you, and what scoring a screen requires: sensitivity and specificity need labels.

Usage

r_contaminated(
  n,
  generator = c("ddm", "rdm"),
  process = c("leading_edge", "delay", "informationless", "mixed"),
  rate = 0.05,
  par = list(),
  st0 = 0,
  anticipation_anchor = c("min", "q10"),
  anticipation_depth = 0.3,
  delay_min = 0,
  delay_max = 2,
  lapse_prop = 0,
  mix = c(leading_edge = 1, delay = 1, informationless = 1)/3,
  dt = 0.001
)

Arguments

n

Number of trials.

generator

The clean core. "ddm" is a first-principles diffusion (Euler–Maruyama); "rdm" is a racing diffusion, two independent single-boundary Wald accumulators with a shared bound, drawn exactly from the inverse Gaussian.

process

Which contaminant process replaces the selected trials:

  • "leading_edge" generates anticipations. Uniform in a band around the observed minimum of the clean trials in this call (see anticipation_anchor), responding at chance. Anchoring to an observed quantity rather than to the generating non-decision time keeps the construct comparable across generators, whose fitted non-decision times can differ substantially at identical observables.

  • "delay" generates delayed start-ups. A genuine trial plus a uniform shift of delay_min to delay_max seconds; the response, and so the intact accuracy, is kept.

  • "informationless" generates lapses of evidence quality. The trial is redrawn from the core with its evidence scaled by lapse_prop while the speed of processing is held fixed: lapse_prop = 0 is a complete lapse (chance accuracy, response times still in the core's range), lapse_prop = 1 the clean process, and values between are partial lapses on the same continuum.

  • "mixed" draws each contaminant trial's process from mix.

rate

Probability that a trial is a contaminant, in [0, 1).

par

Generator parameters as a list. For "ddm": drift, bound, ndt, zr (defaults 1.5, 1.2, 0.30, 0.5). For "rdm": drift as c(correct, error), bound, ndt (defaults c(3, 1.5), 1.5, 0.30). The defaults are illustrative, not calibrated; simulation designs should set parameters from an observable target instead.

st0

Range of across-trial variability in non-decision time, in seconds. The per-trial non-decision time is uniform on ndt ± st0 / 2, the centred parameterization, so the mean response time is invariant and only the leading edge smears.

anticipation_anchor

Observable the anticipation band is centred on: the minimum ("min") or the 10th percentile ("q10") of the clean trials generated in this call.

anticipation_depth

Half-width of the anticipation band as a proportion of the anchor: response times are uniform on anchor * (1 ± anticipation_depth).

delay_min, delay_max

Bounds of the uniform shift, in seconds, for process = "delay".

lapse_prop

Evidence-quality proportion for process = "informationless"; see above.

mix

Relative weights of leading_edge, delay, and informationless for process = "mixed", in that order. Normalized internally.

dt

Euler–Maruyama step size for the "ddm" generator, in seconds. Smaller steps are slower and more faithful: a discretised first-passage sampler misses boundary crossings that happen inside a step, which shifts the response times slightly late. The default of one millisecond was validated against rtdists::rdiffusion() in the package's tests, where the two agree in accuracy and in the response time quantiles.

Value

A data.frame with n rows:

  • rt is the response time in seconds,

  • response is 1 for a correct (upper-boundary / winning-accumulator) response and 0 otherwise,

  • contaminant is the logical ground truth,

  • process is "clean" or the generating contaminant process per trial.

Reproducibility

There is no set.seed() anywhere in rtprep; reproducibility belongs to the caller.

The three processes are clusters, not point definitions

Each named process is one implementation of a psychological cluster, whether premature responding, late starts or disengagement, and the arguments (anticipation_anchor, anticipation_depth, delay_min/delay_max, lapse_prop) sweep within the cluster. Which implementations a simulation uses, and how many per cluster, is a design decision for that simulation to record and defend, not a package default.

References

Ratcliff, R. (1993). Methods for dealing with reaction time outliers. Psychological Bulletin, 114(3), 510–532. doi:10.1037/0033-2909.114.3.510

Ratcliff, R., & Tuerlinckx, F. (2002). Estimating parameters of the diffusion model: Approaches to dealing with contaminant reaction times and parameter variability. Psychonomic Bulletin & Review, 9(3), 438–481. doi:10.3758/bf03196302

See also

rt_screen() for what the flags make of these trials, rt_summary() and ez_ddm() for what the contamination does to the parameters.

Examples

set.seed(1)
d <- r_contaminated(
  200,
  generator = "ddm", process = "leading_edge", rate = 0.1,
  par = list(drift = 1.5, bound = 1.2, ndt = 0.30)
)
table(d$process)
#> 
#>        clean leading_edge 
#>          187           13 
# detection scored against ground truth
scr <- rt_screen(d$rt, rule = rule_sd(2.5))
table(flagged = !scr$.keep, truth = d$contaminant)
#>        truth
#> flagged FALSE TRUE
#>   FALSE   180   13
#>   TRUE      7    0