Draws trials from an evidence accumulation core and replaces a known subset with contaminants generated by one of three psychologically motivated processes. Which trials were replaced, and by which process, comes back with the data. That is what a real data set cannot give you, and what scoring a screen requires: sensitivity and specificity need labels.
Usage
r_contaminated(
n,
generator = c("ddm", "rdm"),
process = c("leading_edge", "delay", "informationless", "mixed"),
rate = 0.05,
par = list(),
st0 = 0,
anticipation_anchor = c("min", "q10"),
anticipation_depth = 0.3,
delay_min = 0,
delay_max = 2,
lapse_prop = 0,
mix = c(leading_edge = 1, delay = 1, informationless = 1)/3,
dt = 0.001
)Arguments
- n
Number of trials.
- generator
The clean core.
"ddm"is a first-principles diffusion (Euler–Maruyama);"rdm"is a racing diffusion, two independent single-boundary Wald accumulators with a shared bound, drawn exactly from the inverse Gaussian.- process
Which contaminant process replaces the selected trials:
"leading_edge"generates anticipations. Uniform in a band around the observed minimum of the clean trials in this call (seeanticipation_anchor), responding at chance. Anchoring to an observed quantity rather than to the generating non-decision time keeps the construct comparable across generators, whose fitted non-decision times can differ substantially at identical observables."delay"generates delayed start-ups. A genuine trial plus a uniform shift ofdelay_mintodelay_maxseconds; the response, and so the intact accuracy, is kept."informationless"generates lapses of evidence quality. The trial is redrawn from the core with its evidence scaled bylapse_propwhile the speed of processing is held fixed:lapse_prop = 0is a complete lapse (chance accuracy, response times still in the core's range),lapse_prop = 1the clean process, and values between are partial lapses on the same continuum."mixed"draws each contaminant trial's process frommix.
- rate
Probability that a trial is a contaminant, in
[0, 1).- par
Generator parameters as a list. For
"ddm":drift,bound,ndt,zr(defaults1.5, 1.2, 0.30, 0.5). For"rdm":driftasc(correct, error),bound,ndt(defaultsc(3, 1.5), 1.5, 0.30). The defaults are illustrative, not calibrated; simulation designs should set parameters from an observable target instead.- st0
Range of across-trial variability in non-decision time, in seconds. The per-trial non-decision time is uniform on
ndt ± st0 / 2, the centred parameterization, so the mean response time is invariant and only the leading edge smears.- anticipation_anchor
Observable the anticipation band is centred on: the minimum (
"min") or the 10th percentile ("q10") of the clean trials generated in this call.- anticipation_depth
Half-width of the anticipation band as a proportion of the anchor: response times are uniform on
anchor * (1 ± anticipation_depth).- delay_min, delay_max
Bounds of the uniform shift, in seconds, for
process = "delay".- lapse_prop
Evidence-quality proportion for
process = "informationless"; see above.- mix
Relative weights of
leading_edge,delay, andinformationlessforprocess = "mixed", in that order. Normalized internally.- dt
Euler–Maruyama step size for the
"ddm"generator, in seconds. Smaller steps are slower and more faithful: a discretised first-passage sampler misses boundary crossings that happen inside a step, which shifts the response times slightly late. The default of one millisecond was validated againstrtdists::rdiffusion()in the package's tests, where the two agree in accuracy and in the response time quantiles.
Value
A data.frame with n rows:
rtis the response time in seconds,responseis1for a correct (upper-boundary / winning-accumulator) response and0otherwise,contaminantis the logical ground truth,processis"clean"or the generating contaminant process per trial.
Reproducibility
There is no set.seed() anywhere in rtprep; reproducibility belongs to
the caller.
The three processes are clusters, not point definitions
Each named process is one implementation of a psychological cluster, whether
premature responding, late starts or disengagement, and the arguments
(anticipation_anchor, anticipation_depth, delay_min/delay_max,
lapse_prop) sweep within the cluster. Which implementations a simulation
uses, and how many per cluster, is a design decision for that simulation to
record and defend, not a package default.
References
Ratcliff, R. (1993). Methods for dealing with reaction time outliers. Psychological Bulletin, 114(3), 510–532. doi:10.1037/0033-2909.114.3.510
Ratcliff, R., & Tuerlinckx, F. (2002). Estimating parameters of the diffusion model: Approaches to dealing with contaminant reaction times and parameter variability. Psychonomic Bulletin & Review, 9(3), 438–481. doi:10.3758/bf03196302
See also
rt_screen() for what the flags make of these trials,
rt_summary() and ez_ddm() for what the contamination does to the
parameters.
Examples
set.seed(1)
d <- r_contaminated(
200,
generator = "ddm", process = "leading_edge", rate = 0.1,
par = list(drift = 1.5, bound = 1.2, ndt = 0.30)
)
table(d$process)
#>
#> clean leading_edge
#> 187 13
# detection scored against ground truth
scr <- rt_screen(d$rt, rule = rule_sd(2.5))
table(flagged = !scr$.keep, truth = d$contaminant)
#> truth
#> flagged FALSE TRUE
#> FALSE 180 13
#> TRUE 7 0
