
Fetch, read and save the trials of a study in one call
Source:R/export-results.R
jatos_export_results.RdRuns the whole pipeline: jatos_results_metadata() for the ids given,
jatos_filter_metadata() with states, worker_types, since,
until and the two exclusion lists,
jatos_download_results() into cache, jatos_url_query(),
jatos_extract_fields() when fields is given, jatos_read_results()
with the metadata joined, and jatos_write_results() to file. Nothing
happens here that those functions do not do on their own; run them
separately when a step needs an argument this wrapper does not pass on
(chunk_size of the download, flatten of the reader). reader,
split, coerce and on_error go to the reader as they are; with
on_error = "skip" a file the reader cannot read is left out of the
trials with a warning, and the final message and the provenance count
only the files that were read. Files that
participants uploaded are not part of a trials dataset; fetch them with
jatos_download_files().
Usage
jatos_export_results(
study_id = NULL,
batch_id = NULL,
...,
cache,
file,
download = TRUE,
fields = NULL,
states = NULL,
worker_types = NULL,
since = NULL,
until = NULL,
exclude_study_result_id = NULL,
exclude_worker_id = NULL,
tz = "UTC",
reader = NULL,
coerce = c("error", "character"),
on_error = c("abort", "skip"),
metadata_cols = default_metadata_cols(),
format = NULL,
split = c("none", "component"),
metadata_file = TRUE,
provenance = TRUE,
archive_study = FALSE,
archive_results = FALSE,
overwrite = FALSE,
conn = jatos_connection()
)Arguments
- study_id, batch_id
One or more ids to select results by; at least one of the two must be given for a download. Passed to
jatos_results_metadata();study_idmay also be one or more uuid strings. Withdownload = FALSEboth may beNULL, which exports the whole cache.- ...
Must be empty.
- cache
The local cache directory, see
jatos_download_results().- file
Path of the trials file to write, see
jatos_write_results(). Its directory must exist.- download
If
FALSE, nothing is requested: the metadata is read from the cache and the dataset is built from the files already there.- fields
JSON keys to extract with
jatos_extract_fields()and join to every trial row. For values that do not sit at the top level of the trial objects, for example inside a survey response; a field that already is a trial column is reported as a clash byjatos_read_results().- states
Study states to keep, for example
"FINISHED".NULLkeeps every state. JATOS states arePRE,STARTED,DATA_RETRIEVED,FINISHED,ABORTEDandFAIL.- worker_types
Worker types to keep, for example
c("PersonalSingle", "GeneralMultiple").NULLkeeps every type; the GUI test runs are"Jatos".- since
Keep study results that started at or after this time: a
POSIXct, aDate, or a string such as"2025-08-24"or"2025-08-24 12:00:00", read in the time zonetz. Rows without a start time are dropped whensinceis given.- until
Keep study results that started before this time, given like
since;sinceanduntiltogether select a half-open interval, sosince = "2025-08-24", until = "2025-08-25"is one day. Rows without a start time are dropped whenuntilis given.- exclude_study_result_id, exclude_worker_id
Study result ids, or worker ids, whose study results are dropped; for the researcher's own runs through a real link, which
worker_typescannot tell apart.- tz
Time zone in which date strings in
sinceanduntilare read, and in which the message shows them."UTC"by default, the zone the server reports times in;Sys.timezone()for local time.- reader
NULLforjatos_read_json(), or a function of one file path returning a data frame with one row per trial.- coerce
What to do with a column whose atomic type differs between files:
"error"stops with the column named;"character"converts that column to character in every file and reports it.- on_error
"abort"stops at the first file the reader cannot read."skip"leaves such files out, reads the rest, and warns once with the files and the first error; a clash between trial and metadata columns still aborts.- metadata_cols
Metadata columns joined to every trial row after the ids, see
jatos_read_results(); thequery_*columns and the extractedfieldsare added to whatever is given here.NULLjoins only those.- format
Format of the trials file, passed to
jatos_write_results();NULLinfers it from the extension offile. The default metadata file shares it; a metadata file given as a path takes its format from its own extension.- split
"none"writes one trials file."component"reads and writes one file per component id (<stem>_component_<id>.<ext>), seejatos_read_results(); the metadata file is written once.- metadata_file
TRUEwrites the study-result table next tofileas<stem>_metadata.<ext>; a path writes it there, in the format of that path's extension (a.csvnext to an.rdstrials file is fine);FALSEskips it.- provenance
If
TRUE, write<stem>_export.jsonwith the host, profile, ids, filters, fields, cache, package version, time, counts and the paths written.- archive_study
If
TRUE, the study archive is downloaded withjatos_export_study()and written next to the dataset: as<stem>_study.jzipfor the one study named instudy_id, as<stem>_study_<id>.jzipper study when several are named or when onlybatch_idis given and the studies come from the metadata. The archive records the experiment that produced the data; the provenance record lists it.- archive_results
If
TRUE, the results archive that the server builds for the same ids (POST /results: everydata.txt, the uploaded files and themetadata.json, seejatos_export_archive()) is written next to the dataset as<stem>_results.zip, and the provenance record lists it with its size and md5 checksum. An archival copy, fetched in full each time; the incremental download intocacheis unchanged.- overwrite
Passed to
jatos_write_results()for every file written. The download keeps its own guard for local files larger than the server's copy.- conn
A
jatos_connection(); not used withdownload = FALSE.
Value
The trials tibble, invisibly (with split = "component" a named
list of tibbles); a message reports the row count and the paths
written.
Details
With download = FALSE the same dataset is built from the cache alone:
the metadata comes from the cache's metadata.json files through
jatos_read_metadata() (restricted to study_id and batch_id when
given, the whole cache otherwise), no request is made and no connection
is needed, and the provenance record says offline: true and lists the
metadata files with their modification times. That is the path for an
analysis machine without credentials, for CI, or for rebuilding a
dataset after the study left the server; rows whose data file is not in
the cache are skipped with a message, as jatos_read_results() does.
archive_study needs the server and is refused offline.
Three files come out of one call, one more each with archive_study = TRUE and archive_results = TRUE. The trials file holds one row per
trial with the ids, batch_id, component_id, worker_id,
worker_type, study_state, study_start_time, the URL query
parameters as query_* columns and the extracted fields in front of
the trial columns. The metadata file (<stem>_metadata.<ext> by default)
holds one row per study result from jatos_study_results(),
which is where exclusions, payments and the participant count of a
methods section come from; it covers the same study results as the
trials file. The provenance file (<stem>_export.json) records host,
ids, filters, package version, time and the counts of the export.
Run it again later and only new or grown results are downloaded; the
files are then rewritten from the whole cache, which is why overwrite
exists. file, format, the target directory, the sidecar paths and
fields are checked before the first request, so a mistake there costs
no download; every path the call will write (with split = "component"
the part files, whose names come from the component ids of the
metadata) is checked once more right after the metadata answer, before
the download. A component result whose download failed, or that the
server's answer did not contain, stops the export before anything is
written, since the dataset would be incomplete; the files fetched so far
stay in the cache for the next run.
See also
jatos_filter_metadata(), jatos_url_query(),
jatos_study_results() for the pieces this adds to the five steps.
Examples
if (FALSE) { # \dontrun{
# needs a JATOS server and a stored API token
trials <- jatos_export_results(
study_id = 12, cache = "JATOS_data", file = "data/study12.rds",
states = "FINISHED"
)
# one batch, as csv, without the researcher's own GUI runs, with a field
# that sits inside a survey response
jatos_export_results(
batch_id = 34, cache = "JATOS_data", file = "data/batch34.csv",
worker_types = c("PersonalSingle", "GeneralMultiple"),
fields = "age", overwrite = TRUE
)
# the same dataset from the cache alone, without credentials
jatos_export_results(
study_id = 12, cache = "JATOS_data", file = "data/study12.rds",
states = "FINISHED", download = FALSE, overwrite = TRUE
)
} # }