This vignette walks through the whole pipeline once: check the token, find the study, fetch result metadata, drop the runs that do not belong in the dataset, download the result files into a local cache, pull a participant id out of every file, read the jsPsych trials into one tibble with the metadata of each run joined, and save that dataset to disk together with the study-result table and a record of what produced it. It ends with study codes.
Every chunk below talks to a JATOS server, so none of them is evaluated when the package is built. The output shown was produced by running the same calls against the package’s test fixtures, a small mocked server with two studies and six component results, so the numbers are small but real.
Credentials
jatos_set_credentials() stores the personal access token
in the credential store of your operating system and the server URL in a
configuration file; both are picked up in every later session. Where a
platform supplies them — CI, a container, a cluster job —
JATOS_HOST and JATOS_TOKEN are read instead
and take precedence. jatos_credentials_sitrep() says which
of those a profile is using. See vignette("credentials")
for the details and for why a connection object holds no token at
all.
jatos_token_info()
#> # A tibble: 1 x 9
#> token_id name username user_id created expires
#> <int> <chr> <chr> <int> <dttm> <dttm>
#> 1 7 analysis-la~ researc~ 3 2025-08-24 01:46:40 NA
#> # i 3 more variables: expired <lgl>, active <lgl>, roles <list>A wrong token gives a 401 error with the server’s message; a wrong host a 404 or, on some installations, an HTML login page, which the package reports as an authentication failure rather than trying to parse it.
Find the study and its batches
jatos_studies() lists every study the token can see.
Components and batches come along as list columns of tibbles.
studies <- jatos_studies()
studies[, c("study_id", "title", "active", "locked")]
#> # A tibble: 2 x 4
#> study_id title active locked
#> <int> <chr> <lgl> <lgl>
#> 1 12 A03_ColorBinding TRUE FALSE
#> 2 13 A01_Registration FALSE TRUE
studies$batches[[1]][, c("batch_id", "title", "active", "allowed_worker_types")]
#> # A tibble: 2 x 4
#> batch_id title active allowed_worker_types
#> <int> <chr> <lgl> <list>
#> 1 34 Default TRUE <chr [2]>
#> 2 35 Prolific wave 2 FALSE <chr [1]>jatos_study(), jatos_components(),
jatos_batches() and jatos_batch() fetch the
same information for one study or batch.
Result metadata
jatos_results_metadata() asks the server what results
exist without downloading any data. It takes any combination of study,
batch, component, study result, component result and group ids and
returns one row per component result. The message summarises what came
back.
meta <- jatos_results_metadata(study_id = 12)
#> i <https://jatos.example.org>: 4 study results, 6 component results (2 not
#> FINISHED).
meta[, c("study_result_id", "component_result_id", "batch_id", "component_state", "data_size")]
#> # A tibble: 6 x 5
#> study_result_id component_result_id batch_id component_state data_size
#> <int> <int> <int> <chr> <dbl>
#> 1 9001 7001 34 FINISHED 2048
#> 2 9002 7002 34 RELOADED 0
#> 3 9002 7003 34 FINISHED 1536
#> 4 9003 7004 34 STARTED 0
#> 5 9004 7005 36 FINISHED 300
#> 6 9004 7006 36 FINISHED 700study_id and component_id also take uuid
strings, which a study keeps when it is exported and imported on another
server while its id changes, so a script keyed by uuid survives the
move. The server takes them in separate fields, and the package sends
them there:
jatos_results_metadata(study_id = "1c2d3e4f-0000-4000-8000-000000000012")
#> i <https://jatos.example.org>: 4 study results, 6 component results (2 not
#> FINISHED).A study result is one run by one participant; a
component result is one component within that run. Every column
name carries its level: study_state and
component_state, study_start_time and
component_start_time, and so on. data_size is
the size of the result data in bytes, which the download below compares
against the local file.
Study result 9002 has two component results for the same component,
one of them a reload with no data. That is the usual reason for more
than one row per study result in a single-component study.
jatos_study_results() collapses the tibble to one row per
study result when that is the level you need:
jatos_study_results(meta)[, c("study_result_id", "batch_id", "study_state", "n_component_results", "data_size")]
#> # A tibble: 4 x 5
#> study_result_id batch_id study_state n_component_results data_size
#> <int> <int> <chr> <int> <dbl>
#> 1 9001 34 FINISHED 1 2048
#> 2 9002 34 FINISHED 2 1536
#> 3 9003 34 STARTED 1 0
#> 4 9004 36 FINISHED 2 1000With a participant key, jatos_study_results() also
counts the runs per participant as n_runs: by default the
query_prolific_pid column that
jatos_url_query() adds (below), or any column named in
participant, such as worker_id or a field
extracted from the data. That is the duplicate check to run before N is
counted.
jatos_study_results(jatos_url_query(meta))[, c("study_result_id", "query_prolific_pid", "n_runs")]
#> # A tibble: 4 x 3
#> study_result_id query_prolific_pid n_runs
#> <int> <chr> <int>
#> 1 9001 <NA> NA
#> 2 9002 <NA> NA
#> 3 9003 <NA> NA
#> 4 9004 p-0004 1The metadata tibble is a plain tibble. Filter it as you like
(meta[meta$component_state == "FINISHED", ], or
dplyr::filter()); every function that takes it only checks
that the columns it uses are still there.
The three exclusions every export makes have a function of their own,
so that the counts are reported next to the call: unfinished runs, the
researcher’s own test runs from the JATOS GUI (worker type
Jatos), and runs before the study went live.
jatos_filter_metadata() works on the study result, so every
component result of a dropped run goes with it.
finished <- jatos_filter_metadata(meta, states = "FINISHED")
#> i Excluded 1 of 4 study results; 3 remain.
#> * 1 by study state (kept "FINISHED")worker_types = c("PersonalSingle", "GeneralMultiple"),
since and until (a half-open interval of start
times, read in the time zone tz, UTC by default),
exclude_study_result_id and exclude_worker_id
(your own test runs through a real link) work the same way, each with
its own count in the message. Nothing is filtered unless you ask for it,
here and in jatos_export_results() below.
Participants who arrive through a Prolific link bring
PROLIFIC_PID, STUDY_ID and
SESSION_ID as URL query parameters, which JATOS stores with
the study result; the url_query list column holds them.
jatos_url_query() widens them into one column per
parameter, prefixed so that STUDY_ID cannot collide with
study_id:
jatos_url_query(meta)[, c("study_result_id", "query_prolific_pid", "query_session_id")]
#> # A tibble: 6 x 3
#> study_result_id query_prolific_pid query_session_id
#> <int> <chr> <chr>
#> 1 9001 <NA> <NA>
#> 2 9002 <NA> <NA>
#> 3 9002 <NA> <NA>
#> 4 9003 <NA> <NA>
#> 5 9004 p-0004 s-0004
#> 6 9004 p-0004 s-0004Download into a cache
jatos_download_results() fetches the
data.txt of every component result in the tibble that is
missing locally or has grown on the server, and stores it under
<path>/batch_<id>/study_result_<id>/comp-result_<id>/.
Results with size 0 are skipped, not requested. Before a long run,
dry_run = TRUE shows what would be fetched without a single
request:
plan <- jatos_download_results(meta, "JATOS_data", dry_run = TRUE)
#> i Batch 34: 2 component results (3.6 kB) to fetch in 1 request.
#> i Batch 36: 2 component results (1.0 kB) to fetch in 1 request.
plan[, c("component_result_id", "data_size", "status")]
#> # A tibble: 6 x 3
#> component_result_id data_size status
#> <int> <dbl> <chr>
#> 1 7001 2048 pending
#> 2 7002 0 empty
#> 3 7003 1536 pending
#> 4 7004 0 empty
#> 5 7005 300 pending
#> 6 7006 700 pendingThe real run shows a progress bar over the requests and reports per batch:
meta <- jatos_download_results(meta, "JATOS_data")
#> v Batch 34: fetched 2 of 2 component results in 1 request.
#> v Batch 36: fetched 2 of 2 component results in 1 request.
meta[, c("component_result_id", "data_size", "file_size", "status")]
#> # A tibble: 6 x 4
#> component_result_id data_size file_size status
#> <int> <dbl> <dbl> <chr>
#> 1 7001 2048 2048 fetched
#> 2 7002 0 NA empty
#> 3 7003 1536 1536 fetched
#> 4 7004 0 NA empty
#> 5 7005 300 300 fetched
#> 6 7006 700 700 fetched
list.files("JATOS_data", recursive = TRUE)
#> batch_34/metadata.json
#> batch_34/study_result_9001/comp-result_7001/data.txt
#> batch_34/study_result_9002/comp-result_7003/data.txt
#> batch_36/metadata.json
#> batch_36/study_result_9004/comp-result_7005/data.txt
#> batch_36/study_result_9004/comp-result_7006/data.txtEach batch directory also receives the server’s own
metadata.json for that batch, written after the batch’s
data, so the cache can be read back without the server later. The run
above cost four requests: two for the data (one per batch) and two for
the metadata files.
Run the same call again and nothing is requested:
meta <- jatos_download_results(meta, "JATOS_data")
table(meta$status)
#>
#> empty unchanged
#> 2 4The comparison is in bytes on both sides. When the server reports a
smaller size than the local file, the local copy stays and the
row gets status shrunk; only overwrite = TRUE
replaces it, because the local file may be the last copy of that
participant’s data. A request that fails does not stop the run: its rows
get status failed, the other chunks are fetched, and one
warning at the end says what went wrong; the next run retries them.
jatos_cache_status() summarises a cache batch by batch,
including data files that sit in a batch directory without a metadata
row:
jatos_cache_status("JATOS_data")
#> # A tibble: 2 x 8
#> batch_id path n_results n_downloaded n_pending n_shrunk n_orphans bytes
#> <int> <chr> <int> <int> <int> <int> <int> <dbl>
#> 1 34 JATOS_data~ 4 2 0 0 0 3584
#> 2 36 JATOS_data~ 2 2 0 0 0 1000The batch_<id>/ layout is the only one the package
reads or writes. A directory in another layout (a
metadata.json at its top level,
JATOS_DATA_<id> folders from smartr) is refused with
a message naming what was found. The check runs before anything in the
directory is read or written, and the message says so: the server is the
source of truth, so download into a fresh directory, or import a results
zip with jatos_import_results() — one exported from the
JATOS GUI, or one fetched with jatos_export_archive() (see
“Work offline” below).
Files participants uploaded
A study that lets participants draw or record uploads files with
jatos.uploadResultFile(); they are not part of the result
data. The metadata lists them per component result in the
files list column, and jatos_result_files()
widens that into one row per file:
jatos_result_files(meta)
#> # A tibble: 1 x 6
#> study_result_id component_result_id component_id batch_id filename size
#> <int> <int> <int> <int> <chr> <dbl>
#> 1 9002 7003 121 34 drawing.png 4096jatos_download_files() fetches them with the same rule
as the result data, applied per file, and stores each one next to the
data.txt of its component result in a files/
folder, which is where the server puts them in its own zips.
dry_run = TRUE plans without a request:
jatos_download_files(meta, "JATOS_data", dry_run = TRUE)
#> i Batch 34: 1 file (4.1 kB) to fetch in 1 request.
files <- jatos_download_files(meta, "JATOS_data")
#> v Batch 34: fetched 1 of 1 file in 1 request.
files[, c("component_result_id", "filename", "size", "file_size", "status")]
#> # A tibble: 1 x 5
#> component_result_id filename size file_size status
#> <int> <chr> <dbl> <dbl> <chr>
#> 1 7003 drawing.png 4096 4096 fetched
list.files("JATOS_data", recursive = TRUE)
#> batch_34/metadata.json
#> batch_34/study_result_9001/comp-result_7001/data.txt
#> batch_34/study_result_9002/comp-result_7003/data.txt
#> batch_34/study_result_9002/comp-result_7003/files/drawing.png
#> batch_36/metadata.json
#> batch_36/study_result_9004/comp-result_7005/data.txt
#> batch_36/study_result_9004/comp-result_7006/data.txtThe status vocabulary is the one of the data download
(fetched, unchanged, empty,
shrunk, missing, failed), and the
files of one component result always travel in one request. A second
call finds nothing to do:
table(jatos_download_files(meta, "JATOS_data")$status)
#>
#> unchanged
#> 1Pull single fields out of the files
Often one value per file is needed before the data is read in full,
for example a participant id to match against a Prolific export.
jatos_extract_fields() reads it by a key-anchored regular
expression, which is fast on thousands of files, and parses every file
where that read did not find exactly one value to compare the two
readings.
meta <- jatos_extract_fields(meta, "participant_id")
meta[, c("component_result_id", "status", "participant_id", "participant_id_status")]
#> # A tibble: 6 x 4
#> component_result_id status participant_id participant_id_status
#> <int> <chr> <chr> <chr>
#> 1 7001 unchanged P1 unique
#> 2 7002 empty <NA> <NA>
#> 3 7003 unchanged P2 unique
#> 4 7004 empty <NA> <NA>
#> 5 7005 unchanged P4 unique
#> 6 7006 unchanged P4 uniqueEvery occurrence of the key in a file is collected, at any depth. The
value is returned when all occurrences agree (unique); a
file where they disagree gets NA and status
conflict, a file without the key absent, and a
file that is not valid JSON unparseable. Every
conflict file, and every absent file as long
as the key was found somewhere, is parsed in full and compared with the
regular expression’s reading (here none, so nothing is reported); a file
with one agreeing value needs no parse, and a key found in no file is
reported absent without parsing. A field that names a
metadata column (batch_id, worker_id,
file, …) is refused, since the extracted values would
replace the server’s, and the tibble remembers which columns came from
an extraction (attribute jatosr_extracted).
jatos_study_results() carries an extracted field up to the
study-result level and flags study results whose components disagree on
it:
jatos_study_results(meta)[, c("study_result_id", "n_component_results", "participant_id")]
#> # A tibble: 4 x 3
#> study_result_id n_component_results participant_id
#> <int> <int> <chr>
#> 1 9001 1 P1
#> 2 9002 2 P2
#> 3 9003 1 <NA>
#> 4 9004 2 P4Read the trials
jatos_read_results() reads every local file of the
tibble and binds the trials. Every row gets the ids of its study result
and component result and, by default, the batch, component, worker,
worker type, study state and start time of that run from the metadata,
so the dataset needs no join later. Files that differ in their columns
are bound with NA filling. A nested object (a survey
response) is stored as a list column, one list per trial,
whatever shape jsonlite gave it in each file, and a message names such
columns; a column whose atomic type differs between files (a number
here, a string there) is an error that names the column, or, with
coerce = "character", is converted to text in every
file.
trials <- jatos_read_results(meta)
#> i 2 rows have no local file.
#> i Stored the nested object column response as list column, one list per
#> trial.
trials[, c("study_result_id", "batch_id", "worker_id", "study_state", "trial_index", "rt")]
#> # A tibble: 6 x 6
#> study_result_id batch_id worker_id study_state trial_index rt
#> <int> <int> <int> <chr> <int> <dbl>
#> 1 9001 34 501 FINISHED 0 512.
#> 2 9001 34 501 FINISHED 1 -1
#> 3 9002 34 502 FINISHED 0 1200
#> 4 9002 34 502 FINISHED 1 950
#> 5 9004 36 504 FINISHED 0 2100
#> 6 9004 36 504 FINISHED 0 8800
names(trials)
#> [1] "study_result_id" "component_result_id" "batch_id"
#> [4] "component_id" "worker_id" "worker_type"
#> [7] "study_state" "study_start_time" "trial_index"
#> [10] "trial_type" "participant_id" "city"
#> [13] "rt" "correct" "filler"
#> [16] "url" "raw" "response"
#> [19] "question_order"metadata_cols chooses the joined columns;
NULL keeps only the ids. A field pulled out with
jatos_extract_fields() is joined when you name it there.
That is meant for a value that does not sit at the trial level, for
example an age answered inside a survey response. Only the
survey component carries it, which the extractor reports as a message,
not a fault:
meta <- jatos_extract_fields(meta, "age")
#> i Cross-checked 3 files without a unique value with jsonlite: 3 parsed, 0
#> unparseable.
#> i 3 files without age (status `absent`).
jatos_read_results(meta, metadata_cols = c("batch_id", "age"))[, c("component_result_id", "batch_id", "age", "trial_index")]
#> i 2 rows have no local file.
#> i Stored the nested object column response as list column, one list per
#> trial.
#> # A tibble: 6 x 4
#> component_result_id batch_id age trial_index
#> <int> <int> <chr> <int>
#> 1 7001 34 <NA> 0
#> 2 7001 34 <NA> 1
#> 3 7003 34 <NA> 0
#> 4 7003 34 <NA> 1
#> 5 7005 36 <NA> 0
#> 6 7006 36 31 0The participant_id extracted earlier is a top-level
field of every trial already (jsPsych.data.addProperties()
puts it there), so the trials carry it without any join. Naming it in
metadata_cols is refused: a trial column is never
overwritten or renamed by the join.
jatos_read_results(meta, metadata_cols = c("batch_id", "participant_id"))
#> i 2 rows have no local file.
#> Error in `jatos_read_results()`:
#> ! The trials of 'JATOS_data/batch_34/study_result_9001/comp-result_7001/data.txt'
#> already carry the column participant_id, which the metadata would overwrite.
#> i Leave it out of `metadata_cols` (or `id_cols`), or rename the metadata
#> column before reading.Three arguments cover studies that are not one jsPsych array per
file. reader takes a function of one file path that returns
a data frame, for example
function(file) utils::read.csv(file) for PsychoJS; the join
and the binding stay the same. on_error = "skip" leaves
files the reader cannot read out of the result and warns once with their
paths. And split = "component" returns one tibble per
component instead of one table, for studies whose components write
different columns:
names(jatos_read_results(meta, split = "component"))
#> i 2 rows have no local file.
#> i Stored the nested object column response as list column, one list per
#> trial.
#> [1] "121" "131" "132"jatos_read_json() reads a single file. A file that holds
several JSON values back to back, which repeated
jatos.appendResultData() calls leave behind (arrays after
arrays, or one object per call), is split into its top-level values and
read as one array.
jatos_read_json(meta$file[1])
#> # A tibble: 2 x 7
#> trial_index trial_type participant_id city rt correct filler
#> <int> <chr> <chr> <chr> <dbl> <lgl> <chr>
#> 1 0 html-keyboard-response P1 Zürich 512. TRUE <NA>
#> 2 1 html-keyboard-response P1 Zürich -1 FALSE xxxxxx~Save the dataset
jatos_write_results() writes the tibble as
.rds, .csv, .csv.gz,
.tsv, .parquet (with the arrow package
installed) or .RData (.rda), taking the format
from the file extension. rds and RData keep
list columns as they are. A csv cannot hold a nested object, so list
columns and nested data-frame columns are serialised cell by cell to
JSON strings, times are written as ISO 8601 in UTC, and NA
as an empty field. parquet keeps a list column when arrow
can give it one type and serialises the others (an object in one trial,
a vector in the next) to JSON strings like csv; a message names
them.
jatos_write_results(trials, "data/study12.rds")
jatos_write_results(trials, "data/study12.csv")
#> i Serialised the list columns response and question_order to JSON strings.
jatos_write_results(trials, "data/study12.RData", object = "study12")
load("data/study12.RData")
#> [1] "study12"The file is written under a temporary name and renamed into place, and an existing file is never replaced unless you say so:
jatos_write_results(trials, "data/study12.rds")
#> Error in `jatos_write_results()`:
#> ! 'data/study12.rds' already exists.
#> i Set `overwrite = TRUE` to replace it.jatos_export_results() runs everything above in one
call: metadata for the ids given, the filters, download into the cache,
the URL query columns, the optional field extraction, the read with the
join, and the write. Run it again after data collection has moved on and
only new or grown results are fetched; the files are then rewritten from
the whole cache.
jatos_export_results(
study_id = 12, cache = "JATOS_data", file = "data/study12.rds",
states = "FINISHED", overwrite = TRUE
)
#> i <https://jatos.example.org>: 4 study results, 6 component results (2 not
#> FINISHED).
#> i Excluded 1 of 4 study results; 3 remain.
#> * 1 by study state (kept "FINISHED")
#> i Reading 4 files.
#> i 1 row has no local file.
#> i Stored the nested object column response as list column, one list per
#> trial.
#> i Writing 6 trials to 'data/study12.rds'.
#> v Wrote 6 trials from 4 component results to 'data/study12.rds'.
#> i Also wrote 'data/study12_metadata.rds' and 'data/study12_export.json'.On a cache that is already current, as here, that costs one request,
the metadata. The trials tibble is returned invisibly; its rows carry
the query_* columns next to the default metadata columns.
until, exclude_study_result_id,
exclude_worker_id and tz pass through to the
filter, reader and metadata_cols to the
reader.
Three files come out of the call. study12.rds holds the
trials. study12_metadata.rds holds one row per study
result, the table from jatos_study_results() with the URL
query columns, which is where exclusions, payments and the participant
count of a methods section come from:
readRDS("data/study12_metadata.rds")[, c("study_result_id", "study_state", "n_component_results", "query_prolific_pid")]
#> # A tibble: 3 x 4
#> study_result_id study_state n_component_results query_prolific_pid
#> <int> <chr> <int> <chr>
#> 1 9001 FINISHED 1 <NA>
#> 2 9002 FINISHED 2 <NA>
#> 3 9004 FINISHED 2 p-0004study12_export.json records what produced the other two,
readable without R:
cat(readLines("data/study12_export.json"), sep = "\n")
#> {
#> "package": "jatosr",
#> "version": "0.1.0",
#> "exported_at": "2026-09-06T11:56:51Z",
#> "host": "https://jatos.example.org",
#> "profile": "default",
#> "study_id": 12,
#> "filters": {
#> "states": "FINISHED"
#> },
#> "cache": "/home/researcher/study12/JATOS_data",
#> "counts": {
#> "on_server": {
#> "study_results": 4,
#> "component_results": 6
#> },
#> "exported": {
#> "study_results": 3,
#> "component_results": 5,
#> "files_read": 4,
#> "trials": 6
#> }
#> },
#> "files": {
#> "trials": "data/study12.rds",
#> "metadata": "data/study12_metadata.rds"
#> }
#> }metadata_file = FALSE and
provenance = FALSE switch the two sidecars off;
split = "component" writes one trials file per component.
Every target is checked before the first request, so a missing directory
or a file in the way costs no download.
The study archive
The dataset says what participants did; the study archive says what
they saw. jatos_export_study() downloads the archive JATOS
itself exports (GET /studies/{id}): a zip with the study’s
properties and components as JSON next to the assets folder that holds
the experiment’s HTML, scripts and stimuli. JATOS imports it as a
.jzip file, so an archive saved next to the data reproduces
the exact experiment that produced them.
jatos_export_study(12, "data/study12_study.jzip")
#> v Wrote the archive of study 12 to 'data/study12_study.jzip' (479 B).
utils::unzip("data/study12_study.jzip", list = TRUE)[, c("Name", "Length")]
#> Name Length
#> 1 A03_ColorBinding.jas 293
#> 2 A03_ColorBinding/index.html 47archive_study = TRUE on
jatos_export_results() does the same as part of the export,
as <stem>_study.jzip next to the other files, and the
provenance record lists it:
jatos_export_results(
study_id = 12, cache = "JATOS_data", file = "data/study12.rds",
states = "FINISHED", archive_study = TRUE, overwrite = TRUE
)
#> i <https://jatos.example.org>: 4 study results, 6 component results (2 not
#> FINISHED).
#> i Excluded 1 of 4 study results; 3 remain.
#> * 1 by study state (kept "FINISHED")
#> i Reading 4 files.
#> i 1 row has no local file.
#> i Stored the nested object column response as list column, one list per
#> trial.
#> v Wrote the archive of study 12 to 'data/study12_study.jzip' (479 B).
#> i Writing 6 trials to 'data/study12.rds'.
#> v Wrote 6 trials from 4 component results to 'data/study12.rds'.
#> i Also wrote 'data/study12_metadata.rds', 'data/study12_export.json', and
#> 'data/study12_study.jzip'.
jsonlite::read_json("data/study12_export.json")$files
#> $trials
#> [1] "data/study12.rds"
#>
#> $metadata
#> [1] "data/study12_metadata.rds"
#>
#> $study_archive
#> [1] "data/study12_study.jzip"When several studies are exported, or when only batch_id
is given and the studies come from the metadata, each archive is named
by its id, <stem>_study_<id>.jzip. The endpoint
needs the user role; a token with the viewer role only can read results
but not the archive.
The results archive
For a data deposit the artefact people trust is the zip JATOS itself
produces, the one the GUI’s “Export Results” downloads: every
data.txt, the uploaded files and a
metadata.json in one archive.
jatos_export_archive() fetches it
(POST /results) for the studies or batches given and writes
it untouched; archive_results = TRUE on
jatos_export_results() does the same next to the dataset,
as <stem>_results.zip, and records its size and md5
in the provenance file. This is an archival copy, fetched in full each
time; the incremental download into the cache is a different thing.
jatos_export_archive(study_id = 12, file = "data/study12_results.zip")
#> v Wrote the results archive to 'data/study12_results.zip' (2.5 kB).
jatos_export_results(
study_id = 12, cache = "JATOS_data", file = "data/study12.rds",
states = "FINISHED", archive_results = TRUE, overwrite = TRUE
)
#> i <https://jatos.example.org>: 4 study results, 6 component results (2 not
#> FINISHED).
#> i Excluded 1 of 4 study results; 3 remain.
#> * 1 by study state (kept "FINISHED")
#> i Reading 4 files.
#> i 1 row has no local file.
#> i Stored the nested object column response as list column, one list per
#> trial.
#> v Wrote the results archive to 'data/study12_results.zip' (2.5 kB).
#> i Writing 6 trials to 'data/study12.rds'.
#> v Wrote 6 trials from 4 component results to 'data/study12.rds'.
#> i Also wrote 'data/study12_metadata.rds', 'data/study12_export.json', and
#> 'data/study12_results.zip'.
jsonlite::read_json("data/study12_export.json")$files$results_archive
#> $file
#> [1] "data/study12_results.zip"
#>
#> $bytes
#> [1] 2487
#>
#> $md5
#> [1] "d0efe5d9a324dbbc16402b6a17380ac1"Raw files per participant
Labs that share raw JSON per participant want one file per person
rather than the cache tree. jatos_write_raw() copies each
data.txt byte for byte to
<path>/<name>.json, named by a metadata column
or an extracted field; a study result with several files gets the
component result id appended, and a name shared by two study results is
an error that lists them.
jatos_write_raw(meta, "raw", name_by = "participant_id")
#> i 2 rows have no local file.
#> v Wrote 4 files to 'raw'.
list.files("raw")
#> [1] "P1.json" "P2.json" "P4_7005.json" "P4_7006.json"Work offline
A cache built by jatos_download_results() is
self-describing. jatos_read_metadata() rebuilds the
metadata tibble from the cached metadata.json files and
attaches the local file paths and sizes, without any request, so an
analysis script can start from the cache alone.
offline <- jatos_read_metadata("JATOS_data")
offline[, c("component_result_id", "batch_id", "data_size", "file_size")]
#> # A tibble: 6 x 4
#> component_result_id batch_id data_size file_size
#> <int> <int> <dbl> <dbl>
#> 1 7001 34 2048 2048
#> 2 7002 34 0 NA
#> 3 7003 34 1536 1536
#> 4 7004 34 0 NA
#> 5 7005 36 300 300
#> 6 7006 36 700 700The sizes in this tibble are those of the cached
metadata.json, so passing it to
jatos_download_results() finds nothing new; fetch fresh
metadata from the server when you want to look for new results.
jatos_export_results(download = FALSE) runs the same
pipeline from the cache: the metadata comes from the cached files, the
filters, the join and the write are the same, no request is made and no
credentials are needed. That is the call for an analysis machine without
a token, for continuous integration, or for rebuilding the dataset after
the study has left the server. The provenance record then says
offline: true and lists the metadata files it was built
from.
Without ids the whole cache is exported (online at least one id is
required, since the server asks for one); with study_id or
batch_id the cached metadata is filtered exactly by
those.
jatos_export_results(
cache = "JATOS_data", file = "data/study12.rds",
states = "FINISHED", download = FALSE, overwrite = TRUE
)
#> i 'JATOS_data': 4 study results, 6 component results in the cached metadata; no
#> request made.
#> i Excluded 1 of 4 study results; 3 remain.
#> * 1 by study state (kept "FINISHED")
#> i Reading 4 files.
#> i 1 row has no local file.
#> i Stored the nested object column response as list column, one list per
#> trial.
#> i Writing 6 trials to 'data/study12.rds'.
#> v Wrote 6 trials from 4 component results to 'data/study12.rds'.
#> i Also wrote 'data/study12_metadata.rds' and 'data/study12_export.json'.Data that did not come through the package, a zip exported from the
JATOS GUI (“Export Results”) or by jatos_export_archive(),
enters a cache with jatos_import_results(), which unpacks
it into the same batch_<id>/ layout with one
metadata.json per batch; the reader, the status report and
later incremental downloads then treat it like any other cache.
jatos_import_results("data/study12_results.zip", "JATOS_import")
#> v Imported batches 34 and 36 into 'JATOS_import': 6 component results listed, 4
#> data files and 1 uploaded file written.
jatos_cache_status("JATOS_import")
#> # A tibble: 2 x 8
#> batch_id path n_results n_downloaded n_pending n_shrunk n_orphans bytes
#> <int> <chr> <int> <int> <int> <int> <int> <dbl>
#> 1 34 JATOS_impo~ 4 2 0 0 0 3584
#> 2 36 JATOS_impo~ 2 2 0 0 0 1000Study codes
A study code is the last part of a study link
(<host>/publix/<code>).
jatos_create_study_codes() generates codes for a batch; for
the personal worker types the server creates n new codes,
for the general types it returns the batch’s single code.
codes <- jatos_create_study_codes(12, batch_id = 34, n = 3, type = "PersonalMultiple", comment = "wave 2")
#> v Received 3 PersonalMultiple study codes for study 12.
codes
#> # A tibble: 3 x 6
#> study_code study_id batch_id type comment study_link
#> <chr> <int> <int> <chr> <chr> <chr>
#> 1 code0001 12 34 PersonalMultiple wave 2 https://jatos.example.o~
#> 2 code0002 12 34 PersonalMultiple wave 2 https://jatos.example.o~
#> 3 code0003 12 34 PersonalMultiple wave 2 https://jatos.example.o~
codes$study_link
#> [1] "https://jatos.example.org/publix/code0001"
#> [2] "https://jatos.example.org/publix/code0002"
#> [3] "https://jatos.example.org/publix/code0003"The server answers with the codes only.
jatos_study_code() fetches the full properties of a code,
including whether it is active, and
jatos_deactivate_study_code() closes codes that should
admit nobody anymore:
jatos_study_code("code0001")
#> # A tibble: 1 x 7
#> study_code batch_id type comment active study_link study_entry_link
#> <chr> <int> <chr> <chr> <lgl> <chr> <chr>
#> 1 code0001 34 PersonalSingle pilot 1 TRUE https://ja~ https://jatos.e~
jatos_deactivate_study_code(codes$study_code)[, c("study_code", "type", "active")]
#> # A tibble: 3 x 3
#> study_code type active
#> <chr> <chr> <lgl>
#> 1 code0001 PersonalMultiple FALSE
#> 2 code0002 PersonalMultiple FALSE
#> 3 code0003 PersonalMultiple FALSEjatos_study_links() builds the run URLs from codes
without a request, for example for codes copied out of the JATOS GUI. It
needs the host, not the token, so this chunk does run:
jatosr::jatos_study_links(c("8kw0pFV5M1e", "Qm3xYtb9Lc2"), host = "https://jatos.example.org")
#> [1] "https://jatos.example.org/publix/8kw0pFV5M1e"
#> [2] "https://jatos.example.org/publix/Qm3xYtb9Lc2"