The problem with the handoff
You have an R analysis that runs on your laptop. Maybe it takes a while. Maybe you need to run it many times, once per species, once per county, once per simulation parameter, once per experimental condition. Maybe both.
CHTC’s high-throughput computing infrastructure can run many independent jobs across a large pool of compute resources. The barrier is rarely the value of the computing. The barrier is the handoff: turning a local analysis into something a scheduler can run somewhere else.
That handoff requires several pieces to line up at once. Your R code needs to run without relying on the interactive session where you developed it. Your software environment needs to be portable. Your files need to move to a submit node. HTCondor needs a submit file. The execute node needs a shell script. Your results need to come back.
submitr is designed to make that handoff easier. It generates the HTCondor submit file, generates the executable script, wraps the SSH and SCP commands that move files to and from the submit node, submits the job, checks status, and downloads results, all from R.
If you are new to CHTC, submitr gives you a guided path to your first successful submission. If you already use CHTC, submitr reduces repetitive setup work and makes common submission patterns easier to reproduce, review, and share.
When to use submitr
Use submitr when you are:
- sending a containerized R analysis to CHTC for the first time;
- teaching researchers the structure of an HTCondor job;
- moving from a single local analysis to many independent HTC jobs;
- standardizing a submit-file and executable-script pattern across projects;
- reducing repeated SSH, SCP, and
condor_submitcommand-line work; - making CHTC submissions easier to review, rerun, and share.
submitr is useful on its own if your project is already organized and containerized. It also fits into a broader workflow for moving from a literate analysis document to a portable, scalable computation.
The toolero family
submitr is the third step in the From the Notebook to the Cluster package family:
toolero organize, scaffold, split
└─ containr freeze the software environment in a container
└─ submitr send the analysis to CHTC and retrieve results
Each package is useful on its own. Together, they form a path from a local R project to a completed high-throughput computing run.
-
toolerohelps you start with a maintainable project structure, use Quarto as a source of truth, and split data into job-sized pieces. -
containrhelps you build a container image from yourrenv.lockso the software environment can travel with the analysis. -
submitrhelps you send the containerized analysis to CHTC, monitor the job, and bring results back.
You can adopt these packages one at a time. submitr does not require toolero, and toolero does not require submitr. The family exists so that each step prepares cleanly for the next when your project is ready to scale.
The folder names, path conventions, and shared vocabulary used consistently across all three packages are collected in one place, CONVENTIONS.md, maintained in the toolero repository, since that is where those conventions are authored. Two of them show up throughout this README: analysis outputs go in output/, and a derived analysis script lives at R/analysis.R. Following them is what lets the same script run unchanged on your laptop and on an execute node.
Before you start
submitr assumes your project is already organized and containerized. Before using it, confirm that:
- your R script runs with
Rscript R/analysis.Routside RStudio; - your container image is pushed to a registry CHTC can access;
- you have SSH access to a CHTC submit node such as
ap2002.chtc.wisc.edu.
Set up SSH connection reuse before anything else. Every submitr function that touches CHTC opens an SSH connection, which can trigger a Duo MFA prompt. Setting up ControlMaster caches your authenticated session and makes the entire workflow significantly smoother. The setup takes two minutes and is worth doing before your first htc_config() call. Full instructions appear after Step 1 below.
Installation
Install the development version from GitHub:
# install.packages("pak")
pak::pak("erwinlares/submitr")A first workflow
library(submitr)
# 1. Start the session (reads htc.cfg, stores config for all calls)
htc_start()
# 2. Generate the submit file
# r_script names the results tarball and, via input_files, tells
# HTCondor to transfer the analysis script itself -- it is not baked
# into the container image
htc_gen_submit(
output_file = "analysis.sub",
container_image = "registry.doit.wisc.edu/your.netid/my-analysis:1.0.0",
executable = "analysis.sh",
r_script = "R/analysis.R",
input_files = "R/analysis.R",
resources = "small",
comments = TRUE
)
# 3. Generate the executable script
htc_gen_executable(
r_script = "R/analysis.R",
output_file = "analysis.sh",
comments = TRUE
)
# 4. Upload files to the submit node
# With no files argument, submitr sends what the submission state recorded
htc_upload()
# 5. Submit the job
cluster_id <- htc_submit(submit_file = "analysis.sub")
# 6. Check progress
htc_status(cluster_id = cluster_id, watch = TRUE)
# 7. Download results
htc_download()Steps 4 and 7 take no arguments because submitr keeps track of what it has generated. The section on the submission state below explains how, and why it still works if you close R between submitting and collecting.
Core workflow functions
htc_start()
htc_start() reads your project’s htc.cfg and stores the connection config for the rest of the R session. All subsequent htc_*() calls use it automatically, with no need to pass config = cfg on every call.
htc_start()
#> v Session started: "your.netid"@"ap2002.chtc.wisc.edu"If this is your first time, htc_start() prompts for your NetID and submit node, writes htc.cfg, and displays ControlMaster setup instructions. On subsequent calls it reads the existing config and validates the connection.
You can still pass config explicitly to any function to override the session config:
other_cfg <- htc_config(path = "other-project/")
htc_upload(files = "job.sub", config = other_cfg)
htc_config()
htc_config() is the lower-level function that reads or creates htc.cfg. Most researchers should use htc_start() instead, which calls htc_config() and stores the result for the session. Use htc_config() directly when you need to manage multiple configs or pass a config to a single call without starting a session.
cfg <- htc_config()
#> Reading HTC config from ./htc.cfg
#> v Connected to "ap2002.chtc.wisc.edu" as "your.netid".That second line comes from a short SSH connection made to tell you whether the server is reachable before you rely on the config, which is useful at a prompt and pointless in a script. Turn it off with check_server = FALSE, or for a whole session with the submitr.check_server option. A companion option, submitr.verbose, silences the progress messages while leaving warnings and errors intact. Both are documented under ?htc_config.
Setting up SSH connection reuse
Before continuing, take two minutes to set up ControlMaster. The quickest way is htc_ssh_setup(), which writes the block below and creates the directory it references without leaving R:
htc_ssh_setup()
#> v Added a ControlMaster block for "*.chtc.wisc.edu" to "~/.ssh/config"That writes the same block you would otherwise add to ~/.ssh/config by hand:
and creates the directory used by ControlPath:
htc_ssh_setup() leaves the file alone if a matching Host block is already there, so it is safe to call again later, and dry_run = TRUE previews the change first if you would rather see it before it is written.
With ControlMaster in place, all subsequent SSH connections, whether uploads, submits, status checks or downloads, reuse the same authenticated session without prompting for Duo MFA. Full documentation is at https://chtc.cs.wisc.edu/uw-research-computing/configure-ssh.
htc_gen_submit()
Generates the HTCondor .sub submit file. It tells HTCondor which container to use, which executable to run, which files to transfer, what resources to request, and what output files to expect.
htc_gen_submit(
output_file = "analysis.sub",
container_image = "docker://registry.doit.wisc.edu/your.netid/my-analysis:1.0.0",
executable = "analysis.sh",
r_script = "R/analysis.R",
input_files = c("R/analysis.R", "data.csv"),
resources = "small",
comments = TRUE
)Use comments = TRUE on a first submission. The generated file includes explanations of each section, making it useful both as a working submit file and as a learning document.
htc_gen_submit() itself only uses r_script to name the results tarball, so that the name matches the one htc_gen_executable() tells the job to build; pass output_files yourself if you want a different name. But your analysis script is not baked into the container image, so it still has to reach the execute node somehow – list it in input_files, as above, so htc_upload() sends it and HTCondor transfers it alongside the executable. Give input_files the paths as they are on your machine (R/analysis.R); htc_upload() needs those to find the files, and the submit file lists them by basename (transfer_input_files = analysis.R), since htc_upload() sends everything flat into one directory on the submit node. For the same reason, two input files that share a name (R/utils.R and scripts/utils.R) are an error. htc_gen_submit() also warns if r_script is supplied but missing from input_files.
executable does not have to be typed here if you are about to call htc_gen_executable() next (or already have): the two generators share the submission state, so whichever one runs second picks up the executable script’s name from whichever one ran first. Passing executable explicitly to both still works as before, and if the two ever disagree, both generators warn rather than silently picking one – your explicit value is still used, so the warning is a nudge to check, not a blocker.
If your project has a project config, _toolero.yml (written by toolero::init_project()), pass it through htc_config(project_config = ) and on to config here, and queue_from in multiple-job mode can be left out entirely – it defaults to the job manifest, manifest.csv, inside config$project$conventions$split_dir, since toolero::write_by_group() always uses that filename:
cfg <- htc_config(project_config = "_toolero.yml")
htc_gen_submit(
mode = "multiple",
config = cfg,
r_script = "R/analysis.R",
input_files = "R/analysis.R"
)Resource presets:
| preset | cpus | memory | disk | when to use |
|---|---|---|---|---|
| small | 1 | 4 GB | 4 GB | first test jobs, lightweight scripts, quick summaries |
| medium | 4 | 16 GB | 15 GB | moderate analyses, multiple input files, model fitting |
| large | 8 | 64 GB | 32 GB | memory-intensive work, large datasets, parallel computation |
Start with "small" for a first test regardless of what your eventual job will need. The HTCondor log file reports actual resource usage after each run, which is the best guide for tuning future submissions. Requesting too little causes jobs to fail; requesting much more than you need makes jobs harder to match with available resources. The log is the ground truth.
htc_gen_executable()
Generates the .sh script that HTCondor runs inside the container. The generated script changes to HTCondor’s scratch directory, creates output/, runs your R script with Rscript, and archives output/ as a .tar.gz for transfer back to the submit node.
htc_gen_executable(
r_script = "R/analysis.R",
output_file = "analysis.sh",
comments = TRUE
)Only output/ itself is created. If your analysis writes to output/figures/, the R script has to create that subfolder, which toolero::save_output() does and a bare ggsave() does not.
If the R script fails, the job still packs output/ and sends the tarball back, with whatever the analysis wrote before it stopped, and then exits with R’s own exit status, so HTCondor still reports the job as failed. The reason is in that job’s .err file, and the partial results are there to look at rather than lost. (Without this, a failed job would leave no tarball, and HTCondor would put it on hold for a missing output file instead.)
Your analysis script is not baked into the container image. It travels to the execute node as an uploaded job input file, the same way data.csv would, so the generated script runs it by bare name (Rscript analysis.R) rather than by an absolute, in-container path. Because HTCondor’s file transfer does not preserve subdirectories, r_script = "R/analysis.R" still resolves to analysis.R at the execute node, which is also the name htc_gen_submit() gives it in transfer_input_files. Data files passed via data_files are the opposite case: those are baked into the image under home_dir at build time, and the script reads them by absolute path. Editing your analysis script therefore only requires re-uploading it and resubmitting – no container rebuild, no registry push.
With comments = TRUE, each section of the generated script is preceded by an explanation of the line beneath it. containr annotates its Dockerfiles the same way round, so a reader moving between the two files reads them the same way: explanation first, instruction second.
output_file here shares the same defaulting through the submission state as executable in htc_gen_submit(): if you called htc_gen_submit() first, its executable value is already recorded, so output_file can be left off. results_folder follows the same config pattern as queue_from above – with a _toolero.yml passed through config, it defaults to config$project$conventions$output_dir instead of the hardcoded "output":
cfg <- htc_config(project_config = "_toolero.yml")
htc_gen_executable(r_script = "R/analysis.R", config = cfg)
htc_check()
Runs a preflight check, locally and in seconds, for the things that otherwise only surface an hour later as a held or failed job on the cluster: a missing input or data file, a "multiple"-mode job whose subset files no longer match subdatasets.csv, a resource request that looks implausible, and a container_image tagged latest or carrying no tag at all. Every argument resolves from the submission state, so the common case is no arguments at all, run right after the two generators:
htc_check()
#> v Preflight check passed -- no issues found.
# Or catch problems before they happen
htc_check()
#> ! Preflight check found 1 error and 1 warning.
#> x ERROR [input_files]: 1 input file(s) not found: R/analysis.R
#> ! WARNING [container_image]: container_image resolves to the "latest" tagIt returns a tibble of issues (zero rows means nothing was found), each tagged "error" (the job will not run without it) or "warning" (worth a second look, not necessarily wrong). When podman or docker is on your PATH, it also asks that tool whether the image can be pulled, which contacts the registry; pass check_image = FALSE to skip that when you are offline. htc_upload(check = TRUE) runs this automatically and aborts on an "error"; warnings never block anything.
htc_upload()
Copies files to the CHTC submit node via scp. Called with no files argument, it sends what the submission state recorded: the submit file, the executable, any shared input files, and in multiple-job mode the subsets and subdatasets.csv. On success, the remote directory it uploaded to is written back to the submission state, so htc_submit() and htc_download() can pick it up without it being retyped.
# Automatic -- uses the submission state built by the two generators
htc_upload()
# Preview the command before running it
htc_upload(dry_run = TRUE)
#> v Dry run -- command that would be executed:
#> `scp analysis.sub analysis.sh R/analysis.R your.netid@ap2002.chtc.wisc.edu:~/`
# Or name the files yourself, which bypasses the submission state entirely
htc_upload(files = c("analysis.sub", "analysis.sh", "R/analysis.R", "data.csv"))
# Uploading somewhere other than ~/ is remembered for later steps
htc_upload(remote_path = "~/projects/penguins/")
# Run htc_check() first and abort the upload if it finds a real problem
htc_upload(check = TRUE)check = TRUE runs htc_check() (described just above) before the transfer and aborts if it finds a missing file or a similar problem worth catching before it becomes a held job on the cluster. It is FALSE by default so existing calls behave exactly as before; turn it on while you are still shaking out a new job.
htc_submit()
Runs condor_submit on the remote server via SSH and returns the cluster ID. Both submit_file and remote_path default to NULL and resolve from the submission state – the submit file htc_gen_submit() wrote and the directory htc_upload() sent it to – so a call with no arguments at all works once those two steps have run:
cluster_id <- htc_submit(verbose = TRUE)
#> Submitting "analysis.sub" on "ap2002.chtc.wisc.edu"...
#> 1 job(s) submitted to cluster 6302860.
#> v Job submitted successfully.The cluster ID, the submit file, and the directory the job was submitted from are all written to the submission state, so none of them has to be repeated when you come back to collect the results.
htc_status()
Runs condor_q on the remote server. Use watch = TRUE to poll until all jobs in the cluster leave the queue. cluster_id defaults to NULL and resolves from the submission state – the ID htc_submit() just returned – so you rarely have to pass it explicitly right after submitting:
# One-shot check
htc_status(cluster_id = cluster_id)
# Watch until complete
htc_status(cluster_id = cluster_id, watch = TRUE)
# Or let it resolve cluster_id from the submission state
htc_status(watch = TRUE)When any jobs in the cluster are held, htc_status() automatically runs a follow-up query and prints the hold reason, so you do not have to leave R to find out why:
htc_status(cluster_id = cluster_id)
#> ! Held job(s) detected. Hold reason(s):
#> 6302860.3 Error from slot1@execute-node: Failed to access user logSet show_hold_reason = FALSE to skip that extra query.
htc_cancel() and htc_release()
Job control from R: htc_cancel() removes a submitted cluster with condor_rm, and htc_release() puts held jobs back into the queue with condor_release. Both resolve cluster_id from the submission state, the same way htc_status() does:
# Cancel the cluster htc_submit() just returned
htc_cancel(cluster_id = cluster_id, reason = "wrong container image")
# Release jobs that HTCondor put on hold
htc_release(cluster_id = cluster_id)
# Or resolve cluster_id from the submission state, same as htc_status()
htc_cancel()Unlike htc_status(), which shows every job in the queue when cluster_id is omitted and nothing can be resolved, htc_cancel() and htc_release() refuse to proceed in that situation rather than falling back to acting on everything – removing or releasing every job you have queued is a much larger mistake than an unfiltered status check. Both support dry_run to preview the condor_rm/condor_release command first.
htc_download()
Copies files back from the submit node via scp. After a full workflow, htc_download() knows which files to retrieve and which remote directory to take them from:
# Automatic -- uses the submission state built during the workflow
htc_download()
# Or specify the cluster ID explicitly
htc_download(cluster_id = "6590895")
# Or specify files directly with glob patterns
htc_download(files = "*.tar.gz", local_path = "downloads/")For a single job it retrieves the results tarball and the three HTCondor log files. For a multiple-job run it retrieves one tarball per subset and one set of logs per process.
If your analysis renders a Quarto document, set embed-resources: true in its YAML header. Both toolero templates already do. Without it a rendered .qmd produces an .html file plus a _files/ directory of supporting assets, and only what you named in output/ comes home; with it the report arrives as a single self-contained file. This removes a whole category of “my figures did not come back”.
htc_collect()
htc_download() leaves you with a folder of tarballs and log files. htc_collect() unpacks each tarball into its own subfolder and returns the job index: a tibble with one row per job, saying whether its results came back, where they landed, what files they hold, and where that job’s HTCondor logs are. Like the other pipeline functions, it resolves what it needs from the submission state:
index <- htc_collect()
index[, c("group_id", "proc_id", "extracted", "n_files", "output_dir")]
#> # A tibble: 3 x 5
#> group_id proc_id extracted n_files output_dir
#> <chr> <int> <lgl> <int> <chr>
#> 1 adelie 0 TRUE 2 ./adelie/output
#> 2 chinstrap 1 FALSE NA NA
#> 3 gentoo 2 TRUE 2 ./gentoo/outputA job whose R script fails still sends its tarball back, holding whatever it wrote before it stopped (see htc_gen_executable() above). A job that never got that far – the container did not start, or the job was removed – sends nothing, so it shows up as a row with extracted = FALSE rather than stopping the collection, and htc_collect() warns once about all such jobs together. The err column points at that job’s .err file, which usually says what went wrong:
lapply(index$err[!index$extracted], readLines)htc_collect() works whatever your script wrote into output/, and it never opens those files, since their types vary by analysis. files is a list column of paths relative to output_dir, so reading every saved model back is one line:
If the analysis used toolero::save_output() and toolero::generate_manifest(), each job’s results folder also holds an output record (project-manifest.json), and has_record says which jobs have one. htc_collect() only checks that the file is there; interpreting it is toolero’s job.
The submission state
Several calls above take no arguments at all, and they are not guessing. As you work, submitr writes what it learns to htc-manifest.yml, a small file that sits in your project beside htc.cfg. The family calls this file the submission state: submitr’s working memory for the job in progress. Its name predates that term and is kept for compatibility, but it is not a job manifest – that is toolero::write_by_group()’s manifest.csv, the list of subsets a multiple-job run reads once, at generation time. The vocabulary section of CONVENTIONS.md lists all four terms.
Each step contributes what it knows, and each step after the first reads back what an earlier one wrote. htc_gen_submit() records the submit file, the mode, the derived results name, and in multiple-job mode the list of subsets. htc_gen_executable() records the analysis script and the executable it wrote. htc_upload() records the remote directory it sent files to. htc_submit() resolves the submit file and remote directory from those two records when you do not pass them, and records the cluster ID HTCondor assigned. htc_status() resolves the cluster ID the same way. By the time you call htc_download(), the submission state holds everything needed to work out which files to ask for.
The reason it is a file rather than something held in memory is the shape of the work. A CHTC job worth sending to CHTC is usually one that takes a while, so you submit it in one sitting and collect it in another, and somewhere in between you close RStudio or your laptop sleeps. Submission state that lived only in the R session would be gone by then, and with it any chance of htc_download() knowing what to retrieve. Because it is on disk, restarting R costs you nothing, and htc_start() leaves it alone.
You can read it at any time. It is ordinary YAML, and looking at it is often the quickest way to see what submitr thinks the state of your job is.
submit_file: analysis.sub
executable_file: analysis.sh
r_script: R/analysis.R
script_stem: analysis
mode: single
output_files: analysis-results.tar.gz
cluster_id: '6302860'
remote_path: ~/By default htc-manifest.yml lives in your project root (path = "."), regardless of where output points – generating files into a subdirectory does not move it along with them. If you do write generated files elsewhere, pass the same path to every function that touches the submission state, so that all five are reading and writing the same file:
htc_gen_submit(r_script = "R/analysis.R", input_files = "R/analysis.R",
output = "jobs/", path = "jobs/")
htc_gen_executable(r_script = "R/analysis.R", output = "jobs/", path = "jobs/")
htc_upload(path = "jobs/")
htc_submit(submit_file = "analysis.sub", path = "jobs/")
htc_download(path = "jobs/")Scaling to many jobs
Once a single job works, scaling up is mostly a matter of changing the queue. Use toolero::write_by_group() to split your dataset and produce a job manifest, then switch to multiple-job mode:
htc_gen_submit(
output_file = "analysis.sub",
container_image = "docker://registry.doit.wisc.edu/your.netid/my-analysis:1.0.0",
executable = "analysis.sh",
r_script = "R/analysis.R",
input_files = "R/analysis.R",
mode = "multiple",
queue_from = "data/jobs/manifest.csv",
resources = "medium",
comments = TRUE
)
htc_gen_executable(
r_script = "R/analysis.R",
output_file = "analysis.sh",
mode = "multiple",
comments = TRUE
)In multiple-job mode, HTCondor passes each subset filename to your R script as a positional argument. Your script should read that argument explicitly:
args <- commandArgs(trailingOnly = TRUE)
input_file <- args[[1]]
data <- readr::read_csv(input_file)One thing to watch when you split more than one dataset. Uploaded files land in a single flat directory on the access point, so two datasets split on the same grouping column produce the same subset filenames and the second set overwrites the first. toolero::write_by_group() takes a prefix argument for exactly this reason; use it whenever a project splits more than one dataset, and the subset names stay distinct all the way through to the tarballs that come back.
A note on the results naming change
Two conventions changed in the development version, and they change the names of files your jobs produce. If you have results sitting on the submit node from an earlier version, download them before upgrading, because htc_download() will now look for names those jobs never created.
The folder your job writes into is now output/ rather than results/. This brings submitr in line with toolero and containr, which already use output/. One folder name across all three packages means an analysis script that runs on your laptop writes to the same place when it runs on an execute node, so toolero::save_output() behaves identically in both, and you do not have to remember which package is in charge of a given directory.
Results tarballs are now named after the script and, in multiple-job mode, the subset the job handled:
| Before | Now | |
|---|---|---|
| Single job | analysis-results.tar.gz |
analysis-results.tar.gz |
| One subset of many | adelie.csv-results.tar.gz |
analysis-adelie-results.tar.gz |
The single-job name is unchanged. The multiple-job name gains the script stem and loses the subset’s file extension, and both halves are deliberate. Adding the script stem means two different analyses splitting the same dataset no longer overwrite each other’s results in the flat namespace of your home directory on the access point. Dropping the extension avoids the awkward adelie.csv-results.tar.gz, in which .csv describes a file that is not a CSV and is not the file being named. Directories are stripped from both stems, which is what lets you follow the family convention of keeping a derived script at R/analysis.R without the job trying to write its tarball into a directory the execute node does not have.
One consequence worth knowing: htc_gen_submit() now needs r_script in order to derive that name. It never opens the executable script and never reads the Dockerfile, so the name of your analysis script is genuinely not something it can work out for itself.
Quick function reference
| Function | What it does |
|---|---|
htc_start() |
Start a session, reading config and storing it for all calls |
htc_config() |
Create or read htc.cfg, optionally checking the server |
htc_ssh_setup() |
Write the ControlMaster block for SSH connection reuse |
htc_gen_submit() |
Generate the HTCondor .sub submit file |
htc_gen_executable() |
Generate the .sh executable script |
htc_check() |
Preflight-check files, resources, and image before upload |
htc_upload() |
Copy files to the submit node via scp
|
htc_submit() |
Run condor_submit on the submit node |
htc_status() |
Check job progress via condor_q, including hold reasons |
htc_cancel() |
Remove a submitted cluster via condor_rm
|
htc_release() |
Release held jobs back into the queue via condor_release
|
htc_download() |
Copy results back from the submit node |
htc_collect() |
Unpack downloaded tarballs and index them, one row per job |
htc_upload(), htc_submit(), htc_status(), htc_cancel(), htc_release(), htc_download(), and htc_collect() can all be called with no arguments (or close to it) once the steps before them have run. All functions that read or write the submission state take a path argument naming the directory it lives in, and all default it to "." independent of output – see the submission state.
What submitr does not do
submitr reduces friction. It does not replace understanding.
- It does not decide whether your workload is appropriate for CHTC.
- It does not manage large input files greater than 1 GB. Those belong in CHTC’s staging area and require a different transfer pattern.
- It does not validate that your container image is correct or that your analysis script will run successfully inside it. Test both locally before submitting to CHTC.
- It does not replace CHTC consultation for complex workloads, custom scheduling requirements, or non-standard resource requests.
The CHTC facilitation team is the right resource for complex workflow questions.
Learn more
The package vignette walks through a complete first submission step by step, with annotated output at each stage:
Related packages
submitr is part of the From the Notebook to the Cluster package family:
- toolero, which organizes and scaffolds the project, uses Quarto as the source of truth, and splits datasets for parallel jobs
- containr, which containerizes the software environment
- submitr, which submits to CHTC and retrieves results (this package)
