Skip to contents

submitr hex sticker

The promise of a first CHTC job

Many research coding projects begin in a notebook-style workflow: an RStudio project, a Quarto document, a few scripts, a folder of input files, and enough local experimentation to understand what the analysis needs to do.

That is a good place to start. A laptop is often the right place to explore data, write early code, make plots, and decide what question the analysis is actually answering. But at some point, the same local workflow can become the wrong place to keep pushing.

Maybe the analysis takes too long. Maybe you need to run the same model across hundreds of parameter combinations. Maybe you need one job per participant, county, simulation, image, genome, or bootstrap sample. Maybe you simply want a workflow that will be easier to rerun six months from now.

That is where high-throughput computing becomes useful.

The UW-Madison Center for High Throughput Computing (CHTC) gives researchers access to large pools of computing capacity. Instead of asking one computer to do everything in sequence, you can break work into independent jobs and let the HTC system run those jobs when resources are available.

submitr helps you take the final step in the From the Notebook to the Cluster workflow: sending a prepared R project to CHTC. It is designed for researchers who know R but may not yet be comfortable with HTCondor submit files, executable shell scripts, ssh, scp, or the rhythm of working on a remote submit node. It is also useful for regular CHTC users who want to reduce repetitive setup work and make job submission easier to reproduce, review, and share.

The goal is not to hide CHTC from you. The goal is to make the standard path visible, repeatable, and less fragile.


The larger idea: make the right choice easy

submitr is part of a small family of R packages for research computing workflows:

local R project
  └─ toolero: organize the project and prepare job-sized inputs
      └─ containr: capture the R software environment in a container image
          └─ submitr: send the containerized job to CHTC

You can use each package on its own.

Use toolero if you want a better project skeleton, cleaner data-loading habits, Quarto scaffolding, or a simple way to split a dataset into many job-sized files.

Use containr if you already have a project with an renv.lock file and want to build a container image that can run somewhere other than your laptop.

Use submitr if your project is already organized and containerized, and you are ready to submit it to CHTC.

Used together, the packages support a practical arc: start with a project that is easier to understand, make its software environment portable, then send it to CHTC with fewer command-line hurdles.

The folder names and path conventions the three packages share are written down in CONVENTIONS.md, in the toolero repository. Two of them appear throughout this vignette: a derived analysis script lives at R/analysis.R, and analysis outputs go in output/. Following them is what lets one script run unchanged on your laptop and on an execute node.


What submitr does

A CHTC job needs a few pieces of information:

  • what code to run;
  • what input files to transfer;
  • what container image to use;
  • how much CPU, memory, and disk to request;
  • what output files to retrieve;
  • how many jobs to queue.

In HTCondor, that information is split across two main files.

The submit file tells HTCondor how to run the job. It describes the executable script, container image, input files, output files, log files, resource requests, and queue instructions.

The executable script tells the job what to do after it starts. For an R analysis, that usually means creating an output folder, running Rscript, and packaging results.

submitr helps you generate those files and then use them:

submitr::htc_start()          # configure and store your submit-node connection
submitr::htc_gen_submit()     # generate the HTCondor submit file
submitr::htc_gen_executable() # generate the executable shell script
submitr::htc_upload()         # copy files to the submit node
submitr::htc_submit()         # submit the job
submitr::htc_status()         # check progress
submitr::htc_download()       # copy results back

Two of those calls take no arguments in a normal workflow. As you go, submitr records what it has generated in a small file called htc-manifest.yml, the submission state, and htc_upload() and htc_download() read it to work out which files to move. There is a section on that below.


Before you submit anything

A successful CHTC submission starts before condor_submit.

Before using submitr, confirm that:

  • your R script runs with Rscript R/analysis.R outside RStudio;
  • your container image is pushed to a registry CHTC can access;
  • you have SSH access to a CHTC submit node such as ap2002.chtc.wisc.edu.

The most important check is simple: your analysis should run outside RStudio.

Rscript R/analysis.R

If that command fails locally, the same analysis is likely to fail on CHTC. Fix that first. CHTC will not know about objects in your Global Environment, local RStudio settings, manually clicked files, or packages that happen to be installed on your laptop.

Set up SSH connection reuse now, before anything else. Every submitr function that touches CHTC opens an SSH connection, which can trigger a Duo MFA prompt. ControlMaster caches your authenticated session so all subsequent calls, whether uploads, submits, status checks or downloads, reuse the same connection without prompting again. The setup takes two minutes and is worth doing before your first htc_start() call. Full instructions appear right after Step 1.

For a first submission, choose something small and intentionally boring. The goal is not to prove that your full analysis can scale yet. The goal is to prove that the pathway works.


A small example analysis

Suppose your project has this shape:

my-analysis/
├── R/
│   └── analysis.R
├── data.csv
├── renv.lock
└── output/

Your R/analysis.R script might look like this:

library(readr)
library(dplyr)

input <- read_csv("data.csv")

summary <- input |>
  group_by(group) |>
  summarise(
    mean_value = mean(value, na.rm = TRUE),
    n          = dplyr::n(),
    .groups    = "drop"
  )

if (!dir.exists("output")) {
  dir.create("output")
}

write_csv(summary, "output/summary.csv")

This script is deliberately modest. A first CHTC job should be easy to inspect. Once the small version works, you can scale the pattern with more confidence.

Note the relative path. "output/" resolves correctly in RStudio, under quarto render, and on an execute node, because the generated executable script moves to HTCondor’s writable scratch directory before calling Rscript. The script itself never has to know where it is running.


Step 1: configure your CHTC connection

Load submitr and start a session:

library(submitr)

htc_start()
#> v Session started: "your.netid"@"ap2002.chtc.wisc.edu"

On first use, htc_start() prompts for your NetID and submit node. It writes an htc.cfg file to the project directory so later calls can reuse the same connection information, and it displays ControlMaster setup instructions. It then stores the configuration for the rest of the session, which is why none of the calls that follow need a config argument.

A later call should look something like this:

htc_start()
#> Reading HTC config from ./htc.cfg
#> v Connected to "ap2002.chtc.wisc.edu" as "your.netid".
#> v Session started: "your.netid"@"ap2002.chtc.wisc.edu"

That middle line comes from a short SSH connection made to tell you whether the server is reachable before you rely on the configuration. It is useful at a prompt and pointless in a script, so it can be turned off with check_server = FALSE, or for a whole session with the submitr.check_server option.

This configuration file is deliberately project-local. Different projects may need different submit nodes, paths, or connection settings.


Setting up SSH connection reuse

Before continuing, take two minutes to configure ControlMaster. Add this block to ~/.ssh/config:

Host *.chtc.wisc.edu
  ControlMaster auto
  ControlPersist 2h
  ControlPath ~/.ssh/connections/%r@%h:%p

Then create the directory used by ControlPath:

mkdir -p ~/.ssh/connections

With ControlMaster in place, all subsequent SSH connections reuse the same authenticated session. You authenticate once when the connection is first established; everything that follows, including file uploads, job submission, status checks and result downloads, happens without prompting for Duo MFA again. Full documentation is at https://chtc.cs.wisc.edu/uw-research-computing/configure-ssh.

The rest of this vignette assumes ControlMaster is in place.


Step 2: generate the submit file

The submit file is the main HTCondor instruction file. It answers the question: what should the HTC system run, and what does it need?

htc_gen_submit(
  output_file     = "analysis.sub",
  container_image = "docker://registry.doit.wisc.edu/your.netid/my-analysis:1.0.0",
  executable      = "analysis.sh",
  r_script        = "R/analysis.R",
  input_files     = c("R/analysis.R", "data.csv"),
  resources       = "small",
  comments        = TRUE,
  output          = "."
)

For a first submission, keep comments = TRUE. The generated file includes explanations of the main sections, making it easier to inspect, learn from, and share with a collaborator or consultant.

r_script is worth a word, because nothing here runs it. htc_gen_submit() uses it only to work out what to call the results tarball, so that the name matches the one htc_gen_executable() will tell the job to build. The function never opens the executable script and never reads the Dockerfile, so the name of your analysis script is genuinely not something it can infer. Supply output_files yourself if you want a different name.

The resources argument uses presets. For a first test, always start with "small" regardless of what your eventual job will need:

preset cpus memory disk when to use
small 1 4 GB 4 GB first test jobs, lightweight scripts
medium 4 16 GB 15 GB moderate analyses, model fitting
large 8 64 GB 32 GB memory-intensive work, large datasets

The HTCondor log file reports actual resource usage after each run. That log is the ground truth for tuning future submissions, not guesswork. Requesting too little causes jobs to fail; requesting much more than you need makes jobs harder to match with available resources.


Step 3: generate the executable script

The executable script answers a different question: once the job starts, what commands should run?

htc_gen_executable(
  r_script    = "R/analysis.R",
  output_file = "analysis.sh",
  comments    = TRUE
)

The generated script handles a standard sequence:

  1. move to HTCondor’s writable scratch directory;
  2. create the output/ folder;
  3. run the R script with Rscript, by its bare name – your analysis script is not baked into the container image, it travels to the execute node as an uploaded job input file (the input_files you listed in Step 2), the same way data.csv did;
  4. archive the folder as analysis-results.tar.gz.

Because HTCondor’s file transfer does not preserve subdirectories, the script lands flat in the scratch directory regardless of the R/ prefix on r_script, so the generated line reads Rscript analysis.R, not Rscript R/analysis.R. This is also why editing your analysis script only means re-uploading it and resubmitting the job – no container rebuild, no registry push. Reference data that genuinely belongs in the image (large or unchanging files you do not expect to edit) is a different case: pass it to htc_gen_executable()’s data_files argument instead, and it keeps the absolute, home_dir-prefixed path that a baked-in file needs.

That last name is derived from r_script, with the directory and extension stripped, which is why R/analysis.R gives analysis-results.tar.gz rather than something containing a R/ that does not exist on the execute node.

Only output/ itself is created. If your analysis writes to output/figures/, the R script has to create that subfolder, which toolero::save_output() does and a bare ggsave() does not.

That sequence is not complicated, but it is exactly the kind of glue code that can become a barrier for researchers who are new to shell scripts. submitr writes the standard version so you can focus on the analysis.


Step 4: preview and upload files

Before copying files to the submit node, do a dry run:

htc_upload(dry_run = TRUE)
#> v Dry run -- command that would be executed:
#>   `scp analysis.sub analysis.sh R/analysis.R your.netid@ap2002.chtc.wisc.edu:~/`

A dry run is a safety habit. It lets you see the command before it changes anything on the remote system. Once the command looks right, upload the files:

Neither call names any files. htc_upload() reads the submission state that the two generators just wrote, and sends the submit file, the executable, and any shared input files you declared. You can still pass files explicitly, which bypasses the submission state entirely:

htc_upload(files = c("analysis.sub", "analysis.sh", "R/analysis.R", "data.csv"))

Step 5: submit the job

cluster_id <- htc_submit(
  submit_file = "analysis.sub",
  verbose     = TRUE
)
#> Submitting "analysis.sub" on "ap2002.chtc.wisc.edu"...
#> 1 job(s) submitted to cluster 6302860.
#> v Job submitted from "~/analysis.sub" on "ap2002.chtc.wisc.edu".

The cluster ID is the handle for this submission. Store it in an object so you can check the job later without having to look it up. htc_submit() also writes it to the submission state, along with the remote directory it submitted from, so the next two steps can find the job without being told.


Step 6: check progress

# One-shot status check
htc_status(cluster_id = cluster_id)

# Watch until the job completes
htc_status(cluster_id = cluster_id, watch = TRUE)

For a small test job, watch = TRUE is useful. For larger workloads, occasional one-shot checks are usually a better fit than keeping an R session occupied.


Step 7: download results

When the job is complete:

That single call retrieves the results tarball and the three HTCondor log files, 6302860-0-job.log, .err and .out. It knows the tarball’s name because htc_gen_submit() recorded it, knows the cluster ID because htc_submit() recorded it, and knows which remote directory to look in for the same reason.

You can still be explicit when you want to be:

htc_download(cluster_id = "6302860")
htc_download(files = "*.tar.gz", local_path = "downloads/")

The logs are not just for failures. They record what happened when the job ran, including actual resource usage, which informs future resource requests.

If your analysis renders a Quarto document, set embed-resources: true in its YAML header. Both toolero templates already do. Without it a rendered .qmd produces an .html file plus a _files/ directory of supporting assets, and only what you named in output/ comes home; with it the report arrives as a single self-contained file.


The submission state

Steps 4 and 7 took no arguments, and it is worth understanding why, because it also explains something that matters when a job runs long.

As you work, submitr writes what it learns to htc-manifest.yml, a small file that sits in your project beside htc.cfg. The family calls it the submission state, to keep it distinct from the job manifest that toolero::write_by_group() writes; the file name predates that distinction and is kept for compatibility. htc_gen_submit() records the submit file, the mode and the derived results name. htc_gen_executable() records the analysis script and the executable it wrote. htc_upload() records the remote directory it sent files to. htc_submit() resolves the submit file and remote directory from those records when you do not supply them, and records the cluster ID HTCondor assigned. By the time you reach htc_download(), the submission state holds everything needed to work out what to ask the submit node for.

It is ordinary YAML, and reading it is often the quickest way to see what submitr thinks the state of your job is:

submit_file: analysis.sub
executable_file: analysis.sh
r_script: R/analysis.R
script_stem: analysis
mode: single
output_files: analysis-results.tar.gz
cluster_id: '6302860'
remote_path: ~/

The reason it is a file rather than something held in the R session is the shape of the work. A job worth sending to CHTC usually takes a while, so you submit it one day and collect it another, and somewhere in between you close RStudio or your laptop sleeps. Anything held in memory would be gone. Because the submission state is on disk, you can restart R, come back on Thursday, call htc_download() with no arguments, and it still knows what to fetch. htc_start() leaves it alone for exactly that reason.


From one test job to many HTC jobs

A first job proves that the path works. The next step is to think like an HTC user: how can the analysis be divided into many independent pieces?

Common patterns include one job per simulation replicate, model specification, input file, county, participant, sample, parameter set, or bootstrap iteration. This is where toolero::write_by_group() helps upstream. It splits a data frame into separate CSV files and writes a job manifest describing those files. Then submitr queues one job per row of the job manifest:

htc_gen_submit(
  output_file     = "analysis.sub",
  container_image = "docker://registry.doit.wisc.edu/your.netid/my-analysis:1.0.0",
  executable      = "analysis.sh",
  r_script        = "R/analysis.R",
  input_files     = "R/analysis.R",
  mode            = "multiple",
  queue_from      = "data/manifest.csv",
  resources       = "medium",
  comments        = TRUE
)

htc_gen_executable(
  r_script    = "R/analysis.R",
  output_file = "analysis.sh",
  mode        = "multiple",
  comments    = TRUE
)

In multiple-job mode, the generated executable passes the per-job input file to your R script as the first command-line argument. Your script should read that argument explicitly:

args       <- commandArgs(trailingOnly = TRUE)
input_file <- args[[1]]

data <- readr::read_csv(input_file)

This is a key pattern. The script stays the same; each job receives a different input. The companion vignette, Single Jobs vs Multiple Jobs on HTCondor, works through both modes side by side.


Where containr fits

CHTC needs to know what software environment your job should use. Your laptop may have the right R packages installed, but the execute node will not automatically have the same setup. A container image solves that problem by packaging the R version, packages, and system libraries needed to run the analysis.

containr handles that step. Notice that generate_dockerfile() is not given your analysis script here – as Step 3 explained, the script travels to CHTC as an uploaded job input file rather than as part of the image, so the image only needs to carry the R version and packages your renv.lock records:

containr::generate_dockerfile(
  r_version = "4.4.0",
  output    = "."
)
containr::build_image(verbose = TRUE)

imgs <- containr::list_images()

containr::push_image(
  image_id = imgs$image_id[1],
  netid    = "your.netid",
  project  = "my-analysis",
  tag      = "1.0.0"
)

After the image is pushed to a registry CHTC can access, submitr can refer to it in container_image.

Use explicit image tags such as "1.0.0" rather than "latest". A versioned tag makes it unambiguous which software environment was used for a particular analysis.


A practical first-submission checklist

Before scaling up, confirm that the small job works end to end:

  • The script runs locally with Rscript R/analysis.R.
  • ControlMaster is configured and the session is authenticated.
  • The container image is pushed to a registry CHTC can access.
  • The image tag is explicit, not "latest".
  • The submit file lists the correct executable and input files, including your analysis script – it is not baked into the image.
  • r_script is the same in both generator calls, so the tarball the job builds is the one the submit file asks for.
  • The dry-run upload shows the expected files.
  • htc_start() connects to the submit node without error.
  • The resource request is reasonable for a test job.
  • The job produces logs and a result archive.

Once this works, you have something valuable: a known-good pathway from local R project to CHTC.


What submitr does not do

submitr reduces friction, but it does not remove the need to make sound research-computing decisions.

It does not:

  • decide whether your workload is a good fit for CHTC;
  • make interactive R code safe for batch execution;
  • guarantee that your container image contains every system dependency;
  • manage restricted or sensitive data;
  • replace CHTC documentation or consultation for complex workflows.

That boundary is intentional. Good tools should make the common path easier while still leaving the important decisions visible.

The CHTC facilitation team is the right resource for complex workflow questions.


A good first goal

Do not make your first submission your largest analysis.

Make your first goal smaller: send one boring job to CHTC, watch it run, and download one result file.

After that, the cluster becomes less mysterious. You can inspect the generated files, adjust resources, split work into many jobs, and grow the workflow with more confidence. That is the role of submitr: to help you take the first successful step from local R code to high-throughput research computing.