
From the Notebook to the Cluster: Your First CHTC Job with submitr
Created 2026-04-30 | Last updated 2026-10-01
Source:vignettes/from-notebook-to-cluster.Rmd
from-notebook-to-cluster.Rmd
The promise of a first CHTC job
Many research coding projects begin in a notebook-style workflow: an RStudio project, a Quarto document, a few scripts, a folder of input files, and enough local experimentation to understand what the analysis needs to do.
That is a good place to start. A laptop is often the right place to explore data, write early code, make plots, and decide what question the analysis is actually answering. But at some point, the same local workflow can become the wrong place to keep pushing.
Maybe the analysis takes too long. Maybe you need to run the same model across hundreds of parameter combinations. Maybe you need one job per participant, county, simulation, image, genome, or bootstrap sample. Maybe you simply want a workflow that will be easier to rerun six months from now.
That is where high-throughput computing becomes useful.
The UW-Madison Center for High Throughput Computing (CHTC) gives researchers access to large pools of computing capacity. Instead of asking one computer to do everything in sequence, you can break work into independent jobs and let the HTC system run those jobs when resources are available.
submitr helps you take the final step in the
From the Notebook to the Cluster workflow: sending a
prepared R project to CHTC. It is designed for researchers who know R
but may not yet be comfortable with HTCondor submit files, executable
shell scripts, ssh, scp, or the rhythm of
working on a remote submit node. It is also useful for regular CHTC
users who want to reduce repetitive setup work and make job submission
easier to reproduce, review, and share.
The goal is not to hide CHTC from you. The goal is to make the standard path visible, repeatable, and less fragile.
The larger idea: make the right choice easy
submitr is part of a small family of R packages for
research computing workflows:
local R project
└─ toolero: organize the project and prepare job-sized inputs
└─ containr: capture the R software environment in a container image
└─ submitr: send the containerized job to CHTC
You can use each package on its own.
Use toolero if you want a better project skeleton,
cleaner data-loading habits, Quarto scaffolding, or a simple way to
split a dataset into many job-sized files.
Use containr if you already have a project with an
renv.lock file and want to build a container image that can
run somewhere other than your laptop.
Use submitr if your project is already organized and
containerized, and you are ready to submit it to CHTC.
Used together, the packages support a practical arc: start with a project that is easier to understand, make its software environment portable, then send it to CHTC with fewer command-line hurdles.
The folder names and path conventions the three packages share are
written down in CONVENTIONS.md,
in the toolero repository. Two of them appear throughout
this vignette: a derived analysis script lives at
R/analysis.R, and analysis outputs go in
output/. Following them is what lets one script run
unchanged on your laptop and on an execute node.
What submitr does
A CHTC job needs a few pieces of information:
- what code to run;
- what input files to transfer;
- what container image to use;
- how much CPU, memory, and disk to request;
- what output files to retrieve;
- how many jobs to queue.
In HTCondor, that information is split across two main files.
The submit file tells HTCondor how to run the job. It describes the executable script, container image, input files, output files, log files, resource requests, and queue instructions.
The executable script tells the job what to do after
it starts. For an R analysis, that usually means creating an output
folder, running Rscript, and packaging results.
submitr helps you generate those files and then use
them:
submitr::htc_start() # configure and store your submit-node connection
submitr::htc_gen_submit() # generate the HTCondor submit file
submitr::htc_gen_executable() # generate the executable shell script
submitr::htc_upload() # copy files to the submit node
submitr::htc_submit() # submit the job
submitr::htc_status() # check progress
submitr::htc_download() # copy results backTwo of those calls take no arguments in a normal workflow. As you go,
submitr records what it has generated in a small file
called htc-manifest.yml, the submission state, and
htc_upload() and htc_download() read it to
work out which files to move. There is a section on that below.
Before you submit anything
A successful CHTC submission starts before
condor_submit.
Before using submitr, confirm that:
- your R script runs with
Rscript R/analysis.Routside RStudio; - your container image is pushed to a registry CHTC can access;
- you have SSH access to a CHTC submit node such as
ap2002.chtc.wisc.edu.
The most important check is simple: your analysis should run outside RStudio.
If that command fails locally, the same analysis is likely to fail on CHTC. Fix that first. CHTC will not know about objects in your Global Environment, local RStudio settings, manually clicked files, or packages that happen to be installed on your laptop.
Set up SSH connection reuse now, before anything
else. Every submitr function that touches CHTC
opens an SSH connection, which can trigger a Duo MFA prompt.
ControlMaster caches your authenticated session so all subsequent calls,
whether uploads, submits, status checks or downloads, reuse the same
connection without prompting again. The setup takes two minutes and is
worth doing before your first htc_start() call. Full
instructions appear right after Step 1.
For a first submission, choose something small and intentionally boring. The goal is not to prove that your full analysis can scale yet. The goal is to prove that the pathway works.
A small example analysis
Suppose your project has this shape:
my-analysis/
├── R/
│ └── analysis.R
├── data.csv
├── renv.lock
└── output/
Your R/analysis.R script might look like this:
library(readr)
library(dplyr)
input <- read_csv("data.csv")
summary <- input |>
group_by(group) |>
summarise(
mean_value = mean(value, na.rm = TRUE),
n = dplyr::n(),
.groups = "drop"
)
if (!dir.exists("output")) {
dir.create("output")
}
write_csv(summary, "output/summary.csv")This script is deliberately modest. A first CHTC job should be easy to inspect. Once the small version works, you can scale the pattern with more confidence.
Note the relative path. "output/" resolves correctly in
RStudio, under quarto render, and on an execute node,
because the generated executable script moves to HTCondor’s writable
scratch directory before calling Rscript. The script itself
never has to know where it is running.
Step 1: configure your CHTC connection
Load submitr and start a session:
On first use, htc_start() prompts for your NetID and
submit node. It writes an htc.cfg file to the project
directory so later calls can reuse the same connection information, and
it displays ControlMaster setup instructions. It then stores the
configuration for the rest of the session, which is why none of the
calls that follow need a config argument.
A later call should look something like this:
htc_start()
#> Reading HTC config from ./htc.cfg
#> v Connected to "ap2002.chtc.wisc.edu" as "your.netid".
#> v Session started: "your.netid"@"ap2002.chtc.wisc.edu"That middle line comes from a short SSH connection made to tell you
whether the server is reachable before you rely on the configuration. It
is useful at a prompt and pointless in a script, so it can be turned off
with check_server = FALSE, or for a whole session with the
submitr.check_server option.
This configuration file is deliberately project-local. Different projects may need different submit nodes, paths, or connection settings.
Setting up SSH connection reuse
Before continuing, take two minutes to configure ControlMaster. Add
this block to ~/.ssh/config:
Then create the directory used by ControlPath:
With ControlMaster in place, all subsequent SSH connections reuse the same authenticated session. You authenticate once when the connection is first established; everything that follows, including file uploads, job submission, status checks and result downloads, happens without prompting for Duo MFA again. Full documentation is at https://chtc.cs.wisc.edu/uw-research-computing/configure-ssh.
The rest of this vignette assumes ControlMaster is in place.
Step 2: generate the submit file
The submit file is the main HTCondor instruction file. It answers the question: what should the HTC system run, and what does it need?
htc_gen_submit(
output_file = "analysis.sub",
container_image = "docker://registry.doit.wisc.edu/your.netid/my-analysis:1.0.0",
executable = "analysis.sh",
r_script = "R/analysis.R",
input_files = c("R/analysis.R", "data.csv"),
resources = "small",
comments = TRUE,
output = "."
)For a first submission, keep comments = TRUE. The
generated file includes explanations of the main sections, making it
easier to inspect, learn from, and share with a collaborator or
consultant.
r_script is worth a word, because nothing here runs it.
htc_gen_submit() uses it only to work out what to call the
results tarball, so that the name matches the one
htc_gen_executable() will tell the job to build. The
function never opens the executable script and never reads the
Dockerfile, so the name of your analysis script is genuinely not
something it can infer. Supply output_files yourself if you
want a different name.
The resources argument uses presets. For a first test,
always start with "small" regardless of what your eventual
job will need:
| preset | cpus | memory | disk | when to use |
|---|---|---|---|---|
| small | 1 | 4 GB | 4 GB | first test jobs, lightweight scripts |
| medium | 4 | 16 GB | 15 GB | moderate analyses, model fitting |
| large | 8 | 64 GB | 32 GB | memory-intensive work, large datasets |
The HTCondor log file reports actual resource usage after each run. That log is the ground truth for tuning future submissions, not guesswork. Requesting too little causes jobs to fail; requesting much more than you need makes jobs harder to match with available resources.
Step 3: generate the executable script
The executable script answers a different question: once the job starts, what commands should run?
htc_gen_executable(
r_script = "R/analysis.R",
output_file = "analysis.sh",
comments = TRUE
)The generated script handles a standard sequence:
- move to HTCondor’s writable scratch directory;
- create the
output/folder; - run the R script with
Rscript, by its bare name – your analysis script is not baked into the container image, it travels to the execute node as an uploaded job input file (theinput_filesyou listed in Step 2), the same waydata.csvdid; - archive the folder as
analysis-results.tar.gz.
Because HTCondor’s file transfer does not preserve subdirectories,
the script lands flat in the scratch directory regardless of the
R/ prefix on r_script, so the generated line
reads Rscript analysis.R, not
Rscript R/analysis.R. This is also why editing your
analysis script only means re-uploading it and resubmitting the job – no
container rebuild, no registry push. Reference data that genuinely
belongs in the image (large or unchanging files you do not expect to
edit) is a different case: pass it to
htc_gen_executable()’s data_files argument
instead, and it keeps the absolute, home_dir-prefixed path
that a baked-in file needs.
That last name is derived from r_script, with the
directory and extension stripped, which is why R/analysis.R
gives analysis-results.tar.gz rather than something
containing a R/ that does not exist on the execute
node.
Only output/ itself is created. If your analysis writes
to output/figures/, the R script has to create that
subfolder, which toolero::save_output() does and a bare
ggsave() does not.
That sequence is not complicated, but it is exactly the kind of glue
code that can become a barrier for researchers who are new to shell
scripts. submitr writes the standard version so you can
focus on the analysis.
Step 4: preview and upload files
Before copying files to the submit node, do a dry run:
htc_upload(dry_run = TRUE)
#> v Dry run -- command that would be executed:
#> `scp analysis.sub analysis.sh R/analysis.R your.netid@ap2002.chtc.wisc.edu:~/`A dry run is a safety habit. It lets you see the command before it changes anything on the remote system. Once the command looks right, upload the files:
Neither call names any files. htc_upload() reads the
submission state that the two generators just wrote, and sends the
submit file, the executable, and any shared input files you declared.
You can still pass files explicitly, which bypasses the
submission state entirely:
htc_upload(files = c("analysis.sub", "analysis.sh", "R/analysis.R", "data.csv"))Step 5: submit the job
cluster_id <- htc_submit(
submit_file = "analysis.sub",
verbose = TRUE
)
#> Submitting "analysis.sub" on "ap2002.chtc.wisc.edu"...
#> 1 job(s) submitted to cluster 6302860.
#> v Job submitted from "~/analysis.sub" on "ap2002.chtc.wisc.edu".The cluster ID is the handle for this submission. Store it in an
object so you can check the job later without having to look it up.
htc_submit() also writes it to the submission state, along
with the remote directory it submitted from, so the next two steps can
find the job without being told.
Step 6: check progress
# One-shot status check
htc_status(cluster_id = cluster_id)
# Watch until the job completes
htc_status(cluster_id = cluster_id, watch = TRUE)For a small test job, watch = TRUE is useful. For larger
workloads, occasional one-shot checks are usually a better fit than
keeping an R session occupied.
Step 7: download results
When the job is complete:
That single call retrieves the results tarball and the three HTCondor
log files, 6302860-0-job.log, .err and
.out. It knows the tarball’s name because
htc_gen_submit() recorded it, knows the cluster ID because
htc_submit() recorded it, and knows which remote directory
to look in for the same reason.
You can still be explicit when you want to be:
htc_download(cluster_id = "6302860")
htc_download(files = "*.tar.gz", local_path = "downloads/")The logs are not just for failures. They record what happened when the job ran, including actual resource usage, which informs future resource requests.
If your analysis renders a Quarto document, set
embed-resources: true in its YAML header. Both
toolero templates already do. Without it a rendered
.qmd produces an .html file plus a
_files/ directory of supporting assets, and only what you
named in output/ comes home; with it the report arrives as
a single self-contained file.
The submission state
Steps 4 and 7 took no arguments, and it is worth understanding why, because it also explains something that matters when a job runs long.
As you work, submitr writes what it learns to
htc-manifest.yml, a small file that sits in your project
beside htc.cfg. The family calls it the submission
state, to keep it distinct from the job manifest that
toolero::write_by_group() writes; the file name predates
that distinction and is kept for compatibility.
htc_gen_submit() records the submit file, the mode and the
derived results name. htc_gen_executable() records the
analysis script and the executable it wrote. htc_upload()
records the remote directory it sent files to. htc_submit()
resolves the submit file and remote directory from those records when
you do not supply them, and records the cluster ID HTCondor assigned. By
the time you reach htc_download(), the submission state
holds everything needed to work out what to ask the submit node for.
It is ordinary YAML, and reading it is often the quickest way to see
what submitr thinks the state of your job is:
submit_file: analysis.sub
executable_file: analysis.sh
r_script: R/analysis.R
script_stem: analysis
mode: single
output_files: analysis-results.tar.gz
cluster_id: '6302860'
remote_path: ~/The reason it is a file rather than something held in the R session
is the shape of the work. A job worth sending to CHTC usually takes a
while, so you submit it one day and collect it another, and somewhere in
between you close RStudio or your laptop sleeps. Anything held in memory
would be gone. Because the submission state is on disk, you can restart
R, come back on Thursday, call htc_download() with no
arguments, and it still knows what to fetch. htc_start()
leaves it alone for exactly that reason.
From one test job to many HTC jobs
A first job proves that the path works. The next step is to think like an HTC user: how can the analysis be divided into many independent pieces?
Common patterns include one job per simulation replicate, model
specification, input file, county, participant, sample, parameter set,
or bootstrap iteration. This is where
toolero::write_by_group() helps upstream. It splits a data
frame into separate CSV files and writes a job manifest describing those
files. Then submitr queues one job per row of the job
manifest:
htc_gen_submit(
output_file = "analysis.sub",
container_image = "docker://registry.doit.wisc.edu/your.netid/my-analysis:1.0.0",
executable = "analysis.sh",
r_script = "R/analysis.R",
input_files = "R/analysis.R",
mode = "multiple",
queue_from = "data/manifest.csv",
resources = "medium",
comments = TRUE
)
htc_gen_executable(
r_script = "R/analysis.R",
output_file = "analysis.sh",
mode = "multiple",
comments = TRUE
)In multiple-job mode, the generated executable passes the per-job input file to your R script as the first command-line argument. Your script should read that argument explicitly:
args <- commandArgs(trailingOnly = TRUE)
input_file <- args[[1]]
data <- readr::read_csv(input_file)This is a key pattern. The script stays the same; each job receives a different input. The companion vignette, Single Jobs vs Multiple Jobs on HTCondor, works through both modes side by side.
Where containr fits
CHTC needs to know what software environment your job should use. Your laptop may have the right R packages installed, but the execute node will not automatically have the same setup. A container image solves that problem by packaging the R version, packages, and system libraries needed to run the analysis.
containr handles that step. Notice that
generate_dockerfile() is not given your analysis script
here – as Step 3 explained, the script travels to CHTC as an uploaded
job input file rather than as part of the image, so the image only needs
to carry the R version and packages your renv.lock
records:
containr::generate_dockerfile(
r_version = "4.4.0",
output = "."
)
containr::build_image(verbose = TRUE)
imgs <- containr::list_images()
containr::push_image(
image_id = imgs$image_id[1],
netid = "your.netid",
project = "my-analysis",
tag = "1.0.0"
)After the image is pushed to a registry CHTC can access,
submitr can refer to it in
container_image.
Use explicit image tags such as "1.0.0" rather than
"latest". A versioned tag makes it unambiguous which
software environment was used for a particular analysis.
A practical first-submission checklist
Before scaling up, confirm that the small job works end to end:
- The script runs locally with
Rscript R/analysis.R. - ControlMaster is configured and the session is authenticated.
- The container image is pushed to a registry CHTC can access.
- The image tag is explicit, not
"latest". - The submit file lists the correct executable and input files, including your analysis script – it is not baked into the image.
-
r_scriptis the same in both generator calls, so the tarball the job builds is the one the submit file asks for. - The dry-run upload shows the expected files.
-
htc_start()connects to the submit node without error. - The resource request is reasonable for a test job.
- The job produces logs and a result archive.
Once this works, you have something valuable: a known-good pathway from local R project to CHTC.
What submitr does not do
submitr reduces friction, but it does not remove the
need to make sound research-computing decisions.
It does not:
- decide whether your workload is a good fit for CHTC;
- make interactive R code safe for batch execution;
- guarantee that your container image contains every system dependency;
- manage restricted or sensitive data;
- replace CHTC documentation or consultation for complex workflows.
That boundary is intentional. Good tools should make the common path easier while still leaving the important decisions visible.
The CHTC facilitation team is the right resource for complex workflow questions.
A good first goal
Do not make your first submission your largest analysis.
Make your first goal smaller: send one boring job to CHTC, watch it run, and download one result file.
After that, the cluster becomes less mysterious. You can inspect the
generated files, adjust resources, split work into many jobs, and grow
the workflow with more confidence. That is the role of
submitr: to help you take the first successful step from
local R code to high-throughput research computing.