
Single Jobs vs Multiple Jobs on HTCondor
Created 2026-05-20 | Last updated 2026-10-01
Source:vignettes/single-vs-multiple-jobs.Rmd
single-vs-multiple-jobs.RmdWhy two modes?
Not every analysis scales the same way. Sometimes you have one dataset and one script and you want to run it once on hardware you don’t have locally. Other times you have the same analysis but need to run it independently across many subsets of the data: once per species, once per site, once per simulation parameter, once per experimental condition.
HTCondor handles both cases, but the job setup is different. The container image, the submit file, the executable script, and the file transfer strategy all change depending on whether you are submitting one job or many. Getting the setup wrong is the most common source of job failures, and the errors are not always obvious.
This vignette walks through both modes side by side using the same analysis, a summary and visualization of the Palmer Penguins dataset. In single mode, the analysis runs once over the full dataset. In multiple mode, it runs once per species, producing independent results for Adelie, Chinstrap, and Gentoo penguins.
The R script is identical in both cases. What changes is how the data gets to the script and how the surrounding infrastructure is configured.
The analysis
The analysis is simple by design so the focus stays on the
infrastructure. The R script loads a CSV file, computes a grouped
summary of body mass and flipper length, produces a scatterplot, and
writes both outputs to an output/ folder.
The script uses toolero::detect_execution_context() to
resolve the input file path. In an interactive RStudio session, the path
is hardcoded for convenience. When rendered via Quarto, it comes from
the YAML params block. When run via Rscript on
HTCondor, it comes from the first command-line argument:
context <- toolero::detect_execution_context()
input_file <- switch(context,
interactive = "data-raw/sample.csv",
quarto = params$input_file,
rscript = commandArgs(trailingOnly = TRUE)[1]
)
penguins <- toolero::read_clean_csv(input_file)The output section writes results to a relative path:
output_dir <- "output"
dir.create(output_dir, showWarnings = FALSE, recursive = TRUE)
toolero::write_clean_csv(p_stats, file.path(output_dir, "results.csv"),
overwrite = TRUE)
ggplot2::ggsave(filename = file.path(output_dir, "plot.png"), plot = p_plot)This is portable. "output/" resolves correctly in
RStudio, in quarto render, and on HTCondor, because the
executable script sets the working directory to HTCondor’s writable
scratch space before calling Rscript. output/
is the folder name all three packages in the family use, which is what
lets the same line of R work in all three places.
The job manifest and the submission state
Two files in the submitr workflow have “manifest” in their name. They serve different purposes and should not be confused, so the family gives each its own name.
The job manifest is a CSV file produced by
toolero::write_by_group(manifest = TRUE). It lists the
subset data files created when splitting a dataset by a grouping column.
htc_gen_submit() reads it via the queue_from
argument to produce the subdatasets.csv that HTCondor uses
to dispatch one job per subset. The job manifest is only relevant in
multiple mode.
The submission state is kept in
htc-manifest.yml, written by submitr into your
project beside htc.cfg. It accumulates metadata as you work
through the pipeline, and each step after the first also reads back what
an earlier one wrote. htc_gen_submit() records the submit
file, the mode, the results name, and in multiple mode the subset names;
htc_gen_executable() records the analysis script and the
executable; htc_upload() records the remote directory it
sent files to; htc_submit() resolves the submit file and
remote directory from those records when you do not supply them, and
records the cluster ID HTCondor assigned. At the end,
htc_download() reads all of it to work out which files to
move, with no glob patterns or manual file lists required.
Because the submission state is a file rather than something held in
the R session, it survives restarting R. That matters more than it may
sound. A job worth sending to CHTC usually takes a while, so you submit
it in one sitting and collect it in another, and in between you close
RStudio or your laptop sleeps. Calling htc_start() again
does not disturb it.
| Job manifest | Submission state | |
|---|---|---|
| What is it | A CSV file (manifest.csv) |
A YAML file (htc-manifest.yml) |
| Created by | toolero::write_by_group() |
htc_gen_submit(), htc_gen_executable(),
htc_submit()
|
| Contains | Subset filenames and row counts | Mode, script stem, results name, subset names, cluster ID, remote path |
| Used by | htc_gen_submit(queue_from = ...) |
htc_upload() and htc_download()
|
| Lives | On disk in the project directory | On disk beside htc.cfg
|
| Survives a restart | Yes | Yes |
| Relevant in | Multiple mode only | Both modes |
Single mode: one dataset, one job
In single mode, the full dataset is baked into the container image at
build time, but the R script is not – it travels to the execute node as
an uploaded job input file, the same file every time, so editing the
analysis only means re-uploading and resubmitting rather than rebuilding
and repushing the image. HTCondor runs the job, the script writes
results to output/, the executable script tars them up, and
HTCondor transfers the tarball back to the submit node.
Building the container
The Dockerfile includes the data file via the data_file
argument. It does not need a code_file argument, since the
analysis script is not baked in:
containr::generate_dockerfile(
r_version = "4.5.0",
data_file = "data-raw/sample.csv",
comments = TRUE,
verbose = TRUE
)This produces a COPY instruction that preserves the
local directory structure inside the container for the data file
only:
Build and push the image:
containr::build_image(
tag = "registry.doit.wisc.edu/your.netid/penguins-analysis:1.0.0",
tool_preference = "docker"
)
containr::push_image(
image_id = "abc123",
netid = "your.netid",
project = "penguins-analysis",
tag = "1.0.0",
tool_preference = "docker",
check_login = FALSE
)Generating the submit file and executable
submitr::htc_gen_submit(
output_file = "analysis.sub",
container_image = "registry.doit.wisc.edu/your.netid/penguins-analysis:1.0.0",
executable = "analysis.sh",
r_script = "R/analysis.R",
input_files = "R/analysis.R"
)
submitr::htc_gen_executable(
r_script = "R/analysis.R",
output_file = "analysis.sh",
data_files = "data-raw/sample.csv"
)r_script appears in both calls and has a different job
in each. htc_gen_executable() uses it to write the
Rscript line. htc_gen_submit() uses it to
derive the results tarball name, since it never opens the executable
script and so cannot infer what the job will produce, and warns if its
basename is missing from input_files. Passing the same
value to both generator calls is what guarantees the name the job builds
is the name the submit file asks for; passing it to
input_files as well is what gets the script itself onto the
execute node, since it is not part of the image.
The generated .sh file:
#!/bin/bash
set -euo pipefail
cd "${_CONDOR_SCRATCH_DIR:-$PWD}"
mkdir -p output
Rscript analysis.R /home/data-raw/sample.csv
tar -czf analysis-results.tar.gz outputThe script changes to the scratch directory, which is writable,
creates the output folder, and runs the R script. Notice the two
different kinds of path on that Rscript line:
analysis.R is bare, because it was transferred in as a job
input file and lands flat in the scratch directory regardless of its
R/ prefix on your machine;
/home/data-raw/sample.csv is absolute, because the data
file actually is baked into the image under home_dir. The
script then tars the results. The tarball lands in the scratch directory
where HTCondor expects to find it. Its name comes from
r_script with the directory and extension stripped, so
R/analysis.R gives
analysis-results.tar.gz.
The generated .sub file:
container_image = docker://registry.doit.wisc.edu/your.netid/penguins-analysis:1.0.0
universe = container
executable = analysis.sh
should_transfer_files = YES
when_to_transfer_output = ON_EXIT
transfer_input_files = R/analysis.R
transfer_output_files = analysis-results.tar.gz
request_cpus = 1
request_memory = 4GB
request_disk = 4GB
queue 1
transfer_input_files = R/analysis.R is what gets the
analysis script onto the execute node – the data file does not need an
entry here because it is already inside the image.
Submitting and downloading
submitr::htc_start()
submitr::htc_upload()
job <- submitr::htc_submit("analysis.sub")
submitr::htc_status(cluster_id = job, watch = TRUE)
submitr::htc_download()Neither htc_upload() nor htc_download()
takes arguments here, because the submission state has everything they
need. htc_gen_submit() recorded that this is a single-mode
job with analysis-results.tar.gz as its output, which files
were generated, and that R/analysis.R needs to travel as an
input file, so htc_upload() sends the submit file, the
executable, and the analysis script together. htc_submit()
recorded the cluster ID and the remote directory. From those,
htc_download() constructs the file list: one tarball plus
three log files, {cluster_id}-0-job.log, .err
and .out.
You can also be explicit if you prefer:
submitr::htc_download(cluster_id = job)
submitr::htc_download(files = "*.tar.gz", local_path = "downloads/")Multiple mode: one analysis, three species
In multiple mode, the same R script runs independently on each subset of the data. As in single mode, the container holds only the software environment – neither the R script nor the data is baked in. The script is a shared input file uploaded once and transferred to every job; each subset file is transferred at runtime by HTCondor as that job’s own input.
The motivation is straightforward: rather than analyzing all three penguin species together, you want to analyze each one independently. Maybe the analysis is computationally expensive, or maybe the subsets come from different sources and should be processed in isolation. HTCondor runs one job per subset, in parallel, across available compute resources.
Splitting the data
Use toolero::write_by_group() to split the dataset by
species and produce a job manifest:
penguins <- toolero::read_clean_csv("data-raw/sample.csv")
toolero::write_by_group(
penguins,
group_col = "species",
output_dir = "data/jobs",
manifest = TRUE
)This produces three CSV files (adelie.csv,
chinstrap.csv, gentoo.csv) and a job manifest
(manifest.csv) listing them.
If a project splits more than one dataset, pass prefix
as well. Uploaded files land in a single flat directory on the access
point, so two datasets split on the same grouping column would produce
the same subset filenames and the second set would overwrite the
first.
Building the container
The container needs neither the R script nor the data – both are transferred at runtime:
containr::generate_dockerfile(
r_version = "4.5.0",
comments = TRUE,
verbose = TRUE
)No code_file or data_file argument. The
Dockerfile carries only R and your packages, with nothing
project-specific copied in.
Build and push as before.
Generating the submit file and executable
submitr::htc_gen_submit(
output_file = "analysis.sub",
container_image = "registry.doit.wisc.edu/your.netid/penguins-analysis:2.0.0",
executable = "analysis.sh",
r_script = "R/analysis.R",
input_files = "R/analysis.R",
mode = "multiple",
queue_from = "data/jobs/manifest.csv"
)
submitr::htc_gen_executable(
r_script = "R/analysis.R",
output_file = "analysis.sh",
mode = "multiple"
)htc_gen_submit() reads the job manifest via
queue_from, extracts the subset filenames, and writes
subdatasets.csv alongside the submit file. It also records
the mode, subset names, script stem, input_files, and the
full local paths of the subsets in the submission state, which is how
htc_upload() later knows to send them and
htc_download() knows what to bring back.
The generated .sh file:
#!/bin/bash
set -euo pipefail
cd "${_CONDOR_SCRATCH_DIR:-$PWD}"
mkdir -p output
Rscript analysis.R ${1}
tar -czf analysis-${1%.*}-results.tar.gz outputAs in single mode, analysis.R is a bare relative name:
the script is a transferred input file, not a baked-in one, so it lands
flat in the scratch directory alongside this executable and
${1}, the per-job subset file.
${1} is the subset filename passed by HTCondor, for
example adelie.csv. The R script receives it as
commandArgs(trailingOnly = TRUE)[1]. The tarball name
combines the script stem with the subset stem, so the Adelie job
produces analysis-adelie-results.tar.gz.
${1%.*} is the shell’s way of stripping the extension:
without it the name would contain .csv, describing a file
that is not a CSV.
The generated .sub file:
container_image = docker://registry.doit.wisc.edu/your.netid/penguins-analysis:2.0.0
universe = container
executable = analysis.sh
arguments = $(file)
should_transfer_files = YES
when_to_transfer_output = ON_EXIT
transfer_input_files = R/analysis.R, $(file)
transfer_output_files = analysis-$Fn(file)-results.tar.gz
request_cpus = 1
request_memory = 4GB
request_disk = 4GB
queue file from subdatasets.csv
Key differences from single mode: arguments = $(file)
passes the subset filename to the executable,
transfer_input_files now carries two kinds of file –
R/analysis.R, shared across every job, and
$(file), substituted per job from
subdatasets.csv – and
queue file from subdatasets.csv submits one job per line in
the file.
$Fn(file) is worth pausing on. It is an HTCondor submit
macro that expands a variable to its file name with the directory and
extension removed, so adelie.csv becomes
adelie. It is the submit language’s counterpart to the
shell’s ${1%.*}, and both are here for the same reason: the
shell script has to create the tarball under a name, and the
submit file has to declare the same name, or HTCondor looks for
a file that was never produced.
Submitting and downloading
submitr::htc_start()
submitr::htc_upload()
job <- submitr::htc_submit("analysis.sub")
submitr::htc_status(cluster_id = job, watch = TRUE)
submitr::htc_download()This is where the submission state pays off.
htc_upload() sends the submit file, the executable, the
analysis script, subdatasets.csv, and all three subset
files without being told, because htc_gen_submit() recorded
where they are. On the way back, there are 12 files to retrieve across
three species and three log types per job, and
htc_download() constructs the full list automatically: it
reads the subset names and script stem from the submission state,
composes analysis-adelie-results.tar.gz,
analysis-chinstrap-results.tar.gz and
analysis-gentoo-results.tar.gz, and generates the nine log
file names from the cluster ID and process count.
One call, no globs, no guessing:
# These are all equivalent
submitr::htc_download()
submitr::htc_download(cluster_id = job)Side-by-side comparison
What lives in the container
| Single mode | Multiple mode | |
|---|---|---|
| R + packages + system libs | Yes | Yes |
| R script | No | No |
| Data files | Yes | No |
The R script is never baked into the image, in either mode – it always travels as an uploaded job input file. Only reference data that genuinely belongs with the image, such as the full dataset in single mode, is baked in.
The .sub file
| Directive | Single mode | Multiple mode |
|---|---|---|
container_image |
docker://... |
docker://... |
universe |
container |
container |
executable |
analysis.sh |
analysis.sh |
arguments |
(absent) | $(file) |
transfer_input_files |
R/analysis.R |
R/analysis.R, $(file) |
transfer_output_files |
analysis-results.tar.gz |
analysis-$Fn(file)-results.tar.gz |
queue |
queue 1 |
queue file from subdatasets.csv |
The .sh file
| Line | Single mode | Multiple mode |
|---|---|---|
| Working directory | cd "${_CONDOR_SCRATCH_DIR:-$PWD}" |
cd "${_CONDOR_SCRATCH_DIR:-$PWD}" |
| Output folder | mkdir -p output |
mkdir -p output |
| Run script | Rscript analysis.R /home/data-raw/sample.csv |
Rscript analysis.R ${1} |
| Package results | tar -czf analysis-results.tar.gz output |
tar -czf analysis-${1%.*}-results.tar.gz output |
Notice the run-script line mixes two kinds of path in single mode:
analysis.R bare, because it was transferred as a job input,
and /home/data-raw/sample.csv absolute, because that file
really is baked into the image.
The R script
The R script is identical in both modes. The only difference is where the input file path comes from:
| Context | Single mode | Multiple mode |
|---|---|---|
| Interactive (RStudio) | Hardcoded path | Hardcoded path |
| Quarto render | params$input_file |
params$input_file |
| Rscript (HTCondor) | Absolute path from .sh
|
${1} from HTCondor |
In both cases, detect_execution_context() resolves the
right source and commandArgs(trailingOnly = TRUE)[1] picks
up whatever the .sh script passes. The R script doesn’t
know whether it’s running in single or multiple mode.
Files uploaded to the submit node
| Single mode | Multiple mode | |
|---|---|---|
.sub file |
Yes | Yes |
.sh file |
Yes | Yes |
| R script | Yes | Yes |
subdatasets.csv |
No | Yes |
| Subset data files | No | Yes |
| Full dataset | No | No |
The R script is uploaded in both modes, since it is never baked into
the image; the full dataset is uploaded in neither, since single mode
bakes it into the image and multiple mode transfers only the per-job
subsets instead. In both modes htc_upload() works this out
from the submission state, so the column above describes what it sends
rather than what you have to type.
Files downloaded after the job
| Single mode | Multiple mode | |
|---|---|---|
| Result tarballs | 1 (analysis-results.tar.gz) |
1 per subset (3 total) |
| Log files | 3 ({cluster}-0-job.{log,err,out}) |
3 per subset (9 total) |
| Total files | 4 | 12 |
How htc_download() resolves files
In both modes, htc_download() reads the submission state
that was built automatically during the workflow. No file lists or glob
patterns are needed.
| Single mode | Multiple mode | |
|---|---|---|
| Submission state knows | Tarball name, cluster ID, remote path | Script stem, subset names, cluster ID, remote path |
| Files resolved | 1 tarball + 3 logs = 4 files | 3 tarballs + 9 logs = 12 files |
| Researcher types | htc_download() |
htc_download() |
The same zero-argument call works for both modes because the submission state captures the difference.
How the two files relate
The job manifest and the submission state are connected but distinct.
The job manifest feeds into the submission state: when
htc_gen_submit() reads the job manifest via
queue_from, it extracts the subset filenames and stores
them in the submission state. From that point on, the submission state
carries the subset names forward so that htc_upload() can
send the right files and htc_download() can reconstruct the
tarball names without re-reading anything.
toolero::write_by_group()
|
v
manifest.csv (job manifest, a CSV on disk)
|
v
htc_gen_submit(queue_from = "manifest.csv")
|
+---> subdatasets.csv (sent to HTCondor)
+---> htc-manifest.yml (submission state: subsets, script stem, mode)
|
v
htc_gen_executable()
+---> htc-manifest.yml (executable and script recorded)
|
v
htc_submit()
+---> htc-manifest.yml (cluster ID, remote path added)
|
v
htc_download() (resolves all files automatically)
A note on the results naming change
If you followed an earlier version of this vignette, the tarball
names have changed. A multiple-mode job over adelie.csv now
produces analysis-adelie-results.tar.gz where it used to
produce adelie.csv-results.tar.gz, and the output folder is
output/ rather than results/. Single-job names
are unchanged.
This is a clean break rather than a compatibility layer. If you have
results sitting on the access point from a job submitted with an earlier
version, download them before upgrading, because
htc_download() will look for names those jobs never
created. The README has the full rationale.
When to use which mode
Use single mode when your analysis processes one dataset as a unit. The full dataset is baked into the image, and only the analysis script itself is transferred at runtime. This is the simplest path and the right starting point for a first CHTC job.
Use multiple mode when you need to run the same analysis independently across subsets of the data. The data files are transferred at runtime, one per job. This scales naturally: adding more subsets means more jobs, not more configuration.
Start with single mode to confirm the analysis runs correctly on CHTC. Switch to multiple mode when you are confident the container, script, and results pipeline are working.