
Getting started with toolero
Erwin Lares
Created 2026-04-30 | Last updated 2026-10-01
Source:vignettes/toolero-intro.Rmd
toolero-intro.Rmd
Background and motivation
toolero grew out of a recurring observation made while
teaching and supporting researchers at UW-Madison: the habits that make
a project reproducible, shareable, and maintainable are easiest to adopt
at the very beginning – and hardest to retrofit once a project is
already underway.
The package is heavily influenced by the workflows taught in
workshops run by The Carpentries
and the UW-Madison
Libraries. Those workshops emphasize consistent project
organization, version control, and reproducible data practices as
foundational skills – not advanced topics. toolero tries to
operationalize those principles into a small set of functions that
reduce the friction of doing the right thing from the start.
toolero is also the first package in From the
Notebook to the Cluster, a three-package family that covers the
full arc from local project setup to high-throughput computing:
toolero organizes and scaffolds, containr
freezes the software environment into a container, and
submitr sends the work to CHTC and brings results back.
This vignette covers toolero on its own. You do not need
the other two packages to benefit from it, but a project started with
toolero is already most of the way to being containerizable
and submittable when that time comes.
toolero can optionally apply UW-Madison branding to a
project’s Quarto output – logo, header, footer, and stylesheet – but
branding is opt-in and off by default. If you are not at UW-Madison, or
would rather use your own look, the rest of the package works exactly
the same way without it.
Who is this for?
toolero is designed for researchers and analysts
who:
- Work primarily in R and use RStudio as their IDE
- Write reports or analyses in Quarto
- Want consistent, reproducible project structure without having to think about it every time
- Split a dataset into pieces and apply the same analysis to each
- Need to record what an analysis produced, and whether each output write succeeded
- May need to publish content to the UW-Madison Knowledge Base
The package is intentionally small. It does not try to be comprehensive. It tries to make the right defaults easy to reach for from the first line of code.
Installation
You can install toolero from CRAN:
install.packages("toolero")Or install the development version from GitHub:
pak::pak("erwinlares/toolero")Project setup: init_project() and
create_qmd()
These two functions are designed to be used together, in order.
init_project() creates the scaffold;
create_qmd() populates it with a working Quarto
document.
Starting with init_project()
Starting a new R project usually means the same manual steps every
time: create a folder, set up an RStudio project, create subdirectories
for data and scripts, initialize renv, initialize
git. None of these steps is hard on its own, but skipping
any of them – especially early on – tends to create friction later.
init_project() handles all of this in a single call:
library(toolero)
init_project(path = "~/Documents/my-project")This creates a new RStudio project at the specified path with the following folder structure already in place:
my-project/
├── data-raw/ # inputs as they arrived, never edited in place
├── data/ # analysis-ready data, produced from data-raw/
├── R/ # code, including scripts derived from .qmd documents
├── scripts/ # hand-written, standalone utility scripts
├── output/
│ ├── figures/ # generated visualizations
│ └── tables/ # generated tables
├── reports/ # rendered .qmd/.html output meant to be shared
├── _toolero.yml # records the folder set and naming conventions
└── .here # marks the project root for here::here()
Why this structure? The folder layout is opinionated but not arbitrary. Separating
data/fromdata-raw/makes it clear which files are original and which have been processed. KeepingR/distinct fromscripts/encourages moving reusable logic into functions over time:R/is where derived scripts land, whether fromqmd_to_r()or from the post-render purl hook, whilescripts/is for utilities you write and maintain by hand.
Every folder init_project() creates that is still empty
when the call finishes also gets a zero-byte .gitkeep, so
the layout survives an opening git commit even before anything has been
written into it. One folder cannot be suppressed, by
custom_folders or by a config file:
R/. usethis::create_project() creates it
unconditionally, so it is present in every project regardless of the
folder set you asked for.
By default, init_project() also sets up
renv and git. This means the project is
reproducible and version-controlled from the first commit.
Why
renvandgitby default?renvensures that the packages your project depends on are recorded and reproducible.gitprovides a full history of changes. Both are much easier to set up at the start than to retrofit later.init_project()usesrenv::scaffold()rather thanrenv::init(), so setting uprenvin the new project does not disturb the R session you called it from. Nothing is snapshotted at creation time – a brand-new project has no code in it yet – so runrenv::snapshot()yourself once the analysis exists and before containerizing.
Setting up renv this way has a quieter side effect. R
reads only one .Rprofile per session – the project’s own if
the working directory has one, ~/.Rprofile otherwise – and
the .Rprofile renv::scaffold() writes becomes
that one. Whatever you keep in your personal ~/.Rprofile (a
handful of options, a favorite helper function) stops loading for this
project, with nothing telling you it happened.
use_rprofile = TRUE appends a guarded block to the
project’s .Rprofile, after renv’s own activation line, that
sources ~/.Rprofile if it exists:
init_project(path = "~/Documents/my-project", use_rprofile = TRUE)It defaults to FALSE, because it is genuinely in tension
with what renv is for: a project that quietly re-sources
your personal environment is not fully isolated from it anymore. Turn it
on when losing your own .Rprofile would cost you more than
that isolation is worth.
If your project needs folders beyond the defaults, or you want to
drop one of them, custom_folders adds and removes without
having to restate the whole set. A "-" prefix removes a
folder from the standard set:
init_project(
path = "~/Documents/my-project",
custom_folders = c("notebooks", "presentations", "-output/figures")
)For a layout that differs enough from the defaults that restating it
on every call gets tedious, generate_project_config()
writes a skeleton YAML config you can edit once and reuse:
generate_project_config("linguistics-project.yml", path = "~")
init_project(
path = "~/Documents/my-project",
config = "~/linguistics-project.yml"
)Whichever route you take, init_project() records the
resolved folder set and naming conventions it ended up with in
the project config, _toolero.yml, at the project
root. Commit that file. It is what lets check_project(),
and later containr and submitr, know how your
project is laid out without being handed the same configuration
again.
To apply UW-Madison branding assets to the project:
init_project(
path = "~/Documents/my-project",
branding = "uw-madison"
)This creates an assets/ folder and populates it with
logo.png, favicon.png,
header.html, footer.html, and
styles.css under UW-Madison RCI branding. Pass
branding = TRUE instead for generic placeholder assets
under the same standardized names, or leave branding at its
default, "none", to skip the assets/ folder
entirely. All three modes are interchangeable from
create_qmd()’s point of view, since it always looks for
those five standardized filenames.
Adding a Quarto document with create_qmd()
Once the project exists, create_qmd() adds a working
Quarto document to it. filename is the one argument you
always have to supply; path defaults to the current
directory. The function has two modes controlled by
include_examples, and several optional features that can be
mixed and matched.
With examples (the default)
When include_examples = TRUE (the default),
create_qmd() scaffolds a complete, runnable analysis
project:
create_qmd("analysis.qmd", path = "~/Documents/my-project")This creates:
-
analysis.qmd– a Quarto document with a fully populated YAML header (including aparams:block), a grouped summary, a scatterplot, and a results-saving section. Input resolution usesresolve_input_path(), so the document is ready to render immediately and works unchanged whether you run it interactively, render it with Quarto, or extract it to a script. -
data-raw/sample.csv– a subset of the Palmer Penguins dataset to develop against. Theparamsblock in the YAML header points at this file. -
assets/logo.png– a placeholder logo that reads “your logo goes here,” unless a logo already exists (for instance frominit_project(branding = )), in which case it is left alone.
The idea is that you can render the document as-is, see results, and then progressively replace the sample analysis with your own. The sample data, the analysis blocks, and the results-saving pattern are all working examples you can study before modifying.
Without examples
When include_examples = FALSE, create_qmd()
creates a minimal skeleton with no sample data and no pre-filled
analysis:
create_qmd(
"analysis.qmd",
path = "~/Documents/my-project",
include_examples = FALSE
)This creates a Quarto document with the YAML header (title, author,
format settings) and a setup chunk that loads
library(toolero). The body has a single
## Introduction heading and an HTML comment prompting you
to add your content. No params block, no analysis code, no
references to sample data. No data-raw/ folder is created,
and no placeholder logo is placed in assets/. The document
is a blank canvas with just enough structure to render.
Use this mode when you already know what your analysis looks like and don’t need the worked example as a starting point.
Custom styling
The use_style argument controls whether CSS and
header/footer assets are wired into the YAML. It works independently of
include_examples:
# Blank document with UW branding (assumes init_project(branding = "uw-madison"))
create_qmd(
"report.qmd",
path = "~/Documents/my-project",
include_examples = FALSE,
use_style = TRUE
)
# Blank document with custom branding from a different directory
create_qmd(
"report.qmd",
path = "~/Documents/my-project",
include_examples = FALSE,
use_style = "my-branding/"
)When use_style = TRUE, the function looks in
assets/ for files by their standardized names –
styles.css, header.html,
footer.html – and wires up whichever are present:
styles.css as css:, header.html
as include-before-body:, footer.html as
include-after-body:. When use_style is a
directory path, it looks there instead. Any subset of the three may be
present; only files that exist are injected.
Styling assets themselves come from
init_project(branding = ), not from
create_qmd(). create_qmd() only wires up what
is already in assets/.
The purl hook
create_qmd() can also scaffold a post-render hook that
extracts R code from the rendered document into a companion
.R file on every render, kept in step with the
.qmd automatically. This is opt-in, via
use_purl = TRUE:
create_qmd(
"analysis.qmd",
path = "~/Documents/my-project",
use_purl = TRUE
)Turning it on:
- Stamps the document’s own YAML header with
purl: true(orpurl: falsewhenuse_purl = FALSE, so a document can positively confirm it should be skipped rather than merely lacking an opinion). - Scaffolds
R/purl.R, which purls every document in the project whose own header carriespurl: true. - Wires
R/purl.Rinto_quarto.yml’s post-render hook, creating_quarto.ymlfrom the package template if it does not exist yet, or merging the hook into an existing file’sproject:block if it does – unless that file already declares awebsite,book, ormanuscriptproject, in which case the automatic wiring is skipped with a warning explaining how to add it by hand.
This is useful for sharing the analysis as a script, running it on a
remote cluster via submitr, or archiving the code
independently of the document. use_purl defaults to
FALSE, so a call with no other arguments creates a plain
.qmd and nothing else.
Pre-populating the YAML header
The header_defaults argument accepts a path to a YAML
file whose top-level keys overwrite the corresponding keys in the
template. Keys not present in the file are left exactly as the template
wrote them – including their quoting, indentation, and any comments,
since every header edit create_qmd() makes is line-based
rather than a parse-and-rewrite. (This argument was named
yaml_data before v0.6.0; the old name still works but is
deprecated and will be removed in v0.7.0 – update to
header_defaults.)
create_qmd(
"analysis.qmd",
path = "~/Documents/my-project",
header_defaults = "~/my-metadata.yml"
)Where my-metadata.yml might look like:
Rather than hand-writing that file for every project,
generate_profile() writes a starting-point YAML skeleton –
pre-filled with placeholders and explanatory comments for the personal
information and formatting preferences that tend to stay identical
across every document you create:
generate_profile(filename = "~/my-metadata.yml")Fill in the placeholders once, and reuse the same file as
header_defaults across projects; filename has
no default, so keeping more than one profile – a personal one and a work
one, say – under different filenames is a normal thing to do, not a
workaround.
This works with both include_examples = TRUE and
FALSE, and composes with use_style and
use_purl.
Summary of what gets created
| File | include_examples = TRUE |
include_examples = FALSE |
|---|---|---|
analysis.qmd |
Full example with analysis blocks and params
|
Skeleton with YAML and empty body |
data-raw/sample.csv |
Yes | No |
assets/logo.png |
Yes, unless one already exists | No |
_quarto.yml and R/purl.R
|
Only when use_purl = TRUE
|
Only when use_purl = TRUE
|
purl: stamp in the YAML header |
Only when use_purl is supplied |
Only when use_purl is supplied |
| CSS/header/footer in YAML | Only when use_style is set |
Only when use_style is set |
Working with data
These functions address common friction points in day-to-day data
work, and the split-apply pair scales an analysis from a single dataset
to many independent pieces. They are general-purpose utilities – useful
in any R project, not just ones set up with toolero.
Reading and writing clean data: read_clean_csv() and
write_clean_csv()
read_clean_csv() combines
readr::read_csv(), janitor::clean_names(), and
optionally tidyr::drop_na() into a single call. The goal is
to get from a raw CSV to a clean, analysis-ready tibble in one step.
data <- read_clean_csv(
"data-raw/my-file.csv",
na = c("", "NA", "N/A", "."),
drop_na = TRUE,
summary = TRUE
)Column names are automatically converted to lowercase with
underscores. The na argument lets you name additional
missing-value codes beyond the readr defaults.
drop_na accepts TRUE (drop any incomplete
row), FALSE (keep everything, the default), or a character
vector of columns to check. summary = TRUE prints row and
column counts, how many names were cleaned, and how many values are
missing.
write_clean_csv() is the natural counterpart,
reinforcing the convention that data-raw/ holds original
inputs and data/ holds analysis-ready outputs. If the data
frame’s names are not already clean, it applies
janitor::clean_names() before writing and warns about the
affected columns, so the output file always has consistent names:
write_clean_csv(data, "data/clean.csv")Splitting and applying: write_by_group() and
run_by_group()
When an analysis needs to run once per group – once per species, once
per site, once per participant – write_by_group() and
run_by_group() form the split-apply pair at the heart of
that workflow. Splitting and applying are deliberately separate steps,
so you can iterate on the analysis function without re-splitting the
data each time.
write_by_group() partitions a data frame by one or more
grouping columns, writes one CSV per group with sanitized filenames, and
optionally writes a job manifest recording each group’s value, row
count, and file path:
write_by_group(
data = penguins,
group_col = "species",
output_dir = "data/jobs",
manifest = TRUE
)group_col also accepts more than one column, writing one
file per combination that actually appears in the data
(adelie--female.csv, and so on), and prefix
prepends a namespace to every filename – useful before a high-throughput
run, where submitr reduces each path in the job manifest to
its basename(), and short group names from two different
datasets could otherwise collide in one flat directory.
If a project was scaffolded with init_project(),
output_dir does not need to be typed out at all.
write_by_group() accepts a config argument – a
path to the project config, _toolero.yml – and resolves
output_dir from its split_dir convention when
output_dir is not supplied directly. An explicit
output_dir always wins; config only fills in
what you didn’t already specify:
write_by_group(
data = penguins,
group_col = "species",
config = "_toolero.yml",
manifest = TRUE
)run_by_group() reads that job manifest (or a named list
of data frames already in memory), applies your function to each subset,
and assembles the results into a single tibble:
summarise_species <- function(data) {
dplyr::summarise(
data,
n = dplyr::n(),
mean_mass = mean(body_mass_g, na.rm = TRUE),
mean_flipper = mean(flipper_length_mm, na.rm = TRUE)
)
}
results <- run_by_group(
manifest = "data/jobs/manifest.csv",
.f = summarise_species
)If .f returns a data frame, the results come back as one
flat tibble with a group-id column prepended; anything else – a model, a
plot, a file path – comes back as a nested tibble with a list-column.
For analyses that are slow or independent across groups,
workers runs them in parallel via furrr. A
bare column name passed through ...
(x = flipper_length_mm) works sequentially without any
special handling, but has to be moved inside .f to survive
workers > 1, since parallel execution has to serialize
every argument to send it to a worker session, and a symbol that only
means something inside the data has nothing to serialize.
Recording what an analysis produced: save_output() and
generate_manifest()
Once an analysis has results, save_output() writes each
object to disk via a function you supply, and appends a row to a
project-level accumulator recording the path, the object’s class, the
function used, and whether the write succeeded:
save_output(
results,
here::here("output", "results.rds"),
.f = saveRDS
)The path is built with here::here(), which starts at the
project root (the folder holding the .here file
init_project() writes) rather than at wherever the code
happens to be running. That matters as soon as a document lives in
reports/: rendering runs its code with
reports/ as the working directory, so a bare
"output/results.rds" would land in
reports/output/. On a cluster, where there is no project
root to find, here::here() uses the job’s working
directory, which is exactly where the results folder is.
save_output() records the path relative to the project
root, as output/results.rds.
At the end of the run, generate_manifest() reads that
accumulator, collapses it to one row per output file, and writes the
output record, output/project-manifest.json – a
record of what the analysis actually produced, which is useful on its
own and becomes essential once a job is running unattended on a cluster.
(The file and the function keep their historical names; the family calls
the file the output record so that it is never confused with the job
manifest write_by_group() writes.)
With no output_dir, both functions use
output/ under the project root, so they always meet at the
same accumulator.
Like write_by_group(), both save_output()
and generate_manifest() accept a config
argument that resolves output_dir from the project config
when you don’t supply it directly.
generate_manifest() also records commit in
the output record: the git commit checked out in git_root
(default ".") at the moment the record was written.
renv.lock already answers which package versions were in
play; commit answers which revision of the analysis script
produced this particular set of outputs – the one piece of provenance
package versions alone can’t supply:
generate_manifest(git_root = ".")commit comes back NULL when the project
isn’t a git repository, has no commits yet, or git isn’t
installed, so recording it is always safe to leave on.
Reading output records back: read_output_records()
The output record is a file for machines, but the question it answers
is one you will ask yourself: what did that run produce, and did
anything fail? read_output_records() reads the record and
returns a tibble with one row per artifact:
records <- read_output_records()
records[records$status == "failure", c("file_path", "error_message")]With no arguments it reads output/ under the project
root, the same folder save_output() and
generate_manifest() use. Given several output folders, it
reads each and stacks the results, with a source column
saying which folder a row came from. That is how a multi-job run comes
back from a cluster: one results folder per job. With
submitr (0.2.0 or later) it takes two calls, one to bring
the results back and one to read them:
jobs <- submitr::htc_collect()
jobs <- jobs[!is.na(jobs$output_dir), ]
records <- read_output_records(setNames(jobs$output_dir, jobs$group_id))htc_collect() knows where each job’s results landed;
read_output_records() knows what the record inside means,
because toolero wrote it. When a job crashed before
generate_manifest() ran, its accumulator is usually still
there, and read_output_records() reads that instead, with a
warning, so the rows for that job still show what was saved before it
stopped.
Auditing a project: check_project()
check_project() audits an existing project directory –
one created by init_project() or any other R project –
against the expected folder structure, an .Rproj file,
renv.lock, git, a README, and a few other common
reproducibility checks. It reads the project config
(_toolero.yml) when one exists, so a customized project
does not need to be handed the same configuration again:
Among those checks is a stale-purled-script check: every
.qmd whose header declares purl: true (see The purl hook above) gets its own row
comparing it against the .R script it should have derived.
A .qmd is the source of truth for its purled script, so a
script missing entirely, or older than the document it came from, is
reported as a warning – it means an edit was made and not yet
re-rendered, and a container or cluster job that bakes in the stale
.R file would run the old analysis with nothing to say
so.
Execution context: detect_execution_context() and
resolve_input_path()
R code often needs to behave differently depending on where it is
running – interactively in RStudio, during a quarto render,
or as a batch Rscript job on a remote cluster.
detect_execution_context() identifies which of these three
environments is active and returns one of "interactive",
"quarto", or "rscript".
For the specific, recurring case of finding the input data file,
resolve_input_path() builds on it directly: it picks the
branch that matches the current context, checks that the result is
usable, and explains what to fix when it is not.
input_file <- resolve_input_path(
interactive = "data-raw/sample.csv",
quarto = params$input_file,
rscript = commandArgs(trailingOnly = TRUE)[1]
)Only the branch matching the current context is ever evaluated, so
the params reference above is safe even under
Rscript, where params does not exist at all.
Arguments can also be omitted: the rscript branch defaults
to the first command line argument, and interactive and
quarto both fall back to the document’s own
params$input_file, so a document whose header declares
input_file can call resolve_input_path() with
no arguments at all.
In the interactive and Quarto contexts, a relative path is read from
the project root, like here::here(), so
input_file: data-raw/sample.csv works for a document in
reports/ as well as for one at the root. The
rscript branch is left as it is, since on a cluster the
path arrives as an argument that already points at the file.
This pattern is built into the template scaffolded by
create_qmd(), so you get it for free without having to
write it yourself. See vignette("detect-execution-context")
for a more detailed treatment of the problem this solves and why it is
worth solving in one place.
Knowledge Base export: generate_kb_xml()
This section is relevant only if you publish content to the UW-Madison Knowledge Base. If you do not, you can safely skip it.
The UW-Madison Knowledge Base requires content to be submitted as XML
with all visual assets embedded in the HTML body.
generate_kb_xml() automates this process entirely.
generate_kb_xml(
html_path = "docs/analysis.html",
output_dir = "exports"
)The function:
- Infers the source
.qmdfrom the HTML path (or accepts it explicitly viaqmd_path) - Re-renders the document with
embed-resources: trueso all CSS, images, and JavaScript are self-contained - Extracts metadata from the
.qmdYAML header –title→kb_title,description→kb_summary,categories→kb_keywords - Produces a
.xmlfile ready for direct KB import
This is why the description and categories
fields in the create_qmd() template matter – they flow
through automatically into the KB article metadata without any extra
work.
When importing into the KB, check the Decode HTML entity in body content option.
Citation metadata: generate_citation()
A repository’s citation metadata tends to get retyped by hand, and
drifts from a person’s actual name, affiliation, and ORCID as a result.
generate_citation() writes a CITATION.cff
skeleton – the Citation File Format
GitHub and other tools use to render a “Cite this repository” button –
optionally pre-filled from a profile written by
generate_profile() (see Pre-populating the YAML
header above):
generate_citation(profile = "~/my-metadata.yml")title, version,
repository-code, url, and license
are facts about the project rather than the person, so they are left as
placeholders (some commented out) regardless of whether
profile is supplied; date-released is filled
in with today’s date. The given-names/family-names split the Citation
File Format requires is done by breaking a profile’s name
field on its last space, which handles the ordinary two-word case but
not multi-word family names, single-word names, or family-name-first
orderings – review the generated file’s given-names and
family-names fields before relying on them.
Syntactic trees: arborize()
Outside the general research workflow, toolero also
ships arborize(), which renders a syntactic tree
description to a standalone PNG via Quarto’s Typst engine – useful for
course handouts, papers, and slides without a full LaTeX installation.
It has its own vignette, vignette("arborize"), since the
notation and sizing options deserve room of their own.
Quick reference
| Function | Brief description |
|---|---|
init_project() |
Creates a new RStudio project with a reproducible folder structure,
renv, git, a README, and optional branding
assets. Records the resolved structure in the project config,
_toolero.yml. |
generate_project_config() |
Writes a skeleton YAML project configuration file, pre-filled with the standard folders and conventions, for projects whose layout should be reused or shared. |
check_project() |
Audits an existing project against the expected folder structure,
common reproducibility hygiene checks, and stale purled .R
scripts. |
create_qmd() |
Creates a Quarto document scaffold. Can generate a full worked
example or a minimal skeleton, pre-fill YAML metadata via
header_defaults, add styling, and opt into the purl
post-render hook. |
generate_profile() |
Writes a reusable YAML skeleton of personal information and document
formatting preferences, meant to be passed as
create_qmd(header_defaults = ). |
qmd_to_r() |
Extracts the R code from a rendered Quarto document into a
standalone .R script. |
read_clean_csv() / write_clean_csv()
|
Reads or writes a CSV with janitor::clean_names()
column names, optional missing-value handling, and an optional ingest
summary. |
write_by_group() |
Splits a data frame by one or more grouping columns and writes one
CSV file per group, optionally with a job manifest.
output_dir can be resolved from _toolero.yml
via config. |
run_by_group() |
Applies a function to each group from a job manifest or a named list, sequentially or in parallel. |
save_output() |
Writes an object to disk via a user-supplied function and records
the write in a project-level accumulator. output_dir can be
resolved from _toolero.yml via config. |
generate_manifest() |
Reads the accumulator and writes the output record, a deduplicated
project-manifest.json describing everything the analysis
produced, including the git commit checked out at the time (when
available). |
read_output_records() |
Reads one or more output records back as a tibble, one row per artifact, falling back to the accumulator when a record is missing. Experimental. |
detect_execution_context() |
Detects whether code is running interactively, during
quarto render, or as an Rscript job. |
resolve_input_path() |
Resolves the input data path for the current execution context, and explains what went wrong when it cannot. |
generate_kb_xml() |
Converts a rendered Quarto HTML document into UW-Madison Knowledge
Base-ready XML with embedded resources and metadata extracted from the
source .qmd. |
generate_citation() |
Writes a CITATION.cff skeleton, optionally pre-filled
with author information from a generate_profile()
file. |
arborize() |
Renders a syntactic tree description as a standalone PNG image via Quarto’s Typst engine. |