Building an RNASeq Pipeline Image
In the Dockerfiles section, we built containers using apt-get and pip as package managers. For serious bioinformatics work, most of the tools we need live in Bioconda. The best way to access those packages inside a Docker image is to build on a base that already has a fast Conda solver built in.
In this section we will build a single, self-contained image that holds every tool needed for a standard RNA-seq pipeline:
| Tool | Purpose | Version |
|---|---|---|
| FastQC | Raw read quality control | 0.12.1 |
| fastp | Adapter trimming & QC | 0.23.4 |
| STAR | Splice-aware alignment | 2.7.11a |
| featureCounts (subread) | Read quantification | 2.0.6 |
| MultiQC | Aggregated QC report | 1.21 |
This image will later be used in the Snakemake workshop as a drop-in replacement for per-rule Conda environments.
1. Choosing a Base Image
Rather than starting from a plain Ubuntu image and manually installing Conda, we use mambaorg/micromamba. This official image ships with Micromamba, a tiny standalone reimplementation of the Conda package manager that resolves environments significantly faster than classic conda. It is the recommended base for reproducible bioinformatics images.
2. Writing the Dockerfile
Create a new directory for this project and open a blank Dockerfile:
Add the following recipe:
# Start from the official Micromamba image for fast Conda installs
FROM mambaorg/micromamba:1.5.8 # (1)!
# Copy the environment specification into the image
COPY env.yaml /tmp/env.yaml # (2)!
# Install all bioinformatics tools in the base environment
RUN micromamba install -y -n base -f /tmp/env.yaml && micromamba clean -a -y # (3)!
# Bake the conda bin directory into PATH so tools are found in any runtime
ENV PATH="/opt/conda/bin:$PATH" # (4)!
# Activate the base environment for every subsequent command
ARG MAMBA_DOCKERFILE_ACTIVATE=1 # (5)!
# Set the default working directory
WORKDIR /data # (6)!
- Pinning the
micromambaimage version (1.5.8) ensures the same Conda solver is used on every build, preventing subtle differences in dependency resolution over time. - We copy a separate
env.yamlfile into the image instead of listing packages inline. This mirrors a real-world best practice: the environment specification is version-controlled separately from the build instructions. micromamba clean -a -yremoves all downloaded package archives and index caches from inside the image in the sameRUNlayer, keeping the final image size as small as possible.- The Micromamba entrypoint activates the conda environment when using
docker run, but tools like Apptainer/Singularity bypass the entrypoint and callbashdirectly. Adding/opt/conda/bintoPATHhere ensures tools are found regardless of how the container is launched. - Setting this build argument tells the Micromamba base image to automatically activate the
baseenvironment in every subsequentRUNstep and in the final container shell. - Any mounted volume will land here by default, so tool commands can reference relative file paths without needing full paths.
Now create the environment file alongside the Dockerfile:
channels:
- bioconda
- conda-forge
- defaults
dependencies:
- fastqc=0.12.1
- fastp=0.23.4
- star=2.7.11a
- subread=2.0.6
- multiqc=1.21
Why a separate env.yaml?
Keeping the package list in a standalone YAML file has two advantages. First, the file is human-readable and easy to update. Second, Docker's layer-caching system will only re-run the micromamba install step when env.yaml actually changes, saving minutes on incremental rebuilds.
3. Building the Image
Run the build from inside the ~/rnaseq_image directory (where both Dockerfile and env.yaml live):
Expected build time
STAR is a large package (~500 MB uncompressed). The first build typically takes 5–10 minutes depending on your connection. Subsequent builds are fast because Docker caches each layer.
Watch for the final output line:
Confirm the image appears in your local registry:
4. Testing the Image
Before using the image in a real pipeline, quickly verify that every tool is accessible and reports the expected version:
docker run --rm rnaseq-pipeline:v1 fastqc --version
docker run --rm rnaseq-pipeline:v1 fastp --version
docker run --rm rnaseq-pipeline:v1 STAR --version
docker run --rm rnaseq-pipeline:v1 featureCounts -v
docker run --rm rnaseq-pipeline:v1 multiqc --version
Each command should return a short version string with no error. If any command fails with command not found, check that your env.yaml spelling matches the exact Conda package names.
featureCounts reports its version to stderr
featureCounts -v writes the version string to standard error. If you want to capture it: docker run --rm rnaseq-pipeline:v1 featureCounts -v 2>&1
5. Pushing to Docker Hub
You now have a portable, version-locked image that encapsulates an entire RNA-seq software stack. Push it to Docker Hub now so it is available on any machine for collaborators and cluster environments. Use the same tagging and pushing steps from the previous Sharing Images section, substituting rnaseq-pipeline for the image name:
docker tag rnaseq-pipeline:v1 yourusername/rnaseq-pipeline:v1
docker push yourusername/rnaseq-pipeline:v1
Once the push completes, this image is ready to use in the Snakemake workshop, where it replaces per-rule Conda environments via a single container: directive.