Benchmarks
Performance
One file (7.3 GB Illumina WES), Apple M1 Pro, from the full report. Java FastQC v0.13.0 is about 3x faster than v0.12.1 at its default four threads, thanks to parallel processing contributed upstream from this project. Speedups are relative to v0.12.1.
Default threads
Single thread (-t 1)
Peak memory (RSS), default threads
Rust v1.0 is v1.0.1 and Rust v1.1 is v1.1.0. Defaults are one thread per
file for Java v0.12.1 and Rust v1.0, -t 4 for Java v0.13.0 and up to 6 threads for
Rust v1.1. Java ran on OpenJDK 21. Measured October 2026; Java v0.13.0 at -t 1 is the best of two separate runs, as the report doesn't include it. Rust v1.1 at -t 1 and the thread scaling below were measured after the gzip decoder began counting towards -t.
Where the speed comes from
Each change was found by CPU profiling and checked for byte-identical output.
Reading input
-
.fastq.gzis decompressed by a zlib-rs decoder (viarapidgzip-core) on its own thread, so it overlaps with the analysis. - FASTQ lines are parsed straight into reused record buffers with a SIMD newline search. No per-read allocation and no garbage collection.
- A 1 MiB read buffer keeps hand-offs from the decoder thread rare.
Per-read analysis
-
Every adapter is found in a single SIMD pass over each read
(
aho-corasick's Teddy), 3.6–5.2x faster than one search per adapter. - Base counts and per-read quality minimums run as SIMD loops, and a 256-byte lookup table replaces branch chains in the per-base modules.
- In-place uppercase conversion and allocation-free tile ID parsing.
Parallelism
- Each file's QC modules are split across worker threads, so they process the same reads side by side. Output is identical at any thread count.
- Modules are assigned to workers by measured cost, so heavy ones don't share a worker.
- Batches are capped by bytes as well as reads, keeping memory flat on long reads. See Multi-threading.
Multi-threading
When a file gets two or more threads, it runs through a pipeline: a reader thread parses
records and hands batches to analysis workers, each of which owns a subset of the QC
modules. Every module still sees every read in order on one thread, so output is the same
at any thread count. Without -t, FastQC-Rust uses the available CPUs, up to 6.
Wall clock time by -t, one 7.3 GB WES file (M1 Pro)
Each file's gzip decoder counts towards -t, so -t 1 decodes
and analyses on a single thread. A single file stops getting faster at around
-t 4, where the analysis outpaces the file's one gzip decoder. Past that point, extra cores are better spent on
more files at once.
Given two or more threads, .fastq.gz decompression runs on its own thread via
rapidgzip-core with a
zlib-rs inflate, so it overlaps with the
analysis. rapidgzip can split a file across decoders (--decompress-threads),
but a single-member .gz can only be split speculatively, which costs several
times the CPU and hundreds of MB for little gain, so one decoder per file is the default.
Test dataset
The full report covers 10 files (98 GB): Illumina WES, Element AVITI, ONT PromethION long reads and BAM, each run through Java v0.12.1, Java v0.13.0, Rust v1.0 and Rust v1.1. It has per-file wall-clock time, CPU time and memory, output correctness checks, a CPU profile, and whether these optimisations could be back-ported to Java.