FastQC: Rust vs Java Benchmark Report
Date: October 2026
Machine: Apple M1 Pro (8 performance + 2 efficiency cores, 16 GB RAM), macOS. Data on an external SSD (~880 MB/s, never the bottleneck).
Implementations:
- Java v0.12.1: the FastQC release most pipelines ran until October 2026. One thread per file, whatever
-tis set to. - Java v0.13.0: adds a parallel pipeline within each file, contributed upstream from this project. Default
-t 4. - Rust v1.0: FastQC-Rust v1.0.1, the previous release. One thread per file.
- Rust v1.1: FastQC-Rust v1.1.0 (#7). Parallel analysis within each file, parallel-capable gzip decoding, default of the available CPUs up to 6. Also measured at
-t 2: one analysis thread plus its gzip decoder thread. (These runs were made as-t 1, before the decoder counted towards-t; it is the same configuration.)
Each tool runs at its default settings, one file at a time, one run per file. Java ran on OpenJDK 21 with the launcher's defaults, except that the long-read files got --memory 4096: the default 512 MB heap is too small for them, as FastQC's own help notes. Speedups are relative to Java v0.12.1 unless stated. macOS Spotlight indexing was using about one core in the background throughout; WES results matched earlier runs on a quiet machine to within 2%.
Summary
| Data | Java v0.12.1 | Java v0.13.0 | Rust v1.0 | Rust v1.1 | Rust v1.1 -t 2 |
|---|---|---|---|---|---|
| Illumina WES (short-read, 126 bp) | 1.0x | 2.7x | 2.6x | 12.5x | 6.6x |
| Element AVITI (short-read) | 1.0x | 2.5x | 2.5x | 9.7x | 6.4x |
| ONT PromethION (long-read) | 1.0x | 2.4x | 2.1x | 9.5x | 7.3x |
| BAM | 1.0x | 2.2x | 1.9x | 3.2x | 2.3x |
Average wall-clock speedup over Java v0.12.1 per file, each tool at its default threads.
Test Dataset
| File | Size | Platform |
|---|---|---|
| SRR7890918_WES_HCC1395-EA_tumor_1.fastq.gz | 7.3 GB | Illumina WES, short-read |
| SRR7890918_WES_HCC1395-EA_tumor_2.fastq.gz | 7.5 GB | Illumina WES, short-read |
| SRR7890919_WES_HCC1395BL-EA_normal_1.fastq.gz | 6.2 GB | Illumina WES, short-read |
| SRR7890919_WES_HCC1395BL-EA_normal_2.fastq.gz | 6.3 GB | Illumina WES, short-read |
| ERR16944282_1.fastq.gz | 487 MB | Element AVITI, short-read |
| ERR16944282_2.fastq.gz | 630 MB | Element AVITI, short-read |
| ERR16944299_1.fastq.gz | 454 MB | Element AVITI, short-read |
| ERR16944299_2.fastq.gz | 597 MB | Element AVITI, short-read |
| ERR16962265.fastq.gz (not benchmarked, see below) | 21.2 GB | ONT PromethION, long-read (wheat WGS) |
| SRR37915503.fastq.gz | 57.6 GB | ONT PromethION, long-read (snake WGS) |
| GM12878_REP1.markdup.sorted.bam | 11.3 GB | Illumina, aligned BAM |
| Total | 119.6 GB | 11 files across 3 platforms + BAM |
Output Correctness
fastqc_data.txt and summary.txt from Rust v1.1 were compared module by module against Java v0.13.0, and between Rust v1.1's -t 2 and default runs. Every file is byte-identical in both comparisons.
| File | Rust v1.1 -t 6 vs Java v0.13.0 | Rust v1.1 -t 2 vs -t 6 |
|---|---|---|
| tumor_1.fastq.gz | identical | identical |
| tumor_2.fastq.gz | identical | identical |
| normal_1.fastq.gz | identical | identical |
| normal_2.fastq.gz | identical | identical |
| ERR16944282_1.fastq.gz | identical | identical |
| ERR16944282_2.fastq.gz | identical | identical |
| ERR16944299_1.fastq.gz | identical | identical |
| ERR16944299_2.fastq.gz | identical | identical |
| SRR37915503.fastq.gz | identical | identical |
| GM12878.bam | identical | identical |
Per-File Results
Wall Clock Time
| File | Java v0.12.1 | Java v0.13.0 | Rust v1.0 | Rust v1.1 | Rust v1.1 vs Java v0.12.1 | Rust v1.1 vs Java v0.13.0 | |
|---|---|---|---|---|---|---|---|
| Illumina WES (short-read, 126 bp) | |||||||
| tumor_1.fastq.gz | 7m 32s | 2m 47s | 2m 52s | 37s | 12.4x | 4.6x | |
| tumor_2.fastq.gz | 7m 32s | 2m 47s | 2m 53s | 37s | 12.4x | 4.6x | |
| normal_1.fastq.gz | 6m 27s | 2m 22s | 2m 28s | 30s | 12.8x | 4.7x | |
| normal_2.fastq.gz | 6m 26s | 2m 21s | 2m 28s | 31s | 12.5x | 4.6x | |
| Element AVITI (short-read) | |||||||
| ERR16944282_1.fastq.gz | 31s | 12s | 12s | 3.0s | 10.1x | 3.9x | |
| ERR16944282_2.fastq.gz | 33s | 13s | 13s | 3.6s | 9.1x | 3.5x | |
| ERR16944299_1.fastq.gz | 31s | 12s | 12s | 2.9s | 10.7x | 4.1x | |
| ERR16944299_2.fastq.gz | 31s | 13s | 13s | 3.5s | 8.8x | 3.7x | |
| ONT PromethION (long-read) | |||||||
| SRR37915503.fastq.gz | 49m 03s | 20m 11s | 23m 33s | 5m 10s | 9.5x | 3.9x | |
| BAM | |||||||
| GM12878.bam | 7m 54s | 3m 37s | 4m 07s | 2m 28s | 3.2x | 1.5x | |
CPU Time (user + system)
| File | Java v0.12.1 | Java v0.13.0 | Rust v1.0 | Rust v1.1 | Rust v1.1 vs Java v0.12.1 | Rust v1.1 vs Java v0.13.0 | |
|---|---|---|---|---|---|---|---|
| Illumina WES (short-read, 126 bp) | |||||||
| tumor_1.fastq.gz | 7m 31s | 9m 21s | 2m 52s | 1m 55s | 3.9x less | 4.9x less | |
| tumor_2.fastq.gz | 7m 32s | 9m 22s | 2m 53s | 1m 55s | 3.9x less | 4.9x less | |
| normal_1.fastq.gz | 6m 26s | 7m 53s | 2m 27s | 1m 35s | 4.1x less | 5.0x less | |
| normal_2.fastq.gz | 6m 25s | 7m 54s | 2m 27s | 1m 44s | 3.7x less | 4.5x less | |
| Element AVITI (short-read) | |||||||
| ERR16944282_1.fastq.gz | 30s | 34s | 12s | 8.5s | 3.5x less | 4.0x less | |
| ERR16944282_2.fastq.gz | 33s | 35s | 13s | 8.9s | 3.7x less | 3.9x less | |
| ERR16944299_1.fastq.gz | 30s | 33s | 12s | 8.2s | 3.7x less | 4.1x less | |
| ERR16944299_2.fastq.gz | 31s | 34s | 13s | 8.8s | 3.5x less | 3.9x less | |
| ONT PromethION (long-read) | |||||||
| SRR37915503.fastq.gz | 1h 03m | 1h 02m | 23m 31s | 12m 12s | 5.2x less | 5.2x less | |
| BAM | |||||||
| GM12878.bam | 7m 53s | 9m 36s | 4m 07s | 4m 15s | 1.9x less | 2.3x less | |
CPU time is what a cluster or cloud bill charges for. Java v0.13.0 and Rust v1.1 spend some extra CPU to finish sooner; Rust v1.1 still uses a fraction of the CPU of either Java version.
Peak Memory (RSS)
| File | Java v0.12.1 | Java v0.13.0 | Rust v1.0 | Rust v1.1 | Rust v1.1 vs Java v0.12.1 | Rust v1.1 vs Java v0.13.0 | |
|---|---|---|---|---|---|---|---|
| Illumina WES (short-read, 126 bp) | |||||||
| tumor_1.fastq.gz | 435 MB | 502 MB | 25 MB | 45 MB | 9.7x less | 11.2x less | |
| tumor_2.fastq.gz | 436 MB | 495 MB | 26 MB | 40 MB | 10.8x less | 12.3x less | |
| normal_1.fastq.gz | 423 MB | 501 MB | 29 MB | 44 MB | 9.6x less | 11.4x less | |
| normal_2.fastq.gz | 424 MB | 527 MB | 25 MB | 40 MB | 10.5x less | 13.1x less | |
| Element AVITI (short-read) | |||||||
| ERR16944282_1.fastq.gz | 424 MB | 475 MB | 50 MB | 71 MB | 6.0x less | 6.7x less | |
| ERR16944282_2.fastq.gz | 431 MB | 479 MB | 50 MB | 68 MB | 6.3x less | 7.0x less | |
| ERR16944299_1.fastq.gz | 419 MB | 478 MB | 50 MB | 133 MB | 3.1x less | 3.6x less | |
| ERR16944299_2.fastq.gz | 414 MB | 461 MB | 57 MB | 155 MB | 2.7x less | 3.0x less | |
| ONT PromethION (long-read) | |||||||
| SRR37915503.fastq.gz | 3044 MB | 4096 MB | 737 MB | 945 MB | 3.2x less | 4.3x less | |
| BAM | |||||||
| GM12878.bam | 405 MB | 510 MB | 28 MB | 32 MB | 12.8x less | 16.1x less | |
Rust v1.1 uses more memory than Rust v1.0, because batches of reads are in flight between its threads, but still 3–13x less than either Java version on short reads. Long reads need far more memory in every implementation, because the per-position arrays grow with read length. Java's long-read figures reflect the 4 GB heap it was given: Java v0.13.0 used all of it.
Throughput by Platform
| Platform | Java v0.12.1 | Java v0.13.0 | Rust v1.0 | Rust v1.1 | Rust v1.1 -t 2 |
|---|---|---|---|---|---|
| Illumina WES (short-read, 126 bp) | 17 MB/s | 48 MB/s | 46 MB/s | 218 MB/s | 116 MB/s |
| Element AVITI (short-read) | 18 MB/s | 46 MB/s | 45 MB/s | 174 MB/s | 115 MB/s |
| ONT PromethION (long-read) | 21 MB/s | 51 MB/s | 44 MB/s | 200 MB/s | 154 MB/s |
| BAM | 26 MB/s | 56 MB/s | 49 MB/s | 82 MB/s | 59 MB/s |
Throughput is compressed file size divided by wall-clock time.
Thread Scaling (Rust v1.1)
One 7.3 GB WES file at each -t value (one run each, measured separately from the table above, after the gzip decoder began counting towards -t):
-t | Wall time | Speedup | CPU time | Peak RSS |
|---|---|---|---|---|
| 1 | 101.7s | 1.00x | 101s | 32 MB |
| 2 | 66.6s | 1.53x | 102s | 50 MB |
| 3 | 54.1s | 1.88x | 106s | 61 MB |
| 4 | 36.2s | 2.81x | 108s | 45 MB |
| 6 (default) | 35.8s | 2.84x | 109s | 43 MB |
At -t 1 the file is decoded and analysed on a single thread; from -t 2 the decoder gets a thread of its own.
A single short-read .fastq.gz stops getting faster at about -t 4: from there, the file's one gzip decoder is the slowest stage. One zlib-rs decoder produces about 1.1 GB/s of decompressed data, the same as system gzip. rapidgzip can split a file across several decoders (--decompress-threads), but a normal single-member .gz can only be split speculatively: 4 decoders were 1.35x faster for 2.8x the CPU and ~900 MB of memory, so one decoder per file is the default. Spare cores are better spent on more files at once.
BAM is decoded on the reader thread, which caps it at about 1.6x however many threads it gets.
Where the Time Goes (Rust v1.1)
CPU profile of Rust v1.1 on the 7.3 GB WES file at -t 2: one analysis thread plus its gzip decoder thread, sampled for 40 seconds mid-run with macOS sample (39,386 active samples across both threads, idle waits excluded).
| Component | Share of CPU |
|---|---|
| Gzip decompression (zlib-rs, on its own thread) | 32% |
| FASTQ parsing (including UTF-8 validation) | 13% |
| Adapter Content (Teddy search) | 12% |
| Per base sequence content | 9% |
| Per base quality + per tile quality | 9% |
| Per base N content | 7% |
| Overrepresented sequences + duplication (hashing) | 4% |
| Per sequence GC + quality, length distribution | 3% |
| Basic Statistics | 3% |
| Memory, threading, other | 9% |
A third of the CPU is gzip decompression, which runs on its own thread. That is why a single file stops scaling at about -t 4: once the analysis is spread across workers, the one decoder is the slowest stage. The analysis side has no single dominant cost any more. The largest remaining item that isn't analysis is UTF-8 validation of each record (about 6%), a possible next step.
What Makes the Rust Version Faster
None of this comes from Rust being inherently faster than Java. Each change was found by CPU profiling and checked for byte-identical output.
1. Parallel analysis within a file
A reader thread parses records and hands batches to worker threads, each of which owns a subset of the QC modules. Every module still sees every read, in order, on one thread, so no results need merging and the output is identical at any thread count. Modules are assigned to workers by their measured cost, so the expensive ones don't share a worker. Batches are capped by bytes as well as by read count, which keeps memory flat on long reads (552 MB → 66 MB on 10 kb reads at -t 8). The same design was ported to Java and shipped in FastQC v0.13.0.
2. Gzip decoding on its own thread
.fastq.gz is decompressed by rapidgzip-core with a zlib-rs inflate, on a separate thread from the parser, so decompression overlaps with the analysis from -t 2. At -t 1 it runs on the reading thread. A 1 MiB read buffer keeps hand-offs between the two rare. Everything is pure Rust: there is no system zlib to link.
3. One SIMD pass for every adapter
FastQC looks for six adapter sequences in every read. Java calls String.indexOf() once per adapter. Rust v1.0 used SIMD memchr per adapter; Rust v1.1 finds all of them in a single pass with aho-corasick's Teddy algorithm, 3.6–5.2x faster again. Hits are tallied once per read at the first match position and summed at the end, instead of incrementing every later position.
4. Vectorised counting
Base counts and each read's lowest quality are computed in loops the compiler turns into SIMD instructions. In the per-position modules, a 256-byte lookup table maps each base straight to a counter index, replacing branch chains the CPU can't predict on random DNA. Together these cut CPU by about a third on 300 bp reads.
5. No per-read allocation
Java creates several objects for every read: a Sequence, a new String from toUpperCase(), String[] arrays from split(":") for tile parsing. Rust parses FASTQ lines straight into reused buffers with a SIMD newline search, uppercases in place, and reads tile IDs without allocating. With no garbage collector, there are no GC pauses.
Could these improvements be back-ported to Java?
Parallel analysis: yes, and it has been. FastQC v0.13.0 ships the same reader-plus-module-workers pipeline (s-andrews/FastQC#197) and removes per-sequence char[] allocations in seven modules (#199). That is where 0.13.0's speedup over 0.12.1 comes from.
The per-read optimisations were tried in Java in April 2026 and didn't carry over:
- SIMD adapter search: no. Java has no portable way to call NEON or SSE substring search, and
String.indexOf()is already a HotSpot intrinsic. A hand-writtenchar[]loop was slower (16.0s vs 12.9s). The Vector API is still incubating and has no substring search. - Lookup-table base counting: marginal. HotSpot already compiles small
switchstatements on bytes into jump tables. An explicit array lookup measured about 3% faster. - Zero-allocation processing: structurally limited. Java strings are immutable and heap-allocated, so
toUpperCase()andsplit()must allocate. Workarounds measured negligible gains, since HotSpot's escape analysis and generational GC already handle short-lived objects well.
Single-threaded, Java v0.13.0 is about 14% slower than 0.12.1 on the WES file (522s vs 456s, reproduced on a quiet machine): the cost of the pipeline when it has only one thread to run on.
Build
FastQC-Rust v1.1 has a single build: pure Rust, with no C toolchain and no system libraries, so binaries are fully static on every platform. The native-zlib feature and the separate pure-Rust variant from 1.0 are gone. BAM (via noodles) and Fast5 decode through the same zlib-rs backend.
Overall Summary
| Metric | Java v0.12.1 | Java v0.13.0 | Rust v1.0 | Rust v1.1 |
|---|---|---|---|---|
| Short-read speed (avg) | 18 MB/s | 47 MB/s | 46 MB/s | 196 MB/s |
| Long-read speed (avg) | 21 MB/s | 51 MB/s | 44 MB/s | 200 MB/s |
| BAM speed | 26 MB/s | 56 MB/s | 49 MB/s | 82 MB/s |
| Peak memory (short-read) | 414 MB–436 MB | 461 MB–527 MB | 25 MB–57 MB | 40 MB–155 MB |
| Peak memory (long-read) | 3.0 GB | 4.0 GB | 737 MB | 945 MB |
| Default threads per file | 1 | 4 | 1 | up to 6 |
| Runtime dependencies | JRE + Perl | JRE + Python | libz (or none) | none |
Real-World Impact
To put these numbers in context, we analysed FastQC task execution data from Seqera Platform Cloud, which runs Nextflow pipelines at scale for bioinformatics teams worldwide.
Scale of FastQC usage (Seqera Platform Cloud only)
| Metric | Value |
|---|---|
| Period analysed | Jan 2025 – Mar 2026 (15 months) |
| Total FastQC tasks | 3.6 million |
| Total CPU hours consumed | 1.1 million hours |
| Annualised | ~2.9 million tasks / ~877,000 CPU hours |
Projected savings
Cloud and cluster costs follow CPU time, not wall time. On the short-read files, Rust v1.1 uses 3.7x less CPU than Java v0.12.1, the version behind that usage data. Applied to the annualised Seqera Cloud usage alone:
| Metric | Java v0.12.1 | Rust v1.1 | Saved |
|---|---|---|---|
| Annual CPU hours | 877,000 | 234,000 | 643,000 hours |
| Annual cloud compute cost* | $26,300 | $7,000 | $19,300 |
| Annual CO2 emissions** | 61 tonnes | 16 tonnes | 45 tonnes |
* Estimated at $0.03/CPU-hour (typical cloud spot pricing for compute-optimised instances).
** Estimated using global average grid carbon intensity (0.35 kgCO2/kWh) and ~200 W per CPU core.
Beyond Seqera Platform
These figures cover a single cloud platform. FastQC runs in university clusters, hospital genomics labs, national sequencing centres and cloud pipelines worldwide, so the total compute spent on it is likely orders of magnitude larger.
The memory reduction matters too: a short-read FastQC task that needed ~500 MB fits in ~50 MB, so it can run on smaller instances and packs more densely in schedulers like Kubernetes and AWS Batch.