FastQC: Rust vs Java Benchmark Report

Date: October 2026
Machine: Apple M1 Pro (8 performance + 2 efficiency cores, 16 GB RAM), macOS. Data on an external SSD (~880 MB/s, never the bottleneck).

Implementations:

Each tool runs at its default settings, one file at a time, one run per file. Java ran on OpenJDK 21 with the launcher's defaults, except that the long-read files got --memory 4096: the default 512 MB heap is too small for them, as FastQC's own help notes. Speedups are relative to Java v0.12.1 unless stated. macOS Spotlight indexing was using about one core in the background throughout; WES results matched earlier runs on a quiet machine to within 2%.

Summary

DataJava v0.12.1Java v0.13.0Rust v1.0Rust v1.1Rust v1.1 -t 2
Illumina WES (short-read, 126 bp)1.0x2.7x2.6x12.5x6.6x
Element AVITI (short-read)1.0x2.5x2.5x9.7x6.4x
ONT PromethION (long-read)1.0x2.4x2.1x9.5x7.3x
BAM1.0x2.2x1.9x3.2x2.3x

Average wall-clock speedup over Java v0.12.1 per file, each tool at its default threads.

Test Dataset

FileSizePlatform
SRR7890918_WES_HCC1395-EA_tumor_1.fastq.gz7.3 GBIllumina WES, short-read
SRR7890918_WES_HCC1395-EA_tumor_2.fastq.gz7.5 GBIllumina WES, short-read
SRR7890919_WES_HCC1395BL-EA_normal_1.fastq.gz6.2 GBIllumina WES, short-read
SRR7890919_WES_HCC1395BL-EA_normal_2.fastq.gz6.3 GBIllumina WES, short-read
ERR16944282_1.fastq.gz487 MBElement AVITI, short-read
ERR16944282_2.fastq.gz630 MBElement AVITI, short-read
ERR16944299_1.fastq.gz454 MBElement AVITI, short-read
ERR16944299_2.fastq.gz597 MBElement AVITI, short-read
ERR16962265.fastq.gz (not benchmarked, see below)21.2 GBONT PromethION, long-read (wheat WGS)
SRR37915503.fastq.gz57.6 GBONT PromethION, long-read (snake WGS)
GM12878_REP1.markdup.sorted.bam11.3 GBIllumina, aligned BAM
Total119.6 GB11 files across 3 platforms + BAM
ERR16962265.fastq.gz is listed for completeness but left out of the results: the local copy is an incomplete download (the first 22.7 GB of the 66.6 GB file on ENA), so the gzip stream ends early and it can't be read to the end. The April 2026 report's figures for this file came from the same partial copy.

Output Correctness

fastqc_data.txt and summary.txt from Rust v1.1 were compared module by module against Java v0.13.0, and between Rust v1.1's -t 2 and default runs. Every file is byte-identical in both comparisons.

FileRust v1.1 -t 6 vs Java v0.13.0Rust v1.1 -t 2 vs -t 6
tumor_1.fastq.gzidenticalidentical
tumor_2.fastq.gzidenticalidentical
normal_1.fastq.gzidenticalidentical
normal_2.fastq.gzidenticalidentical
ERR16944282_1.fastq.gzidenticalidentical
ERR16944282_2.fastq.gzidenticalidentical
ERR16944299_1.fastq.gzidenticalidentical
ERR16944299_2.fastq.gzidenticalidentical
SRR37915503.fastq.gzidenticalidentical
GM12878.bamidenticalidentical

Per-File Results

Wall Clock Time

Java v0.12.1
Java v0.13.0
Rust v1.0
Rust v1.1
FileJava v0.12.1Java v0.13.0Rust v1.0Rust v1.1Rust v1.1 vs
Java v0.12.1
Rust v1.1 vs
Java v0.13.0
Illumina WES (short-read, 126 bp)
tumor_1.fastq.gz7m 32s2m 47s2m 52s37s12.4x4.6x
tumor_2.fastq.gz7m 32s2m 47s2m 53s37s12.4x4.6x
normal_1.fastq.gz6m 27s2m 22s2m 28s30s12.8x4.7x
normal_2.fastq.gz6m 26s2m 21s2m 28s31s12.5x4.6x
Element AVITI (short-read)
ERR16944282_1.fastq.gz31s12s12s3.0s10.1x3.9x
ERR16944282_2.fastq.gz33s13s13s3.6s9.1x3.5x
ERR16944299_1.fastq.gz31s12s12s2.9s10.7x4.1x
ERR16944299_2.fastq.gz31s13s13s3.5s8.8x3.7x
ONT PromethION (long-read)
SRR37915503.fastq.gz49m 03s20m 11s23m 33s5m 10s9.5x3.9x
BAM
GM12878.bam7m 54s3m 37s4m 07s2m 28s3.2x1.5x

CPU Time (user + system)

Java v0.12.1
Java v0.13.0
Rust v1.0
Rust v1.1
FileJava v0.12.1Java v0.13.0Rust v1.0Rust v1.1Rust v1.1 vs
Java v0.12.1
Rust v1.1 vs
Java v0.13.0
Illumina WES (short-read, 126 bp)
tumor_1.fastq.gz7m 31s9m 21s2m 52s1m 55s3.9x less4.9x less
tumor_2.fastq.gz7m 32s9m 22s2m 53s1m 55s3.9x less4.9x less
normal_1.fastq.gz6m 26s7m 53s2m 27s1m 35s4.1x less5.0x less
normal_2.fastq.gz6m 25s7m 54s2m 27s1m 44s3.7x less4.5x less
Element AVITI (short-read)
ERR16944282_1.fastq.gz30s34s12s8.5s3.5x less4.0x less
ERR16944282_2.fastq.gz33s35s13s8.9s3.7x less3.9x less
ERR16944299_1.fastq.gz30s33s12s8.2s3.7x less4.1x less
ERR16944299_2.fastq.gz31s34s13s8.8s3.5x less3.9x less
ONT PromethION (long-read)
SRR37915503.fastq.gz1h 03m1h 02m23m 31s12m 12s5.2x less5.2x less
BAM
GM12878.bam7m 53s9m 36s4m 07s4m 15s1.9x less2.3x less

CPU time is what a cluster or cloud bill charges for. Java v0.13.0 and Rust v1.1 spend some extra CPU to finish sooner; Rust v1.1 still uses a fraction of the CPU of either Java version.

Peak Memory (RSS)

Java v0.12.1
Java v0.13.0
Rust v1.0
Rust v1.1
FileJava v0.12.1Java v0.13.0Rust v1.0Rust v1.1Rust v1.1 vs
Java v0.12.1
Rust v1.1 vs
Java v0.13.0
Illumina WES (short-read, 126 bp)
tumor_1.fastq.gz435 MB502 MB25 MB45 MB9.7x less11.2x less
tumor_2.fastq.gz436 MB495 MB26 MB40 MB10.8x less12.3x less
normal_1.fastq.gz423 MB501 MB29 MB44 MB9.6x less11.4x less
normal_2.fastq.gz424 MB527 MB25 MB40 MB10.5x less13.1x less
Element AVITI (short-read)
ERR16944282_1.fastq.gz424 MB475 MB50 MB71 MB6.0x less6.7x less
ERR16944282_2.fastq.gz431 MB479 MB50 MB68 MB6.3x less7.0x less
ERR16944299_1.fastq.gz419 MB478 MB50 MB133 MB3.1x less3.6x less
ERR16944299_2.fastq.gz414 MB461 MB57 MB155 MB2.7x less3.0x less
ONT PromethION (long-read)
SRR37915503.fastq.gz3044 MB4096 MB737 MB945 MB3.2x less4.3x less
BAM
GM12878.bam405 MB510 MB28 MB32 MB12.8x less16.1x less

Rust v1.1 uses more memory than Rust v1.0, because batches of reads are in flight between its threads, but still 3–13x less than either Java version on short reads. Long reads need far more memory in every implementation, because the per-position arrays grow with read length. Java's long-read figures reflect the 4 GB heap it was given: Java v0.13.0 used all of it.

Throughput by Platform

PlatformJava v0.12.1Java v0.13.0Rust v1.0Rust v1.1Rust v1.1 -t 2
Illumina WES (short-read, 126 bp)17 MB/s48 MB/s46 MB/s218 MB/s116 MB/s
Element AVITI (short-read)18 MB/s46 MB/s45 MB/s174 MB/s115 MB/s
ONT PromethION (long-read)21 MB/s51 MB/s44 MB/s200 MB/s154 MB/s
BAM26 MB/s56 MB/s49 MB/s82 MB/s59 MB/s

Throughput is compressed file size divided by wall-clock time.

Thread Scaling (Rust v1.1)

One 7.3 GB WES file at each -t value (one run each, measured separately from the table above, after the gzip decoder began counting towards -t):

-tWall timeSpeedupCPU timePeak RSS
1101.7s1.00x101s32 MB
266.6s1.53x102s50 MB
354.1s1.88x106s61 MB
436.2s2.81x108s45 MB
6 (default)35.8s2.84x109s43 MB

At -t 1 the file is decoded and analysed on a single thread; from -t 2 the decoder gets a thread of its own.

A single short-read .fastq.gz stops getting faster at about -t 4: from there, the file's one gzip decoder is the slowest stage. One zlib-rs decoder produces about 1.1 GB/s of decompressed data, the same as system gzip. rapidgzip can split a file across several decoders (--decompress-threads), but a normal single-member .gz can only be split speculatively: 4 decoders were 1.35x faster for 2.8x the CPU and ~900 MB of memory, so one decoder per file is the default. Spare cores are better spent on more files at once.

BAM is decoded on the reader thread, which caps it at about 1.6x however many threads it gets.

Where the Time Goes (Rust v1.1)

CPU profile of Rust v1.1 on the 7.3 GB WES file at -t 2: one analysis thread plus its gzip decoder thread, sampled for 40 seconds mid-run with macOS sample (39,386 active samples across both threads, idle waits excluded).

ComponentShare of CPU
Gzip decompression (zlib-rs, on its own thread)32%
FASTQ parsing (including UTF-8 validation)13%
Adapter Content (Teddy search)12%
Per base sequence content9%
Per base quality + per tile quality9%
Per base N content7%
Overrepresented sequences + duplication (hashing)4%
Per sequence GC + quality, length distribution3%
Basic Statistics3%
Memory, threading, other9%

A third of the CPU is gzip decompression, which runs on its own thread. That is why a single file stops scaling at about -t 4: once the analysis is spread across workers, the one decoder is the slowest stage. The analysis side has no single dominant cost any more. The largest remaining item that isn't analysis is UTF-8 validation of each record (about 6%), a possible next step.

What Makes the Rust Version Faster

None of this comes from Rust being inherently faster than Java. Each change was found by CPU profiling and checked for byte-identical output.

1. Parallel analysis within a file

A reader thread parses records and hands batches to worker threads, each of which owns a subset of the QC modules. Every module still sees every read, in order, on one thread, so no results need merging and the output is identical at any thread count. Modules are assigned to workers by their measured cost, so the expensive ones don't share a worker. Batches are capped by bytes as well as by read count, which keeps memory flat on long reads (552 MB → 66 MB on 10 kb reads at -t 8). The same design was ported to Java and shipped in FastQC v0.13.0.

2. Gzip decoding on its own thread

.fastq.gz is decompressed by rapidgzip-core with a zlib-rs inflate, on a separate thread from the parser, so decompression overlaps with the analysis from -t 2. At -t 1 it runs on the reading thread. A 1 MiB read buffer keeps hand-offs between the two rare. Everything is pure Rust: there is no system zlib to link.

3. One SIMD pass for every adapter

FastQC looks for six adapter sequences in every read. Java calls String.indexOf() once per adapter. Rust v1.0 used SIMD memchr per adapter; Rust v1.1 finds all of them in a single pass with aho-corasick's Teddy algorithm, 3.6–5.2x faster again. Hits are tallied once per read at the first match position and summed at the end, instead of incrementing every later position.

4. Vectorised counting

Base counts and each read's lowest quality are computed in loops the compiler turns into SIMD instructions. In the per-position modules, a 256-byte lookup table maps each base straight to a counter index, replacing branch chains the CPU can't predict on random DNA. Together these cut CPU by about a third on 300 bp reads.

5. No per-read allocation

Java creates several objects for every read: a Sequence, a new String from toUpperCase(), String[] arrays from split(":") for tile parsing. Rust parses FASTQ lines straight into reused buffers with a SIMD newline search, uppercases in place, and reads tile IDs without allocating. With no garbage collector, there are no GC pauses.

Could these improvements be back-ported to Java?

Parallel analysis: yes, and it has been. FastQC v0.13.0 ships the same reader-plus-module-workers pipeline (s-andrews/FastQC#197) and removes per-sequence char[] allocations in seven modules (#199). That is where 0.13.0's speedup over 0.12.1 comes from.

The per-read optimisations were tried in Java in April 2026 and didn't carry over:

Single-threaded, Java v0.13.0 is about 14% slower than 0.12.1 on the WES file (522s vs 456s, reproduced on a quiet machine): the cost of the pipeline when it has only one thread to run on.

Build

FastQC-Rust v1.1 has a single build: pure Rust, with no C toolchain and no system libraries, so binaries are fully static on every platform. The native-zlib feature and the separate pure-Rust variant from 1.0 are gone. BAM (via noodles) and Fast5 decode through the same zlib-rs backend.

Overall Summary

MetricJava v0.12.1Java v0.13.0Rust v1.0Rust v1.1
Short-read speed (avg)18 MB/s47 MB/s46 MB/s196 MB/s
Long-read speed (avg)21 MB/s51 MB/s44 MB/s200 MB/s
BAM speed26 MB/s56 MB/s49 MB/s82 MB/s
Peak memory (short-read)414 MB–436 MB461 MB–527 MB25 MB–57 MB40 MB–155 MB
Peak memory (long-read)3.0 GB4.0 GB737 MB945 MB
Default threads per file141up to 6
Runtime dependenciesJRE + PerlJRE + Pythonlibz (or none)none

Real-World Impact

To put these numbers in context, we analysed FastQC task execution data from Seqera Platform Cloud, which runs Nextflow pipelines at scale for bioinformatics teams worldwide.

Scale of FastQC usage (Seqera Platform Cloud only)

MetricValue
Period analysedJan 2025 – Mar 2026 (15 months)
Total FastQC tasks3.6 million
Total CPU hours consumed1.1 million hours
Annualised~2.9 million tasks / ~877,000 CPU hours

Projected savings

Cloud and cluster costs follow CPU time, not wall time. On the short-read files, Rust v1.1 uses 3.7x less CPU than Java v0.12.1, the version behind that usage data. Applied to the annualised Seqera Cloud usage alone:

MetricJava v0.12.1Rust v1.1Saved
Annual CPU hours877,000234,000643,000 hours
Annual cloud compute cost*$26,300$7,000$19,300
Annual CO2 emissions**61 tonnes16 tonnes45 tonnes

* Estimated at $0.03/CPU-hour (typical cloud spot pricing for compute-optimised instances).
** Estimated using global average grid carbon intensity (0.35 kgCO2/kWh) and ~200 W per CPU core.

Beyond Seqera Platform

These figures cover a single cloud platform. FastQC runs in university clusters, hospital genomics labs, national sequencing centres and cloud pipelines worldwide, so the total compute spent on it is likely orders of magnitude larger.

The memory reduction matters too: a short-read FastQC task that needed ~500 MB fits in ~50 MB, so it can run on smaller instances and packs more densely in schedulers like Kubernetes and AWS Batch.