ctJpeg writes exactly the same JPEG files as Google’s jpegli,
byte for byte, and encodes a 39-megapixel photo 6.9 times as fast. Unlike an audio
encoder it is not a stream of frames but a sequence of stages: each parallel stage hands
a few hundred independent jobs to one worker per thread, and the next stage starts only
when all of them are done. These nine pictures show how the stages fanned out, one
after another, on ThinkMeta ConcurrentTasks.
Picture 1 of 9
The reference: jpegli
jpegli is the JPEG encoder from the JPEG XL team: at the same visual
quality its files are smaller than those of mozjpeg and libjpeg-turbo
– but it encodes on one core. Reading, encoding and writing follow
one another on a single thread.
These bytes are the yardstick: every following picture
has to write exactly the same JPEG file.
Where the time goes
For a 39-megapixel photo, tokenizing the progressive scans takes
44 % of the time, the pixel phase (colour, adaptive
quantization, DCT, quantization) 34 %, writing the bitstream
6 % and Huffman 1 %. For a real speed-up, at least the
pixel phase and the tokenizer must run in parallel.
Run time · one photo, 39 MP · development
jpegli: 0.75 s
Picture 2 of 9
Asynchronous reading and writing
Reading and writing move out of the encoder: a reading task on an I/O
scheduler fetches the photo ahead in large reads, and the encoder already
computes while the rest is still on its way – every strip of the image
starts as soon as its own rows have arrived. A writing task writes the JPEG
while the encoder cleans up.
In practice a photo usually comes from the disk, not from the file cache.
From the SSD, the 39-megapixel photo is encoded 28 % faster.
Measured
Run time relative to reading the whole file first, median of alternating pairs, file from the NVMe SSD:
39 MP: 0.72 with 16 threads, 0.75 with 4, 0.92 on one thread
11 MP: 0.90 / 0.67 / 0.92
When the file is already in the cache, it is copied once more: with
all threads 1 to 23 % slower, with 4 threads or on one thread the
same. Every variant was bit-identical.
Real overlap
Reading and computing only overlap when every strip waits for its own
rows alone and the reading task hands each block over at once.
Writing alongside
The Huffman tables stand in front of the scans, so the file can only
be written once the whole JPEG is assembled. Writing then runs on the
I/O thread while the encoder releases its memory – for the large
photo that saves another 2 to 3 %.
Run time · 39 MP photo from the SSD, 16 threads
0.185 → 0.133 s
Picture 3 of 9
The port: same flow, our own code
First bit identity with a readable port, then parallelism. jpegli is
ported to C++ by hand and split into phases: setup, pixel phase,
entropy coding. Between them sits a buffer of coefficients. The pixel
phase already works in bands of block rows and computes the context rows
it needs itself – ready for parallel work, but still on one thread.
For now about 1.7 times slower than jpegli, but bit-identical in every test case.
Two kinds of identical
jpegli does not write the same bytes on every CPU: its SIMD
library uses fused multiply-add on AVX2 but not on SSE, and sums
vector lanes in an order that depends on the vector width. The
port reproduces both classes – byte for byte jpegli on SSE2,
and byte for byte jpegli on AVX2.
Still in use today
This flow is still in the code: ctJpeg --sequential
runs the whole encoder as one task – the reference path for
comparisons.
Run time · one photo, 39 MP · development
1.27 s · jpegli 0.75 s
Picture 4 of 9
The pixel phase goes parallel
The bands of the pixel phase become jobs, about four per worker. Each
worker fetches the next band from a shared counter; the main task waits
until all are done. The pixel phase drops from 0.76 to 0.12 s
– 6.4 times as fast on 16 threads.
The limit now: the serial entropy coding, 0.48 s.
Why the bands are independent
The adaptive quantization only looks at small local windows around
each block. When a band computes a few rows of context itself, its
coefficients no longer depend on any other band.
Run time · one photo, 39 MP · development
0.70 s · 1.07× jpegli
Picture 5 of 9
Tokenizing and writing go parallel
The entropy block breaks apart. The progressive scans are independent of
each other, one job each; the large scans are additionally split into
bands. A short serial Stitch joins the bands exactly as
jpegli would have coded them in one go. Huffman stays serial: its tables
need the symbol counts of all scans.
Every scan ends on a byte boundary, so each one is written into its own
buffer in parallel; Assemble puts them together.
What the stitch sews
Within a scan, runs of empty blocks are coded as one end-of-band
run that crosses block rows – and jpegli caps a run at 32,767
blocks and splits refinement runs after 255 correction bits. A band
starts with a fresh state and reports its beginning compactly; the
stitch replays it with the real state.
Only half right
The profiler said then: limited by the 8 physical cores, not by
memory. Picture 7 shows that this was only half true – memory
became the limit as soon as each thread got faster.
Run time · one photo, 39 MP · development
0.31 s · 2.4× jpegli
Picture 6 of 9
Less work per node
The graph stays the same; every node gets faster. SIMD kernels in the
pixel phase follow jpegli’s arithmetic exactly, and the tokenizer works on bit masks per block instead of looping over every coefficient.
On one thread ctJpeg is now 1.4 to 1.5 times as fast as jpegli. But the
pixel phase only scales about threefold.
Nothing measurable
Ring buffers alone, which shrank each worker’s working set
from 8.5 to 2 MB, brought nothing measurable. Why the pixel
phase stopped scaling only became clear in the next picture.
Run time · one photo, 39 MP · development
0.24 s · 3.1× jpegli
Picture 7 of 9
Smaller jobs that fit the cache
In a fork-join, the last job of a stage decides when it
ends. A job trace showed which one it was for every stage: the work was
spread correctly, but single jobs were too large or ran into memory. So
the jobs got smaller: the pixel phase works on tiles about 1,024 pixels
wide, the DC scan is split into bands, and large scans are written in
chunks – Count works out where each chunk starts,
Splice joins the chunks bit-exactly.
The total time now falls all the way to 16 threads; before, the optimum was 6 to 8.
What the job trace found
The DC scan ran as one 32 ms job and finished last → DC bands: tokenizing 69 → 56 ms.
One worker wrote each large refinement scan → chunks: writing 28 → 10–13 ms.
The last AC bands set the end → finer bands: tokenizing 39 → 33 ms.
Misleading clue
Before the tiles, bands of a single block row were faster with 8
threads. That looked like a load-balancing problem but was a cache
effect – only the same image at a quarter of the width made it
visible.
Run time · the same photo · tuning runs
0.14 s · 5.4× jpegli
Picture 8 of 9
Cleaning up goes parallel, too
With 16 threads about 0.05 s still lay outside the stages –
most of it handing used memory back to the operating system. Now that
happens in jobs, as early as possible: the input image as the first job
of Tokenize, the coefficients during Write (nothing reads them any more),
the rest in a new stage, Release.
Freeing memory drops from 26 to 7–8 ms. Over 14 test photos ctJpeg is now 4.8 to 5 times as fast as jpegli.
Why not simply skip it?
Without freeing, the cost only moves to the end of the process
– measured: same total time. Spread over the workers it
shrinks. On the side, committed memory fell from 1,241 to
408 MB.
Tried and dropped
Freeing the coefficient buffer in parallel pieces: the pieces of
one memory region slow each other down, freeing went up from 11 to
14 ms.
Run time · the same photo · tuning runs
0.143 → 0.109 s · 6.9× jpegli
Picture 9 of 9
The complete graph
Ten stages, three of them serial and short: Setup, Huffman, Assemble.
Every other stage fans out to all cores and joins again before the next
one starts. The 39-megapixel photo takes 0.108 s instead of
0.750 s – bit-identical to jpegli, and 3.4 to 4.7 times as fast
as libjpeg-turbo on large images, with jpegli’s smaller files.
Why tasks here
At its core ctJpeg is fork-join: no job waits for another in the
middle of its work. The encoder itself runs as a task and waits for
each stage; on ThinkMeta ConcurrentTasks that wait suspends a fiber instead of blocking
a thread, and the encoder reads as plain sequential code.
Other CPUs
Six photos, 99 megapixels in all:
Intel Core i9-11900K, 8 cores: 5.3×
Intel Xeon E3-1275 v6, 4 cores: 5.6×
Intel Core Ultra 7 155H, notebook: 4.2×
Bit-identical to jpegli on every machine.
Structure for free
The graph now lives in one place in the code. Measured against the
build before: 0.96 to 1.04 in every case – within the noise.
The structure costs nothing.
Speed-up over jpegli · benchmark package, 39 MP photo
6.9×
Try it yourself
ctJpeg, the jpegli reference and the benchmark for your own photos are on
GitHub.
Have a workload that should use every core?
Get in touch.