ctJpeg writes exactly the same JPEG files as Google’s jpegli, byte for byte, and encodes a 39-megapixel photo 6.9 times as fast. Unlike an audio encoder it is not a stream of frames but a sequence of stages: each parallel stage hands a few hundred independent jobs to one worker per thread, and the next stage starts only when all of them are done. These nine pictures show how the stages fanned out, one after another, on ThinkMeta ConcurrentTasks.

Read, Write, cjpegliReadsynchronousWritesynchronouscjpegliall phases in turn Read, Write, cjpegli1 thread · one thing after anotherReadsynchronousWritesynchronouscjpegliall phases in turn

Picture 1 of 9

The reference: jpegli

jpegli is the JPEG encoder from the JPEG XL team: at the same visual quality its files are smaller than those of mozjpeg and libjpeg-turbo – but it encodes on one core. Reading, encoding and writing follow one another on a single thread.

These bytes are the yardstick: every following picture has to write exactly the same JPEG file.

Where the time goes

For a 39-megapixel photo, tokenizing the progressive scans takes 44 % of the time, the pixel phase (colour, adaptive quantization, DCT, quantization) 34 %, writing the bitstream 6 % and Huffman 1 %. For a real speed-up, at least the pixel phase and the tokenizer must run in parallel.

Run time · one photo, 39 MP · development

jpegli: 0.75 s

Read, Write, EncodeReadasynchronousWriteasynchronousEncodestrip by strip Read, Write, EncodeI/O schedulercompute schedulerI/O schedulerReadasynchronousWriteasynchronousEncodestrip by strip

Picture 2 of 9

Asynchronous reading and writing

Reading and writing move out of the encoder: a reading task on an I/O scheduler fetches the photo ahead in large reads, and the encoder already computes while the rest is still on its way – every strip of the image starts as soon as its own rows have arrived. A writing task writes the JPEG while the encoder cleans up.

In practice a photo usually comes from the disk, not from the file cache. From the SSD, the 39-megapixel photo is encoded 28 % faster.

Measured

Run time relative to reading the whole file first, median of alternating pairs, file from the NVMe SSD:

  • 39 MP: 0.72 with 16 threads, 0.75 with 4, 0.92 on one thread
  • 11 MP: 0.90 / 0.67 / 0.92

When the file is already in the cache, it is copied once more: with all threads 1 to 23 % slower, with 4 threads or on one thread the same. Every variant was bit-identical.

Real overlap

Reading and computing only overlap when every strip waits for its own rows alone and the reading task hands each block over at once.

Writing alongside

The Huffman tables stand in front of the scans, so the file can only be written once the whole JPEG is assembled. Writing then runs on the I/O thread while the encoder releases its memory – for the large photo that saves another 2 to 3 %.

Run time · 39 MP photo from the SSD, 16 threads

0.185 → 0.133 s

Read, Write, Setup, Pixel phase, coefficients, EntropyReadasynchronousWriteasynchronousSetuptables, scan scriptPixel phasein bands, serialcoefficientsEntropytokens, Huffman, bits Read, Write, Setup, Pixel phase, coefficients, EntropyI/O scheduler1 thread · split into phasesI/O schedulerReadasynchronousWriteasynchronousSetuptables, scan scriptPixel phasein bands, serialcoefficientsEntropytokens, Huffman, bits

Picture 3 of 9

The port: same flow, our own code

First bit identity with a readable port, then parallelism. jpegli is ported to C++ by hand and split into phases: setup, pixel phase, entropy coding. Between them sits a buffer of coefficients. The pixel phase already works in bands of block rows and computes the context rows it needs itself – ready for parallel work, but still on one thread.

For now about 1.7 times slower than jpegli, but bit-identical in every test case.

Two kinds of identical

jpegli does not write the same bytes on every CPU: its SIMD library uses fused multiply-add on AVX2 but not on SSE, and sums vector lanes in an order that depends on the vector width. The port reproduces both classes – byte for byte jpegli on SSE2, and byte for byte jpegli on AVX2.

Still in use today

This flow is still in the code: ctJpeg --sequential runs the whole encoder as one task – the reference path for comparisons.

Run time · one photo, 39 MP · development

1.27 s · jpegli 0.75 s

Read, Write, Setup, Pixel, coefficients, Entropyjob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixelbandsPixelbandsPixelbandscoefficientsEntropystill serial Read, Write, Setup, Pixel, coefficients, EntropyI/O schedulerscheduler · 1 worker per threadI/O schedulerjob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixelbandsPixelbandsPixelbandscoefficientsEntropystill serial

Picture 4 of 9

The pixel phase goes parallel

The bands of the pixel phase become jobs, about four per worker. Each worker fetches the next band from a shared counter; the main task waits until all are done. The pixel phase drops from 0.76 to 0.12 s – 6.4 times as fast on 16 threads.

The limit now: the serial entropy coding, 0.48 s.

Why the bands are independent

The adaptive quantization only looks at small local windows around each block. When a band computes a few rows of context itself, its coefficients no longer depend on any other band.

Run time · one photo, 39 MP · development

0.70 s · 1.07× jpegli

Read, Write, Setup, Pixel, coefficients, Tokenize, Tokens, Stitch, Huffman, Write, Assemblejob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixelbandsPixelbandsPixelbandscoefficientsTokenizescans + AC bandsTokenizescans + AC bandsTokenizescans + AC bandsTokensStitchper scanHuffmanWriteper scanWriteper scanWriteper scanAssemble Read, Write, Setup, Pixel, coefficients, Tokenize, Tokens, Stitch, Huffman, Write, AssembleI/O schedulerscheduler · 1 worker per threadI/O schedulerjob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixelbandsPixelbandsPixelbandscoefficientsTokenizescans + AC bandsTokenizescans + AC bandsTokenizescans + AC bandsTokensStitchper scanHuffmanWriteper scanWriteper scanWriteper scanAssemble

Picture 5 of 9

Tokenizing and writing go parallel

The entropy block breaks apart. The progressive scans are independent of each other, one job each; the large scans are additionally split into bands. A short serial Stitch joins the bands exactly as jpegli would have coded them in one go. Huffman stays serial: its tables need the symbol counts of all scans.

Every scan ends on a byte boundary, so each one is written into its own buffer in parallel; Assemble puts them together.

What the stitch sews

Within a scan, runs of empty blocks are coded as one end-of-band run that crosses block rows – and jpegli caps a run at 32,767 blocks and splits refinement runs after 255 correction bits. A band starts with a fresh state and reports its beginning compactly; the stitch replays it with the real state.

Only half right

The profiler said then: limited by the 8 physical cores, not by memory. Picture 7 shows that this was only half true – memory became the limit as soon as each thread got faster.

Run time · one photo, 39 MP · development

0.31 s · 2.4× jpegli

Read, Write, Setup, Pixel, coefficients, Tokenize, Tokens, Stitch, Huffman, Write, Assemblejob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixel+ SIMDPixel+ SIMDPixel+ SIMDcoefficientsTokenize+ bit masksTokenize+ bit masksTokenize+ bit masksTokensStitchHuffmanWriteWriteWriteAssemble Read, Write, Setup, Pixel, coefficients, Tokenize, Tokens, Stitch, Huffman, Write, AssembleI/O schedulerscheduler · 1 worker per threadI/O schedulerjob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixel+ SIMDPixel+ SIMDPixel+ SIMDcoefficientsTokenize+ bit masksTokenize+ bit masksTokenize+ bit masksTokensStitchHuffmanWriteWriteWriteAssemble

Picture 6 of 9

Less work per node

The graph stays the same; every node gets faster. SIMD kernels in the pixel phase follow jpegli’s arithmetic exactly, and the tokenizer works on bit masks per block instead of looping over every coefficient.

On one thread ctJpeg is now 1.4 to 1.5 times as fast as jpegli. But the pixel phase only scales about threefold.

Nothing measurable

Ring buffers alone, which shrank each worker’s working set from 8.5 to 2 MB, brought nothing measurable. Why the pixel phase stopped scaling only became clear in the next picture.

Run time · one photo, 39 MP · development

0.24 s · 3.1× jpegli

Read, Write, Setup, Pixel, coefficients, Tokenize, Tokens, Stitch, Huffman, Count, cursors, Write, chunk bits, Splice, Assemblejob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixel+ tilesPixel+ tilesPixel+ tilescoefficientsTokenize+ DC bandsTokenize+ DC bandsTokenize+ DC bandsTokensStitchStitchStitchHuffmanCount+ newCount+ newCount+ newcursorsWrite+ in chunksWrite+ in chunksWrite+ in chunkschunk bitsSplice+ newSplice+ newSplice+ newAssemble Read, Write, Setup, Pixel, coefficients, Tokenize, Tokens, Stitch, Huffman, Count, cursors, Write, chunk bits, Splice, AssembleI/O schedulerscheduler · 1 worker per threadI/O schedulerjob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixel+ tilesPixel+ tilesPixel+ tilescoefficientsTokenize+ DC bandsTokenize+ DC bandsTokenize+ DC bandsTokensStitchStitchStitchHuffmanCount+ newCount+ newCount+ newcursorsWrite+ in chunksWrite+ in chunksWrite+ in chunkschunk bitsSplice+ newSplice+ newSplice+ newAssemble

Picture 7 of 9

Smaller jobs that fit the cache

In a fork-join, the last job of a stage decides when it ends. A job trace showed which one it was for every stage: the work was spread correctly, but single jobs were too large or ran into memory. So the jobs got smaller: the pixel phase works on tiles about 1,024 pixels wide, the DC scan is split into bands, and large scans are written in chunks – Count works out where each chunk starts, Splice joins the chunks bit-exactly.

The total time now falls all the way to 16 threads; before, the optimum was 6 to 8.

What the job trace found

  • The DC scan ran as one 32 ms job and finished last → DC bands: tokenizing 69 → 56 ms.
  • One worker wrote each large refinement scan → chunks: writing 28 → 10–13 ms.
  • 16 workers × 2 MB overflowed the 16 MB L3 cache → tiles: pixel phase 89 → 38–46 ms.
  • The last AC bands set the end → finer bands: tokenizing 39 → 33 ms.

Misleading clue

Before the tiles, bands of a single block row were faster with 8 threads. That looked like a load-balancing problem but was a cache effect – only the same image at a quarter of the width made it visible.

Run time · the same photo · tuning runs

0.14 s · 5.4× jpegli

Read, Write, Setup, Pixel, coefficients, Tokenize, Tokens, Stitch, Huffman, Count, cursors, Write, chunk bits, Splice, Assemble, Releasejob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixelPixelPixelcoefficientsTokenize+ frees the imageTokenize+ frees the imageTokenize+ frees the imageTokensStitchStitchStitchHuffmanCountCountCountcursorsWrite+ frees coefficientsWrite+ frees coefficientsWrite+ frees coefficientschunk bitsSpliceSpliceSpliceAssembleRelease+ newRelease+ newRelease+ new Read, Write, Setup, Pixel, coefficients, Tokenize, Tokens, Stitch, Huffman, Count, cursors, Write, chunk bits, Splice, Assemble, ReleaseI/O schedulerscheduler · 1 worker per threadI/O schedulerjob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixelPixelPixelcoefficientsTokenize+ frees the imageTokenize+ frees the imageTokenize+ frees the imageTokensStitchStitchStitchHuffmanCountCountCountcursorsWrite+ frees coefficientsWrite+ frees coefficientsWrite+ frees coefficientschunk bitsSpliceSpliceSpliceAssembleRelease+ newRelease+ newRelease+ new

Picture 8 of 9

Cleaning up goes parallel, too

With 16 threads about 0.05 s still lay outside the stages – most of it handing used memory back to the operating system. Now that happens in jobs, as early as possible: the input image as the first job of Tokenize, the coefficients during Write (nothing reads them any more), the rest in a new stage, Release.

Freeing memory drops from 26 to 7–8 ms. Over 14 test photos ctJpeg is now 4.8 to 5 times as fast as jpegli.

Why not simply skip it?

Without freeing, the cost only moves to the end of the process – measured: same total time. Spread over the workers it shrinks. On the side, committed memory fell from 1,241 to 408 MB.

Tried and dropped

Freeing the coefficient buffer in parallel pieces: the pieces of one memory region slow each other down, freeing went up from 11 to 14 ms.

Run time · the same photo · tuning runs

0.143 → 0.109 s · 6.9× jpegli

Read, Write, Setup, Pixel, coefficients + DC plane, Tokenize, Tokens, Stitch, Huffman, Count, cursors, Write, chunk bits, Splice, Assemble, Releasejob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixelPixelPixelcoefficients + DC planeTokenizeTokenizeTokenizeTokensStitchStitchStitchHuffmanCountCountCountcursorsWriteWriteWritechunk bitsSpliceSpliceSpliceAssembleReleaseReleaseRelease Read, Write, Setup, Pixel, coefficients + DC plane, Tokenize, Tokens, Stitch, Huffman, Count, cursors, Write, chunk bits, Splice, Assemble, ReleaseI/O schedulerscheduler · 1 worker per threadI/O schedulerjob k−1job kjob k+1ReadasynchronousWriteasynchronousSetupPixelPixelPixelcoefficients + DC planeTokenizeTokenizeTokenizeTokensStitchStitchStitchHuffmanCountCountCountcursorsWriteWriteWritechunk bitsSpliceSpliceSpliceAssembleReleaseReleaseRelease

Picture 9 of 9

The complete graph

Ten stages, three of them serial and short: Setup, Huffman, Assemble. Every other stage fans out to all cores and joins again before the next one starts. The 39-megapixel photo takes 0.108 s instead of 0.750 s – bit-identical to jpegli, and 3.4 to 4.7 times as fast as libjpeg-turbo on large images, with jpegli’s smaller files.

Why tasks here

At its core ctJpeg is fork-join: no job waits for another in the middle of its work. The encoder itself runs as a task and waits for each stage; on ThinkMeta ConcurrentTasks that wait suspends a fiber instead of blocking a thread, and the encoder reads as plain sequential code.

Other CPUs

Six photos, 99 megapixels in all:

  • Intel Core i9-11900K, 8 cores: 5.3×
  • Intel Xeon E3-1275 v6, 4 cores: 5.6×
  • Intel Core Ultra 7 155H, notebook: 4.2×

Bit-identical to jpegli on every machine.

Structure for free

The graph now lives in one place in the code. Measured against the build before: 0.96 to 1.04 in every case – within the noise. The structure costs nothing.

Speed-up over jpegli · benchmark package, 39 MP photo

6.9×

Try it yourself

ctJpeg, the jpegli reference and the benchmark for your own photos are on GitHub. Have a workload that should use every core? Get in touch.

An unhandled error has occurred. Reload