From 4e471e974c3ebcecb51bac8fe13a6eecff376551 Mon Sep 17 00:00:00 2001 From: Jason Rhinelander Date: Mon, 21 Sep 2026 18:43:44 -0300 Subject: [PATCH 1/3] CLAUDE.md: format with utils/format.sh before committing It pins clang-format-19; the system clang-format is a different version and produces different output, which fails the CI format check. --- CLAUDE.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/CLAUDE.md b/CLAUDE.md index 5be834ff..ca26240f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -79,5 +79,9 @@ Many headers come in pairs: `foo.h` (C API for FFI use) and `foo.hpp` (C++ API). ## Code Style +- **Format with `./utils/format.sh` before every commit**, and check it with `./utils/format.sh verify` + (exit 0 means clean). Do not run `clang-format` yourself: the script pins **clang-format-19** and + formats the whole source list, and other versions — including whatever `clang-format` points at — + produce different output that fails CI. - **Prefer DRY code**: when logic is duplicated across two or more call sites, extract a shared helper. Do this proactively when writing new code, not only when asked. - **Specify the shape upfront**: when asked to implement something that overlaps with existing code, identify and extract the shared piece before writing the new code, so duplication never appears in the first place. From a8d244420e6c97132752512e46c83eecbb29331a Mon Sep 17 00:00:00 2001 From: Jason Rhinelander Date: Mon, 21 Sep 2026 17:43:07 -0300 Subject: [PATCH 2/3] Add session::image::thumbhash image placeholders ThumbHash (https://github.com/evanw/thumbhash, MIT) encodes a ~20-25 byte placeholder that a receiving client renders as a blurred preview while the real attachment downloads. It is a better fit than BlurHash on every axis we care about: natively binary rather than base83, alpha support, and it spends its bits like a codec (7x7 DCT on luminance, 3x3 per chroma axis, 5x5 on alpha) instead of splitting them evenly across R, G and B. Measured against 13 photos, scored by RMSE against the original downsampled to 32px: thumbhash 20.5 bytes / 29.25, raw 4x4 pixels 48 bytes / 30.87, BlurHash 4x3 28 bytes / 36.68. Smallest and best of everything tried. There is no upstream C or C++ implementation, so this is a port, with three deliberate departures: - The DCT basis goes through src/image/det_trig.hpp rather than std::cos, and every a*b+c is an explicit std::fma. Neither std::cos's accuracy nor whether the compiler contracts a multiply-add is fixed by the standard, and both change the emitted hash: V8's cos differs from glibc's by 1 ulp on ~3.5% of the arguments this DCT uses, which alters ~7% of hashes. Since a thumbhash is sent to other people, that would make it a weak fingerprint of the sending platform. Exact integer argument reduction also happens to be 27x more accurate than the reference formulation, which hands libm a thrice-rounded angle. Verified identical across 7 GCC and 2 Clang configurations spanning -O0..-O3, -march=native, -mfma and -ffp-contract both fast and off. - The basis tables are hoisted out of the innermost loops; the reference rebuilds them per (cx, cy) when encoding and per pixel when decoding. Accumulation order is untouched, so this is bit-identical, and about 3x faster each way: encode 0.11ms at 32x32, decode 0.085ms at 32x24. - Decoding is resolution-independent, so there is no upscale-and-blur step and no image library needed on the receive path. Keep the decode small and let the UI scale it: at 32px, upscaling 10x with any linear filter differs from a full-size decode by RMSE 0.87, which is imperceptible. The hash therefore differs from upstream encoders by at most 1 in an individual 4-bit AC coefficient. Decoders remain fully interoperable; nothing requires two encoders to agree. THUMBHASH_REFERENCE_COS switches the basis back to the upstream std::cos formulation, which is what the algorithm is defined as and what its published vectors were produced with; it is there to keep that legible and measurable, and must not be defined in production. Two tests guard the parts that would otherwise fail silently. The pinned hash vectors catch a build that reintroduces a platform dependency, rather than letting it leak. The det_trig case checks the hand-written Taylor coefficients against libm across the DCT's whole argument set, because a mistranscribed factorial still yields smooth, plausible output -- just with the wrong values in every hash. The API is shaped so the safe path is the obvious one. decode() takes explicit dimensions, because a hash carries no usable record of the source's shape: what it stores is the DCT component counts, which a decoder needs in order to know how many luminance AC coefficients precede the chroma terms. That parameter tracks the shape loosely, so upstream exposes it as an aspect ratio, but one side is always pinned at the per-channel maximum and the complete set of results is 7/n for n in 1..7 and reciprocals (13 values), or 5/n for n in 1..5 with alpha. It also saturates, so a 1000x100 banner reads as 7.0 rather than 10.0 and a 1x100 sliver is wrong by a factor of fourteen. It is therefore named component_aspect_ratio(), and the overload that guesses an output shape from it is decode_unsized(); both are documented as diagnostics for a hash that arrives with no metadata, not as part of a display path. valid() and expected_size() are for a carrier storing peer-supplied values. A hash's length is not merely bounded but fully determined by its own header -- only six lengths are reachable at all -- so a receive path can check exact structural validity for the cost of a few bit extractions, which is strictly stronger than a length cap that any blob of the right size would pass. Note for the libvips work: encode() takes RGBA at <=100x100, so the caller does the downscaling, and the scaler is now the weakest link for reproducibility -- two clients that downscale the same photo differently produce different hashes. --- include/session/image/thumbhash.hpp | 144 +++++++++++ src/CMakeLists.txt | 1 + src/image/det_trig.hpp | 113 ++++++++ src/image/thumbhash.cpp | 385 ++++++++++++++++++++++++++++ tests/CMakeLists.txt | 1 + tests/test_image_thumbhash.cpp | 273 ++++++++++++++++++++ 6 files changed, 917 insertions(+) create mode 100644 include/session/image/thumbhash.hpp create mode 100644 src/image/det_trig.hpp create mode 100644 src/image/thumbhash.cpp create mode 100644 tests/test_image_thumbhash.cpp diff --git a/include/session/image/thumbhash.hpp b/include/session/image/thumbhash.hpp new file mode 100644 index 00000000..bca1f3d2 --- /dev/null +++ b/include/session/image/thumbhash.hpp @@ -0,0 +1,144 @@ +#pragma once + +#include +#include +#include +#include +#include + +namespace session::image::thumbhash { + +/// The largest input dimension `encode` accepts. ThumbHash's own limit: the output is at most a +/// 7x7 DCT, so a larger input costs time without adding information. +inline constexpr uint32_t max_input_dimension = 100; + +/// The largest hash `encode` can produce (7x7 luminance + 3x3 P + 3x3 Q + 5x5 alpha). +inline constexpr size_t max_hash_size = 25; + +/// A decoded placeholder: RGBA8, row-major, *not* alpha-premultiplied. +struct image { + uint32_t width = 0; + uint32_t height = 0; + std::vector rgba; +}; + +/// API: image/thumbhash/encode +/// +/// Encodes a small RGBA image into a ThumbHash: a ~20-25 byte placeholder that a receiving client +/// can render as a blurred preview before the real attachment arrives. +/// +/// The caller is responsible for scaling the source image down; `width` and `height` must each be +/// between 1 and `max_input_dimension`. Note that the scaler matters for reproducibility: two +/// clients that downscale the same photo differently will produce different hashes. See the note +/// on cross-platform reproducibility in the implementation. +/// +/// Inputs: +/// - `rgba` -- the pixels, row-major, 4 bytes per pixel, *not* alpha-premultiplied. Must be +/// exactly `width * height * 4` bytes. +/// - `width`, `height` -- the image dimensions, each in [1, `max_input_dimension`]. +/// +/// Outputs: +/// - the hash, between 5 and `max_hash_size` bytes. +/// +/// Throws `std::invalid_argument` if the dimensions are out of range or `rgba` is the wrong size. +std::vector encode(std::span rgba, uint32_t width, uint32_t height); + +/// API: image/thumbhash/decode +/// +/// Decodes a ThumbHash to RGBA pixels at the requested resolution. This is the function clients +/// should use: pass the attachment's real dimensions, scaled down small. +/// +/// The output shape is entirely yours to choose -- the DCT basis is continuous, so any grid is +/// valid and this simply samples it at the points you ask for. A hash carries no usable record +/// of the source's shape (see `component_aspect_ratio`), so the aspect ratio must come from the +/// attachment metadata, not from here. +/// +/// Keep the resolution small and let the UI scale the result. The hash holds at most 7 cycles +/// across the image, so a 32px decode captures essentially everything: upscaling it 10x with any +/// linear filter differs from a full-size decode by an RMSE of under 1, which is imperceptible, +/// and decode cost grows with the *output* pixel count (0.085ms at 32x24, 4.6ms at 240x180). Do +/// not scale it with nearest-neighbour filtering, which will look blocky. +/// +/// Inputs: +/// - `hash` -- the bytes produced by `encode`. +/// - `width`, `height` -- the output dimensions, each at least 1. +/// +/// Outputs: +/// - the decoded RGBA8 image. +/// +/// Throws `std::invalid_argument` if the hash is malformed or the dimensions are zero. +image decode(std::span hash, uint32_t width, uint32_t height); + +/// API: image/thumbhash/expected_size +/// +/// The exact number of bytes a well-formed hash starting with these bytes must occupy, or nullopt +/// if the leading bytes are not a usable header. +/// +/// A hash's length is not merely bounded, it is fully determined: the alpha flag and the 3-bit +/// component count fix how many coefficient nibbles follow. Only six lengths are reachable at +/// all -- 17, 19, 21, 23 or 24 bytes without alpha, and 23 or 25 with it. Reads the header only: +/// no DCT, no allocation. +/// +/// Note that `decode` is more permissive than this, requiring only that enough bytes are present. +/// Use `valid` at a trust boundary, where "exactly well-formed" is what you mean. +std::optional expected_size(std::span hash); + +/// API: image/thumbhash/valid +/// +/// Whether `hash` is structurally well-formed: a usable header, and exactly the length that +/// header implies. +/// +/// This is what a receive path should test before storing a value a remote peer supplied. It is +/// strictly stronger than a length cap, which a blob of the right size but arbitrary content +/// would pass, and it costs a handful of bit extractions. +bool valid(std::span hash); + +/// API: image/thumbhash/average_rgba +/// +/// Returns the average colour of the original image as RGBA8 (not alpha-premultiplied), read +/// straight out of the hash's DC terms. Much cheaper than decoding, and enough for a flat-colour +/// placeholder. +/// +/// Throws `std::invalid_argument` if the hash is malformed. +std::array average_rgba(std::span hash); + +// ─── Diagnostics and last resorts ──────────────────────────────────────────── +// +// The two functions below exist for inspecting a hash and for the case where one turns up with no +// accompanying metadata at all. Neither belongs on a normal display path. + +/// API: image/thumbhash/component_aspect_ratio +/// +/// Returns the ratio of the hash's DCT component counts. This is NOT the image's aspect ratio +/// and must never be used to lay one out. +/// +/// A ThumbHash does not record the source's shape. It records how many luminance AC coefficients +/// are present, because a decoder cannot parse the hash without knowing where the chroma terms +/// begin. The encoder picks that count roughly in proportion to the source, so a wide image gets +/// more horizontal detail than vertical, and dividing the two out gives something that correlates +/// with the shape. Upstream exposes it as `thumbHashToApproximateAspectRatio`; this is the same +/// value under a name that does not invite misuse. +/// +/// It is quantised to a handful of values, because one side is always pinned at the per-channel +/// maximum. The complete set of results is 7/n for n in 1..7 and their reciprocals without alpha +/// (13 values: 1/7 ... 1 ... 7), or 5/n for n in 1..5 and reciprocals with it (9 values). It also +/// saturates, so every image wider than 7:1 reports exactly 7.0: a 1000x100 banner reads as 7.0 +/// rather than 10.0, and a 1x100 sliver is wrong by a factor of fourteen. +/// +/// Throws `std::invalid_argument` if the hash is malformed. +double component_aspect_ratio(std::span hash); + +/// API: image/thumbhash/decode_unsized +/// +/// Decodes a hash when its dimensions are genuinely unknown, guessing the output shape from +/// `component_aspect_ratio` with `size` pixels on the longer edge. +/// +/// The guess is as coarse as that function is -- it can be wrong by more than a factor of ten -- +/// so a placeholder sized from the result will visibly jump when the real image loads. Use +/// `decode` with the attachment's real dimensions instead; reach for this only when there are +/// none, such as when inspecting a hash in isolation. +/// +/// Throws `std::invalid_argument` if the hash is malformed or `size` is zero. +image decode_unsized(std::span hash, uint32_t size = 32); + +} // namespace session::image::thumbhash diff --git a/src/CMakeLists.txt b/src/CMakeLists.txt index 70b42085..3be79243 100644 --- a/src/CMakeLists.txt +++ b/src/CMakeLists.txt @@ -53,6 +53,7 @@ endmacro() add_libsession_util_library(util file.cpp + image/thumbhash.cpp logging.cpp util.cpp ) diff --git a/src/image/det_trig.hpp b/src/image/det_trig.hpp new file mode 100644 index 00000000..28236a2e --- /dev/null +++ b/src/image/det_trig.hpp @@ -0,0 +1,113 @@ +#pragma once + +// Bit-reproducible trigonometry for the ThumbHash DCT. +// +// std::cos is not required by IEEE-754 to be correctly rounded, and implementations disagree: V8's +// fdlibm-derived cos differs from glibc's by 1 ulp on ~3.5% of the arguments this DCT uses, which +// is enough to change the emitted hash for ~7% of images. Since a thumbhash is sent to other +// people, that would make it a weak fingerprint of which platform produced it. +// +// Reproducibility here rests on two things: +// +// - Argument reduction is exact integer arithmetic, so no high-precision pi is needed and no +// rounding happens before the series. (This is also *more* accurate than the reference +// formulation, which hands libm an angle that has already been rounded three times: worst error +// against a high-precision reference is 1.5e-16 here versus 4.0e-15 there.) +// +// - Every remaining step is an IEEE-754 operation that the standard requires to be correctly +// rounded: +, -, * and fusedMultiplyAdd. The polynomials are written with explicit std::fma +// rather than `a + z * b`, because the latter is contractible: a compiler may or may not fuse +// it depending on -ffp-contract and on whether the target has an FMA instruction, and the fused +// and unfused results differ in the last bit. Spelling the fusion out makes the result +// independent of the build rather than dependent on a flag. +// +// What this still assumes: the default round-to-nearest-even rounding mode, and no x87 excess +// precision (i.e. SSE2 math on x86). Both hold on every platform we target. +// tests/test_image_thumbhash.cpp pins hash vectors so that a build which breaks either fails +// loudly rather than silently leaking. + +#include +#include + +namespace session::image::detail { + +namespace trig { + + // cos(u) and sin(u) by Taylor series on |u| <= pi/4, where the series is well conditioned. + // Terms run past u^18/18! ~= 3e-18, comfortably under a double's 2.2e-16 relative resolution. + // Evaluated by Horner from the smallest term up, each step a single fused multiply-add. + inline double cos_series(double z) { // z = u*u + double c = -1.0 / 6402373705728000; // -1/18! + c = std::fma(c, z, 1.0 / 20922789888000); + c = std::fma(c, z, -1.0 / 87178291200); + c = std::fma(c, z, 1.0 / 479001600); + c = std::fma(c, z, -1.0 / 3628800); + c = std::fma(c, z, 1.0 / 40320); + c = std::fma(c, z, -1.0 / 720); + c = std::fma(c, z, 1.0 / 24); + c = std::fma(c, z, -1.0 / 2); + return std::fma(c, z, 1.0); + } + + inline double sin_series(double u, double z) { // z = u*u + double s = 1.0 / 355687428096000; // 1/17! + s = std::fma(s, z, -1.0 / 1307674368000); + s = std::fma(s, z, 1.0 / 6227020800); + s = std::fma(s, z, -1.0 / 39916800); + s = std::fma(s, z, 1.0 / 362880); + s = std::fma(s, z, -1.0 / 5040); + s = std::fma(s, z, 1.0 / 120); + s = std::fma(s, z, -1.0 / 6); + s = std::fma(s, z, 1.0); + return u * s; + } + + // The double nearest pi. Only ever multiplies an already-reduced ratio in [0, 1/4], so its + // error contributes at most ~1e-16 to the angle. + inline constexpr double pi = 3.14159265358979323846; + +} // namespace trig + +/// cos(pi * n / d), for d > 0. Identical on every IEEE-754 platform, in any build configuration. +inline double cos_pi(int64_t n, int64_t d) { + // Reduce to one period exactly, in integers: cos has period 2 in n/d. + int64_t p = 2 * d; + n %= p; + if (n < 0) + n += p; + if (n > d) + n = p - n; // cos(2pi - t) == cos(t); now 0 <= n <= d + double sign = 1.0; + if (2 * n > d) { + n = d - n; // cos(pi - t) == -cos(t); now t <= pi/2 + sign = -1.0; + } + if (4 * n > d) { + // t > pi/4: swap to the sine of the complement, pi/2 - t == pi*(d-2n)/(2d), which keeps + // the series argument inside [0, pi/4]. + double u = trig::pi * (double(d - 2 * n) / double(2 * d)); + return sign * trig::sin_series(u, u * u); + } + double u = trig::pi * (double(n) / double(d)); + return sign * trig::cos_series(u * u); +} + +/// The DCT basis value cos(pi/n * k * (i + 0.5)) that ThumbHash needs. +/// +/// Upstream writes this literally as `cos(pi / n * k * (i + 0.5))`, which rounds the angle three +/// times before libm ever sees it. Restating it as an exact ratio lets cos_pi reduce it in +/// integers instead, which is both reproducible and rather more accurate. +/// +/// Defining THUMBHASH_REFERENCE_COS selects that upstream formulation instead. It is the +/// definition of the algorithm, and is kept so the computation we are approximating stays legible +/// and so the deterministic path can be measured against it, but it is not reproducible -- it +/// depends on the platform's std::cos. Production builds must not define it. +inline double cos_dct(int k, int i, int n) { +#ifdef THUMBHASH_REFERENCE_COS + return std::cos(trig::pi / n * k * (i + 0.5)); +#else + return cos_pi(int64_t(k) * (2 * i + 1), 2 * int64_t(n)); +#endif +} + +} // namespace session::image::detail diff --git a/src/image/thumbhash.cpp b/src/image/thumbhash.cpp new file mode 100644 index 00000000..43acf629 --- /dev/null +++ b/src/image/thumbhash.cpp @@ -0,0 +1,385 @@ +#include "session/image/thumbhash.hpp" + +#include +#include +#include + +#include "det_trig.hpp" + +// Port of ThumbHash (https://github.com/evanw/thumbhash, MIT), with two deliberate departures from +// the reference implementation, both verified to leave the output bit-identical except where +// noted: +// +// - The DCT basis is evaluated through session::image::detail (see det_trig.hpp), so that the +// result does not depend on the platform's libm or on FMA contraction. This *does* change the +// hash relative to the reference, by at most 1 in an individual 4-bit AC coefficient, in +// exchange for being identical everywhere. Decoders remain fully interoperable; nothing +// requires two encoders to agree. +// +// - The basis tables are hoisted out of the innermost loops. The reference rebuilds them inside +// the (cx, cy) loop when encoding and inside the per-pixel loop when decoding, recomputing each +// value many times over. Accumulation order is untouched, so this is a pure strength +// reduction: bit-identical, and roughly 3x faster in both directions. + +namespace session::image::thumbhash { + +namespace { + + using detail::cos_dct; + + // ECMAScript Math.round: round half toward +Infinity. Not the same as floor(v + 0.5), which + // rounds 0.49999999999999994 up to 1 because the addition itself rounds to 1.0. + int iround(double v) { + double f = std::floor(v); + return int(v - f >= 0.5 ? f + 1 : f); + } + + struct channel { + double dc = 0; + std::vector ac; + double scale = 0; + }; + + // Largest component count any channel uses: luminance goes up to 7x7, and max(3, lx) never + // exceeds that. + constexpr int max_components = 7; + + // Basis table indexed [c * n + i] for cos(pi/n * c * (i + 0.5)). Built once per image and + // shared by every channel. + std::vector basis(int n) { + std::vector t(size_t(max_components) * n); + for (int c = 0; c < max_components; c++) + for (int i = 0; i < n; i++) + t[size_t(c) * n + i] = cos_dct(c, i, n); + return t; + } + + // DCT over the w*h channel, keeping the triangular set of (cx, cy) terms the format keeps: + // cx * ny < nx * (ny - cy). + channel encode_channel( + const std::vector& ch, + int w, + int h, + int nx, + int ny, + const std::vector& fxt, + const std::vector& fyt) { + channel out; + for (int cy = 0; cy < ny; cy++) { + for (int cx = 0; cx * ny < nx * (ny - cy); cx++) { + double f = 0; + const double* fx = &fxt[size_t(cx) * w]; + const double* fyr = &fyt[size_t(cy) * h]; + for (int y = 0; y < h; y++) { + double fy = fyr[y]; + for (int x = 0; x < w; x++) + f = std::fma(ch[size_t(x) + size_t(y) * w] * fx[size_t(x)], fy, f); + } + f /= double(w) * h; + if (cx || cy) { + out.ac.push_back(f); + out.scale = std::max(out.scale, std::fabs(f)); + } else { + out.dc = f; + } + } + } + if (out.scale) { + double inv = 0.5 / out.scale; + for (auto& v : out.ac) + v = std::fma(inv, v, 0.5); + } + return out; + } + + uint8_t byte_at(std::span h, size_t i) { + return std::to_integer(h[i]); + } + + // The part of the header that determines the hash's layout: which channels are present and + // how many coefficients each carries. Everything after this point in the hash is a function + // of these, which is what lets `expected_size` work without decoding. + struct shape { + bool has_alpha; + int lx, ly; + size_t ac_start; + }; + + shape read_shape(std::span hash) { + shape s{}; + uint32_t h16 = byte_at(hash, 3) | (uint32_t(byte_at(hash, 4)) << 8); + s.has_alpha = (byte_at(hash, 2) & 0x80) != 0; + bool landscape = (h16 >> 15) != 0; + s.lx = std::max(3, landscape ? (s.has_alpha ? 5 : 7) : int(h16 & 7)); + s.ly = std::max(3, landscape ? int(h16 & 7) : (s.has_alpha ? 5 : 7)); + s.ac_start = s.has_alpha ? 6 : 5; + return s; + } + + // Number of AC coefficients a channel contributes: the triangular set the format keeps, minus + // the DC term. Must stay in step with the loop in `decode_at`'s decode_channel. + size_t ac_count(int nx, int ny) { + size_t n = 0; + for (int cy = 0; cy < ny; cy++) + for (int cx = cy ? 0 : 1; cx * ny < nx * (ny - cy); cx++) + n++; + return n; + } + + // Everything a decoder needs from the fixed-size part of a hash. + struct header { + double l_dc, p_dc, q_dc, a_dc; + double l_scale, p_scale, q_scale, a_scale; + bool has_alpha; + int lx, ly; + size_t ac_start; + }; + + header read_header(std::span hash) { + if (hash.size() < 5) + throw std::invalid_argument{"thumbhash: too short"}; + auto s = read_shape(hash); + header hd{}; + uint32_t h24 = byte_at(hash, 0) | (uint32_t(byte_at(hash, 1)) << 8) | + (uint32_t(byte_at(hash, 2)) << 16); + uint32_t h16 = byte_at(hash, 3) | (uint32_t(byte_at(hash, 4)) << 8); + hd.l_dc = (h24 & 63) / 63.0; + hd.p_dc = ((h24 >> 6) & 63) / 31.5 - 1; + hd.q_dc = ((h24 >> 12) & 63) / 31.5 - 1; + hd.l_scale = ((h24 >> 18) & 31) / 31.0; + hd.p_scale = ((h16 >> 3) & 63) / 63.0; + hd.q_scale = ((h16 >> 9) & 63) / 63.0; + hd.has_alpha = s.has_alpha; + hd.lx = s.lx; + hd.ly = s.ly; + hd.ac_start = s.ac_start; + if (hd.has_alpha && hash.size() < 6) + throw std::invalid_argument{"thumbhash: truncated before alpha DC"}; + hd.a_dc = hd.has_alpha ? (byte_at(hash, 5) & 15) / 15.0 : 1.0; + hd.a_scale = hd.has_alpha ? (byte_at(hash, 5) >> 4) / 15.0 : 0.0; + return hd; + } + +} // namespace + +std::optional expected_size(std::span hash) { + if (hash.size() < 5) + return std::nullopt; + auto s = read_shape(hash); + if (s.has_alpha && hash.size() < 6) + return std::nullopt; + size_t nibbles = ac_count(s.lx, s.ly) + 2 * ac_count(3, 3) + (s.has_alpha ? ac_count(5, 5) : 0); + return s.ac_start + (nibbles + 1) / 2; +} + +bool valid(std::span hash) { + auto n = expected_size(hash); + return n && *n == hash.size(); +} + +std::vector encode(std::span rgba, uint32_t width, uint32_t height) { + if (width < 1 || height < 1 || width > max_input_dimension || height > max_input_dimension) + throw std::invalid_argument{ + "thumbhash: dimensions must be between 1 and " + + std::to_string(max_input_dimension)}; + if (rgba.size() != size_t(width) * height * 4) + throw std::invalid_argument{"thumbhash: rgba buffer size does not match dimensions"}; + + int w = int(width), h = int(height); + auto px = [&](size_t i) { return double(std::to_integer(rgba[i])); }; + + double avg_r = 0, avg_g = 0, avg_b = 0, avg_a = 0; + for (size_t i = 0, j = 0; i < size_t(w) * h; i++, j += 4) { + double alpha = px(j + 3) / 255.0; + double s = alpha / 255; + avg_r = std::fma(s, px(j), avg_r); + avg_g = std::fma(s, px(j + 1), avg_g); + avg_b = std::fma(s, px(j + 2), avg_b); + avg_a += alpha; + } + if (avg_a) { + avg_r /= avg_a; + avg_g /= avg_a; + avg_b /= avg_a; + } + + bool has_alpha = avg_a < double(w) * h; + int l_limit = has_alpha ? 5 : 7; // fewer luminance components when alpha needs the space + int lx = std::max(1, iround(double(l_limit) * w / std::max(w, h))); + int ly = std::max(1, iround(double(l_limit) * h / std::max(w, h))); + + // Convert to LPQA, compositing over the average colour so that transparent regions do not drag + // the DCT toward black. + size_t n = size_t(w) * h; + std::vector l(n), p(n), q(n), a(n); + for (size_t i = 0, j = 0; i < n; i++, j += 4) { + double alpha = px(j + 3) / 255.0; + double s = alpha / 255, inv = 1 - alpha; + double r = std::fma(avg_r, inv, s * px(j)); + double g = std::fma(avg_g, inv, s * px(j + 1)); + double b = std::fma(avg_b, inv, s * px(j + 2)); + l[i] = (r + g + b) / 3; + p[i] = (r + g) / 2 - b; + q[i] = r - g; + a[i] = alpha; + } + + auto fxt = basis(w), fyt = basis(h); + auto lc = encode_channel(l, w, h, std::max(3, lx), std::max(3, ly), fxt, fyt); + auto pc = encode_channel(p, w, h, 3, 3, fxt, fyt); + auto qc = encode_channel(q, w, h, 3, 3, fxt, fyt); + channel ac_a; + if (has_alpha) + ac_a = encode_channel(a, w, h, 5, 5, fxt, fyt); + + bool landscape = w > h; + uint32_t header24 = uint32_t(iround(63 * lc.dc)) | + (uint32_t(iround(std::fma(31.5, pc.dc, 31.5))) << 6) | + (uint32_t(iround(std::fma(31.5, qc.dc, 31.5))) << 12) | + (uint32_t(iround(31 * lc.scale)) << 18) | (uint32_t(has_alpha) << 23); + uint32_t header16 = uint32_t(landscape ? ly : lx) | (uint32_t(iround(63 * pc.scale)) << 3) | + (uint32_t(iround(63 * qc.scale)) << 9) | (uint32_t(landscape) << 15); + + std::vector hash; + hash.reserve(max_hash_size); + auto push = [&hash](int v) { hash.push_back(std::byte(static_cast(v))); }; + push(header24 & 255); + push((header24 >> 8) & 255); + push((header24 >> 16) & 255); + push(header16 & 255); + push((header16 >> 8) & 255); + if (has_alpha) + push(iround(15 * ac_a.dc) | (iround(15 * ac_a.scale) << 4)); + + size_t ac_start = hash.size(), ac_index = 0; + auto put = [&](const std::vector& v) { + for (double f : v) { + size_t byte = ac_start + (ac_index >> 1); + if (byte >= hash.size()) + hash.resize(byte + 1, std::byte{0}); + // Two 4-bit coefficients per byte, low nibble first. + hash[byte] |= std::byte(static_cast(iround(15 * f) << ((ac_index & 1) << 2))); + ac_index++; + } + }; + put(lc.ac); + put(pc.ac); + put(qc.ac); + if (has_alpha) + put(ac_a.ac); + return hash; +} + +double component_aspect_ratio(std::span hash) { + if (hash.size() < 5) + throw std::invalid_argument{"thumbhash: too short"}; + bool has_alpha = byte_at(hash, 2) & 0x80; + bool landscape = byte_at(hash, 4) & 0x80; + int lx = landscape ? (has_alpha ? 5 : 7) : (byte_at(hash, 3) & 7); + int ly = landscape ? (byte_at(hash, 3) & 7) : (has_alpha ? 5 : 7); + if (ly == 0) + throw std::invalid_argument{"thumbhash: invalid component count"}; + return double(lx) / ly; +} + +std::array average_rgba(std::span hash) { + auto hd = read_header(hash); + double b = std::fma(-(2.0 / 3.0), hd.p_dc, hd.l_dc); + double r = (std::fma(3.0, hd.l_dc, -b) + hd.q_dc) / 2; + double g = r - hd.q_dc; + auto to8 = [](double v) { return std::byte(uint8_t(std::max(0.0, 255 * std::min(1.0, v)))); }; + return {to8(r), to8(g), to8(b), to8(hd.a_dc)}; +} + +image decode(std::span hash, uint32_t width, uint32_t height) { + if (width < 1 || height < 1) + throw std::invalid_argument{"thumbhash: output dimensions must be non-zero"}; + auto hd = read_header(hash); + int w = int(width), h = int(height); + + size_t ac_index = 0; + auto decode_channel = [&](int nx, int ny, double scale) { + std::vector ac; + for (int cy = 0; cy < ny; cy++) + for (int cx = cy ? 0 : 1; cx * ny < nx * (ny - cy); cx++) { + size_t byte = hd.ac_start + (ac_index >> 1); + if (byte >= hash.size()) + throw std::invalid_argument{"thumbhash: truncated in AC terms"}; + int nib = (byte_at(hash, byte) >> ((ac_index & 1) << 2)) & 15; + ac_index++; + ac.push_back((nib / 7.5 - 1) * scale); + } + return ac; + }; + // The format boosts chroma by 1.25x on decode to offset quantisation loss. + auto l_ac = decode_channel(hd.lx, hd.ly, hd.l_scale); + auto p_ac = decode_channel(3, 3, hd.p_scale * 1.25); + auto q_ac = decode_channel(3, 3, hd.q_scale * 1.25); + std::vector a_ac; + if (hd.has_alpha) + a_ac = decode_channel(5, 5, hd.a_scale); + + image out{width, height, std::vector(size_t(w) * h * 4)}; + + // The basis separates: fx depends only on (cx, x) and fy only on (cy, y). + int nx = std::max(hd.lx, hd.has_alpha ? 5 : 3) + 1; + int ny = std::max(hd.ly, hd.has_alpha ? 5 : 3) + 1; + std::vector fxt(size_t(w) * nx), fyt(size_t(h) * ny); + for (int x = 0; x < w; x++) + for (int cx = 0; cx < nx; cx++) + fxt[size_t(x) * nx + cx] = cos_dct(cx, x, w); + for (int y = 0; y < h; y++) + for (int cy = 0; cy < ny; cy++) + fyt[size_t(y) * ny + cy] = cos_dct(cy, y, h); + + for (int y = 0, i = 0; y < h; y++) { + const double* fy = &fyt[size_t(y) * ny]; + for (int x = 0; x < w; x++, i += 4) { + double l = hd.l_dc, p = hd.p_dc, q = hd.q_dc, a = hd.a_dc; + const double* fx = &fxt[size_t(x) * nx]; + + for (int cy = 0, j = 0; cy < hd.ly; cy++) { + double fy2 = fy[size_t(cy)] * 2; + for (int cx = cy ? 0 : 1; cx * hd.ly < hd.lx * (hd.ly - cy); cx++, j++) + l = std::fma(l_ac[size_t(j)] * fx[size_t(cx)], fy2, l); + } + for (int cy = 0, j = 0; cy < 3; cy++) { + double fy2 = fy[size_t(cy)] * 2; + for (int cx = cy ? 0 : 1; cx < 3 - cy; cx++, j++) { + double f = fx[size_t(cx)] * fy2; + p = std::fma(p_ac[size_t(j)], f, p); + q = std::fma(q_ac[size_t(j)], f, q); + } + } + if (hd.has_alpha) + for (int cy = 0, j = 0; cy < 5; cy++) { + double fy2 = fy[size_t(cy)] * 2; + for (int cx = cy ? 0 : 1; cx < 5 - cy; cx++, j++) + a = std::fma(a_ac[size_t(j)] * fx[size_t(cx)], fy2, a); + } + + double b = std::fma(-(2.0 / 3.0), p, l); + double r = (std::fma(3.0, l, -b) + q) / 2; + double g = r - q; + auto to8 = [](double v) { + return std::byte(uint8_t(std::max(0.0, 255 * std::min(1.0, v)))); + }; + out.rgba[size_t(i)] = to8(r); + out.rgba[size_t(i) + 1] = to8(g); + out.rgba[size_t(i) + 2] = to8(b); + out.rgba[size_t(i) + 3] = to8(a); + } + } + return out; +} + +image decode_unsized(std::span hash, uint32_t size) { + if (size < 1) + throw std::invalid_argument{"thumbhash: output size must be non-zero"}; + double ratio = component_aspect_ratio(hash); + int w = iround(ratio > 1 ? size : size * ratio); + int h = iround(ratio > 1 ? size / ratio : size); + return decode(hash, uint32_t(std::max(1, w)), uint32_t(std::max(1, h))); +} + +} // namespace session::image::thumbhash diff --git a/tests/CMakeLists.txt b/tests/CMakeLists.txt index ed631566..a3a9942c 100644 --- a/tests/CMakeLists.txt +++ b/tests/CMakeLists.txt @@ -43,6 +43,7 @@ set(LIB_SESSION_UTESTS_SOURCES test_group_info.cpp test_group_members.cpp test_hash.cpp + test_image_thumbhash.cpp #test_logging.cpp # Handled separately, see below test_multi_encrypt.cpp test_mnemonics.cpp diff --git a/tests/test_image_thumbhash.cpp b/tests/test_image_thumbhash.cpp new file mode 100644 index 00000000..bf3e11cb --- /dev/null +++ b/tests/test_image_thumbhash.cpp @@ -0,0 +1,273 @@ +#include + +#include +#include +#include +#include +#include + +#include "../src/image/det_trig.hpp" +#include "session/image/thumbhash.hpp" + +using namespace session::image; +using namespace std::literals; + +namespace { + +// A synthetic image with enough structure in every channel that all the DCT terms are exercised: +// a diagonal luminance ramp, opposing red/blue gradients, and a green blob. +std::vector test_image(int w, int h, bool alpha) { + std::vector px(size_t(w) * h * 4); + for (int y = 0; y < h; y++) + for (int x = 0; x < w; x++) { + double fx = double(x) / std::max(1, w - 1), fy = double(y) / std::max(1, h - 1); + double blob = std::exp(-8 * ((fx - 0.3) * (fx - 0.3) + (fy - 0.7) * (fy - 0.7))); + auto set = [&](int c, double v) { + px[(size_t(y) * w + x) * 4 + c] = + std::byte(uint8_t(std::clamp(v, 0.0, 1.0) * 255 + 0.5)); + }; + set(0, 0.9 * fx + 0.05); + set(1, blob); + set(2, 0.9 * (1 - fy) + 0.05); + // A soft circular cutout, so the alpha DCT has something to encode. + set(3, alpha ? std::clamp(1.6 - 3.0 * std::hypot(fx - 0.5, fy - 0.5), 0.0, 1.0) : 1.0); + } + return px; +} + +} // namespace + +// det_trig replaces std::cos with a hand-written Taylor series, so the coefficient table is a +// transcription that nothing else would catch if it were wrong: a bad term still produces smooth, +// plausible-looking output, just with the wrong values baked into every hash. Check it against +// libm over exactly the arguments the DCT uses. The tolerance is far looser than the ~4e-15 the +// two genuinely differ by (libm is handed a thrice-rounded angle) and far tighter than any +// consequential typo, which would show up at 1e-5 or worse. +TEST_CASE("det_trig cosine agrees with libm", "[image][thumbhash][det_trig]") { + using session::image::detail::cos_dct; + using session::image::detail::cos_pi; + + double worst = 0; + for (int n = 1; n <= 100; n++) + for (int k = 0; k <= 6; k++) + for (int i = 0; i < n; i++) + worst = std::max( + worst, std::abs(cos_dct(k, i, n) - std::cos(M_PI / n * k * (i + 0.5)))); + CHECK(worst < 1e-13); + + // Exact values the reduction must land on, covering every branch: each octant, the sign flip + // past pi/2, and wrap-around of the integer reduction. + CHECK(cos_pi(0, 7) == 1.0); + CHECK(cos_pi(7, 7) == -1.0); + CHECK(cos_pi(14, 7) == 1.0); // 2pi + CHECK(cos_pi(-7, 7) == -1.0); // negative numerator + CHECK(cos_pi(700, 7) == 1.0); // many periods + CHECK(cos_pi(1, 2) == Catch::Approx(0.0).margin(1e-16)); // pi/2 + CHECK(cos_pi(1, 3) == Catch::Approx(0.5)); // pi/3 + CHECK(cos_pi(2, 3) == Catch::Approx(-0.5)); // 2pi/3 + CHECK(cos_pi(1, 4) == Catch::Approx(std::sqrt(0.5))); // pi/4, the branch boundary + CHECK(cos_pi(3, 4) == Catch::Approx(-std::sqrt(0.5))); + CHECK(cos_pi(1, 6) == Catch::Approx(std::sqrt(3.0) / 2)); // pi/6, cos branch + CHECK(cos_pi(5, 6) == Catch::Approx(-std::sqrt(3.0) / 2)); +} + +TEST_CASE("thumbhash round-trip", "[image][thumbhash]") { + for (auto alpha : {false, true}) { + for (auto [w, h] : {std::pair{32, 32}, {64, 48}, {48, 64}, {100, 100}, {7, 3}, {1, 1}}) { + auto px = test_image(w, h, alpha); + auto hash = thumbhash::encode(px, w, h); + + CHECK(hash.size() >= 5); + CHECK(hash.size() <= thumbhash::max_hash_size); + // Alpha is only carried when some pixel is actually translucent. + CHECK(((std::to_integer(hash[2]) & 0x80) != 0) == alpha); + + auto img = thumbhash::decode_unsized(hash, 32); + CHECK(img.rgba.size() == size_t(img.width) * img.height * 4); + CHECK(img.width >= 1); + CHECK(img.height >= 1); + CHECK(std::max(img.width, img.height) == 32); + + auto exact = thumbhash::decode(hash, 20, 10); + CHECK(exact.width == 20); + CHECK(exact.height == 10); + CHECK(exact.rgba.size() == 20 * 10 * 4); + } + } +} + +namespace { + +// Per-channel mean of a decoded image. +std::array channel_means(const thumbhash::image& im) { + std::array sum{}; + size_t n = size_t(im.width) * im.height; + for (size_t i = 0; i < n; i++) + for (int c = 0; c < 4; c++) + sum[c] += std::to_integer(im.rgba[i * 4 + c]); + for (auto& v : sum) + v /= double(n); + return sum; +} + +} // namespace + +TEST_CASE("thumbhash decode is resolution independent", "[image][thumbhash]") { + auto hash = thumbhash::encode(test_image(64, 48, false), 64, 48); + + // The DCT basis is continuous, so decoding at different resolutions samples the same + // underlying surface: the overall colour should not shift with the output grid. + auto coarse = channel_means(thumbhash::decode(hash, 16, 12)); + auto fine = channel_means(thumbhash::decode(hash, 160, 120)); + for (int c = 0; c < 4; c++) + CHECK(std::abs(coarse[c] - fine[c]) < 2.0); +} + +TEST_CASE("thumbhash average colour matches the decoded mean", "[image][thumbhash]") { + for (auto alpha : {false, true}) { + auto hash = thumbhash::encode(test_image(40, 40, alpha), 40, 40); + auto avg = thumbhash::average_rgba(hash); + // average_rgba reads the DC terms, which are the mean of the image; the AC terms very + // nearly integrate away over the whole grid, so the decoded mean should land close by. + // Clamping and the 1.25x chroma boost keep it from being exact. + auto mean = channel_means(thumbhash::decode(hash, 32, 32)); + for (int c = 0; c < 4; c++) + CHECK(std::abs(double(std::to_integer(avg[c])) - mean[c]) <= 24.0); + } +} + +TEST_CASE("thumbhash aspect ratio", "[image][thumbhash]") { + auto wide = thumbhash::encode(test_image(96, 32, false), 96, 32); + auto tall = thumbhash::encode(test_image(32, 96, false), 32, 96); + auto square = thumbhash::encode(test_image(48, 48, false), 48, 48); + CHECK(thumbhash::component_aspect_ratio(wide) > 1.5); + CHECK(thumbhash::component_aspect_ratio(tall) < 0.67); + CHECK(thumbhash::component_aspect_ratio(square) == 1.0); + + // decode_unsized() should follow the component ratio. + auto w = thumbhash::decode_unsized(wide, 32); + CHECK(w.width == 32); + CHECK(w.height < 32); +} + +// component_aspect_ratio() reports the DCT component counts, not the image's shape, so it can only +// ever return 7/n for n in 1..7 (or 5/n with alpha) and saturates beyond that. Pinned here so the +// limitation stays visible rather than being discovered in a client: a 1000x637 photo scaled to +// 100x64 by the sender comes back as 7/4, and anything wider than 7:1 comes back as exactly 7. +TEST_CASE("thumbhash component ratio is not an aspect ratio", "[image][thumbhash]") { + struct { + int w, h; + double expected; + } const cases[] = { + {100, 64, 7.0 / 4}, // a 1000x637 source: true 1.570 + {100, 56, 7.0 / 4}, // 1920x1080: true 1.778 + {100, 75, 7.0 / 5}, // 4032x3024: true 1.333 + {100, 25, 7.0 / 2}, // 2000x500: true 4.0 + {100, 100, 1.0}, // + {64, 100, 4.0 / 7}, // portrait + {100, 10, 7.0}, // 1000x100 banner: true 10.0, saturated + {100, 1, 7.0}, // 100:1, still 7.0 + }; + for (const auto& c : cases) { + auto hash = thumbhash::encode(test_image(c.w, c.h, false), c.w, c.h); + INFO(c.w << "x" << c.h); + CHECK(thumbhash::component_aspect_ratio(hash) == Catch::Approx(c.expected)); + } + + // The whole representable set, so a change to it cannot slip through silently. + std::set seen; + for (int w = 1; w <= 100; w++) + for (int h = 1; h <= 100; h++) + seen.insert(thumbhash::component_aspect_ratio( + thumbhash::encode(test_image(w, h, false), w, h))); + CHECK(seen.size() == 13); + CHECK(*seen.begin() == Catch::Approx(1.0 / 7)); + CHECK(*seen.rbegin() == Catch::Approx(7.0)); + + // Callers who know the real size should bypass it entirely. + auto hash = thumbhash::encode(test_image(100, 64, false), 100, 64); + auto img = thumbhash::decode(hash, 32, 20); // the true 1000x637 shape + CHECK(img.width == 32); + CHECK(img.height == 20); +} + +// A hash's length is a pure function of its header, so a carrier storing peer-supplied values can +// check exact structural validity rather than a loose cap -- a cap would pass a blob of the right +// size but arbitrary content. +TEST_CASE("thumbhash validity is exact, not a length cap", "[image][thumbhash]") { + std::set lengths; + for (auto alpha : {false, true}) + for (int w = 1; w <= 100; w += 7) + for (int h = 1; h <= 100; h += 7) { + auto hash = thumbhash::encode(test_image(w, h, alpha), w, h); + INFO(w << "x" << h << (alpha ? " rgba" : " rgb")); + CHECK(thumbhash::valid(hash)); + CHECK(thumbhash::expected_size(hash) == hash.size()); + CHECK(hash.size() <= thumbhash::max_hash_size); + lengths.insert(hash.size()); + + // One byte short or one byte long is not the length its own header demands. + CHECK_FALSE(thumbhash::valid(std::span{hash}.first(hash.size() - 1))); + auto padded = hash; + padded.push_back(std::byte{0}); + CHECK_FALSE(thumbhash::valid(padded)); + } + + // Only eight lengths are reachable at all. + CHECK(lengths == std::set{17, 19, 21, 23, 24, 25}); + + // Too short to hold a header at all. + for (size_t n = 0; n < 5; n++) + CHECK(thumbhash::expected_size(std::vector(n)) == std::nullopt); + // Header says alpha, but the alpha DC byte is missing. + std::vector alpha_hdr(5, std::byte{0}); + alpha_hdr[2] = std::byte{0x80}; + CHECK(thumbhash::expected_size(alpha_hdr) == std::nullopt); +} + +TEST_CASE("thumbhash rejects bad input", "[image][thumbhash]") { + auto px = test_image(8, 8, false); + CHECK_THROWS_AS(thumbhash::encode(px, 8, 7), std::invalid_argument); // size mismatch + CHECK_THROWS_AS(thumbhash::encode(px, 0, 8), std::invalid_argument); // zero dimension + CHECK_THROWS_AS(thumbhash::encode(test_image(101, 1, false), 101, 1), std::invalid_argument); + + auto hash = thumbhash::encode(px, 8, 8); + auto prefix = [&hash](size_t n) { return std::span{hash}.first(n); }; + CHECK_THROWS_AS(thumbhash::decode(prefix(4), 8, 8), std::invalid_argument); + CHECK_THROWS_AS(thumbhash::decode({}, 8, 8), std::invalid_argument); + CHECK_THROWS_AS(thumbhash::decode(hash, 0, 4), std::invalid_argument); + CHECK_THROWS_AS(thumbhash::decode_unsized(hash, 0), std::invalid_argument); + CHECK_THROWS_AS(thumbhash::decode_unsized(prefix(4)), std::invalid_argument); + // Truncated in the middle of the AC nibbles. + CHECK_THROWS_AS(thumbhash::decode(prefix(6), 8, 8), std::invalid_argument); +} + +// These vectors pin the floating-point behaviour of the encoder. A thumbhash is sent to other +// people, so if it varied with the platform's libm or with whether the compiler contracted a*b+c +// into an FMA, it would leak a bit about which client produced it. src/image/det_trig.hpp removes +// both sources; this test is the backstop that makes a build which reintroduces one fail loudly +// instead of silently fingerprinting users. +// +// If this fails on a new platform, do not regenerate the vectors: find out which assumption broke +// (most likely the rounding mode, or x87 excess precision from a non-SSE2 x86 build). +TEST_CASE("thumbhash is bit-reproducible", "[image][thumbhash]") { + struct { + int w, h; + bool alpha; + std::string_view expected; + } const vectors[] = { + {32, 32, false, "1b67067f262062763f9a885289d879678789670777909a09"sv}, + {64, 48, false, "1b67067da62062763f9a885289d87976767007a999"sv}, + {48, 64, false, "1b67067d2620623f9a2895d879769878767007a999"sv}, + {32, 32, true, "9d3782250a2617b138d871efbd777007b99808688888808968"sv}, + {64, 48, true, "9d3782248c3717a138d871ef7d0777908b8980868808988806"sv}, + {100, 100, false, "1b67067f261061763f9a885189e879678789670777909a09"sv}, + {17, 5, true, "9d378219883506b037d883777007b98828788888808a78"sv}, + {1, 1, false, "d5102ad70708f808888888808f8088f80888808ff8088800"sv}, + }; + for (const auto& v : vectors) { + auto hash = thumbhash::encode(test_image(v.w, v.h, v.alpha), v.w, v.h); + INFO(v.w << "x" << v.h << (v.alpha ? " rgba" : " rgb")); + CHECK(oxenc::to_hex(hash) == v.expected); + } +} From d610c1949c12fabf9d212df74ba5b90d6806d692 Mon Sep 17 00:00:00 2001 From: Jason Rhinelander Date: Mon, 21 Sep 2026 21:42:12 -0300 Subject: [PATCH 3/3] thumbhash: fix decode overflow, tighten valid(), pin the decoder Review fixes for #172. decode() indexed its output with an `int` that wrapped at 2^29 output pixels, after which the stores went outside the buffer: decode(hash, 32768, 16384) is a signed overflow under UBSan, and decode(hash, 32768, 20480) segfaults. The index is now size_t, the byte count is computed in uint64_t and rejected if it would not fit a size_t (which is what would otherwise under-allocate on a 32-bit target and overflow the same way at a much lower pixel count), and the dimensions are bounded to int32 so the `int` loop counters and det_trig's 2*i+1 cannot overflow either. No policy cap: a big decode is merely slow, as intended -- the header now says to decode at 100x100 or less and scale the result up with libvips, which suits a blurred image well. valid() accepted a stored component count of 0. No conforming encoder emits one -- both this encoder and upstream clamp to max(1, ...) -- but it is a bit-field a peer controls, and while the decoder's max(3, ...) clamp waves it through, component_aspect_ratio reads the field unclamped and rejects it. A carrier that stored a value on the strength of valid() could then get an exception out of the documented no-metadata path. expected_size() now rejects it, so the trust boundary agrees with the rest of the API. Nothing tested the decoder. The pinned vectors cover the encoder; the round-trip test checks buffer sizes but no pixels, resolution-independence compares the decoder against itself, and the average-colour test compares two quantities that both derive from the same DC terms. A decoder producing consistently wrong pixels passed the entire suite. Its output is now pinned at 8x6 for the same eight vectors. Also: extract lpqa_to_rgba8, which decode() and average_rgba() had character-for-character in common, so the flat-colour placeholder cannot drift from the image it stands in for; fix a comment naming decode_at, which was renamed to decode; and correct a test comment saying "eight lengths" about a set of six. The encoder is untouched: its pinned vectors are unchanged across -O0..-O3, -march=native, -mfma and -ffp-contract fast/off. --- include/session/image/thumbhash.hpp | 20 ++++-- src/image/thumbhash.cpp | 67 ++++++++++++------ tests/test_image_thumbhash.cpp | 104 +++++++++++++++++++++++++++- 3 files changed, 163 insertions(+), 28 deletions(-) diff --git a/include/session/image/thumbhash.hpp b/include/session/image/thumbhash.hpp index bca1f3d2..06778f4e 100644 --- a/include/session/image/thumbhash.hpp +++ b/include/session/image/thumbhash.hpp @@ -53,20 +53,26 @@ std::vector encode(std::span rgba, uint32_t width, u /// of the source's shape (see `component_aspect_ratio`), so the aspect ratio must come from the /// attachment metadata, not from here. /// -/// Keep the resolution small and let the UI scale the result. The hash holds at most 7 cycles -/// across the image, so a 32px decode captures essentially everything: upscaling it 10x with any -/// linear filter differs from a full-size decode by an RMSE of under 1, which is imperceptible, -/// and decode cost grows with the *output* pixel count (0.085ms at 32x24, 4.6ms at 240x180). Do -/// not scale it with nearest-neighbour filtering, which will look blocky. +/// **Decode small -- 100x100 or less -- and scale the result up.** There is no enforced maximum, +/// but decoding evaluates the DCT basis per output pixel, so cost grows with the output area and +/// nothing is gained by it: the hash holds at most 7 cycles across the image, so a 32px decode +/// already captures essentially everything it contains. Upscaling that 10x with any linear filter +/// differs from a full-size decode by an RMSE under 1, which is imperceptible, and costs 0.085ms +/// against 4.6ms for decoding 240x180 directly. To fill a larger area, decode small and scale +/// with libvips (or the platform's own scaler), which handles a blurred image like this very well. +/// Do not scale with nearest-neighbour filtering, which will look blocky. /// /// Inputs: /// - `hash` -- the bytes produced by `encode`. -/// - `width`, `height` -- the output dimensions, each at least 1. +/// - `width`, `height` -- the output dimensions, each at least 1. There is no upper limit beyond +/// what the arithmetic can represent, but see above: these are meant to be small. /// /// Outputs: /// - the decoded RGBA8 image. /// -/// Throws `std::invalid_argument` if the hash is malformed or the dimensions are zero. +/// Throws `std::invalid_argument` if the hash is malformed, the dimensions are zero, or they are +/// large enough that the output size would not fit in a `size_t`. A merely very large request is +/// not rejected, and will throw `std::bad_alloc` or `std::length_error` if it cannot be allocated. image decode(std::span hash, uint32_t width, uint32_t height); /// API: image/thumbhash/expected_size diff --git a/src/image/thumbhash.cpp b/src/image/thumbhash.cpp index 43acf629..aec5df43 100644 --- a/src/image/thumbhash.cpp +++ b/src/image/thumbhash.cpp @@ -2,6 +2,7 @@ #include #include +#include #include #include "det_trig.hpp" @@ -101,6 +102,7 @@ namespace { // of these, which is what lets `expected_size` work without decoding. struct shape { bool has_alpha; + int stored; // the raw 3-bit count, before the decoder's max(3, ...) clamp int lx, ly; size_t ac_start; }; @@ -109,15 +111,32 @@ namespace { shape s{}; uint32_t h16 = byte_at(hash, 3) | (uint32_t(byte_at(hash, 4)) << 8); s.has_alpha = (byte_at(hash, 2) & 0x80) != 0; + s.stored = int(h16 & 7); bool landscape = (h16 >> 15) != 0; - s.lx = std::max(3, landscape ? (s.has_alpha ? 5 : 7) : int(h16 & 7)); - s.ly = std::max(3, landscape ? int(h16 & 7) : (s.has_alpha ? 5 : 7)); + // The DCT needs at least three components per axis, so the decoder works with a clamped + // count even where the encoder stored 1 or 2. component_aspect_ratio deliberately reports + // the unclamped value instead, which is what upstream does and what makes 7:1 reachable. + s.lx = std::max(3, landscape ? (s.has_alpha ? 5 : 7) : s.stored); + s.ly = std::max(3, landscape ? s.stored : (s.has_alpha ? 5 : 7)); s.ac_start = s.has_alpha ? 6 : 5; return s; } + // LPQA -> RGBA8. Used both per-pixel by `decode` and once by `average_rgba` for the DC terms; + // sharing it is what keeps the flat-colour placeholder from drifting away from the decoded + // image it stands in for. + std::array lpqa_to_rgba8(double l, double p, double q, double a) { + double b = std::fma(-(2.0 / 3.0), p, l); + double r = (std::fma(3.0, l, -b) + q) / 2; + double g = r - q; + auto to8 = [](double v) { + return std::byte(uint8_t(std::max(0.0, 255 * std::min(1.0, v)))); + }; + return {to8(r), to8(g), to8(b), to8(a)}; + } + // Number of AC coefficients a channel contributes: the triangular set the format keeps, minus - // the DC term. Must stay in step with the loop in `decode_at`'s decode_channel. + // the DC term. Must stay in step with the loop in `decode`'s decode_channel. size_t ac_count(int nx, int ny) { size_t n = 0; for (int cy = 0; cy < ny; cy++) @@ -168,6 +187,12 @@ std::optional expected_size(std::span hash) { auto s = read_shape(hash); if (s.has_alpha && hash.size() < 6) return std::nullopt; + // No conforming encoder emits a component count of 0 -- both this one and upstream clamp to + // max(1, ...). The decoder's max(3, ...) would silently accept it, but component_aspect_ratio + // reads the field unclamped and rejects it, so treating it as well-formed here would let a + // value through the trust boundary that the rest of the API refuses. + if (s.stored == 0) + return std::nullopt; size_t nibbles = ac_count(s.lx, s.ly) + 2 * ac_count(3, 3) + (s.has_alpha ? ac_count(5, 5) : 0); return s.ac_start + (nibbles + 1) / 2; } @@ -284,16 +309,20 @@ double component_aspect_ratio(std::span hash) { std::array average_rgba(std::span hash) { auto hd = read_header(hash); - double b = std::fma(-(2.0 / 3.0), hd.p_dc, hd.l_dc); - double r = (std::fma(3.0, hd.l_dc, -b) + hd.q_dc) / 2; - double g = r - hd.q_dc; - auto to8 = [](double v) { return std::byte(uint8_t(std::max(0.0, 255 * std::min(1.0, v)))); }; - return {to8(r), to8(g), to8(b), to8(hd.a_dc)}; + return lpqa_to_rgba8(hd.l_dc, hd.p_dc, hd.q_dc, hd.a_dc); } image decode(std::span hash, uint32_t width, uint32_t height) { if (width < 1 || height < 1) throw std::invalid_argument{"thumbhash: output dimensions must be non-zero"}; + // No policy cap on the size -- a big decode is merely slow -- but the arithmetic below has to + // survive one. Both the byte count and the pixel index must fit in size_t, which they do not + // on a 32-bit target for anything over ~1G pixels, and `int` counters would overflow at 2^29 + // pixels on any target. + uint64_t bytes = uint64_t(width) * height * 4; + if (bytes > std::numeric_limits::max() || + uint64_t(std::max(width, height)) > uint64_t(std::numeric_limits::max())) + throw std::invalid_argument{"thumbhash: output dimensions too large"}; auto hd = read_header(hash); int w = int(width), h = int(height); @@ -319,7 +348,7 @@ image decode(std::span hash, uint32_t width, uint32_t height) { if (hd.has_alpha) a_ac = decode_channel(5, 5, hd.a_scale); - image out{width, height, std::vector(size_t(w) * h * 4)}; + image out{width, height, std::vector(size_t(bytes))}; // The basis separates: fx depends only on (cx, x) and fy only on (cy, y). int nx = std::max(hd.lx, hd.has_alpha ? 5 : 3) + 1; @@ -332,8 +361,11 @@ image decode(std::span hash, uint32_t width, uint32_t height) { for (int cy = 0; cy < ny; cy++) fyt[size_t(y) * ny + cy] = cos_dct(cy, y, h); - for (int y = 0, i = 0; y < h; y++) { + // `i` is size_t, not int: at 2^29 output pixels a 32-bit counter wraps and the stores below + // land outside the buffer. + for (int y = 0; y < h; y++) { const double* fy = &fyt[size_t(y) * ny]; + size_t i = size_t(y) * w * 4; for (int x = 0; x < w; x++, i += 4) { double l = hd.l_dc, p = hd.p_dc, q = hd.q_dc, a = hd.a_dc; const double* fx = &fxt[size_t(x) * nx]; @@ -358,16 +390,11 @@ image decode(std::span hash, uint32_t width, uint32_t height) { a = std::fma(a_ac[size_t(j)] * fx[size_t(cx)], fy2, a); } - double b = std::fma(-(2.0 / 3.0), p, l); - double r = (std::fma(3.0, l, -b) + q) / 2; - double g = r - q; - auto to8 = [](double v) { - return std::byte(uint8_t(std::max(0.0, 255 * std::min(1.0, v)))); - }; - out.rgba[size_t(i)] = to8(r); - out.rgba[size_t(i) + 1] = to8(g); - out.rgba[size_t(i) + 2] = to8(b); - out.rgba[size_t(i) + 3] = to8(a); + auto rgba = lpqa_to_rgba8(l, p, q, a); + out.rgba[i] = rgba[0]; + out.rgba[i + 1] = rgba[1]; + out.rgba[i + 2] = rgba[2]; + out.rgba[i + 3] = rgba[3]; } } return out; diff --git a/tests/test_image_thumbhash.cpp b/tests/test_image_thumbhash.cpp index bf3e11cb..d1543d0c 100644 --- a/tests/test_image_thumbhash.cpp +++ b/tests/test_image_thumbhash.cpp @@ -213,7 +213,7 @@ TEST_CASE("thumbhash validity is exact, not a length cap", "[image][thumbhash]") CHECK_FALSE(thumbhash::valid(padded)); } - // Only eight lengths are reachable at all. + // Only six lengths are reachable at all. CHECK(lengths == std::set{17, 19, 21, 23, 24, 25}); // Too short to hold a header at all. @@ -223,6 +223,19 @@ TEST_CASE("thumbhash validity is exact, not a length cap", "[image][thumbhash]") std::vector alpha_hdr(5, std::byte{0}); alpha_hdr[2] = std::byte{0x80}; CHECK(thumbhash::expected_size(alpha_hdr) == std::nullopt); + + // A component count of 0 is not something any encoder emits, but it is one bit-field a peer + // controls, and the decoder's max(3, ...) clamp would otherwise wave it through while + // component_aspect_ratio -- which reads the field unclamped -- rejects it. valid() has to + // agree with the rest of the API about what it will accept, or it is not a trust boundary. + for (auto [w, h] : {std::pair{100, 43}, {43, 100}}) { + auto hash = thumbhash::encode(test_image(w, h, false), w, h); + REQUIRE(thumbhash::valid(hash)); + hash[3] &= std::byte{0xf8}; + INFO(w << "x" << h << " with the component count zeroed"); + CHECK_FALSE(thumbhash::valid(hash)); + CHECK(thumbhash::expected_size(hash) == std::nullopt); + } } TEST_CASE("thumbhash rejects bad input", "[image][thumbhash]") { @@ -271,3 +284,92 @@ TEST_CASE("thumbhash is bit-reproducible", "[image][thumbhash]") { CHECK(oxenc::to_hex(hash) == v.expected); } } + +// The vectors above pin the encoder only. Everything else that touches the decoder checks it +// against itself -- resolution independence compares two decodes, and the average-colour test +// compares two quantities that both derive from the same DC terms -- so a decoder that produced +// consistently wrong pixels would pass the whole suite. Pin its actual output too. +// +// Decoded at 8x6 so the expected bytes stay readable; that is enough pixels to exercise every +// branch of the reconstruction, including the alpha channel on the RGBA entries. +TEST_CASE("thumbhash decoder output is pinned", "[image][thumbhash]") { + struct { + int w, h; + bool alpha; + std::string_view expected; + } const vectors[] = { + {32, + 32, + false, + "1200ebff2d0cfaff4819faff6715f4ff9001efffbc00ecffed00f5ffff00ffff0017beff1a33d1ff" + "3840d5ff583aceff8122c9ffad00c5ffda00cafff800d1ff005f8eff1b80a8ff3c8cadff5a7da3ff" + "805c99ffaa3092ffd50b93fff20099ff009859ff26bd79ff4dc983ff6ab377ff8f8869ffb5515cff" + "df255bfffd1061ff01ad28ff31d24aff5adb54ff7ac149ffa0913bffc6532dfff0232bffff0d32ff" + "008a00ff12aa05ff39af0eff5d9506ff8a6900ffb82f00ffe70100ffff0004ff"sv}, + {64, + 48, + false, + "1100eaff2c0af8ff4617f9ff6614f3ff9001efffbd00edffec00f4ffff00feff001ac2ff1e36d5ff" + "3b44d8ff5a3cd0ff8223c9ffae00c5ffdc00ccfffb00d4ff005a89ff177ca3ff3989abff577ba1ff" + "7d5996ffa72d8fffd30891fff00097ff009c5dff2ac17cff50cc86ff6eb67aff928c6dffb95560ff" + "e1285efffe1162ff00ac27ff2fcf47ff56d751ff76bd46ff9d8e39ffc4522bffee222affff0c31ff" + "008a00ff13ab06ff3bb110ff5f9708ff8b6a00ffb82f00ffe70200ffff0005ff"sv}, + {48, + 64, + false, + "1400edff2b09f7ff491afcff6715f4ff8d00ecffbe00eeffef00f7ffff00feff0018bfff1730ceff" + "3a43d7ff5a3cd0ff7e20c6ffad00c4ffdb00cbfff800d0ff006190ff167ba3ff3e8eafff5d80a7ff" + "7e5a97ffa92f91ffd60c94fff20099ff009b5cff21b873ff4dc983ff6db67aff8c8666ffb6515dff" + "e1275dfffb0f60ff04b12cff2ccd45ff5bdc56ff7cc34cff9c8e38ffc7542efff3262effff0b30ff" + "008b00ff0da501ff3cb211ff609809ff856400ffb82f00ffeb0500ffff0002ff"sv}, + {32, + 32, + true, + "6c559e096d559c2d71549a49784f9850824696538e3b9649983196169e2a9700696195396a619366" + "6e609190755b8fa27f528ea98a478e9b953c8e629b358f2365738649677385816b7183b9726b81d3" + "7c6280da885680c6924b8185984582426681774d678176886c7f74c2737973db7d6f73de896373c6" + "935774859a5075426a896c2d6c886b6570856999777f68ab827568ab8d686997985c6a5e9e556b22" + "6e8b66006f8a650f7487643d7b81634b8676634b9269643d9c5d650ea3566600"sv}, + {64, + 48, + true, + "6b569e006d559c2171539a4c794e98568543965b9336964b9f299600a62197006664953368639378" + "6d6191b7755b8fd181508edb8e438ec69b368e71a22e8f126177864b6377859f687483f4706e81ff" + "7c6280ff8a5480ff964781a69e3f824160877751628676aa678374ff6f7c73ff7c7073ff8a6273ff" + "975474a59e4c7541648f6c22668e6b766b8a69c4738368df807768df8e6769c09b596a6ba3516b11" + "689166006a9065006f8c643a7884634e8477634e93686439a05a6500a7516600"sv}, + {100, + 100, + false, + "0e00e7ff2d0bf9ff4a1bfdff6a18f7ff9002efffbb00ebffeb00f3ffff00ffff0011b9ff1831cfff" + "3942d6ff5a3bd0ff8022c8ffaa00c2ffd700c7fff600cfff005c8bff1b80a7ff3f90b1ff5d81a7ff" + "815d9bffa92f91ffd40a92fff20099ff009758ff29c07cff53d089ff71b97eff938c6dffb8535eff" + "e1285dffff1464ff00ab26ff32d34bff5ee059ff7fc64effa2933effc6542dfff0232bffff0f33ff" + "008200ff0ea601ff38af0eff5d9506ff886600ffb32a00ffe20000ffff0001ff"sv}, + {17, + 5, + true, + "6e5b96006f5b940c725a931878569120804f902689469018913e9000963890006e628f266f628e49" + "72618c73775d8a927f568a9c894d898391448a42963f8a066d6e84486e6d8381716c81c9776780fb" + "7f607fff885780dc914e80879648813e6d79794e6e78788b727677d7777276ff806a75ff896075dd" + "915776869651773e6e8170156f807047727e6e8078796da081716da08a676e7d935d6f3c98587005" + "6e856c0070846b0073826a07797c6919817469188b6a6a0693606b00995a6c00"sv}, + {1, + 1, + false, + "ffff00ffffff00ff000000ff9a8affff8475ffff000000ffffff00ff4a7600ff707b00ff000000ff" + "000000ff0000ffff0000ffff000000ff000000ff293b00ff5547d8ff0000ffff0000ffffffffffff" + "ffffffff0000ffff0000ffff1007f6ff3a2fe7ff0000ffff0000ffffffffffffffffffff0000ffff" + "0000ffff0f07f6ff546500ff000000ff000000ff0000ffff0000ffff000000ff000000ff273900ff" + "517d00fff5ff00ff000000ff5d51ffff5c51ffff000000ffedff00ff426f00ff"sv}, + }; + for (const auto& v : vectors) { + auto hash = thumbhash::encode(test_image(v.w, v.h, v.alpha), v.w, v.h); + auto img = thumbhash::decode(hash, 8, 6); + REQUIRE(img.width == 8); + REQUIRE(img.height == 6); + REQUIRE(img.rgba.size() == 8 * 6 * 4); + INFO(v.w << "x" << v.h << (v.alpha ? " rgba" : " rgb")); + CHECK(oxenc::to_hex(img.rgba) == v.expected); + } +}